Record the lightnovelworld series-identity decision (#77)

A Series identity is discovered from the Site's own links, never derived
from an address. ADR-0008 carries the decision and the evidence behind
it: 3 of 41 sampled novels serve chapters under a slug that differs from
their series slug, one novel serves chapters under two, and neither slug
is computable from the other.

CONTEXT.md gains Chapter Slug as a term, defines it as plural by nature
and never an identity, and pins Series identity to the canonical slug the
Site publishes. Latest Chapter is restated as the highest-numbered
chapter rather than the newest-dated one, settling #79.

The research note is corrected in place where later probes refuted it:
the ul.clstyle container it named is the hidden empty "Latest Reading"
template rather than the chapter list, and its caveat about the comment
region understated the risk, since that region is writable by any
visitor and the scan takes an unbounded maximum into a shared row.

Docs only; no code. Implementation is specified in #80.
This commit is contained in:
2026-08-11 09:28:26 +07:00
parent 1ee5eb67ea
commit 0fe07d12f8
3 changed files with 692 additions and 8 deletions
+18 -8
View File
@@ -7,17 +7,24 @@ so progress survives across sites and devices.
## Language ## Language
**Series**: **Series**:
One ongoing work — a manga or a novel — as published by a Site. Identified by its One ongoing work — a manga or a novel — as published by a Site. Identified by the canonical
stable slug on that Site, never by its title. A Series exists once and is shared by slug the Site itself publishes for it, never by its title and never by a Chapter Slug. A
every Reader who bookmarks it; it owns the facts that are true regardless of who is Series exists once and is shared by every Reader who bookmarks it; it owns the facts that
reading — title, cover, Latest Chapter. A Reader cannot change them; they describe the are true regardless of who is reading — title, cover, Latest Chapter. A Reader cannot
Series, not anyone's relationship to it. change them; they describe the Series, not anyone's relationship to it.
_Avoid_: manga, title, book, comic _Avoid_: manga, title, book, comic
**Site**: **Site**:
One third-party source a Series is published on. A Series on two Sites is two Series. One third-party source a Series is published on. A Series on two Sites is two Series.
_Avoid_: source, host, provider, domain _Avoid_: source, host, provider, domain
**Chapter Slug**:
A slug a Site builds its chapter addresses from. Not an identity: one Series may have
several, any of them may differ from the slug that identifies the Series, and none is
computable from another. Only the Site's own links say which ones a Series uses, so a
Chapter Slug is always discovered, never derived.
_Avoid_: series slug, url slug, permalink, chapter path
**Cover**: **Cover**:
The image that stands for a Series wherever it is listed. A fact about the Series like The image that stands for a Series wherever it is listed. A fact about the Series like
its title — one Cover per Series, shared by every Reader, never per-Reader. Defined by its title — one Cover per Series, shared by every Reader, never per-Reader. Defined by
@@ -49,9 +56,12 @@ is real activity, so only Progress reorders the list.
_Avoid_: position, bookmark (the noun is taken), last read _Avoid_: position, bookmark (the noun is taken), last read
**Latest Chapter**: **Latest Chapter**:
The newest chapter a Site has published for a Series, discovered without the reader The highest-numbered chapter a Site has published for a Series, discovered without the
present. Distinct from Progress in every way that matters: it is a fact about the Site, reader present. The number is what ranks it, never a date and never the Site's own
not about the reader, and it must never reorder the list. "newest chapter" banner — where a Site disagrees with itself, its list of chapters is
the record and its summary of that list is not. Distinct from Progress in every way
that matters: it is a fact about the Site, not about the reader, and it must never
reorder the list.
_Avoid_: newest, current chapter, update _Avoid_: newest, current chapter, update
**Poll**: **Poll**:
@@ -0,0 +1,102 @@
# ADR-0008: A Series identity is discovered from the Site's links, never derived from an address
Date: 2026-08-11
Status: accepted
## Decision
On lightnovelworld, a Series is identified by the slug in its `/novel/<slug>/`
address, and the userscript obtains that address by reading the chapter page's
`a[aria-label="All Chapter"]` anchor. It no longer constructs the address by
string manipulation of the chapter path. When neither that anchor nor the
microdata breadcrumb is present, the page resolves to `type: "other"` and no
Bookmark is offered.
A Chapter Slug — the slug a chapter address is built from — is not an identity
and is not stored. The backend finds chapters by matching
`lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/` against the series page
body **truncated at the first `wpd-threads`**.
## Why
Measured on lightnovelworld between 2026-08-10 and 2026-08-11.
The chapter slug and the series slug are two independent facts. In a 41-novel
sample, 3 diverged (~7%). `/my-longevity-simulation-chapter-1/` returns 200
while `/novel/my-longevity-simulation/` returns 404 with no redirect, and that
novel's real address is `/novel/immortality-simulator/`. Divergence runs in both
directions: `the-sword-illuminates-the-great-wilderness` is served by chapter
slug `radiant-blade-of-the-wilderness`. Neither slug is computable from the
other, and the Site publishes no alternative-names field, so the mapping exists
only in the chapter page's own markup.
Storing the Chapter Slug beside the identity does not work, because a Series may
have more than one. `/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/`
serves chapters 1–99 under `…-one-skill-not` and 100–423 under
`…-one-skill-not-them-all`. Both resolve, and both chapter pages point back at
the same Series.
Deriving the identity from the chapter path also made one Series produce two
rows: bookmarking from the series page yielded `lightnovelworld:immortality-simulator`,
and from a chapter page `lightnovelworld:my-longevity-simulation`.
The pointer is reliable. Across 8 chapter pages — chapter 1, chapter 1200, the
latest chapter, both slugs of the split novel, two divergent novels, and a novel
with a number in its title — the `All Chapter` anchor and breadcrumb position 2
were both present and agreed every time, including on the old-slug pages. Three
narrower selectors were rejected on evidence: `a[href*="/novel/"]` matches the
header nav index first; matching the text "All Chapter" false-matches the novel
titled "…Not Them All Chapter 200"; and the JSON-LD breadcrumb's position 2 is
the chapter, not the Series.
The scan is truncated because a series page server-renders a wpdiscuz comment
thread below the chapter list, and comment bodies are HTML that can carry an
anchor. Verified on `/novel/the-sword-illuminates-the-great-wilderness/`:
comment `#wpd-comm-358_0` rendered in the initial HTML, corroborated by that
page's comment RSS feed. The scanner takes the maximum chapter number with no
upper bound, and a Series row is shared by every Reader (ADR-0003), so one
comment containing a link to a high-numbered chapter would pin that Series'
Latest Chapter for everyone. `wpd-threads` occurs exactly once per page and
follows every chapter anchor on all 4 series pages measured. `wpdcom` and
`wpdiscuz` are unusable: they occur 111 to 143 times per page, including in
`<head>` before the chapter list.
## Considered options
**Correct `series_url` only, leaving the Chapter Slug as the identity.** This
repairs the Poll with no migration, because the Poll reads the stored address
rather than the identity. Rejected: it keeps an identity that the Site does not
guarantee to be stable, and leaves the duplicate-row hazard in place.
**Scope the match to the chapter-list container.** Rejected on measurement. The
prior research named `ul.clstyle`; on all 4 pages sampled that is the hidden,
empty "Latest Reading" template, and the real list is a classless `<ul>` in
`div.eplister.eplisterfull`. The backend scans raw HTML with no parser, so
container extraction means a second regex against class names, which is more
fragile than the one-off truncation and protects nothing extra.
**Repair the stored rows with a SQL migration.** Rejected as impossible: the
database holds no source for the correct slug. The `series` table has no chapter
address, and no endpoint on the Site maps one slug to the other.
## Consequences
Existing Bookmarks on divergent novels stop matching their own chapter pages,
because `keyOf` changes. The userscript therefore migrates a row in place when
it sees the mismatch: it rewrites the row's key, identity and address in the
cache, the retry-queue entry, and the last-checked map, then syncs. This heals
only when the Reader next opens a chapter page of that novel.
A migrated row leaves its old `series` row behind. Nothing deletes it, but
`DueForLatestCheck` is an inner join against bookmarks, so a Series with no
Bookmarks is never polled again. The old row is permanently stored and
permanently inert.
A Bookmark whose Reader never opens a chapter page of that novel does not heal.
It continues to return 404 at every cooldown, as it does today.
If the truncation marker disappears, the scan is skipped and logged rather than
run against the whole page. An env-gated live test, `TestSmokeLnwCommentBoundary`,
asserts that `wpd-threads` still occurs exactly once and still follows the last
chapter anchor. It skips when its environment variable is unset, matching the
existing `TestSmokeKagane*` convention.
@@ -0,0 +1,572 @@
# lightnovelworld.net — chapter slug vs. series slug
Research note for Gitea issue #77. All pages fetched live on **2026-08-11** with
plain `curl` and a desktop UA. **No Cloudflare challenge was encountered on any
request** — every fetch below returned real HTML on the first try, so the
Playwright fallback was never needed.
Every claim carries the URL it came from. Nothing here is inferred from the
existing code; where a claim is an interpretation rather than an observation it
is marked `[INFERENCE]`.
---
## 1. Summary answer table
| Question | Answer | Evidence |
|---|---|---|
| Does the chapter slug equal the series slug? | **No, not reliably.** 3 of 41 sampled novels diverge at their latest chapter (~7%); a 4th diverges historically. | §4 |
| Best chapter-page → series-URL pointer | `a[aria-label='All Chapter']` and the microdata breadcrumb `a[itemprop=item]` at `position 2`. Both carry the full absolute series URL. | §3 |
| Is JSON-LD `BreadcrumbList` usable? | **No.** The Rank Math JSON-LD breadcrumb on a chapter page contains only *Home* and *the chapter itself* — the series is absent. | §3.1 |
| Is `<link rel=canonical>` usable? | **No.** It points at the chapter page itself. | §3.2 |
| Are `og:` tags usable? | **No.** `og:url` is the chapter page; no `og:` tag names the series. | §3.3 |
| Does the series page list chapters in the initial HTML? | **Yes — all of them.** 1232 unique chapters (1…1232) present in one document. | §5 |
| Is there a paginated chapter endpoint? | **No.** `/novel/<slug>/chapters/` and `/novel/<slug>/chapters/page-2` both **404**. | §5 |
| Does the series page carry *other* novels' chapter anchors? | **No.** Across 30 series pages, every `-chapter-N/` URL belonged to the page's own novel. | §6 |
| Unscoped regex `lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/` on a series page | **SAFE** | §6 |
| Does `/novel/<chapter-slug>/` 404? | **Yes**, 404, no redirect. | §7 |
| Is there a chapter-slug → series-slug endpoint that avoids parsing a chapter page? | **No endpoint found.** But `/<chapter-slug>/` (no `-chapter-N`) **302s to chapter 1**, which still requires parsing a chapter page. | §7 |
| Is there an "Alternative names" field explaining divergence? | **No such field exists** on lightnovelworld series pages. | §4.3 |
---
## 2. The two reference pages
| URL | Status |
|---|---|
| `https://lightnovelworld.net/my-longevity-simulation-chapter-1/` | **200**, 0 redirects |
| `https://lightnovelworld.net/novel/immortality-simulator/` | **200**, 0 redirects |
| `https://lightnovelworld.net/novel/my-longevity-simulation/` | **404**, 0 redirects |
Confirmed: chapter slug `my-longevity-simulation`, series slug
`immortality-simulator`, and the naive `/novel/<chapter-slug>/` construction is
a hard 404.
---
## 3. Chapter page → series URL: every in-page pointer, in priority order
All snippets below are verbatim from
`https://lightnovelworld.net/my-longevity-simulation-chapter-1/`.
### Priority 1 — `a[aria-label='All Chapter']` ✅ WORKS
Single occurrence in the document, inside the chapter navigation bar:
```html
<div class="nvs nvsc"><a href='https://lightnovelworld.net/novel/immortality-simulator/' aria-label='All Chapter'><i class="fas fa-list-ul"></i> All Chapter</a></div>
```
Note the **single quotes** on both attributes — a regex written for `href="` will
miss it. This is the most narrowly-targeted pointer: exactly one element on the
page has `aria-label='All Chapter'`.
### Priority 2 — microdata `BreadcrumbList` / `a[itemprop=item]` ✅ WORKS
At line 363–379 of the served HTML. `position 2` is the series:
```html
<div itemscope="" itemtype="http://schema.org/BreadcrumbList">
<span itemprop="itemListElement" itemscope="" itemtype="http://schema.org/ListItem">
<a itemprop="item" href="https://lightnovelworld.net/"><span itemprop="name">Home</span></a>
<meta itemprop="position" content="1">
</span>
›
<span itemprop="itemListElement" itemscope="" itemtype="http://schema.org/ListItem">
<a itemprop="item" href="https://lightnovelworld.net/novel/immortality-simulator/"><span itemprop="name">Immortality Simulator</span></a>
<meta itemprop="position" content="2">
</span>
›
<span itemprop="itemListElement" itemscope="" itemtype="http://schema.org/ListItem">
<a itemprop="item" href="https://lightnovelworld.net/my-longevity-simulation-chapter-1/"><span itemprop="name">My Longevity Simulation Chapter 1</span></a>
<meta itemprop="position" content="3">
</span>
</div>
```
Bonus: `<span itemprop="name">Immortality Simulator</span>` gives the clean
series **title** as well, which the userscript currently derives by stripping
`Chapter <n>` off `h1.entry-title`.
Caveat: `itemprop="item"` is **not** unique to the breadcrumb — the site header
nav also uses `itemprop="url"` on anchors, and unrelated `itemprop="item"` usage
could appear. Scope any selector to the enclosing
`[itemtype="http://schema.org/BreadcrumbList"]`.
### Priority 3 — corroborating `/novel/` anchors ⚠️ AMBIGUOUS
The chapter page contains exactly **8** distinct `/novel/…` hrefs. Two are
navigational (`/novel/` index), one is the series, and **five are a
"recommended" strip of unrelated novels**:
```
href="/novel/"
href="https://lightnovelworld.net/novel/"
href="https://lightnovelworld.net/novel/dreadful-radio-game/"
href="https://lightnovelworld.net/novel/evil-god-average/"
href="https://lightnovelworld.net/novel/immortality-simulator/"
href="https://lightnovelworld.net/novel/omniscient-first-persons-viewpoint/"
href="https://lightnovelworld.net/novel/reverend-insanity/"
href="https://lightnovelworld.net/novel/surviving-in-a-romance-fantasy-novel/"
```
So **"grab the first `/novel/<slug>/` link on a chapter page" is unsafe** — the
recommendation strip contaminates it. Use Priority 1 or 2.
### ❌ NOT usable — JSON-LD `BreadcrumbList` (§3.1)
There is exactly one `application/ld+json` block on the page (Rank Math SEO).
Verbatim:
```html
<script type="application/ld+json" class="rank-math-schema">{"@context":"https://schema.org","@graph":[{"@type":"BreadcrumbList","@id":"https://lightnovelworld.net/my-longevity-simulation-chapter-1/#breadcrumb","itemListElement":[{"@type":"ListItem","position":"1","item":{"@id":"https://lightnovelworld.net","name":"Home"}},{"@type":"ListItem","position":"2","item":{"@id":"https://lightnovelworld.net/my-longevity-simulation-chapter-1/","name":"My Longevity Simulation Chapter 1"}}]}]}</script>
```
Only two items — Home and the chapter. **The series does not appear.** The
JSON-LD breadcrumb and the visible microdata breadcrumb disagree; the microdata
one is the richer of the two. Do not use the JSON-LD.
### ❌ NOT usable — `<link rel=canonical>` (§3.2)
```html
<link rel="canonical" href="https://lightnovelworld.net/my-longevity-simulation-chapter-1/" />
```
Self-referential.
### ❌ NOT usable — `og:` meta tags (§3.3)
```html
<meta property="og:type" content="article" />
<meta property="og:title" content="Read My Longevity Simulation Chapter 1 Online Free - LightNovelworld" />
<meta property="og:url" content="https://lightnovelworld.net/my-longevity-simulation-chapter-1/" />
<meta property="og:site_name" content="Light Novel World" />
```
`og:url` is the chapter. `og:title` carries the **chapter-slug-derived** title
("My Longevity Simulation"), *not* the series title ("Immortality Simulator") —
so og tags actively reinforce the wrong name.
### Also present — `rel=next` / `rel=prev` chapter navigation
```html
<a aria-label="next" href="https://lightnovelworld.net/my-longevity-simulation-chapter-2/" rel="next">Next <i class="fa fa-angle-right" aria-hidden="true"></i></a>
```
Chapter 1 has no `rel=prev`. Useful for walking chapters, useless for finding
the series.
---
## 4. The reverse direction, and how common divergence is
### 4.1 Sample method
Two independent samples, deduplicated:
- 14 novels taken from the latest-updates listing on `https://lightnovelworld.net/`
(distinct chapter slugs), each chapter page fetched and its
`aria-label='All Chapter'` href read for the series slug.
- 30 series URLs taken from `https://lightnovelworld.net/novel/`, each series
page fetched and its chapter anchors' slug prefix extracted. Two of the 30
(`/novel/feed/` and `/novel/list-mode/`) are not novels and are excluded,
leaving 28.
One novel (`investing-in-my-crippled-wife-every-return-makes-me-stronger`)
appears in both samples. **Distinct novels sampled: 41.**
### 4.2 Divergence results
| Series slug (`/novel/<…>/`) | Chapter slug (`/<…>-chapter-N/`) | Result |
|---|---|---|
| `immortality-simulator` | `my-longevity-simulation` | **DIVERGE** |
| `the-sword-illuminates-the-great-wilderness` | `radiant-blade-of-the-wilderness` | **DIVERGE** |
| `a-villains-will-to-survive` | `the-villain-wants-to-live` | **DIVERGE** |
| `all-jobs-and-classes-i-just-wanted-one-skill-not-them-all` | `all-jobs-and-classes-i-just-wanted-one-skill-not` (ch. 1–99) **and** `all-jobs-and-classes-i-just-wanted-one-skill-not-them-all` (ch. 100–423) | **SPLIT** (latest matches) |
| `86-eighty-six` | same | match |
| `a-dungeon-beneath-my-house-let-me-gain-800-exp` | same | match |
| `a-journey-of-black-and-red` | same | match |
| `a-knight-who-eternally-regresses` | same | match |
| `a-regressors-tale-of-cultivation` | same | match |
| `a-will-eternal` | same | match |
| `absolute-resonance` | same | match |
| `absolute-sword-sense` | same | match |
| `advent-of-the-three-calamities` | same | match |
| `after-changing-to-the-ruthless-way-the-brothers-cried-and-begged-for-forgiveness` | same | match |
| `against-the-gods` | same | match |
| `apocalypse-i-built-the-infinite-train` | same | match |
| `arcane-exfil` | same | match |
| `as-a-mafia-boss-i-refuse-to-be-an-extra` | same | match |
| `ascendance-of-a-bookworm` | same | match |
| `avatar-conquering-the-elements` | same | match |
| `battle-world-ascending-without-limits` | same | match |
| `became-the-patron-of-villains` | same | match |
| `ending-maker` | same | match |
| `got-dropped-into-a-ghost-story-still-gotta-work` | same | match |
| `greed-all-for-what` | same | match |
| `investing-in-my-crippled-wife-every-return-makes-me-stronger` | same | match |
| `lord-of-mysteries-2-circle-of-inevitability` | same | match |
| `lord-of-the-mysteries` | same | match |
| `my-vampire-system` | same | match |
| `cleaver-of-sin` | same | match |
| `greatest-legacy-of-the-magus-universe` | same | match |
| `magus-infinite` | same | match |
| `path-of-the-extra` | same | match |
| `regnum-aetern-dual-rebirth` | same | match |
| `shadow-slave` | same | match |
| `slime-evolution` | same | match |
| `sss-awakening-i-can-class-change-at-will` | same | match |
| `the-dao-of-reincarnation-ascending-through-a-hundred-lives` | same | match |
| `the-gamers-pov` | same | match |
| `the-insane-regressor-throne-of-pride` | same | match |
| `the-villains-pov` | same | match |
**Counts (41 distinct novels):**
- **37 match** (90.2%)
- **3 diverge at the latest chapter** (7.3%) — these are the ones that break the poller *today*
- **1 splits mid-series** but its newest chapters happen to match the series slug (2.4%)
`[INFERENCE]` The true site-wide divergence rate is probably in the same
ballpark, but 41 novels out of a catalogue of thousands is a small sample and
the listing page is alphabetically front-loaded. Treat ~7% as an order of
magnitude, not a precise figure.
### 4.3 Is there a derivable rule? **No.**
The divergence is **not directional**, so you cannot compute one slug from the
other:
- `immortality-simulator` — the *series* carries the polished English title
(`<h1 class="entry-title" itemprop="name">` is absent from the chapter page's
own naming; `og:title` says "My Longevity Simulation"), while the *chapters*
carry the literal translation.
- `the-sword-illuminates-the-great-wilderness` — the reverse: the *series* has
the literal title
(`<h1 class="entry-title" itemprop="name">The Sword Illuminates the Great Wilderness</h1>`,
`<title>Read The Sword Illuminates the Great Wilderness English Online Free - LightNovelworld</title>`)
while the *chapters* carry the polished one (`radiant-blade-of-the-wilderness`).
- `a-villains-will-to-survive` —
`<h1 class="entry-title" itemprop="name">A Villain’s Will to Survive</h1>`,
chapters at `the-villain-wants-to-live-chapter-N/`.
`[INFERENCE]` The consistent explanation is that a novel is retitled after
publication; WordPress updates the series post's slug but leaves the already-published
chapter posts' slugs alone. The direction of the retitle varies per novel, which
is why no rule exists. This is consistent with the split case in §4.4, but the
site exposes no field that states it.
**There is no "Alternative names" / "Associated names" field.** Scanning
`https://lightnovelworld.net/novel/a-villains-will-to-survive/` and
`https://lightnovelworld.net/novel/immortality-simulator/` for such a label
returns nothing; the series info panel exposes only **Author**, **Released**,
**Type**, **Status**. (The only `names?` match in the document is the `>Name*<`
label on the comment form.) So the old title is **not** recoverable from the
series page — the mapping only exists in the chapter anchors themselves.
### 4.4 The split case — a slug can change *mid-series*
`https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/`
carries **two** chapter slugs in one page:
| Chapter slug prefix | Anchors | Chapter range |
|---|---|---|
| `all-jobs-and-classes-i-just-wanted-one-skill-not` | 100 | 1 – 99 |
| `all-jobs-and-classes-i-just-wanted-one-skill-not-them-all` | 325 | 100 – 423 |
Both resolve, and **both point back at the same series**:
```
200 https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-chapter-1/
<a href='https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/' aria-label='All Chapter'>
200 https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-404/
<a href='https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/' aria-label='All Chapter'>
```
**Consequence:** a single stored chapter slug is not a sufficient key even for
one novel. A fix that captures "the chapter slug at bookmark time" and pins the
poller regex to it would still miss chapters published under a later slug. This
is the strongest argument for the unscoped regex over a stored-slug regex.
---
## 5. Series page → chapter list
From `https://lightnovelworld.net/novel/immortality-simulator/` (200, 655 037 bytes):
- **1234** anchors matching `href="https://lightnovelworld.net/<slug>-chapter-<n>/"`.
- **1232 unique** chapter URLs, covering chapter numbers **1 through 1232 with no gaps**.
- The 2 extras are the First/Last shortcut widget, which duplicates two chapters:
```html
<h2>Read Immortality Simulator</h2></div>
<div class="lastend">
<div class="inepcx">
<a href="https://lightnovelworld.net/my-longevity-simulation-chapter-1/">
<span>First Chapter</span>
<span class="epcur epcurfirst">Vol. 1 Ch. 1</span>
</a>
</div>
<div class="inepcx">
<a href="https://lightnovelworld.net/my-longevity-simulation-chapter-1232/">
```
**The entire chapter list is in the initial HTML.** No JS hydration, no
pagination, no separate endpoint. Confirmed by probing the shapes the issue
speculated about:
| URL | Status |
|---|---|
| `https://lightnovelworld.net/novel/immortality-simulator/chapters/` | **404** |
| `https://lightnovelworld.net/novel/immortality-simulator/chapters/page-2` | **404** |
The only `page-numbers` / pagination markup in the document belongs to
`class="wpdiscuz-comment-pagination"` (the comment widget), not the chapter list.
This holds for large novels too: `greed-all-for-what` served 2666 chapter
anchors and `my-vampire-system` 2547, all inline in one response.
**All chapter hrefs are absolute** (`https://lightnovelworld.net/…`). A grep for
relative `href="/<slug>-chapter-N/"` on the series page returned **0** matches,
so a pattern anchored on the literal host is safe.
---
## 6. Is the unscoped chapter regex SAFE on a series page?
### Verdict: **SAFE**
Proposed pattern: `lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/`
**Proof.** For each of the 30 series pages fetched from
`https://lightnovelworld.net/novel/`, every occurrence of `<slug>-chapter-<n>/`
anywhere in the document (href or not) was reduced to its slug prefix and
deduplicated:
- **27 of 28 real series pages had exactly 1 distinct chapter-slug prefix**, equal
to that novel's own chapter slug.
- **1 page had 2 prefixes** — `all-jobs-and-classes-…`, and §4.4 proves *both
belong to that same novel*. Not foreign.
- **0 pages carried any other novel's chapter URL.**
The recommendation and sidebar widgets on a series page link to **series** URLs
only, never chapter URLs. On
`https://lightnovelworld.net/novel/immortality-simulator/` the recommendation
strip renders as:
```html
<a href="https://lightnovelworld.net/novel/reverend-insanity/" itemprop="url" title="Reverend Insanity" class="tip" rel="292">
```
— `/novel/<slug>/`, which the pattern cannot match.
The one widget that *would* carry foreign chapter links, "Latest Reading", is an
**empty client-side template**. Its container is `display:none` and its `<ul>` is
empty in the served HTML; the row markup lives in an inert
`<span id="series-history-tpl" style='display:none'>` with mustache placeholders
and a `href="#/{{number}}"` — no real chapter URL:
```html
<div class="bixbox bxcl" id="series-history" style="display:none;">
<div class="releases"><h2>Latest Reading</h2></div>
<div class="series-history-pool">
<ul class="clstyle" id="series-history-ul"></ul>
</div>
</div>
<span id="series-history-tpl" style='display:none'>
<li data-id="{{id}}" data-num="{{chapter}}" data-vol="{{volume}}" data-title="{{title}}">
<div class="chbox"><div class="eph-num">
<a onclick="…return series_history.redirect({{id}});" href="#/{{number}}">
```
It is populated from the visitor's own local history, so a **server-side fetch
never sees content there** — the poller is immune. A browser-rendered fetch with
a fresh profile is likewise immune (no history to render).
### Caveats to record with the verdict
1. **`[INFERENCE]` This is an empirical guarantee, not a structural one.** Unlike
the asura/novelfull cases where scoping to the slug makes foreign contamination
*impossible*, here we only know that lightnovelworld's series template does not
currently emit foreign chapter anchors. If the theme ever adds a "latest site
updates" strip rendered server-side, the unscoped pattern breaks silently and
in the worst direction (a foreign chapter number *higher* than the real one
wins the maximum and the bookmark shows a phantom update).
**Correction, 2026-08-11.** This caveat understated the risk. A series page
server-renders a wpdiscuz comment thread below the chapter list, and comment
bodies are HTML that can carry an `<a href>`. Verified on
`/novel/the-sword-illuminates-the-great-wilderness/`: `#wpdcom` present,
comment `#wpd-comm-358_0` by "hasbi asy" rendered as `<p>where's everyone</p>`
inside `.wpd-comment-text`, corroborated by that page's comment RSS feed;
`firstLoadWithAjax: 0` confirms the thread is in the initial HTML. The sampled
pages carried 0 or 1 comments and no links, which is why §6 read clean. So the
unscoped pattern does not merely depend on the *theme* staying unchanged — it
reads a region any visitor can write to, and the scanner takes the maximum with
no upper bound. The scan must stop before the comment thread.
2. ~~Consider scoping the match to the chapter-list container rather than the whole
document, which would restore the structural guarantee at low cost. The list
items sit under `ul.clstyle` (`class="bixbox bxcl"`).~~
**Refuted, 2026-08-11**, measured on 4 series pages
(`the-sword-illuminates-the-great-wilderness`, `all-jobs-…-them-all`,
`a-will-eternal`, `immortality-simulator`). The single `ul.clstyle` per page is
the hidden, empty "Latest Reading" template (`#series-history-ul`, see §5), not
the chapter list — scoping to it would match nothing. The real chapter list is a
classless `<ul>` inside `div.eplister.eplisterfull`, in
`div.bixbox.bxcl.epcheck`.
The workable boundary is `wpd-threads`: it occurs exactly once per page, and on
all 4 pages every chapter anchor precedes it and no chapter-shaped href follows
it. `wpdcom` and `wpdiscuz` are unusable as markers — they occur 111 to 143
times per page, including in `<head>` roughly 15 KB *before* the chapter list,
so cutting there would discard the list itself. `id='comments'` also occurs once
but is single-quoted; the canonical `id="comments"` never appears.
3. `[0-9.]+` is fine: no decimal chapter numbers appeared anywhere in the sampled
pages, but the existing `strings.Trim(m[1], ".")` guard already handles them.
4. The current test fixture `lnwSeriesFixture` in
`backend/internal/latest/sites_test.go` includes
`<a href="https://lightnovelworld.net/overgeared-chapter-9999/">` and asserts it
is ignored. **That anchor is not representative of a real series page** — no
sampled page contained a foreign chapter anchor. Relaxing the regex will make
that assertion fail, and the correct response is to fix the fixture, not to
keep the scoping.
---
## 7. Redirects and reverse-lookup endpoints
| URL | Status | Redirects | Final |
|---|---|---|---|
| `https://lightnovelworld.net/novel/my-longevity-simulation/` | **404** | 0 | — |
| `https://lightnovelworld.net/novel/radiant-blade-of-the-wilderness/` | **404** | 0 | — |
| `https://lightnovelworld.net/immortality-simulator-chapter-1/` | **404** | 0 | — |
| `https://lightnovelworld.net/my-longevity-simulation/` | **200** | **1** | `https://lightnovelworld.net/my-longevity-simulation-chapter-1/` |
Two findings:
1. **`/novel/<chapter-slug>/` genuinely 404s.** Confirmed for both
`my-longevity-simulation` and `radiant-blade-of-the-wilderness`. There is no
alias, no 301, no fallback. Any stored series URL built by that construction is
permanently dead.
2. **The chapter slug alone (no `-chapter-N`) 302s to chapter 1** of that novel.
This is the closest thing to a reverse-lookup endpoint, but it lands on a
*chapter page*, so recovering the series URL still requires parsing that page's
`aria-label='All Chapter'` anchor or breadcrumb. **No endpoint maps chapter slug
→ series slug directly.**
3. The inverse construction also fails: `/<series-slug>-chapter-1/` 404s for a
divergent novel, so you cannot probe your way from a series slug to a chapter
slug either. The chapter slug must be read off the series page's anchors.
Also present but not a lookup path: the site is WordPress and exposes
`https://lightnovelworld.net/wp-json/wp/v2/posts/177695` (advertised via
`<link rel="alternate" title="JSON" type="application/json" …>` on the chapter
page). Whether the REST API exposes a chapter→series relation was **not tested**
— out of scope for this note, and it would still require fetching the chapter
page to learn the post ID.
---
## 8. Implications for issue #77
> The issue text itself could not be read: `gh` is not installed in this
> environment, so `issue://77` failed to resolve. The proposed regex quoted below
> is the one supplied in the task brief.
### 8.1 The bug
`backend/internal/latest/sites.go:128-137` derives the chapter-anchor pattern
from the **series slug**:
```go
m := lnwSlugRe.FindStringSubmatch(u.Path) // /novel/<series-slug>/
re = regexp.MustCompile(`lightnovelworld\.net/` + regexp.QuoteMeta(m[1]) + `-chapter-([0-9.]+)/`)
```
For `immortality-simulator` this compiles to a pattern matching
`…net/immortality-simulator-chapter-N/`, which appears **zero times** on the
series page — the anchors all say `my-longevity-simulation`. `latestChapterFrom`
returns `ok == false`, indistinguishable from a Cloudflare challenge or a site
redesign, and the bookmark silently stops tracking updates. Same for
`the-sword-illuminates-the-great-wilderness` and `a-villains-will-to-survive`.
### 8.2 The proposed fix is sound
Dropping the scoping to
`lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/` is sound in shape and fixes
all three currently-broken novels plus the split case in §4.4 — which no
stored-slug approach can fix, since that novel legitimately has two chapter
slugs. It must not be applied to the whole document, however: truncate the body
at `wpd-threads` first, per the corrected caveat 2 in §6. The `ul.clstyle`
suggestion originally recorded here is refuted.
Also worth removing: the `overgeared-chapter-9999` line in `lnwSeriesFixture`
and the `"lightnovelworld takes the max and ignores another series"` test case,
which encode a contamination scenario that §6 shows does not occur on this site.
Replacing them with a fixture that mirrors §4.4's two-slugs-one-novel reality
would be a better regression test.
### 8.3 The userscript has the same bug, and it is worse
`userscript/novel-bookmark.user.js` builds the series URL by string-concatenating
the chapter slug (lines 154 and 169):
```js
seriesUrl: "https://lightnovelworld.net/novel/" + m[1] + "/",
```
For a divergent novel this **writes a permanently-404 series URL into the
database at bookmark time**. Fixing only the backend regex leaves those rows
broken, and `fetchableSeriesURL` will happily accept the dead URL because it only
checks the hostname (`sites.go:321-324`).
The userscript should read the series URL off the page instead of constructing
it. On a chapter page, prefer in this order (§3):
```js
document.querySelector("a[aria-label='All Chapter']")?.href
?? document.querySelector("[itemtype$='BreadcrumbList'] a[itemprop=item][href*='/novel/']")?.href
```
The breadcrumb's `<span itemprop="name">` also supplies the correct series title
directly, replacing the current `h1.entry-title` + strip-`Chapter <n>` heuristic
— which today yields "My Longevity Simulation" where the series is actually
titled "Immortality Simulator".
A backfill for already-stored rows is likely needed: any `lightnovelworld` row
whose `series_url` 404s can be repaired by fetching
`https://lightnovelworld.net/<series_id>-chapter-1/` (or `/<series_id>/`, which
302s there per §7) and reading its `All Chapter` anchor.
### 8.4 Note on `series_id`
Rows are keyed `lightnovelworld:<series_id>` where `series_id` is currently the
**chapter slug** (from the userscript's `m[1]`). If the fix changes `series_id` to
the real series slug, existing divergent bookmarks change key and need migrating.
Whether that is worth it depends on whether `series_id` is used to rebuild URLs
anywhere — `sites.go` uses the stored `series_url`, not `series_id`, for
lightnovelworld, so storing the correct `series_url` may be sufficient without a
re-key.
---
## Reproduction
```sh
UA='Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0 Safari/537.36'
# §2 status codes
curl -sS -o /dev/null -w '%{http_code} %{url_effective}\n' -L -A "$UA" \
https://lightnovelworld.net/my-longevity-simulation-chapter-1/ \
https://lightnovelworld.net/novel/immortality-simulator/ \
https://lightnovelworld.net/novel/my-longevity-simulation/
# §3 the two working pointers
curl -sS -A "$UA" https://lightnovelworld.net/my-longevity-simulation-chapter-1/ \
| grep -E "aria-label='All Chapter'|itemprop=\"item\""
# §6 SAFE proof — distinct chapter-slug prefixes per series page (expect 1)
curl -sS -A "$UA" https://lightnovelworld.net/novel/immortality-simulator/ \
| grep -oE '[a-z0-9-]+-chapter-[0-9.]+/' | sed -E 's/-chapter-[0-9.]+\///' | sort | uniq -c
```