Record the lightnovelworld series-identity decision (#77) (#81)

Docs only. No code, no tests, nothing to run. Implementation is specified in #80.

Outcome of a grilling session on 2026-08-11 against #77, backed by live measurement of lightnovelworld over 2026-08-10/11.

## What changed

**`docs/adr/0008-series-identity-is-discovered-not-derived.md`** (new)

A Series identity is discovered from the Site's own links, never derived from an address.
On lightnovelworld the userscript reads the chapter page's `All Chapter` anchor instead of
building a `/novel/<slug>/` address by string manipulation. A Chapter Slug is not an
identity and is not stored. The backend's chapter scan drops its per-Series scoping and
runs against the body truncated before the visitor comment thread.

Evidence in the ADR: 3 of 41 sampled novels serve chapters under a slug that differs from
their series slug, divergence runs in both directions, one novel serves chapters under two
slugs, and neither slug is computable from the other. The pointer was checked on 8 chapter
pages and agreed every time. Three narrower selectors are recorded as rejected, each with
the measurement that killed it.

Three rejected options are recorded with reasons: correcting the stored address only, which
keeps an identity the Site does not guarantee; scoping the scan to a container, which the
probe refuted; and a SQL migration, which is impossible because the database holds no
source for the correct slug.

**`CONTEXT.md`**

- **Series** - identity is the canonical slug the Site publishes, never the title and never a Chapter Slug.
- **Chapter Slug** - new term. A slug a Site builds its chapter addresses from. Not an identity: one Series may have several, and none is computable from another.
- **Latest Chapter** - now the highest-numbered chapter, explicitly not a date and not the Site's own newest-chapter banner. Settles #79.

**`docs/research/lightnovelworld-chapter-vs-series-slug.md`** (new, committed with its corrections)

The 41-novel survey behind the ADR. Two claims are struck through and corrected in place,
with the date and sample size of the probe that refuted each: the `ul.clstyle` container it
named is the hidden, empty "Latest Reading" template rather than the chapter list, and its
caveat about the comment region understated the risk, because that region is writable by
any visitor while the scan takes an unbounded maximum into a Series row shared by every
Reader (ADR-0003).

## Review notes

Nothing here constrains code that exists today - the ADR describes work not yet written.
The part worth disagreeing with, if any of it is wrong, is the fail-closed rule: a missing
truncation marker means skip the Series and log, never scan the whole page.

Related: #77 (the defect), #80 (the spec), #79 (the numbering anomaly, closed by decision),
#71 (the same size cap seen from the cover side).

Reviewed-on: #81
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
This commit was merged in pull request #81.
This commit is contained in:
2026-08-11 09:34:26 +07:00
committed by sulthan
parent 1ee5eb67ea
commit f1eb7d514c
3 changed files with 692 additions and 8 deletions
@@ -0,0 +1,572 @@
# lightnovelworld.net — chapter slug vs. series slug
Research note for Gitea issue #77. All pages fetched live on **2026-08-11** with
plain `curl` and a desktop UA. **No Cloudflare challenge was encountered on any
request** — every fetch below returned real HTML on the first try, so the
Playwright fallback was never needed.
Every claim carries the URL it came from. Nothing here is inferred from the
existing code; where a claim is an interpretation rather than an observation it
is marked `[INFERENCE]`.
---
## 1. Summary answer table
| Question | Answer | Evidence |
|---|---|---|
| Does the chapter slug equal the series slug? | **No, not reliably.** 3 of 41 sampled novels diverge at their latest chapter (~7%); a 4th diverges historically. | §4 |
| Best chapter-page → series-URL pointer | `a[aria-label='All Chapter']` and the microdata breadcrumb `a[itemprop=item]` at `position 2`. Both carry the full absolute series URL. | §3 |
| Is JSON-LD `BreadcrumbList` usable? | **No.** The Rank Math JSON-LD breadcrumb on a chapter page contains only *Home* and *the chapter itself* — the series is absent. | §3.1 |
| Is `<link rel=canonical>` usable? | **No.** It points at the chapter page itself. | §3.2 |
| Are `og:` tags usable? | **No.** `og:url` is the chapter page; no `og:` tag names the series. | §3.3 |
| Does the series page list chapters in the initial HTML? | **Yes — all of them.** 1232 unique chapters (1…1232) present in one document. | §5 |
| Is there a paginated chapter endpoint? | **No.** `/novel/<slug>/chapters/` and `/novel/<slug>/chapters/page-2` both **404**. | §5 |
| Does the series page carry *other* novels' chapter anchors? | **No.** Across 30 series pages, every `-chapter-N/` URL belonged to the page's own novel. | §6 |
| Unscoped regex `lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/` on a series page | **SAFE** | §6 |
| Does `/novel/<chapter-slug>/` 404? | **Yes**, 404, no redirect. | §7 |
| Is there a chapter-slug → series-slug endpoint that avoids parsing a chapter page? | **No endpoint found.** But `/<chapter-slug>/` (no `-chapter-N`) **302s to chapter 1**, which still requires parsing a chapter page. | §7 |
| Is there an "Alternative names" field explaining divergence? | **No such field exists** on lightnovelworld series pages. | §4.3 |
---
## 2. The two reference pages
| URL | Status |
|---|---|
| `https://lightnovelworld.net/my-longevity-simulation-chapter-1/` | **200**, 0 redirects |
| `https://lightnovelworld.net/novel/immortality-simulator/` | **200**, 0 redirects |
| `https://lightnovelworld.net/novel/my-longevity-simulation/` | **404**, 0 redirects |
Confirmed: chapter slug `my-longevity-simulation`, series slug
`immortality-simulator`, and the naive `/novel/<chapter-slug>/` construction is
a hard 404.
---
## 3. Chapter page → series URL: every in-page pointer, in priority order
All snippets below are verbatim from
`https://lightnovelworld.net/my-longevity-simulation-chapter-1/`.
### Priority 1 — `a[aria-label='All Chapter']` ✅ WORKS
Single occurrence in the document, inside the chapter navigation bar:
```html
<div class="nvs nvsc"><a href='https://lightnovelworld.net/novel/immortality-simulator/' aria-label='All Chapter'><i class="fas fa-list-ul"></i> All Chapter</a></div>
```
Note the **single quotes** on both attributes — a regex written for `href="` will
miss it. This is the most narrowly-targeted pointer: exactly one element on the
page has `aria-label='All Chapter'`.
### Priority 2 — microdata `BreadcrumbList` / `a[itemprop=item]` ✅ WORKS
At line 363–379 of the served HTML. `position 2` is the series:
```html
<div itemscope="" itemtype="http://schema.org/BreadcrumbList">
<span itemprop="itemListElement" itemscope="" itemtype="http://schema.org/ListItem">
<a itemprop="item" href="https://lightnovelworld.net/"><span itemprop="name">Home</span></a>
<meta itemprop="position" content="1">
</span>
›
<span itemprop="itemListElement" itemscope="" itemtype="http://schema.org/ListItem">
<a itemprop="item" href="https://lightnovelworld.net/novel/immortality-simulator/"><span itemprop="name">Immortality Simulator</span></a>
<meta itemprop="position" content="2">
</span>
›
<span itemprop="itemListElement" itemscope="" itemtype="http://schema.org/ListItem">
<a itemprop="item" href="https://lightnovelworld.net/my-longevity-simulation-chapter-1/"><span itemprop="name">My Longevity Simulation Chapter 1</span></a>
<meta itemprop="position" content="3">
</span>
</div>
```
Bonus: `<span itemprop="name">Immortality Simulator</span>` gives the clean
series **title** as well, which the userscript currently derives by stripping
`Chapter <n>` off `h1.entry-title`.
Caveat: `itemprop="item"` is **not** unique to the breadcrumb — the site header
nav also uses `itemprop="url"` on anchors, and unrelated `itemprop="item"` usage
could appear. Scope any selector to the enclosing
`[itemtype="http://schema.org/BreadcrumbList"]`.
### Priority 3 — corroborating `/novel/` anchors ⚠️ AMBIGUOUS
The chapter page contains exactly **8** distinct `/novel/…` hrefs. Two are
navigational (`/novel/` index), one is the series, and **five are a
"recommended" strip of unrelated novels**:
```
href="/novel/"
href="https://lightnovelworld.net/novel/"
href="https://lightnovelworld.net/novel/dreadful-radio-game/"
href="https://lightnovelworld.net/novel/evil-god-average/"
href="https://lightnovelworld.net/novel/immortality-simulator/"
href="https://lightnovelworld.net/novel/omniscient-first-persons-viewpoint/"
href="https://lightnovelworld.net/novel/reverend-insanity/"
href="https://lightnovelworld.net/novel/surviving-in-a-romance-fantasy-novel/"
```
So **"grab the first `/novel/<slug>/` link on a chapter page" is unsafe** — the
recommendation strip contaminates it. Use Priority 1 or 2.
### ❌ NOT usable — JSON-LD `BreadcrumbList` (§3.1)
There is exactly one `application/ld+json` block on the page (Rank Math SEO).
Verbatim:
```html
<script type="application/ld+json" class="rank-math-schema">{"@context":"https://schema.org","@graph":[{"@type":"BreadcrumbList","@id":"https://lightnovelworld.net/my-longevity-simulation-chapter-1/#breadcrumb","itemListElement":[{"@type":"ListItem","position":"1","item":{"@id":"https://lightnovelworld.net","name":"Home"}},{"@type":"ListItem","position":"2","item":{"@id":"https://lightnovelworld.net/my-longevity-simulation-chapter-1/","name":"My Longevity Simulation Chapter 1"}}]}]}</script>
```
Only two items — Home and the chapter. **The series does not appear.** The
JSON-LD breadcrumb and the visible microdata breadcrumb disagree; the microdata
one is the richer of the two. Do not use the JSON-LD.
### ❌ NOT usable — `<link rel=canonical>` (§3.2)
```html
<link rel="canonical" href="https://lightnovelworld.net/my-longevity-simulation-chapter-1/" />
```
Self-referential.
### ❌ NOT usable — `og:` meta tags (§3.3)
```html
<meta property="og:type" content="article" />
<meta property="og:title" content="Read My Longevity Simulation Chapter 1 Online Free - LightNovelworld" />
<meta property="og:url" content="https://lightnovelworld.net/my-longevity-simulation-chapter-1/" />
<meta property="og:site_name" content="Light Novel World" />
```
`og:url` is the chapter. `og:title` carries the **chapter-slug-derived** title
("My Longevity Simulation"), *not* the series title ("Immortality Simulator") —
so og tags actively reinforce the wrong name.
### Also present — `rel=next` / `rel=prev` chapter navigation
```html
<a aria-label="next" href="https://lightnovelworld.net/my-longevity-simulation-chapter-2/" rel="next">Next <i class="fa fa-angle-right" aria-hidden="true"></i></a>
```
Chapter 1 has no `rel=prev`. Useful for walking chapters, useless for finding
the series.
---
## 4. The reverse direction, and how common divergence is
### 4.1 Sample method
Two independent samples, deduplicated:
- 14 novels taken from the latest-updates listing on `https://lightnovelworld.net/`
(distinct chapter slugs), each chapter page fetched and its
`aria-label='All Chapter'` href read for the series slug.
- 30 series URLs taken from `https://lightnovelworld.net/novel/`, each series
page fetched and its chapter anchors' slug prefix extracted. Two of the 30
(`/novel/feed/` and `/novel/list-mode/`) are not novels and are excluded,
leaving 28.
One novel (`investing-in-my-crippled-wife-every-return-makes-me-stronger`)
appears in both samples. **Distinct novels sampled: 41.**
### 4.2 Divergence results
| Series slug (`/novel/<…>/`) | Chapter slug (`/<…>-chapter-N/`) | Result |
|---|---|---|
| `immortality-simulator` | `my-longevity-simulation` | **DIVERGE** |
| `the-sword-illuminates-the-great-wilderness` | `radiant-blade-of-the-wilderness` | **DIVERGE** |
| `a-villains-will-to-survive` | `the-villain-wants-to-live` | **DIVERGE** |
| `all-jobs-and-classes-i-just-wanted-one-skill-not-them-all` | `all-jobs-and-classes-i-just-wanted-one-skill-not` (ch. 1–99) **and** `all-jobs-and-classes-i-just-wanted-one-skill-not-them-all` (ch. 100–423) | **SPLIT** (latest matches) |
| `86-eighty-six` | same | match |
| `a-dungeon-beneath-my-house-let-me-gain-800-exp` | same | match |
| `a-journey-of-black-and-red` | same | match |
| `a-knight-who-eternally-regresses` | same | match |
| `a-regressors-tale-of-cultivation` | same | match |
| `a-will-eternal` | same | match |
| `absolute-resonance` | same | match |
| `absolute-sword-sense` | same | match |
| `advent-of-the-three-calamities` | same | match |
| `after-changing-to-the-ruthless-way-the-brothers-cried-and-begged-for-forgiveness` | same | match |
| `against-the-gods` | same | match |
| `apocalypse-i-built-the-infinite-train` | same | match |
| `arcane-exfil` | same | match |
| `as-a-mafia-boss-i-refuse-to-be-an-extra` | same | match |
| `ascendance-of-a-bookworm` | same | match |
| `avatar-conquering-the-elements` | same | match |
| `battle-world-ascending-without-limits` | same | match |
| `became-the-patron-of-villains` | same | match |
| `ending-maker` | same | match |
| `got-dropped-into-a-ghost-story-still-gotta-work` | same | match |
| `greed-all-for-what` | same | match |
| `investing-in-my-crippled-wife-every-return-makes-me-stronger` | same | match |
| `lord-of-mysteries-2-circle-of-inevitability` | same | match |
| `lord-of-the-mysteries` | same | match |
| `my-vampire-system` | same | match |
| `cleaver-of-sin` | same | match |
| `greatest-legacy-of-the-magus-universe` | same | match |
| `magus-infinite` | same | match |
| `path-of-the-extra` | same | match |
| `regnum-aetern-dual-rebirth` | same | match |
| `shadow-slave` | same | match |
| `slime-evolution` | same | match |
| `sss-awakening-i-can-class-change-at-will` | same | match |
| `the-dao-of-reincarnation-ascending-through-a-hundred-lives` | same | match |
| `the-gamers-pov` | same | match |
| `the-insane-regressor-throne-of-pride` | same | match |
| `the-villains-pov` | same | match |
**Counts (41 distinct novels):**
- **37 match** (90.2%)
- **3 diverge at the latest chapter** (7.3%) — these are the ones that break the poller *today*
- **1 splits mid-series** but its newest chapters happen to match the series slug (2.4%)
`[INFERENCE]` The true site-wide divergence rate is probably in the same
ballpark, but 41 novels out of a catalogue of thousands is a small sample and
the listing page is alphabetically front-loaded. Treat ~7% as an order of
magnitude, not a precise figure.
### 4.3 Is there a derivable rule? **No.**
The divergence is **not directional**, so you cannot compute one slug from the
other:
- `immortality-simulator` — the *series* carries the polished English title
(`<h1 class="entry-title" itemprop="name">` is absent from the chapter page's
own naming; `og:title` says "My Longevity Simulation"), while the *chapters*
carry the literal translation.
- `the-sword-illuminates-the-great-wilderness` — the reverse: the *series* has
the literal title
(`<h1 class="entry-title" itemprop="name">The Sword Illuminates the Great Wilderness</h1>`,
`<title>Read The Sword Illuminates the Great Wilderness English Online Free - LightNovelworld</title>`)
while the *chapters* carry the polished one (`radiant-blade-of-the-wilderness`).
- `a-villains-will-to-survive` —
`<h1 class="entry-title" itemprop="name">A Villain’s Will to Survive</h1>`,
chapters at `the-villain-wants-to-live-chapter-N/`.
`[INFERENCE]` The consistent explanation is that a novel is retitled after
publication; WordPress updates the series post's slug but leaves the already-published
chapter posts' slugs alone. The direction of the retitle varies per novel, which
is why no rule exists. This is consistent with the split case in §4.4, but the
site exposes no field that states it.
**There is no "Alternative names" / "Associated names" field.** Scanning
`https://lightnovelworld.net/novel/a-villains-will-to-survive/` and
`https://lightnovelworld.net/novel/immortality-simulator/` for such a label
returns nothing; the series info panel exposes only **Author**, **Released**,
**Type**, **Status**. (The only `names?` match in the document is the `>Name*<`
label on the comment form.) So the old title is **not** recoverable from the
series page — the mapping only exists in the chapter anchors themselves.
### 4.4 The split case — a slug can change *mid-series*
`https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/`
carries **two** chapter slugs in one page:
| Chapter slug prefix | Anchors | Chapter range |
|---|---|---|
| `all-jobs-and-classes-i-just-wanted-one-skill-not` | 100 | 1 – 99 |
| `all-jobs-and-classes-i-just-wanted-one-skill-not-them-all` | 325 | 100 – 423 |
Both resolve, and **both point back at the same series**:
```
200 https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-chapter-1/
<a href='https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/' aria-label='All Chapter'>
200 https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-404/
<a href='https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/' aria-label='All Chapter'>
```
**Consequence:** a single stored chapter slug is not a sufficient key even for
one novel. A fix that captures "the chapter slug at bookmark time" and pins the
poller regex to it would still miss chapters published under a later slug. This
is the strongest argument for the unscoped regex over a stored-slug regex.
---
## 5. Series page → chapter list
From `https://lightnovelworld.net/novel/immortality-simulator/` (200, 655 037 bytes):
- **1234** anchors matching `href="https://lightnovelworld.net/<slug>-chapter-<n>/"`.
- **1232 unique** chapter URLs, covering chapter numbers **1 through 1232 with no gaps**.
- The 2 extras are the First/Last shortcut widget, which duplicates two chapters:
```html
<h2>Read Immortality Simulator</h2></div>
<div class="lastend">
<div class="inepcx">
<a href="https://lightnovelworld.net/my-longevity-simulation-chapter-1/">
<span>First Chapter</span>
<span class="epcur epcurfirst">Vol. 1 Ch. 1</span>
</a>
</div>
<div class="inepcx">
<a href="https://lightnovelworld.net/my-longevity-simulation-chapter-1232/">
```
**The entire chapter list is in the initial HTML.** No JS hydration, no
pagination, no separate endpoint. Confirmed by probing the shapes the issue
speculated about:
| URL | Status |
|---|---|
| `https://lightnovelworld.net/novel/immortality-simulator/chapters/` | **404** |
| `https://lightnovelworld.net/novel/immortality-simulator/chapters/page-2` | **404** |
The only `page-numbers` / pagination markup in the document belongs to
`class="wpdiscuz-comment-pagination"` (the comment widget), not the chapter list.
This holds for large novels too: `greed-all-for-what` served 2666 chapter
anchors and `my-vampire-system` 2547, all inline in one response.
**All chapter hrefs are absolute** (`https://lightnovelworld.net/…`). A grep for
relative `href="/<slug>-chapter-N/"` on the series page returned **0** matches,
so a pattern anchored on the literal host is safe.
---
## 6. Is the unscoped chapter regex SAFE on a series page?
### Verdict: **SAFE**
Proposed pattern: `lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/`
**Proof.** For each of the 30 series pages fetched from
`https://lightnovelworld.net/novel/`, every occurrence of `<slug>-chapter-<n>/`
anywhere in the document (href or not) was reduced to its slug prefix and
deduplicated:
- **27 of 28 real series pages had exactly 1 distinct chapter-slug prefix**, equal
to that novel's own chapter slug.
- **1 page had 2 prefixes** — `all-jobs-and-classes-…`, and §4.4 proves *both
belong to that same novel*. Not foreign.
- **0 pages carried any other novel's chapter URL.**
The recommendation and sidebar widgets on a series page link to **series** URLs
only, never chapter URLs. On
`https://lightnovelworld.net/novel/immortality-simulator/` the recommendation
strip renders as:
```html
<a href="https://lightnovelworld.net/novel/reverend-insanity/" itemprop="url" title="Reverend Insanity" class="tip" rel="292">
```
— `/novel/<slug>/`, which the pattern cannot match.
The one widget that *would* carry foreign chapter links, "Latest Reading", is an
**empty client-side template**. Its container is `display:none` and its `<ul>` is
empty in the served HTML; the row markup lives in an inert
`<span id="series-history-tpl" style='display:none'>` with mustache placeholders
and a `href="#/{{number}}"` — no real chapter URL:
```html
<div class="bixbox bxcl" id="series-history" style="display:none;">
<div class="releases"><h2>Latest Reading</h2></div>
<div class="series-history-pool">
<ul class="clstyle" id="series-history-ul"></ul>
</div>
</div>
<span id="series-history-tpl" style='display:none'>
<li data-id="{{id}}" data-num="{{chapter}}" data-vol="{{volume}}" data-title="{{title}}">
<div class="chbox"><div class="eph-num">
<a onclick="…return series_history.redirect({{id}});" href="#/{{number}}">
```
It is populated from the visitor's own local history, so a **server-side fetch
never sees content there** — the poller is immune. A browser-rendered fetch with
a fresh profile is likewise immune (no history to render).
### Caveats to record with the verdict
1. **`[INFERENCE]` This is an empirical guarantee, not a structural one.** Unlike
the asura/novelfull cases where scoping to the slug makes foreign contamination
*impossible*, here we only know that lightnovelworld's series template does not
currently emit foreign chapter anchors. If the theme ever adds a "latest site
updates" strip rendered server-side, the unscoped pattern breaks silently and
in the worst direction (a foreign chapter number *higher* than the real one
wins the maximum and the bookmark shows a phantom update).
**Correction, 2026-08-11.** This caveat understated the risk. A series page
server-renders a wpdiscuz comment thread below the chapter list, and comment
bodies are HTML that can carry an `<a href>`. Verified on
`/novel/the-sword-illuminates-the-great-wilderness/`: `#wpdcom` present,
comment `#wpd-comm-358_0` by "hasbi asy" rendered as `<p>where's everyone</p>`
inside `.wpd-comment-text`, corroborated by that page's comment RSS feed;
`firstLoadWithAjax: 0` confirms the thread is in the initial HTML. The sampled
pages carried 0 or 1 comments and no links, which is why §6 read clean. So the
unscoped pattern does not merely depend on the *theme* staying unchanged — it
reads a region any visitor can write to, and the scanner takes the maximum with
no upper bound. The scan must stop before the comment thread.
2. ~~Consider scoping the match to the chapter-list container rather than the whole
document, which would restore the structural guarantee at low cost. The list
items sit under `ul.clstyle` (`class="bixbox bxcl"`).~~
**Refuted, 2026-08-11**, measured on 4 series pages
(`the-sword-illuminates-the-great-wilderness`, `all-jobs-…-them-all`,
`a-will-eternal`, `immortality-simulator`). The single `ul.clstyle` per page is
the hidden, empty "Latest Reading" template (`#series-history-ul`, see §5), not
the chapter list — scoping to it would match nothing. The real chapter list is a
classless `<ul>` inside `div.eplister.eplisterfull`, in
`div.bixbox.bxcl.epcheck`.
The workable boundary is `wpd-threads`: it occurs exactly once per page, and on
all 4 pages every chapter anchor precedes it and no chapter-shaped href follows
it. `wpdcom` and `wpdiscuz` are unusable as markers — they occur 111 to 143
times per page, including in `<head>` roughly 15 KB *before* the chapter list,
so cutting there would discard the list itself. `id='comments'` also occurs once
but is single-quoted; the canonical `id="comments"` never appears.
3. `[0-9.]+` is fine: no decimal chapter numbers appeared anywhere in the sampled
pages, but the existing `strings.Trim(m[1], ".")` guard already handles them.
4. The current test fixture `lnwSeriesFixture` in
`backend/internal/latest/sites_test.go` includes
`<a href="https://lightnovelworld.net/overgeared-chapter-9999/">` and asserts it
is ignored. **That anchor is not representative of a real series page** — no
sampled page contained a foreign chapter anchor. Relaxing the regex will make
that assertion fail, and the correct response is to fix the fixture, not to
keep the scoping.
---
## 7. Redirects and reverse-lookup endpoints
| URL | Status | Redirects | Final |
|---|---|---|---|
| `https://lightnovelworld.net/novel/my-longevity-simulation/` | **404** | 0 | — |
| `https://lightnovelworld.net/novel/radiant-blade-of-the-wilderness/` | **404** | 0 | — |
| `https://lightnovelworld.net/immortality-simulator-chapter-1/` | **404** | 0 | — |
| `https://lightnovelworld.net/my-longevity-simulation/` | **200** | **1** | `https://lightnovelworld.net/my-longevity-simulation-chapter-1/` |
Two findings:
1. **`/novel/<chapter-slug>/` genuinely 404s.** Confirmed for both
`my-longevity-simulation` and `radiant-blade-of-the-wilderness`. There is no
alias, no 301, no fallback. Any stored series URL built by that construction is
permanently dead.
2. **The chapter slug alone (no `-chapter-N`) 302s to chapter 1** of that novel.
This is the closest thing to a reverse-lookup endpoint, but it lands on a
*chapter page*, so recovering the series URL still requires parsing that page's
`aria-label='All Chapter'` anchor or breadcrumb. **No endpoint maps chapter slug
→ series slug directly.**
3. The inverse construction also fails: `/<series-slug>-chapter-1/` 404s for a
divergent novel, so you cannot probe your way from a series slug to a chapter
slug either. The chapter slug must be read off the series page's anchors.
Also present but not a lookup path: the site is WordPress and exposes
`https://lightnovelworld.net/wp-json/wp/v2/posts/177695` (advertised via
`<link rel="alternate" title="JSON" type="application/json" …>` on the chapter
page). Whether the REST API exposes a chapter→series relation was **not tested**
— out of scope for this note, and it would still require fetching the chapter
page to learn the post ID.
---
## 8. Implications for issue #77
> The issue text itself could not be read: `gh` is not installed in this
> environment, so `issue://77` failed to resolve. The proposed regex quoted below
> is the one supplied in the task brief.
### 8.1 The bug
`backend/internal/latest/sites.go:128-137` derives the chapter-anchor pattern
from the **series slug**:
```go
m := lnwSlugRe.FindStringSubmatch(u.Path) // /novel/<series-slug>/
re = regexp.MustCompile(`lightnovelworld\.net/` + regexp.QuoteMeta(m[1]) + `-chapter-([0-9.]+)/`)
```
For `immortality-simulator` this compiles to a pattern matching
`…net/immortality-simulator-chapter-N/`, which appears **zero times** on the
series page — the anchors all say `my-longevity-simulation`. `latestChapterFrom`
returns `ok == false`, indistinguishable from a Cloudflare challenge or a site
redesign, and the bookmark silently stops tracking updates. Same for
`the-sword-illuminates-the-great-wilderness` and `a-villains-will-to-survive`.
### 8.2 The proposed fix is sound
Dropping the scoping to
`lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/` is sound in shape and fixes
all three currently-broken novels plus the split case in §4.4 — which no
stored-slug approach can fix, since that novel legitimately has two chapter
slugs. It must not be applied to the whole document, however: truncate the body
at `wpd-threads` first, per the corrected caveat 2 in §6. The `ul.clstyle`
suggestion originally recorded here is refuted.
Also worth removing: the `overgeared-chapter-9999` line in `lnwSeriesFixture`
and the `"lightnovelworld takes the max and ignores another series"` test case,
which encode a contamination scenario that §6 shows does not occur on this site.
Replacing them with a fixture that mirrors §4.4's two-slugs-one-novel reality
would be a better regression test.
### 8.3 The userscript has the same bug, and it is worse
`userscript/novel-bookmark.user.js` builds the series URL by string-concatenating
the chapter slug (lines 154 and 169):
```js
seriesUrl: "https://lightnovelworld.net/novel/" + m[1] + "/",
```
For a divergent novel this **writes a permanently-404 series URL into the
database at bookmark time**. Fixing only the backend regex leaves those rows
broken, and `fetchableSeriesURL` will happily accept the dead URL because it only
checks the hostname (`sites.go:321-324`).
The userscript should read the series URL off the page instead of constructing
it. On a chapter page, prefer in this order (§3):
```js
document.querySelector("a[aria-label='All Chapter']")?.href
?? document.querySelector("[itemtype$='BreadcrumbList'] a[itemprop=item][href*='/novel/']")?.href
```
The breadcrumb's `<span itemprop="name">` also supplies the correct series title
directly, replacing the current `h1.entry-title` + strip-`Chapter <n>` heuristic
— which today yields "My Longevity Simulation" where the series is actually
titled "Immortality Simulator".
A backfill for already-stored rows is likely needed: any `lightnovelworld` row
whose `series_url` 404s can be repaired by fetching
`https://lightnovelworld.net/<series_id>-chapter-1/` (or `/<series_id>/`, which
302s there per §7) and reading its `All Chapter` anchor.
### 8.4 Note on `series_id`
Rows are keyed `lightnovelworld:<series_id>` where `series_id` is currently the
**chapter slug** (from the userscript's `m[1]`). If the fix changes `series_id` to
the real series slug, existing divergent bookmarks change key and need migrating.
Whether that is worth it depends on whether `series_id` is used to rebuild URLs
anywhere — `sites.go` uses the stored `series_url`, not `series_id`, for
lightnovelworld, so storing the correct `series_url` may be sufficient without a
re-key.
---
## Reproduction
```sh
UA='Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0 Safari/537.36'
# §2 status codes
curl -sS -o /dev/null -w '%{http_code} %{url_effective}\n' -L -A "$UA" \
https://lightnovelworld.net/my-longevity-simulation-chapter-1/ \
https://lightnovelworld.net/novel/immortality-simulator/ \
https://lightnovelworld.net/novel/my-longevity-simulation/
# §3 the two working pointers
curl -sS -A "$UA" https://lightnovelworld.net/my-longevity-simulation-chapter-1/ \
| grep -E "aria-label='All Chapter'|itemprop=\"item\""
# §6 SAFE proof — distinct chapter-slug prefixes per series page (expect 1)
curl -sS -A "$UA" https://lightnovelworld.net/novel/immortality-simulator/ \
| grep -oE '[a-z0-9-]+-chapter-[0-9.]+/' | sed -E 's/-chapter-[0-9.]+\///' | sort | uniq -c
```