Files
mangaBookmark/docs/adr/0008-series-identity-is-discovered-not-derived.md
T
sulthan 0fe07d12f8 Record the lightnovelworld series-identity decision (#77)
A Series identity is discovered from the Site's own links, never derived
from an address. ADR-0008 carries the decision and the evidence behind
it: 3 of 41 sampled novels serve chapters under a slug that differs from
their series slug, one novel serves chapters under two, and neither slug
is computable from the other.

CONTEXT.md gains Chapter Slug as a term, defines it as plural by nature
and never an identity, and pins Series identity to the canonical slug the
Site publishes. Latest Chapter is restated as the highest-numbered
chapter rather than the newest-dated one, settling #79.

The research note is corrected in place where later probes refuted it:
the ul.clstyle container it named is the hidden empty "Latest Reading"
template rather than the chapter list, and its caveat about the comment
region understated the risk, since that region is writable by any
visitor and the scan takes an unbounded maximum into a shared row.

Docs only; no code. Implementation is specified in #80.
2026-08-11 09:28:26 +07:00

103 lines
5.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ADR-0008: A Series identity is discovered from the Site's links, never derived from an address
Date: 2026-08-11
Status: accepted
## Decision
On lightnovelworld, a Series is identified by the slug in its `/novel/<slug>/`
address, and the userscript obtains that address by reading the chapter page's
`a[aria-label="All Chapter"]` anchor. It no longer constructs the address by
string manipulation of the chapter path. When neither that anchor nor the
microdata breadcrumb is present, the page resolves to `type: "other"` and no
Bookmark is offered.
A Chapter Slug — the slug a chapter address is built from — is not an identity
and is not stored. The backend finds chapters by matching
`lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/` against the series page
body **truncated at the first `wpd-threads`**.
## Why
Measured on lightnovelworld between 2026-08-10 and 2026-08-11.
The chapter slug and the series slug are two independent facts. In a 41-novel
sample, 3 diverged (~7%). `/my-longevity-simulation-chapter-1/` returns 200
while `/novel/my-longevity-simulation/` returns 404 with no redirect, and that
novel's real address is `/novel/immortality-simulator/`. Divergence runs in both
directions: `the-sword-illuminates-the-great-wilderness` is served by chapter
slug `radiant-blade-of-the-wilderness`. Neither slug is computable from the
other, and the Site publishes no alternative-names field, so the mapping exists
only in the chapter page's own markup.
Storing the Chapter Slug beside the identity does not work, because a Series may
have more than one. `/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/`
serves chapters 1–99 under `…-one-skill-not` and 100–423 under
`…-one-skill-not-them-all`. Both resolve, and both chapter pages point back at
the same Series.
Deriving the identity from the chapter path also made one Series produce two
rows: bookmarking from the series page yielded `lightnovelworld:immortality-simulator`,
and from a chapter page `lightnovelworld:my-longevity-simulation`.
The pointer is reliable. Across 8 chapter pages — chapter 1, chapter 1200, the
latest chapter, both slugs of the split novel, two divergent novels, and a novel
with a number in its title — the `All Chapter` anchor and breadcrumb position 2
were both present and agreed every time, including on the old-slug pages. Three
narrower selectors were rejected on evidence: `a[href*="/novel/"]` matches the
header nav index first; matching the text "All Chapter" false-matches the novel
titled "…Not Them All Chapter 200"; and the JSON-LD breadcrumb's position 2 is
the chapter, not the Series.
The scan is truncated because a series page server-renders a wpdiscuz comment
thread below the chapter list, and comment bodies are HTML that can carry an
anchor. Verified on `/novel/the-sword-illuminates-the-great-wilderness/`:
comment `#wpd-comm-358_0` rendered in the initial HTML, corroborated by that
page's comment RSS feed. The scanner takes the maximum chapter number with no
upper bound, and a Series row is shared by every Reader (ADR-0003), so one
comment containing a link to a high-numbered chapter would pin that Series'
Latest Chapter for everyone. `wpd-threads` occurs exactly once per page and
follows every chapter anchor on all 4 series pages measured. `wpdcom` and
`wpdiscuz` are unusable: they occur 111 to 143 times per page, including in
`<head>` before the chapter list.
## Considered options
**Correct `series_url` only, leaving the Chapter Slug as the identity.** This
repairs the Poll with no migration, because the Poll reads the stored address
rather than the identity. Rejected: it keeps an identity that the Site does not
guarantee to be stable, and leaves the duplicate-row hazard in place.
**Scope the match to the chapter-list container.** Rejected on measurement. The
prior research named `ul.clstyle`; on all 4 pages sampled that is the hidden,
empty "Latest Reading" template, and the real list is a classless `<ul>` in
`div.eplister.eplisterfull`. The backend scans raw HTML with no parser, so
container extraction means a second regex against class names, which is more
fragile than the one-off truncation and protects nothing extra.
**Repair the stored rows with a SQL migration.** Rejected as impossible: the
database holds no source for the correct slug. The `series` table has no chapter
address, and no endpoint on the Site maps one slug to the other.
## Consequences
Existing Bookmarks on divergent novels stop matching their own chapter pages,
because `keyOf` changes. The userscript therefore migrates a row in place when
it sees the mismatch: it rewrites the row's key, identity and address in the
cache, the retry-queue entry, and the last-checked map, then syncs. This heals
only when the Reader next opens a chapter page of that novel.
A migrated row leaves its old `series` row behind. Nothing deletes it, but
`DueForLatestCheck` is an inner join against bookmarks, so a Series with no
Bookmarks is never polled again. The old row is permanently stored and
permanently inert.
A Bookmark whose Reader never opens a chapter page of that novel does not heal.
It continues to return 404 at every cooldown, as it does today.
If the truncation marker disappears, the scan is skipped and logged rather than
run against the whole page. An env-gated live test, `TestSmokeLnwCommentBoundary`,
asserts that `wpd-threads` still occurs exactly once and still follows the last
chapter anchor. It skips when its environment variable is unset, matching the
existing `TestSmokeKagane*` convention.