lnw identity is read from the chapter page, not derived #89

Closed
opened 2026-08-11 11:28:09 +07:00 by sulthan · 1 comment
Owner

Parent

#80 (spec) — implements part of #77.

What to build

A Reader who bookmarks a lightnovelworld novel from a chapter page can get a Series identity that does not exist. The userscript builds the Series address by string-manipulating the chapter path, and on 7 to 10% of this Site's novels (3 of 41 sampled) the chapter address is built from a different slug than the Series is published under. The resulting Bookmark 404s on every Poll, so it never shows a New Chapter — and if the Reader also bookmarked the same novel from its own page, the list holds the novel twice with only one row updating. The Site gives no hint these novels are different from any other.

Stop guessing. Every chapter page carries a pointer back to its Series: an All Chapter anchor whose href is the Series address, plus a microdata breadcrumb saying the same thing. Read the pointer and take the identity from it. The anchor's aria-label is the primary selector — chosen over three alternatives on measurement: a generic /novel/ href match hits the header navigation index first, a text match false-matches a novel whose title ends in "All Chapter 200", and the JSON-LD breadcrumb's second position is the chapter rather than the Series. The microdata breadcrumb's second crumb is the fallback.

Where neither pointer is present, the page is not one the script understands: it resolves to the "other" page type, so no Bookmark button is offered rather than a Bookmark being created under an invented identity. Across 8 chapter pages measured — chapter 1, chapter 1200, the newest chapter, both slugs of the split novel, two divergent novels, and a novel with a number in its title — both pointers were present and agreed every time.

The Chapter Slug the address was built from is not stored. It is not an identity and a Series may have several. It is carried on the detected page object only, so the repair in the follow-up ticket can recognise a stale row.

The test harness needs a small extension first: its document.querySelector answers a meta-property selector with an element exposing getAttribute, and every other selector with text only. It must answer an attribute selector with an element exposing getAttribute. Extend the stub rather than working around it.

novelfull is untouched, as are all four manga Sites.

Acceptance criteria

  • On a chapter page with the pointer present, the resolved identity comes from the pointer and not from the address
  • With the pointer absent and the breadcrumb present, the identity is the same
  • With both absent, the page resolves to the "other" type and the panel offers no Bookmark button
  • On a divergent novel the two slugs differ and the pointer's slug is the one used — pinned as a regression test, so the derivation cannot be silently reintroduced
  • A novel whose title ends in a number resolves correctly
  • An old chapter address using a retired Chapter Slug resolves to the same Series as a current one
  • The Chapter Slug is available on the detected page object and is written to no store
  • novelfull's adapter is unchanged
  • The documented live URL shapes for this Site reflect that the Series address is discovered, not derived
  • node --check userscript/novel-bookmark.user.js and node --test userscript/test/novel-logic.test.js pass

Blocked by

  • None — can start immediately.
## Parent #80 (spec) — implements part of #77. ## What to build A Reader who bookmarks a lightnovelworld novel from a chapter page can get a Series identity that does not exist. The userscript builds the Series address by string-manipulating the chapter path, and on 7 to 10% of this Site's novels (3 of 41 sampled) the chapter address is built from a different slug than the Series is published under. The resulting Bookmark 404s on every Poll, so it never shows a New Chapter — and if the Reader also bookmarked the same novel from its own page, the list holds the novel twice with only one row updating. The Site gives no hint these novels are different from any other. Stop guessing. Every chapter page carries a pointer back to its Series: an All Chapter anchor whose href is the Series address, plus a microdata breadcrumb saying the same thing. Read the pointer and take the identity from it. The anchor's `aria-label` is the primary selector — chosen over three alternatives on measurement: a generic `/novel/` href match hits the header navigation index first, a text match false-matches a novel whose title ends in "All Chapter 200", and the JSON-LD breadcrumb's second position is the chapter rather than the Series. The microdata breadcrumb's second crumb is the fallback. Where neither pointer is present, the page is not one the script understands: it resolves to the "other" page type, so no Bookmark button is offered rather than a Bookmark being created under an invented identity. Across 8 chapter pages measured — chapter 1, chapter 1200, the newest chapter, both slugs of the split novel, two divergent novels, and a novel with a number in its title — both pointers were present and agreed every time. The Chapter Slug the address was built from is **not** stored. It is not an identity and a Series may have several. It is carried on the detected page object only, so the repair in the follow-up ticket can recognise a stale row. The test harness needs a small extension first: its `document.querySelector` answers a meta-property selector with an element exposing `getAttribute`, and every other selector with text only. It must answer an attribute selector with an element exposing `getAttribute`. Extend the stub rather than working around it. novelfull is untouched, as are all four manga Sites. ## Acceptance criteria - [ ] On a chapter page with the pointer present, the resolved identity comes from the pointer and not from the address - [ ] With the pointer absent and the breadcrumb present, the identity is the same - [ ] With both absent, the page resolves to the "other" type and the panel offers no Bookmark button - [ ] On a divergent novel the two slugs differ and the pointer's slug is the one used — pinned as a regression test, so the derivation cannot be silently reintroduced - [ ] A novel whose title ends in a number resolves correctly - [ ] An old chapter address using a retired Chapter Slug resolves to the same Series as a current one - [ ] The Chapter Slug is available on the detected page object and is written to no store - [ ] novelfull's adapter is unchanged - [ ] The documented live URL shapes for this Site reflect that the Series address is discovered, not derived - [ ] `node --check userscript/novel-bookmark.user.js` and `node --test userscript/test/novel-logic.test.js` pass ## Blocked by - None — can start immediately.
sulthan added the ready-for-agent label 2026-08-11 11:28:09 +07:00
sulthan self-assigned this 2026-08-11 11:34:30 +07:00
Author
Owner

Landed on main as e864770 (merge of ticket/89-lnw-identity-discovered; commits a4da2e2, 8beaa4f).

The lightnovelworld adapter's chapter branch no longer derives the Series address. It reads a[aria-label='All Chapter'], falls back to the microdata breadcrumb's second crumb scoped to the enclosing BreadcrumbList, and resolves to type 'other' when neither is present or the href is not a /novel// address on lightnovelworld.net — no fallback to derivation at any point. The page object carries chapterSlug (null on series pages) and nothing stores it. The harness's querySelector stub now answers attribute selectors with a getAttribute-bearing element. userscript/AGENTS.md records that the Series address is discovered, not derived.

20/20 tests passing, node --check clean. All ten acceptance criteria met, including the divergent-novel regression pin and both Chapter Slugs of the split novel resolving to one Series.

Known gap, deliberately not fixed here: the userscript's own background latest-chapter check still scopes its anchor scan to the stored seriesId slug, so a split-slug novel undercounts client-side. The backend Poll is the authority for Latest Chapter and #87 fixes it there; the client-side scope is filed separately rather than smuggled into this ticket.

Landed on main as e864770 (merge of ticket/89-lnw-identity-discovered; commits a4da2e2, 8beaa4f). The lightnovelworld adapter's chapter branch no longer derives the Series address. It reads a[aria-label='All Chapter'], falls back to the microdata breadcrumb's second crumb scoped to the enclosing BreadcrumbList, and resolves to type 'other' when neither is present or the href is not a /novel/<slug>/ address on lightnovelworld.net — no fallback to derivation at any point. The page object carries chapterSlug (null on series pages) and nothing stores it. The harness's querySelector stub now answers attribute selectors with a getAttribute-bearing element. userscript/AGENTS.md records that the Series address is discovered, not derived. 20/20 tests passing, node --check clean. All ten acceptance criteria met, including the divergent-novel regression pin and both Chapter Slugs of the split novel resolving to one Series. Known gap, deliberately not fixed here: the userscript's own background latest-chapter check still scopes its anchor scan to the stored seriesId slug, so a split-slug novel undercounts client-side. The backend Poll is the authority for Latest Chapter and #87 fixes it there; the client-side scope is filed separately rather than smuggled into this ticket.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: sulthan/mangaBookmark#89