fix(userscript): read comix titles from document.title, parse kagane volume chapters

comix.to is an SPA whose client router rewrites document.title but never
touches the server-rendered og:title. The adapter read og:title, so a
bookmark taken after a cold load got the homepage's title ("Comix - Read
Comics online for free") and one taken after an in-page hop got the
previous series' title. Titles now come from document.title, with the
chapter page's " - Ch.<n>" tail stripped.

comix also serves no og:image at all, which is why every comix bookmark
fell back to the monogram placeholder. The cover is now the img whose alt
matches the cleaned title.

Both fixes need the page to have finished its client-side route change,
and comix fills document.title a beat after the URL changes - later than
the nav watcher's 300ms snapshot. The watcher therefore also re-detects
when the detect() signature changes, not only when the URL does.

kagane reader URLs carry no chapter number, so it comes out of og:title.
Volume-numbered series render "<Series> - Volume <v> Chapter <n>" with no
episode name, a shape the suffix regex did not match. One unmatched title
caused both reported symptoms: the volume tail stayed in the stored title
("SP Baby - Volume 1 Chapter 1"), and chapterNum came back null so no
chapter was ever recorded for the series. The regex now takes an optional
"Volume <v> " segment.

All three page shapes were captured live on 2026-08-08 and are pinned as
regression tests in userscript/test/logic.test.js.
This commit is contained in:
2026-08-08 23:03:09 +07:00
parent 2ef769d421
commit e72765df85
3 changed files with 115 additions and 21 deletions
+19
View File
@@ -44,6 +44,25 @@ Guidance for OpenCode (and Claude Code) working under `userscript/`. See root `A
Encodings (incl. triple-encoded punctuation like `%25252D`) identical
on /manga/ and /title/ pages, so decode-once seriesIds match — verified
2026-07-28.
- **comix.to**: series `/title/<id>-<slug>`, chapter
`/title/<id>-<slug>/<uploadId>-chapter-<n>`. Only the leading `<id>` is
identity — the slug re-renders when a series is renamed (`comixSeriesId`).
An SPA that **never rewrites `og:title`**: the server-rendered head keeps
whatever document loaded first, so on a cold load `og:title` is the homepage's
"Comix — Read Comics online for free" and after an in-page hop it is the
*previous* series' name. `document.title` is the one thing client routing does
update, so titles come from there, with the chapter page's `" · Ch.<n>"` tail
stripped. Covers likewise: `og:image` is absent, so the cover is the `img`
whose `alt` matches the cleaned title — verified live 2026-08-08.
- **kagane.to**: series `/series/<uuid>`, reader
`/series/<uuid>/reader/<bookUuid>`. Reader URLs carry no chapter number, so
the number comes out of `og:title`. Two shapes exist: `"<Series> - Chapter
<n>[ - Episode <n>]"` and, for volume-numbered series, `"<Series> - Volume <v>
Chapter <n>"` with no episode name — both must yield a bare series title, or
the volume tail lands in the bookmark's title. Its covers are challenge- and
CORP-protected, so the web UI proxies them; the userscript still stores the
raw `og:image`. Behind a Cloudflare JS challenge, so the backend polls it
through the headless browser.
- **novelfull.com** (novel script): series `/<slug>.html`, chapter
`/<slug>/chapter-<n>[-<title-slug>].html`. No `og:*` tags at all — title from
`h3.title` (series) or `a.truyen-title` (chapter), cover from