Compare commits

..

24 Commits

Author SHA1 Message Date
sulthan c62c3bb07b chore: drop asuracomic.net from the userscript, CORS allowlist and docs (#97)
Closes #96.

## What

Removes every reference that still invites a Reader onto `asuracomic.net`.
The domain's deep links 301 to the `asurascans.com` **root**, discarding the
path (re-checked 2026-07-25), so a page on it never yields a series document
client-side and a stored address on it never yields a series page server-side.
#95 already pinned each Site to one hostname, so the backend rejects such an
address cleanly; this is the cleanup around that.

| File | Change |
|---|---|
| `userscript/manga-bookmark.user.js` | drops the `@match`, narrows the asura adapter to `/(^\|\.)asurascans\.com$/` |
| `userscript/test/logic.test.js` | new test pinning the narrowed host match |
| `.env.example`, `docker-compose.yml` | origin dropped from the `ALLOWED_ORIGINS` default |
| `DEPLOY.md` | same, and the sample list gains the two novel origins it was missing |
| `backend/api_test.go` | CORS fixtures and round-trip seed move to `asurascans.com` |
| `README.md`, `AGENTS.md` | notes say the host is dropped, not "stays matched" |

## Behaviour

- A Reader landing on `asuracomic.net` gets no userscript UI. Previously the
  script loaded and could do nothing useful — the redirect had already
  discarded the path.
- A request whose `Origin` is `https://asuracomic.net` is no longer reflected
  by a deployment using the shipped defaults.
- No backend logic changed: the CORS rule, the address gate and the poller are
  untouched. `AllowedOrigins` is data, not code.

## Security invariant preserved

CORS still reflects `Origin` only when it appears in `ALLOWED_ORIGINS`, with
`GET,PUT,DELETE,OPTIONS` and a `204` preflight — `TestCORSPreflight` and
`TestCORSDisallowedOrigin` still pin both halves, now against a live origin.
This change only removes a value from the allowlist, which is a narrowing.

## Verification

- `go test ./...` — full backend suite green (real Postgres per package).
- `node --test test/*.test.js` — 66/66 green, up one from the new match test.

## Deploy note (does not happen on merge)

The live allowlist comes from the VPS `.env`, not from these defaults, so the
origin must be dropped there in the same deploy. The one-off row repair for any
stored `asuracomic.net` address is in #96.

Reviewed-on: #97
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-12 05:53:17 +07:00
sulthan 21615be2bd feat: one registry entry per Site, one shared Series-page read (#95)
Closes #94.

## What

Two phases per the spec, in three feature commits plus two review-fix commits:

**Phase one — one registry entry per Site** (`2d134fb`)
The six per-site comparison points that used to live across three files collapse
into one `sites` map in `backend/internal/latest/sites.go`: Latest Chapter parse,
Cover parse, browser-backed list, fetcher route, host pins, and the browser
payload read all become lookups into it. `browserBackedSites()` is derived from
the registry (sorted, deterministic); `fetcherFor` and `fetchableSeriesURL` keep
their signatures and become lookups; `BrowserFetcher.Get` dispatches through the
entries' `Read`/`Done` while the tab lifecycle stays in `BrowserFetcher.run`.

**Phase two — one shared Series-page read** (`f215130`)
`readSeriesPage` (new `read.go`) performs the read the Poll and the Acquisition
have in common: gate, route, fetch, parse Latest Chapter, parse Cover address.
It returns facts only — polling and persistence policies (stamp order, cooldowns,
cover policy) stay with the callers; `acquire.go` gained the comment naming the
deliberate post-fetch stamp order. The poll's legacy cover heal and the
no-chapter byte-count diagnostic were restored after review (`d998f87`) so the
claims "the Poll keeps its own Cover policy" and "pinning is the only
behavioural change" both hold.

## Behaviour

- All six Sites now pin their host exactly; asura/demonic/comix previously
  accepted any https host. For asura this is a strict improvement: its dead old
  domain redirects deep links to the site root and would parse the wrong
  document.
- Everything else is unchanged: existing parse tables, the challenge-body table
  and the gate table pass unmodified except the one deliberate exception — the
  gate table gains the three new pin cases.

## Security invariants preserved

- The address gate is recognisably the same rule, now a single registry lookup:
  `https` + exact hostname match, all callers route through it. No fetch path
  was widened; asura/demonic/comix were narrowed.
- The second host pin inside each browser entry's Read is retained deliberately
  (browser = strong SSRF primitive, `series_url` is client-supplied) and is not
  deduplicated against the shared gate.
- Review hardening: `fetcherFor` now fails closed for unknown site strings
  (previously fell through to the TLS fetcher on an unreachable path), and the
  browser dispatch iterates a sorted list so outcomes cannot depend on map order.
- The security review's log-injection finding was checked against Go's
  `url.Parse` and does not hold: control characters are rejected anywhere in a
  URL, so a client-supplied value in a log line cannot carry a newline.

## Review

Reviewed on three axes (spec, standards, security) by read-only subagents over
`672c16f..f1b26f4`. No blocking findings; all minor/nit findings addressed in
`d998f87` and `700de20`. Verified end to end with `go test ./...` (Docker
Postgres per test package) on every commit.

## Out of scope (tracked separately)

- Dropping asuracomic.net (CORS allowlist, userscript match, API fixtures,
  live env) — separate issue, per spec.

Reviewed-on: #95
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-12 05:43:37 +07:00
sulthan 672c16ffbf Remove the client latest-chapter scan for lightnovelworld (#91) (#93)
Follows the spec published on #91: remove the client-side latest-chapter scan for lightnovelworld rather than porting #87's truncation into a second codebase.

- lightnovelworld.latestChapterFromAnchors deleted, not stubbed: absence is what the background-fetch guard keys off.
- computeLatestChapter tolerates an adapter with no scanner (yields null) and is exported as the test seam.
- backgroundRefreshLatest skips a scanner-less Site before the due filter: no Series page fetched, no freshness timestamp recorded, no batch slot consumed. The on-page path (maybeCaptureLatestOnSeriesPage) routes through the same null-tolerant computation.
- novelfull's scanner, the shared max-chapter helper and all four manga Sites untouched.
- userscript/AGENTS.md records the Poll-only contract for this Site and why.

All seven acceptance criteria from the spec met. node --check clean; novel suite 30/30 (the regression pin fails if a lnw scan is reintroduced, scoped or not); manga suite 35/35, manga userscript byte-for-byte unchanged.

Two-axis code review: no hard standard violations, spec-clean; one follow-up commit matching the sibling adapter guard from the manga script.

Reviewed-on: #93
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-11 20:58:18 +07:00
sulthan e7e22a12a5 lightnovelworld Series identity is read from the chapter page (#80) (#92)
Implements spec #80 / ADR-0008 — Gitea issues #86, #87, #88, #89, #90, all closed.

A Reader bookmarks a novel on lightnovelworld and it never shows a New Chapter, because the Series identity was derived from the chapter address instead of read from the page. One Series can publish under several Chapter Slugs, so the derived key points at a slug that 404s.

- **#89** — the userscript's lnw adapter stops deriving `seriesUrl`/`seriesId` from the path. It reads the page's own pointer (`a[aria-label='All Chapter']`), falling back to the microdata breadcrumb's second crumb, and carries `chapterSlug` on the page object, stored nowhere.
- **#87** — the Poll's lnw chapter scan is unscoped (no stored-slug pattern can cover a Series' whole list) and truncated at the `wpd-threads` comment thread, the one region a visitor can write to. Marker absent means skip and log with the body length, never scan whole. Corrects the `maxBodyBytes` headroom comment to the measured 3.5x.
- **#86** — the scan fixture is now text trimmed from a real, wholly-fetched Series page instead of a hand-written cross-series anchor that no live page carries.
- **#90** — stale stored rows repair themselves on the next chapter visit: a pure transform over cache, queue and last-checked map, silent to the Reader, with progress, favourite and lifecycle bucket preserved when two rows merge.
- **#88** — an env-gated live canary (`SMOKE_LNW_SERIES_URL`) proving the marker still occurs exactly once and still follows the last chapter anchor, asserted against the production symbols themselves.

Verified on the merged branch: `go test ./...` green, `gofmt -l internal/latest/` silent, both userscripts `node --check` clean, 35/35 + 29/29 logic tests. Live canary green (marker once at byte 612,182 of 651,795). #90 verified on device with Playwright.

Open follow-up: **#91** — the userscript's client-side latest-chapter scan is still scoped to the derived slug.

Reviewed-on: #92
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-11 18:21:50 +07:00
sulthan 90d8ab72ad chore: refresh graphify map and tidy AGENTS.md (#84)
graphify update regeneration: semantic hashes now populated in manifest.json, graph rebuilt (1634 nodes, 3179 edges). Track the map (5 curated files + .graphify_root) so a fresh checkout starts with it; cost.json, cache/, dated snapshots and .rebuild.lock stay ignored.

AGENTS.md: drop stale Relevant skills and Notes sections; graphify rule now says to always query the graph first.

Reviewed-on: #84
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-11 11:19:47 +07:00
sulthan c400c91a80 Implement a batch of tickets through per-ticket subagents (#82)
## What this adds

Two files that turn the one-ticket-at-a-time `/implement` loop into an orchestrated batch.

**`.claude/skills/implement-tickets/SKILL.md`** — user-invoked (`disable-model-invocation: true`, so it costs no context until typed). The agent that runs it is an orchestrator, not an implementer:

1. Collect the tickets over `tea`, reading each `Blocked by` line.
2. Plan waves from the blocking edges, three tickets wide, and fix every cross-ticket contract (shared signature, JSON shape, column, token) before anything is dispatched.
3. Present the plan and stop for approval.
4. Per ticket: `git worktree add ../ticket-<n>`, copy the gitignored `.env`, claim the issue, write a brief to `.scratch/`, then dispatch the whole wave as one `task` batch.
5. Land each result — merge `--no-ff`, comment the report, close, remove the worktree. Textual conflicts are the orchestrator's; a semantic clash goes back to whichever ticket owns the contract.
6. Full suite once on the merged base.

**`.omp/agents/ticket-implementer.md`** — the worker. Brief-driven, worktree-bound, and gated on review before it reports: it runs the `code-review` skill over its own diff with `cr-spec` and `cr-standards` on the two axes, fixes Critical and Important findings in at most two rounds, and returns a short status contract (`DONE` / `DONE_WITH_CONCERNS` / `BLOCKED` / `NEEDS_CONTEXT` / `REVIEW_BLOCKED`).

The brief template makes the subagent read `tea issue <n> --comments` for its ticket and for the issue that ticket refers to — the comments carry decisions the body never got updated with — and names the `tdd` skill at each seam where a test comes first. Briefs are written in the ubiquitous language of `CONTEXT.md`; a brief that says "scrape" where the domain says Poll hands the subagent the wrong model of the system.

## Verification

Dispatched a real `ticket-implementer` as a probe. The agent resolved from `.omp/agents`, and it spawned `cr-spec`, which replied. That was the one thing that could have silently killed the design: `task.maxRecursionDepth` defaults to 2, and the chain is session to orchestrator to implementer to reviewer. It clears. If that ever changes, the implementer returns `REVIEW_BLOCKED` and the orchestrator runs the review itself.

Confirmed against the omp binary that `autoloadSkills: code-review, tdd` is split by `parseArrayOrCSV`, not swallowed as one unknown name.

## Notes

- Agents are discovered from `.omp/agents`, never `.claude/agents` — the latter is deliberately skipped by omp because its frontmatter is a different contract.
- No product code changes. `.gitignore` gains `.scratch/`, where briefs and reports live.
- Not included: retry after a failed dispatch, a state file for resuming a crashed wave, a cheap model tier for mechanical tickets. Add them when a real batch needs them.

Reviewed-on: #82
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-11 10:08:09 +07:00
sulthan f1eb7d514c Record the lightnovelworld series-identity decision (#77) (#81)
Docs only. No code, no tests, nothing to run. Implementation is specified in #80.

Outcome of a grilling session on 2026-08-11 against #77, backed by live measurement of lightnovelworld over 2026-08-10/11.

## What changed

**`docs/adr/0008-series-identity-is-discovered-not-derived.md`** (new)

A Series identity is discovered from the Site's own links, never derived from an address.
On lightnovelworld the userscript reads the chapter page's `All Chapter` anchor instead of
building a `/novel/<slug>/` address by string manipulation. A Chapter Slug is not an
identity and is not stored. The backend's chapter scan drops its per-Series scoping and
runs against the body truncated before the visitor comment thread.

Evidence in the ADR: 3 of 41 sampled novels serve chapters under a slug that differs from
their series slug, divergence runs in both directions, one novel serves chapters under two
slugs, and neither slug is computable from the other. The pointer was checked on 8 chapter
pages and agreed every time. Three narrower selectors are recorded as rejected, each with
the measurement that killed it.

Three rejected options are recorded with reasons: correcting the stored address only, which
keeps an identity the Site does not guarantee; scoping the scan to a container, which the
probe refuted; and a SQL migration, which is impossible because the database holds no
source for the correct slug.

**`CONTEXT.md`**

- **Series** - identity is the canonical slug the Site publishes, never the title and never a Chapter Slug.
- **Chapter Slug** - new term. A slug a Site builds its chapter addresses from. Not an identity: one Series may have several, and none is computable from another.
- **Latest Chapter** - now the highest-numbered chapter, explicitly not a date and not the Site's own newest-chapter banner. Settles #79.

**`docs/research/lightnovelworld-chapter-vs-series-slug.md`** (new, committed with its corrections)

The 41-novel survey behind the ADR. Two claims are struck through and corrected in place,
with the date and sample size of the probe that refuted each: the `ul.clstyle` container it
named is the hidden, empty "Latest Reading" template rather than the chapter list, and its
caveat about the comment region understated the risk, because that region is writable by
any visitor while the scan takes an unbounded maximum into a Series row shared by every
Reader (ADR-0003).

## Review notes

Nothing here constrains code that exists today - the ADR describes work not yet written.
The part worth disagreeing with, if any of it is wrong, is the fail-closed rule: a missing
truncation marker means skip the Series and log, never scan the whole page.

Related: #77 (the defect), #80 (the spec), #79 (the numbering anomaly, closed by decision),
#71 (the same size cap seen from the cover side).

Reviewed-on: #81
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-11 09:34:26 +07:00
sulthan 1ee5eb67ea Clear the stale Chrome singleton lock at browser boot (#76)
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 19:11:24 +07:00
sulthan b22ae82897 Restore the novel script's lost module-scope constants (#74) (#75)
> *This was generated by AI during triage.*

Fixes #74.

`novel-bookmark.user.js` was split out of `manga-bookmark.user.js` and lost four module-scope constants. Every use of them is behind a `try/catch` or a fire-and-forget promise, so the `ReferenceError`s were swallowed rather than reported.

| constant | used at | effect while missing |
| --- | --- | --- |
| `LATEST_CHECK_THROTTLE_MS` | `:722` | `backgroundRefreshLatest()` throws before computing `due` — no background latest-check ever runs for novels (the symptom in #74) |
| `LATEST_CHECK_BATCH` | `:724` | same throw |
| `CACHE_KEY` | `:215`, `:224` | `loadCache()` always returns `[]`, `saveCache()` silently no-ops — the local cache never persists |
| `LASTCHECKED_KEY` | `:232`, `:241` | last-checked map never persists, so the throttle would not hold even once the first two are defined |

#74 named only the two throttle constants. The two cache keys are the same lost lines with the same root cause, so they are restored here too — fixing only the pair the issue named would leave `backgroundRefreshLatest()` re-fetching every series on every navigation, because `saveLastChecked()` would still be a no-op.

Values and comments copied verbatim from `manga-bookmark.user.js:37-43`; throttle 4h, batch 1.

## Verification

- `node --check userscript/novel-bookmark.user.js` — clean.
- `node --test userscript/test/logic.test.js userscript/test/novel-logic.test.js` — 47/47 pass.
- New test `every SCREAMING_CASE constant the script uses is declared in it` scans both scripts (comments and string literals stripped first, so prose and SVG path data do not trip it). Confirmed it fails — 1 failing test — when `LATEST_CHECK_BATCH` is deleted again, and passes when restored.

A behavioural test cannot reach this: the storage helpers and the background refresh are exactly the layers the harness does not cover (see the `testing-the-userscript` skill), and the errors are swallowed anyway. A static guard is the only instrument that sees this bug class.

On-device confirmation that the ember now lights for novels is still outstanding — that needs Violentmonkey against a live novelfull/lightnovelworld page.

Reviewed-on: #75
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 18:12:03 +07:00
sulthan 7c7d597019 Delete the kagane-specific cover path (#63) (#73)
Closes #63

Deletes the second way to reach a Cover. Since #62, every Site's cover bytes land in the content-addressed store at creation or on the poll, and the one public route serves them all — nothing needs the kagane proxy anymore.

## What went

- **Template-level rewrite:** `Bookmark.CoverURL()` and both templates' use of it. Cards and chrome now render `.Cover` — the wire value — and nothing else. `Bookmark.CoverSource` was dead once `CoverURL` went, so it and its `bookmarkColumns` entry are gone too.
- **Kagane-only cover route and its identifier validation:** `GET /img/kagane/{id}`, `web.CoverFetcher`, `coverIDRe`, and the whole `internal/web/cover.go`.
- **The proxy's persistence:** `store.KaganeImageID`, `GetKaganeCover`, `PutKaganeCover`, `kaganeCoverSourceURL`, `kaganeCoverRe`.
- **The kagane-shaped branch in the byte-fetch routing:** `fetchCoverBytes` no longer takes a `site` argument and no longer names a Site. The URL shape kagane's API publishes is claimed by the browser module itself — `kaganeImageURLRe` + `browserCoverURL` live in `latest/browser.go` with the rest of the per-Site knowledge — and `BrowserFetcher.Image` is now URL-driven (it validates the URL it will navigate to, same SSRF discipline as before). The no-plain-TLS-fallback rule for a claimed URL is preserved: a claimed address with no browser is an error, never a challenge-page fetch.

## What stayed (deliberately)

- `BrowserFetcher.Image` and the browser-backed acquisition path: kagane genuinely serves cover bytes behind the challenge + `cross-origin-resource-policy: same-origin`, so the sidecar remains the only fetcher for them — it just routes by URL claim now instead of by Site name.
- `fetcherFor`'s per-Site page routing (kagane/novelfull page fetches) — that is the page path, not a cover path.

## Acceptance criteria

- [x] Template-level kagane cover rewrite gone
- [x] Kagane-only cover route and its identifier validation gone
- [x] Tests removed/rewritten against the general route, guarantees kept: unstored + traversal-shaped addresses serve nothing (`TestPublicCoverRejectsUnknownAddress`), non-image content types never echoed (`TestPublicCoverNeverEchoesNonImage` — new; the store-side gate was already pinned by `TestCoverStoreAcceptsAnySourceURL`). Store reopen-persistence and filesystem content-addressing tests rewritten against `PutCover`/`GetCover`, no guarantee lost.
- [x] No Site name in a cover code path outside the acquisition module (`grep kagane backend`: store/web/templates/api are clean; remaining hits are `latest/browser.go` + `latest/sites.go`, tests, docs)
- [x] Web UI and panel render Covers for all six Sites (templates render the wire address; panel renders `b.cover` — untouched, it never had a kagane path)
- [x] `go test ./...` green

## Verification

- `go vet ./...` clean
- `go test ./...` — all packages pass (root 16.9s, latest 12.7s, store 12.7s, web 0.004s)
- `CGO_ENABLED=0 go build` produces the static binary
- Cover-path tests run verbosely: `TestPublicCoverServesStoredBytesUnauthenticated`, `TestPublicCoverRejectsUnknownAddress` (unknown/malformed/traversal/empty), `TestPublicCoverNeverEchoesNonImage`, `TestListRendersAcquiredCover`, `TestAcquireKaganeCoverThroughBrowser`, `TestRunOncePrefetchesKaganeCover`, `TestRunOnceRoutesNonKaganeCoverToPublicFetcher` all pass; the three `SMOKE_*` tests skip without the browser sidecar, as designed

Live browser verification of the "web UI and panel render Covers for all six Sites" criterion is being run separately with Playwright against real Site pages and a locally mocked backend.

Reviewed-on: #73
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 18:02:47 +07:00
sulthan 78234f3c19 Browser-backed Sites join the Cover pipeline (#62) (#72)
Fixes #62

Browser-backed Sites join the Cover pipeline: kagane and novelfull Series now get their Covers at creation, through the same acquisition path as every other Site, instead of waiting for a poll pass.

## What changed

`latest.Acquirer` (creation-time acquisition, fired by the first Bookmark of a Series) previously skipped kagane and novelfull entirely — their pages only yield a Cloudflare challenge to the TLS client, so the request was spent for nothing. It now routes them like the poller does, with the two Sites split exactly as the issue demands:

- **kagane** — page fetched through the browser sidecar, cover URL extracted from the API JSON, bytes fetched through the browser sidecar (the only path that clears the challenge) into the content-addressed store. With no `BROWSER_WS_URL` configured, acquisition is skipped entirely and nothing falls back to a plain fetch.
- **novelfull** — page fetched through the browser sidecar, cover URL extracted from the HTML, bytes fetched over plain TLS through the ordinary gated fetcher (its image paths answer 200 with `access-control-allow-origin: *`, measured 2026-08-09). With no browser configured, the page fetch falls back to the TLS client — novelfull's challenge is a live time-varying fact (AGENTS.md), so when the page body answers, the Cover still lands; when it is challenged, nothing happens.

The byte-routing rule (kagane → browser, every other Site → TLS) is now one shared function (`latest.fetchCoverBytes`) used by both the Poller and the Acquirer, so the two cannot drift apart.

## Acceptance criteria

- [x] kagane cover bytes are fetched through the browser sidecar and stored in the content-addressed store — `TestAcquireKaganeCoverThroughBrowser`
- [x] novelfull cover URLs are extracted from the browser-fetched HTML, and its bytes are fetched over plain TLS — `TestAcquireNovelfullCoverOverPlainTLS`
- [x] With no browser sidecar configured, kagane Covers are absent and nothing falls back to a plain fetch — `TestAcquireKaganeSkippedWithoutBrowser`
- [x] With no browser sidecar configured, novelfull Covers still work if its page body is available — `TestAcquireNovelfullCoverWithoutBrowser`
- [x] Manually verified on-device: a kagane Series shows its Cover in the panel, not a broken-image glyph — being run by a separate manual-verification agent against a mocked scenario (no prod data); not part of this PR
- [x] `go test ./...` is green, with live-network checks gated behind `SMOKE_BROWSER_WS_URL` like the existing kagane image smoke test — new `TestSmokeAcquireKaganeCover` proves the end-to-end acquire path against the real browser when the env var is set

## Verification

- `go test ./...` green across all packages
- New unit tests exercise every routing decision with fakes — no network in the default suite
- Smoke test gated behind `SMOKE_BROWSER_WS_URL`, skipped by default

## Post-review changes (a66491a)

- **One routing rule for pages too** — `fetcherFor` is now a shared function used by both the Poller and the Acquirer; novelfull falls back to the plain-TLS fetcher in *both* when no browser is configured, so pre-existing (client-scraped) novelfull rows get healed by the poll as well, not just Series created after this change (`TestNovelfullUsesTLSWhenNoBrowserFetcher`).
- **Byte-level no-fallback proof** — `TestAcquireKaganeBytesNeverFallBackToPlainTLS` pins that kagane cover bytes never route to the TLS fetcher even when the page came through a browser.
- **Acquirer wired independent of the TLS client** — if `NewTLSFetcher` fails, kagane/novelfull acquisition still works via the sidecar (`main.go`).
- AGENTS.md (root + backend) updated for the novelfull plain-TLS fallback.

Reviewed-on: #72
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 11:06:36 +07:00
sulthan b9220b3dfc The Poll fills blank Covers for every Site and both Libraries (#70)
Closes #61.

## Summary

Permanently-blank Series (the half of #47 that creation-time acquisition cannot reach) heal on the next due poll cycle. The cover path is no longer kagane-only: every Site and both Libraries fill a blank Cover from the series page the chapter poll already fetched, and never replace a Cover that already exists.

## What changed

### `backend/internal/latest/poller.go`

- **`fillBlankCover`** — when `Cover` and `CoverAddress` are both blank, extract a source URL via `coverFrom` from the series-page body and store bytes through `SetSeriesCover`. Skips any Series that already has a source URL (owned by prefetch) or a stored address (never overwrite).
- **`prefetchCover`** — source-URL healing path, now site-uniform. Kagane no longer special-cases into `PutKaganeCover` alone; every Site lands on `SetSeriesCover`, so the wire Cover becomes a content-addressed public URL. Reuses already-stored bytes when present.
- **`storeCover` / `fetchCoverBytes`** — shared fetch+persist. Only kagane routes image bytes through the browser fetcher; every other Site uses plain TLS `CoverBytesFetch`. Failures log with the Series key and never return to the chapter path.
- **`checkOne`** — after a successful series-page fetch, calls `fillBlankCover` once regardless of whether chapter extraction succeeded (cover fill is independent of the chapter signal).

### `backend/internal/latest/poller_test.go`

Extended the existing poller harness (real store, fake fetchers) rather than a new one:

- `TestRunOnceFillsBlankCoverFromSeriesPage` — asura manga, lightnovelworld novel, kagane manga; asserts wire Cover + correct fetcher routing.
- `TestRunOnceDoesNotReplaceExistingCover` — second poll does not refetch.
- `TestRunOnceRetriesFailedBlankCoverOnNextPoll` — failed fill stays blank, next due cycle retries (no separate queue).
- `TestRunOnceBlankCoverFailureDoesNotBlockChapter` — chapter still lands; failure log carries the Series key.
- Kagane prefetch test now also asserts the content-addressed wire Cover.

## Acceptance criteria (#61)

| Criterion | Status |
|---|---|
| Cover prefetch runs for every Site | done |
| Cover prefetch runs for both Libraries | done |
| Poll fills a blank Cover | done |
| Poll never replaces an existing Cover | done |
| Failed cover fetch does not fail/block chapter poll | done |
| Failed cover fetch retried next poll, no separate queue | done |
| Failures logged with the Series | done |
| Existing poller tests extended | done |
| `go test ./...` green | done |
| Manually verified: blank Series gets Cover after a poll cycle | **left for you** |

## Out of scope / not closed

- Does **not** close #47 or #55 (per ticket).
- No migration/backfill script — the Poll walks every Series already.
- No admin refetch (#54).

## Review notes addressed

- Removed the kagane-only `PutKaganeCover` branch from prefetch so source-URL healing also sets `CoverAddress` (wire Cover).
- Guard so `fillBlankCover` does not double-fetch after `prefetchCover` healed the same snapshot.
- Single `fillBlankCover` call site after the series-page fetch.

## Test plan

- [x] `go test ./...` (backend; needs Docker/Postgres via `pgtest`)
- [ ] After deploy: pick a Series that was blank, wait one poll cycle, confirm Cover in web UI and userscript panel

Reviewed-on: #70
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 10:14:27 +07:00
sulthan e2c054e7ce Covers render in the userscript panel, from a public route (#60) (#69)
Closes #60.

Spec: #55. Originating bug: #47. Architecture: `docs/adr/0007-backend-hosts-cover-bytes.md`. Neither #47 nor #55 is closed from here.

## What this branch does

The panel now renders Covers from the deployment's own origin, and both userscripts stop having an opinion about where a Cover lives.

**The public route was already in place.** `GET /covers/{address}` landed with #59 (`92eba07`) and is registered on the bare mux, outside `httpmw.Auth` and outside the web UI's Discord session — `backend/main.go:210-214`, handler `backend/internal/api/handlers.go:142-158`. It reads no cookie and no header, answers `404` for an address that was never stored (and for a row whose file has gone missing — recorded-but-gone is not-found, never a fabricated body), refuses anything that is not `^[0-9a-f]{64}$` *before* the value becomes a path, and sets `Cache-Control: public, max-age=604800, immutable`. Those four properties are asserted by `backend/cover_test.go:231-278`. This branch re-verified them rather than re-implementing them; the only backend line it touches is a comment.

**Both userscripts lose cover scraping entirely.** Every adapter's `cover:` field is gone, along with the two helpers that fed them: the manga script's `coverFromPage()` (the `img[alt]` DOM scan comix needed, because comix publishes no `og:image`) and the novel script's `metaName()` plus the now-callerless module-level `meta()`. Nothing under `userscript/` reads `og:image`, `meta[name=image]`, or `img[alt]` any more.

**Nothing sends a cover either.** `delete body.cover` sits in `apiPut` — `manga-bookmark.user.js:486`, `novel-bookmark.user.js:275` — which is the single chokepoint every write passes through (`pushBookmark`, the retry-queue flush, `toggleFavorite`, `toggleArchive`). It operates on the `Object.assign` copy, so the in-memory row keeps the cover it renders with. This matters beyond tidiness: a Reader upgrading from an older copy has `localStorage` rows carrying third-party scraped URLs, and without the strip those would ride back up on the next write. The handler discards the field regardless (`handlers.go:53-59`) — it is permanently inert, not pending removal.

**Failed loads get the designed empty state, not the broken-image glyph.** `onerror: (e) => e.target.replaceWith(el("div", { class: "cover ph" }))` on the cover `<img>` in both card renderers (`manga:1380-1390`, `novel:1134-1144`). The replacement is byte-identical to the existing no-cover branch on the very next line, so it picks up the `.cover.ph` styling already in the panel CSS — no new tokens, no new rule. `el()` routes any `on*` prop through `addEventListener`, so this is a listener, not an inline attribute string, and the swap is a `createElement` + DOM call with no markup parsing anywhere near it. This is the half of #47 that was visible on kagane.

**The deleted scraping's tests went with it**: the two comix cover cases, the `pageImages` and `namedMetas` fixtures, the `img[alt]` and `meta[name=...]` stub branches, the now-dead `querySelectorAll` stub member, and every stale `og:image` fixture and `p.cover` assertion across both suites. The export lists needed no change and that was checked, not assumed — `coverFromPage` and `metaName` were module-private on `origin/main` and no cover symbol ever appeared in `module.exports`.

Docs that described the deleted behaviour were corrected in the same breath, because leaving them would instruct the next agent to put the scraping back: `userscript/AGENTS.md` (adapter contract + the per-site notes for comix, kagane and novelfull), the README's adapter reference, and the userscript testing skill's stub table.

## Verification

- `go test -count=1 ./...` — green across all nine packages (`backend` 29.8s, `latest`, `store`, `session`, `token`, `userscript`, `web`).
- `node --check` clean on both userscripts; `node --test` on both logic suites — 46 tests, 46 pass.
- `gofmt -l` clean; `go build ./...` clean.
- The `onerror` swap is DOM behaviour and deliberately has no coverage in the Node harness — that harness stubs a browser precisely so it never needs a DOM, and #60 says not to invent coverage for it. It was instead exercised for real: the `el()` helper and the exact render expression were loaded into a headless Chromium with a deliberately unloadable `src`, and the resulting DOM was `<div class="cover ph"></div>`. Ad hoc, not committed.
- **Not done, needs you:** the on-device criterion — a comix Series bookmarked mid-chapter showing its Cover in the panel. That needs a real install against the deployment and is the one box left unticked on #60.

## Reviewed

Both `/code-review` axes ran against `cc0fa92`. Spec found no missed requirement and no scope creep; standards found the diff clean on the four areas it scrutinised (the `delete body.cover` placement, the `onerror` handler's DOM safety, comment quality, dead-code removal). Their combined findings — the dead `querySelectorAll` stub, the stale README and skill text, and the handler comment whose premise this change invalidates — are fixed in `8b58019`.

## Out of scope, deliberately

The kagane-specific cover proxy still exists and still carries its session gate (#63 deletes it). The poll's blank-Cover fill (#61) and browser-backed Sites joining the pipeline (#62) are untouched.

Reviewed-on: #69
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 08:52:57 +07:00
sulthan 92eba07da7 A newly bookmarked Series acquires its Cover at creation (#59) (#68)
Closes #59.

Part of spec #55, and the ticket that fixes the reported bug #47. Architecture: `docs/adr/0007-backend-hosts-cover-bytes.md`. Does not close #47 or #55.

## What changed

A Reader bookmarks a Series nobody holds yet — the exact case in #47 — and within seconds the list shows its artwork instead of a broken image. The first Bookmark to create a Series fires `Store.OnSeriesCreated` after commit, and the new `latest.Acquirer` turns that into **one** series-page fetch that yields both the Latest Chapter and the cover URL. The bytes go through the gated cover fetcher from #57 and are stored content-addressed through #56, so the wire carries an absolute URL on this deployment's own origin — never a third-party address, and never one that 404s.

### Store

- Migration `0009_series_cover_address.sql` adds `series.cover_address`. The two facts are now split: `series.cover` is the third-party source address the bytes came from (the acquisition path's dedupe key), `series.cover_address` is the SHA-256 they are stored under. An empty `cover_address` is precisely what "no Cover yet" means, which is the distinction both the API and the UI depend on.
- `SetSeriesCover` writes the address only after the bytes are on disk, so the wire can never name an object that is not there.
- `CoverWireURL` builds `PUBLIC_BASE_URL + /covers/<sha256>` for every scanned row, and returns `""` for a blank address.
- The cover columns are gone from `Upsert`'s `INSERT` and its `DO UPDATE`. A client-supplied cover cannot reach the shared Series row on any path, not just the creation path.
- `Open` now rejects a base URL that is not an absolute `http(s)` origin: `PUBLIC_BASE_URL=bookmarks.example.com` would otherwise start cleanly and emit addresses no browser can load.

### Acquisition

- `internal/latest/acquire.go`: one fetch, gated by the poller's own `fetchableSeriesURL` (a `series_url` arrives in a client-supplied PUT body, so without the gate a token-holder chooses what the server fetches from its own network position).
- Asynchronous and log-and-drop. The Bookmark, its progress and its Latest Chapter are already committed; a Site that is down or a cover that cannot be produced disturbs none of them.
- Bounded by a two-slot semaphore. A bulk sync creating N Series would otherwise fire N simultaneous requests from one IP — the traffic shape the poller's stagger exists to avoid.
- Cancelled at shutdown (shares the poller's context) and stamps `latest_checked_at`, so the poller does not refetch the same page a tick later.
- Browser-backed Sites (kagane, novelfull) are deliberately skipped: their pages only yield a Cloudflare challenge to the TLS client, so the request would be spent for nothing. They arrive in #62.

### Wire and route

- `GET /covers/{address}` serves the bytes publicly and uncredentialed with `Cache-Control: public, max-age=604800, immutable`. The address is gated by a `^[0-9a-f]{64}$` pattern and cross-checked against a pure function of itself before any filesystem read, so no request shaped like a traversal reaches disk.
- `PUT /bookmarks/{key}` still accepts a `cover` field and discards it, permanently. Rejecting it would break every installed userscript the moment this deploys, and ADR-0004's compatibility argument depends on those scripts continuing to work. The decode site says so in place of a TODO nobody intends to keep.
- `store.CoverContentType` canonicalises comix's non-standard `image/jpg` to `image/jpeg`, so one image cannot land under two spellings. This one was found by the live smoke test, not by reading.

### Config

`PUBLIC_BASE_URL` is new and required (cover URLs must go out absolute — the userscript renders them on third-party origins, where a relative path resolves against the Site). Documented in `.env.example`, `docker-compose.yml` (`:?` so compose fails too), `DEPLOY.md` and `backend/AGENTS.md`.

## Acceptance criteria

All twelve of #59's criteria are met; the checklist on the issue is ticked with the evidence.

## Verification

- `go test ./...` green (Docker-backed Postgres suite).
- Live smoke against a real backend + Postgres: bookmarking `comix:n8we-dungeons-and-crayons` produced `"cover": "http://127.0.0.1:8099/covers/8ce74d80…"` and `"latest_chapter": "Chapter 81"` within seconds of the PUT; `curl` on that address returned `200`, `Content-Type: image/jpeg`, `Cache-Control: public, max-age=604800, immutable`, and a 280x420 JPEG. That run is what surfaced the `image/jpg` content type.
- Mutation-checked the asynchrony test: removing the `go` from `Acquire` turns `TestAcquireDoesNotBlockTheWrite` red.

## Reviewed

Both axes of `/code-review` were run against this diff before commit. Their findings that were actionable here are folded in: the concurrency bound, the shutdown tie, the `PUBLIC_BASE_URL` validation, the missing `latest_checked_at` stamp, and a test that could not fail.

## Known sequencing

A kagane/novelfull Series created between this deploy and #62 has no cover source at all: the acquisition skips those Sites and `Upsert` no longer persists the userscript-scraped address. This is #59's stated boundary rather than a defect, but it is a user-visible gap on two Sites and should order #62 accordingly.

Reviewed-on: #68
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 04:07:53 +07:00
sulthan b6b88bde8a feat(latest): extract per-site covers (#58) (#67)
Closes #58

## Summary

- Add pure per-Site cover extraction beside latest-chapter parsing for all six Sites.
- Read Asura, Demonic, LightNovelWorld, and NovelFull metadata; read the Comix target detail state; read Kagane's browser-fetched `series_covers[].image_id` JSON.
- Preserve published cover URLs, percent-encode Demonic raw spaces, select Comix's smaller published `medium`, and avoid thumbnail rendition URL synthesis.
- Add live-source fixtures plus no-cover and Cloudflare challenge coverage for every Site.

## Correctness

- Scope Comix extraction to the requested series detail key, avoiding recommended posters.
- Parse Kagane's current live API shape and emit its canonical compressed image route from the published image ID; unrelated JSON fields are ignored.
- Validate Kagane image IDs against the existing UUID-shaped route constraint.
- Keep extraction pure; storage, polling, and wire integration remain outside issue #58.

## Acceptance criteria

- [x] Cover extraction exists for all six Sites in the existing latest parser module.
- [x] Each Site has a live-source fixture with source URL and date.
- [x] Comix reads the state blob, not metadata.
- [x] Demonic raw spaces are percent-encoded.
- [x] Comix returns the smaller published rendition.
- [x] No-cover pages return empty.
- [x] Cloudflare challenge pages return empty.
- [x] No thumbnail URL is synthesized by editing a published URL.
- [x] `go test ./...` passes.

## Verification

- `go test ./...`
- `go vet ./...`
- `git diff --check`

Parent issues #47 and #55 remain open as requested.

Reviewed-on: #67
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 01:43:57 +07:00
sulthan 9d6d3bde72 Add gated cover byte fetcher (#66)
## Summary

Adds a plain-TLS cover byte fetcher with a destination-class SSRF gate and wires public cover sources through the content-addressed filesystem store.

## Changes

- Resolve hostnames before connecting; refuse non-HTTPS, loopback, private, link-local, unique-local, CGNAT, credentials, and mixed public/private DNS answers.
- Re-check every redirect and resolve/classify again at dial time to close DNS rebinding.
- Reuse `maxBodyBytes`; reject oversized responses and non-image content types before persistence.
- Add generic `Store.GetCover`/`PutCover` source-URL storage while preserving the browser-backed kagane path.
- Keep cover prefetch failures isolated from chapter polling.
- Add observable tests for TLS, no-connection refusals, all refused address classes, redirect blocking, streaming body caps, non-image rejection, content-addressed persistence, DNS rebinding, and poller routing.

## Verification

- `go test -count=1 ./...`
- `go vet ./...`

Both pass. No test touches the live network.

Closes #57

Reviewed-on: #66
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 00:45:55 +07:00
sulthan e8d1cba6c5 Move cover bytes to content-addressed filesystem storage (#65)
Refs #56

## Summary

Moves Kagane cover bytes out of Postgres bytea storage into an immutable, content-addressed filesystem store. Reader-visible behavior remains unchanged: the existing session-gated route serves stored bytes, missing bytes use the existing browser fetch path, and no browser still returns a missing cover.

## Changes

- Added migration 0008, which drops the legacy `covers` table and recreates it with only `address`, `path`, and `content_type`. Existing byte rows are intentionally dropped.
- Added SHA-256 source-URL addressing with two-level sharding (`ab/cd/<sha256>`). Writes use a temp file plus atomic link; reads validate the stored relative path before opening it.
- Made `COVER_DIR` required in runtime config and Compose. Compose passes it as a Docker build argument and volume target, so custom durable paths keep image ownership, runtime config, and the named `cover-data` volume aligned.
- Updated every `store.Open` caller and documented configuration, deployment, backup, and troubleshooting behavior.
- Added filesystem, restart, migration-drop, no-browser, and content-addressing coverage.

## Verification

- `go test ./...`
- `CGO_ENABLED=0 go build ./...`
- `docker build --build-arg COVER_DIR=/data/covers -t manga-bookmark-cover-check-custom ./backend`
- `docker compose config --format json` confirms custom `COVER_DIR` is the volume target
- `git diff --check origin/main`
- LSP diagnostics clean for touched Go files

Parents #47 and #55 remain open as required by #56.

Reviewed-on: #65
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 00:11:20 +07:00
sulthan 30c57bd39c Define Cover and record hosting its bytes (#47) (#64)
Defines **Cover** in the glossary and records ADR-0007, the decision behind #47's fix.

## Why these two files, and why now

`CONTEXT.md` named Cover inside the **Series** entry — "facts true regardless of who is reading — title, cover, Latest Chapter" — but never said *what* one is. That gap is the bug. Nothing in the model distinguished "an address on a Site" from "an image a Reader's browser can display", so both clients were left to work it out independently, and one of them got it wrong. kagane serves covers with `cross-origin-resource-policy: same-origin`, the web UI rewrote them to a proxy in its templates, the JSON API did not, and the panel rendered a broken-image glyph. The new entry closes the ambiguity: *an address no client can load is not a Cover, it is a missing one.*

ADR-0007 records what follows from that — the backend fetches, stores and serves every Site's cover bytes — plus the alternatives that were rejected and, more importantly, the two places this deliberately departs from existing precedent:

- **Destination-class control instead of a host allowlist.** `fetchableSeriesURL` sets the allowlist precedent for `series_url`, and covers do not follow it. Cover hosts are CDNs that move independently of their Site — demonicscans serves its covers from `readermc.org` — so an allowlist would stop producing Covers the day a Site switched CDN, and that failure would look exactly like #47. The resolve-then-classify step is what actually stops the SSRF.
- **A public cover route where the kagane proxy is session-gated.** An `<img>` cannot send a bearer token, and it cannot be given one either: the panel's shadow root is `mode: "open"`, so the host page's JavaScript can read any `src` the script sets.

Both are security-adjacent departures, which is precisely why they are written down rather than left in a commit message.

## Scope

Documentation only — no code, no schema, no behaviour. The implementation is #56–#63.

## Why this should merge promptly rather than sit

All eight implementation tickets cite `docs/adr/0007-backend-hosts-cover-bytes.md` as the authority for decisions they must not relitigate, and they are written in the vocabulary this glossary entry defines. An agent picking up #56 reads both from `main`. Until this lands they get a 404 and either invent a rationale or stall — so this PR gates the tickets, not the other way round.

Related: #47 (bug), #55 (spec), #54 (deferred admin refetch).
Reviewed-on: #64
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 23:19:37 +07:00
sulthan 8081a0a5d8 Give the Tailscale ACL step a working policy file (#53)
Follow-up to #52, which merged before this landed. Docs only — no code, no compose changes.

`DEPLOY.md` §7 told the operator to "tag the two machines" and showed a bare `acls` fragment. Following it literally does not work and is actively harmful:

- the fragment references `tag:bookmark-api` / `tag:bookmark-browser` without a `tagOwners` section, so the policy is rejected on save;
- it never says how a tag gets onto a device (`tailscale up --advertise-tags=...`, which re-authenticates);
- replacing the tailnet's default allow-all with only that one rule **removes the operator's own SSH access to the browser machine**.

Replaced with a complete, saveable policy file: `tagOwners`, the CDP rule, a second rule preserving own-device access including `:22`, and a `tests` block so a later edit that widens 9222 is rejected rather than silently applied.

Also records two things that were assumed rather than stated:

- **Why tagging is load-bearing.** Tailscale has no `deny`, so restricting 9222 means removing the blanket accept and enumerating what remains. That is only expressible if the browser machine falls outside a selector that still covers your own devices — which is exactly what a tag does, since a tagged device has no user and stops matching `autogroup:member` / `autogroup:self`. Without that, the whole step reads as arbitrary ceremony.
- **Tagging replaces a device's user identity**, so it suits a dedicated box and disrupts a daily driver. Both paths are now written down.

Finally, separates two checks the old text conflated: the existing `curl` runs on the home machine and proves only the **bind**, because node-local traffic is not filtered. Proving the **ACL** needs a third device, so that check is now its own step.

Verified: `tailscale.com/docs/reference/syntax/policy-file` and `/docs/features/tags` (validated Apr 2026 / Dec 2025) for `tagOwners`, `autogroup:self` semantics vs tagged devices, `--advertise-tags` re-auth and key-expiry behaviour. Markdown fences balanced.
Reviewed-on: #53
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 16:12:03 +07:00
sulthan 2a3bb6922d Move the browser off the VPS to its own unit (#46) (#52)
Closes #46 once deployed.

The headless browser leaves the API stack and becomes its own compose unit
(`chrome/docker-compose.yml`) intended for the home machine, reached over the
tailnet. No fallback sidecar is left on the VPS.

The backend needs no code change — `BROWSER_WS_URL` was already the only
coupling. Its default is now empty rather than a pinned Docker IP, so an
unconfigured or unreachable browser degrades exactly as it always has: plain-TLS
libraries unaffected, kagane/novelfull logged and skipped, stored covers still
served.

### What shipped

- `chrome/docker-compose.yml` + `chrome/.env.example` — the browser unit, with
  the CDP port bound to `${BROWSER_BIND_ADDR}` (no default) and the resource
  limits from the epic: 512 MiB / 1 GiB memory+swap, `oom_score_adj 800`,
  halved CPU weight, shm 1 GiB -> 128 MiB.
- API stack drops the service, its `depends_on` and the `browser` network.
- `bookmark-api` gains the `default` network. Dropping `browser` had left it on
  `db` alone, which is `internal: true` — no published port and, worse, no
  egress for the poller at all. Caught by actually bringing the stack up.
- ADR-0006 for the topology; `DEPLOY.md` §7 for first-time setup of the browser
  machine; `REDEPLOY.md` §8 for its independent update cadence; architecture
  diagrams, config tables and troubleshooting rows across README/AGENTS/env.

### Verified locally

- Browser unit builds and runs: Chrome 151, UA carries no `HeadlessChrome`,
  all limits applied as declared.
- **Live smoke passes through the new unit**: `TestSmokeKaganeImage` fetched
  56710 bytes of `image/webp`, `TestSmokeKaganeGet` got a 200 with a real
  chapter list. The challenge cleared under the reduced 128 MiB shm.
- Bind isolation proven: refused on the host's non-loopback address, accepted
  on the configured one.
- 321 MiB peak of the 512 MiB cap after a full solve; 0 restarts, no OOM kill.
- API stack comes up clean, `/healthz` 200; egress confirmed present on
  `default` and absent on `db`.
- `go test ./...`, `go vet`, `gofmt` clean.

### Left to the operator

Provisioning the home machine, the Tailscale ACL, setting `BROWSER_WS_URL` in
production, and observing acceptance criteria 5-7 (covers with the machine off,
several days of zero OOM/restarts, VPS memory improvement). `DEPLOY.md` §7 now
carries the before/after `free -m` reading those need.

Reviewed-on: #52
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 15:28:21 +07:00
sulthan d1800d0707 Prefetch Kagane covers during latest polling (#51)
## Summary
- Add an optional browser-backed cover fetcher to the latest-chapter poller.
- Prefetch missing Kagane covers during the existing due-series cycle and persist them before a Reader opens the web UI.
- Keep chapter polling, cooldown stamping, and on-first-view fallback independent from cover failures.

## Behavior and safety
- Stored Kagane covers are detected before browser work, so later poll cycles do not refetch them.
- Nil cover fetchers and non-Kagane series retain the existing behavior.
- Shared Kagane image-id and content-type validation prevents challenge or non-image responses from poisoning persistent cover storage.
- The browser is wired into both the chapter and cover poller paths from the composition root.

## Verification
- `go test ./...`
- Focused latest, store, and web package tests
- Deterministic tests cover missing covers, cached covers, failed fetches, invalid content types, nil fetchers, and non-Kagane series.

Closes #45

Reviewed-on: #51
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 14:58:59 +07:00
sulthan 84cfd1b2c1 Make browser sidecar on-demand (#44) (#50)
Closes #44. Chrome now starts on first CDP connection, tracks concurrent helpers, reaps after 300 seconds idle, preserves the named profile, and classifies reap interruptions. Shutdown stops Chrome's process group so cookie batches flush. ADR-0005 records the measured constraints and decisions. Verification: docker build, live CDP wake, graceful stop cleanup, sh -n, and go test ./... (7 packages, 3 no tests).

Reviewed-on: #50
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 07:39:07 +07:00
sulthan bfae84c5c3 Persist kagane covers in Postgres (#49)
Closes #43

Persist kagane cover bytes in a dedicated Postgres covers table keyed by image ID. The web handler reads storage before the browser, writes validated fetches through, and no longer keeps an in-process cover cache. Added migration, store persistence tests including reopen, handler coverage for stored/miss/rejected paths, and corrected repository guidance.

Verification:
- go test ./...
- CGO_ENABLED=0 go build ./...

Reviewed-on: #49
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 06:52:02 +07:00
sulthan cd3a7e3d01 feat(latest): split browser poll cooldown (#48)
## Summary

Split latest-chapter polling cooldowns by fetch cost. Browser-backed kagane and novelfull series now rest longer without changing the cadence of plain-TLS sites.

## Behavior

- Plain-TLS series keep the 1h default cooldown.
- Browser-backed series use `LATEST_CHAPTER_POLL_BROWSER_COOLDOWN`, defaulting to 6h.
- Both cooldowns share the existing 15m minimum floor; invalid values retain the existing fallback behavior.
- The poller still selects both classes in one due query per cycle.
- Existing ordering and exclusions remain unchanged: reader-count precedence, least-recently-checked ordering, finished exclusion, archived polling, and orphan exclusion.

## Implementation

- Added the browser cooldown to backend configuration and passed it through production poller construction.
- Added the browser-site list as the single routing source used for both due-query cutoff selection and fetcher choice.
- Kept all query values parameterized; the site list is passed as a bound PostgreSQL array parameter.
- Updated startup logging to report interval, plain cooldown, browser cooldown, batch, and stagger.
- Documented the variable, default, and floor in `README.md`, `.env.example`, `backend/AGENTS.md`, and `docker-compose.yml`.

## Review findings addressed

The first review found that configuration parsing was correct but `startLatestPoller` did not pass `BrowserCooldown` into `latest.Poller`; every browser-backed row would therefore have been due immediately. Production construction now goes through `newLatestPoller`, with a regression test covering both cooldown fields.

The review also identified duplicated browser-site knowledge in fetch routing. `slices.Contains(browserBackedSites, site)` now reuses the same list already supplied to the store query.

## Verification

- Focused backend tests pass: `go test ./internal/latest ./internal/store .`.
- Full suite passes: `go test ./...`.
- `graphify update .` completed.
- Issue #42 was updated and closed.

Reviewed-on: #48
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 06:25:48 +07:00
63 changed files with 59752 additions and 1080 deletions
+134
View File
@@ -0,0 +1,134 @@
---
name: implement-tickets
description: "Orchestrate a batch of tickets: plan the briefs, then hand each ticket to its own implementer subagent in its own worktree."
disable-model-invocation: true
---
# Implement tickets
You are the **orchestrator**. You write briefs, dispatch, land results, and talk
to the tracker. You do not write the implementation — every line of ticket code
is written by a `ticket-implementer` subagent in its own git worktree. Reach for
the editor yourself only for a merge conflict resolution.
Ticket source and `tea` usage: `docs/agents/issue-tracker.md`. Codebase
questions: `graphify query "<question>"` before grepping.
## 1. Collect the tickets
The user's argument is the selector: issue numbers, a label, a parent issue, or
nothing. With nothing, take the open issues labelled `ready-for-agent`.
Fetch each with `tea issue <n> --comments`, and read the **whole** body —
acceptance criteria and the `Blocked by` line are what the rest of this skill
runs on. A ticket whose blockers are still open is out of this batch unless a
blocker is also in it.
## 2. Plan the batch
Explore enough of the codebase to write briefs a fresh context can act on: the
files each ticket lands in, the patterns it must follow, the `AGENTS.md`
invariants it touches.
Then decide three things:
- **Waves.** Blocking edges set the order; tickets with no open blocker inside
the batch share a wave. Cap each wave at **3** concurrent tickets unless the
user set another width.
- **Contracts.** Two tickets in one wave that meet at a function signature, a
JSON shape, a table column, or a token name: you decide the shape now and
write the identical wording into both briefs. A contract left for the
subagents to negotiate is a merge conflict you scheduled.
- **Splits.** A ticket too big for one fresh context window goes into the wave
as two briefs, or back to the user.
## 3. Get the plan approved
Present, and stop:
- the wave list, and for each ticket: number, title, one-line brief summary,
the files or areas it will touch, its verification commands
- every cross-ticket contract, verbatim as it will appear in the briefs
- anything you had to assume
Wait for approval. Apply the user's edits to the plan, do not relitigate them.
## 4. Run a wave
Per ticket, before dispatch:
```bash
git worktree add ../ticket-<n> -b ticket/<n>-<slug> <base> # base = the branch you are on
cp .env ../ticket-<n>/ 2>/dev/null # gitignored, worktrees do not get it
tea issue edit <n> --add-assignees <your gitea username> # tea login list has it
```
Write the brief to `.scratch/<batch-slug>/t<n>-brief.md` using the template
below, in the ubiquitous language of `CONTEXT.md` — a brief that says "scrape"
where the domain says Poll hands the subagent the wrong model of the system.
Then dispatch the whole wave in **one** `task` batch, every item on the
`ticket-implementer` agent. Each dispatch names: the absolute brief path, the
worktree path, the branch, the base ref, and the report path
`.scratch/<batch-slug>/t<n>-report.md`.
<brief-template>
# Ticket #<n> — <title>
**Read first.** `tea issue <n> --comments` for this ticket, then the issue it
refers to — the parent or spec — the same way. The comments carry decisions the
body never got updated with. This brief stays the requirements; those two reads
are the intent behind them.
**Goal.** The end-to-end behaviour this ticket makes work, from the user's side.
**Acceptance criteria.** Verbatim from the ticket.
**Contract.** The exact shared signatures / shapes / names this ticket must
implement or consume, and which sibling ticket is on the other end. Omit when
the ticket touches nothing shared.
**Where it lands.** The files and packages, and the existing pattern to follow
in each.
**Binding invariants.** The `AGENTS.md` rules this change can break — name them.
**TDD seams.** Where a test comes first — run the `tdd` skill at each one and
follow its red → green loop. Or "none — verify after".
**Verify.** The exact commands, e.g. `cd backend && go test ./...`,
`node --test userscript/test/logic.test.js`.
**Out of scope.** What not to touch, especially a sibling ticket's files.
</brief-template>
## 5. Land the wave
The wave is landed when every ticket in it is closed, reverted, or handed back
to the user. Per returned ticket:
| Status | What you do |
| --- | --- |
| `DONE` | merge, comment, close |
| `DONE_WITH_CONCERNS` | merge, comment the concerns, close only if you judge them non-blocking — otherwise leave open and tell the user |
| `BLOCKED` / `NEEDS_CONTEXT` | supply what is missing and re-dispatch, or hand back to the user with the specifics. Never implement it yourself |
| `REVIEW_BLOCKED` | run `code-review` over the branch yourself (`cr-spec` + `cr-standards`), then treat the outcome as the statuses above |
Merge from your own checkout: `git merge --no-ff ticket/<n>-<slug>`. A textual
conflict is yours to resolve (`resolving-merge-conflicts`). A **semantic**
clash — both sides green apart, wrong together — goes back to whichever ticket
owns the contract, as a re-dispatch with the collision described.
Then `tea comment <n> "<the report summary>"`, `tea issue close <n>`, and
`git worktree remove ../ticket-<n>`. Keep the report file.
Only once the whole wave is landed does the next wave start — its briefs may
need what this one changed.
## 6. Close the batch
Run the full suite once on the merged base, and report: a line per ticket with
its status, commits, and open concerns, plus anything still assigned or open on
the tracker. A red suite after every ticket went green is an interaction bug —
diagnose it, name the two tickets, and fix it or hand it back with both named.
@@ -14,7 +14,7 @@ parsers, helpers. UI, network, and storage behaviour are verified on-device.
```bash
node --check userscript/manga-bookmark.user.js # parse check, silent on success
node --test userscript/test/logic.test.js # 14 tests as of 2026-07-28
node --test userscript/test/logic.test.js # 35 tests as of 2026-08-10
```
Run both before every commit that touches the userscript.
@@ -31,7 +31,7 @@ The test file installs four globals **before** requiring the userscript:
|---|---|---|
| `localStorage` | `Map`-backed stub | `loadCache`, `loadQueue`, and the key-migration IIFE touch it at module scope |
| `location` | `{href, hostname, pathname, origin}` | read during boot |
| `document` | `querySelector` for `meta[property="…"]` only, plus a no-op `addEventListener` | adapters read `og:title`/`og:image` |
| `document` | `querySelector` for `meta[property="…"]` only, plus a no-op `addEventListener` | adapters read `og:title` (covers are the backend's, never scraped) |
| `document.body` | **left `undefined`** | this is the whole trick |
`document.body === undefined` sends the userscript's boot block down its `else`
+38 -29
View File
@@ -15,7 +15,7 @@ OWNER_DISCORD_ID=changeme-your-discord-user-id
# Comma-separated origins allowed to call the API (CORS). Both Asura domains
# plus Demonic, Comix, Kagane, and the two novel sites. Add/remove as the
# sites' hostnames change.
ALLOWED_ORIGINS=https://asuracomic.net,https://asurascans.com,https://demonicscans.org,https://comix.to,https://kagane.to,https://novelfull.com,https://lightnovelworld.net
ALLOWED_ORIGINS=https://asurascans.com,https://demonicscans.org,https://comix.to,https://kagane.to,https://novelfull.com,https://lightnovelworld.net
# Password for the bundled Postgres container, and therefore half of the
# DATABASE_URL compose builds for the backend. Generate one:
@@ -25,6 +25,15 @@ POSTGRES_PASSWORD=changeme-generate-a-long-random-password
# Override only to point the backend at a Postgres compose does not run.
# DATABASE_URL=postgres://user:pass@host:5432/bookmarks?sslmode=require
# Directory inside bookmark-api for immutable, content-addressed Cover bytes.
# Compose builds the image and mounts its named volume at this path.
COVER_DIR=/covers
# Public origin this deployment answers on, no trailing slash. Required: Cover
# URLs go out absolute, because the userscript renders them on a Site's own
# origin where a relative path would resolve against the Site (ADR-0007).
PUBLIC_BASE_URL=https://bookmark-api.example.com
# --- Prod override (Traefik) only ---
# Subdomain Traefik routes to this service (required by the prod override).
# BOOKMARK_API_HOST=bookmark-api.example.com
@@ -67,13 +76,15 @@ DISCORD_REDIRECT_URI=
# Set to 0 to turn it off entirely.
# LATEST_CHAPTER_POLL_ENABLED=1
#
# Two independent clocks. COOLDOWN is how long one series rests between checks;
# INTERVAL is how often the poller wakes up and looks for series past that
# cooldown. Shortening INTERVAL cannot shorten a COOLDOWN.
# LATEST_CHAPTER_POLL_COOLDOWN=1h # per series, floor 15m
# LATEST_CHAPTER_POLL_INTERVAL=10m # how often to wake
# LATEST_CHAPTER_POLL_BATCH=14 # series per wake
# LATEST_CHAPTER_POLL_STAGGER=20s # delay between fetches in a batch
# Two independent clocks. COOLDOWN is how long a plain-TLS series rests between
# checks; BROWSER_COOLDOWN is the longer rest for kagane and novelfull. INTERVAL
# is how often the poller wakes up and looks for series past their cooldowns.
# Shortening INTERVAL cannot shorten either cooldown.
LATEST_CHAPTER_POLL_COOLDOWN=1h # plain-TLS per series, floor 15m
LATEST_CHAPTER_POLL_BROWSER_COOLDOWN=6h # browser-backed per series, floor 15m
LATEST_CHAPTER_POLL_INTERVAL=10m # how often to wake
LATEST_CHAPTER_POLL_BATCH=14 # series per wake
LATEST_CHAPTER_POLL_STAGGER=20s # delay between fetches in a batch
#
# Uses a ticker, not an immediate first run: the first poll happens one
# INTERVAL after startup, not at startup. A container restarting more often
@@ -83,29 +94,27 @@ DISCORD_REDIRECT_URI=
# defaults. Beyond that the cadence stretches uniformly rather than breaking;
# raise BATCH or lower INTERVAL. Keep BATCH x STAGGER under INTERVAL.
# Headless-shell CDP endpoint for sites behind a JavaScript challenge (kagane).
# Unset disables browser polling; those sites then rely on the userscript alone.
# Leave commented — the compose files' own default (ws://172.28.0.10:9222) is
# correct. Do NOT set this to the "headless-shell" DNS name: Chrome's DevTools
# HTTP handler 500s any /json/version request whose Host header isn't an IP or
# "localhost", which silently breaks every kagane poll.
# BROWSER_WS_URL=ws://172.28.0.10:9222
# Clock zone the headless browser reports. A UTC clock is itself the bot
# signal — Cloudflare treats it as the datacenter default — and kagane's
# challenge then never clears. Measured 2026-08-08, identical container, one
# Indonesian egress IP: UTC never cleared in 60s (twice); Asia/Jakarta and
# America/New_York both cleared in 4s. So any real zone works; it does not
# have to match the IP's country, it just must not be UTC.
# CDP endpoint of the browser, used for the two sites behind a Cloudflare
# JavaScript challenge (kagane, novelfull) and by the web UI's kagane cover
# proxy. Unset disables browser polling and serves 404 for covers not already
# stored; those sites then rely on the userscript alone. That is also exactly
# how an unreachable browser degrades, so a home machine that is off costs
# chapter freshness and nothing else.
#
# Unset falls back to the host's /etc/timezone, which is a real zone whenever
# the host clock is set to local time. Set this when the host runs UTC — a UTC
# server is exactly the case that fails. Only the browser sidecar reads it —
# the backend's own zone is API_TZ below, and is cosmetic.
# BROWSER_TZ=Asia/Jakarta
# The browser does NOT run in this stack. It is its own compose unit on the
# home machine (chrome/docker-compose.yml, chrome/.env.example) and is reached
# over the tailnet, so set this to that machine's tailnet address:
#
# BROWSER_WS_URL=ws://100.x.y.z:9222
#
# It must be the tailnet **IP**, never a MagicDNS hostname and never the old
# Docker service name: Chrome's DevTools HTTP handler 500s any /json/version
# request whose Host header isn't an IP or "localhost", which silently breaks
# every kagane poll. Left unset here on purpose — a wrong default would poll a
# stranger's address, and "no browser" is a safe, self-announcing state.
# BROWSER_WS_URL=ws://100.x.y.z:9222
# Zone the backend stamps its log lines in. Cosmetic only — it exists so the
# API's logs read on the same clock as the browser sidecar's. Nothing else in
# Zone the backend stamps its log lines in. Cosmetic only. Nothing else in
# the service has a zone: bookmark timestamps are unix ms, and the two real
# time columns are timestamptz. Defaults to Asia/Jakarta; set to UTC for the
# conventional server default.
+6 -1
View File
@@ -5,8 +5,13 @@
backend/server
backend/backend
.playwright-mcp/
graphify-out/
# graphify map is committed; only regenerable/local parts are ignored
graphify-out/cost.json
graphify-out/cache/
graphify-out/[0-9][0-9][0-9][0-9]-[0-9][0-9]-[0-9][0-9]/
graphify-out/.rebuild.lock
plans/
.scratch/
docs/superpowers/
.superpowers/
go.work
+90
View File
@@ -0,0 +1,90 @@
---
name: ticket-implementer
description: Implements one ticket end to end inside its own git worktree - reads a brief file, implements, tests, commits, runs the two-axis code review through cr-spec and cr-standards, fixes findings, writes a report file, returns a short status contract. Dispatched by the implement-tickets skill.
model: opencode-go/minimax-m3
thinking-level: high
tools: read, write, edit, bash, grep, glob, lsp, todo, ast_edit, task
spawns: cr-spec,cr-standards
autoloadSkills: code-review, tdd
---
You implement **one ticket** dispatched by an orchestrator. Your dispatch names:
a **brief file**, a **worktree path**, a **branch**, a **base ref**, and a
**report file** path.
## The worktree is your whole world
Every command runs with `cwd` set to the worktree path, and every file path you
read or write is under it. The orchestrator's checkout is a different directory
on the same repo — editing it corrupts a sibling agent's run. If a command must
run elsewhere, say so in the report instead of doing it.
Your branch is already checked out there. Never `git checkout`, `git switch`,
`git rebase`, or `git worktree` anything.
## Order of work
1. Read the brief file. It is the single source of requirements — use its exact
values verbatim.
2. Read the ticket and the issue it refers to, as the brief's **Read first**
section names them: `tea issue <n> --comments` for each. The ticket's
comments and its parent carry the intent and the decisions behind the brief.
Read no other ticket and no other brief.
3. Read `AGENTS.md` in the worktree, plus the nested `AGENTS.md` for the area
you touch. Its invariants bind you: security rules, design system, comment
policy.
4. Ask before writing code if requirements, acceptance criteria, approach, or
dependencies are unclear. Asking is free; guessing is not.
5. Implement exactly what the brief specifies. At each TDD seam the brief names,
run the `tdd` skill and follow its red → green loop.
Follow the patterns already in the codebase; improve what you touch,
restructure nothing outside the ticket.
6. Verify. Focused tests while iterating, the brief's full verification commands
once at the end. Test output must be pristine.
7. Commit to your branch. Reference the ticket number in the subject.
8. Review (below), fix, re-verify, commit the fixes.
9. Write the report file, then return the status contract.
## Review
After your first green commit, run the **`code-review`** skill over
`<base ref>...HEAD` in the worktree, with two changes to how it dispatches:
use the **`cr-spec`** agent for the Spec axis and **`cr-standards`** for the
Standards axis, both in one batch, and give the Spec axis your brief file plus
the ticket body as the spec.
Fix every Critical and Important finding, then re-run the tests that cover the
amended code. Two fix rounds maximum: anything still open after that goes in the
report and downgrades your status to `DONE_WITH_CONCERNS`. Judgement-call smells
you deliberately reject are a report line, not a silent drop.
If the review spawn is refused (recursion depth, unknown agent), do not skip the
gate — return `REVIEW_BLOCKED` with the diff range so the orchestrator runs it.
## Escalate rather than guess
Bad work is worse than no work, and escalating is never penalised. Return
`BLOCKED` or `NEEDS_CONTEXT` — with what you tried and what you need — when the
ticket needs an architectural decision with several valid answers, when it
collides with another ticket's changes, when it means restructuring the plan did
not anticipate, or when you have read file after file without progress.
## Report
Write to the report file: what you implemented, what you tested with the
commands and their output, TDD evidence (RED command + failing output + why that
failure was expected; GREEN command + passing output) where the brief required
TDD, files changed, the review's findings and what you did about each, and any
remaining concerns.
Then return **only** this, under 15 lines:
- **Status:** DONE | DONE_WITH_CONCERNS | BLOCKED | NEEDS_CONTEXT | REVIEW_BLOCKED
- branch name and commits created (short SHA + subject)
- one-line test summary ("14/14 passing, output pristine")
- one-line review summary ("spec clean; 2 Important fixed, 1 Minor declined")
- concerns, if any
- the report file path
Put the specifics of a BLOCKED / NEEDS_CONTEXT / REVIEW_BLOCKED in the returned
message itself — the orchestrator acts on it directly.
+13 -18
View File
@@ -6,7 +6,7 @@ Guidance for OpenCode (and Claude Code) working in this repo.
Read-progress tracker for two libraries — manga and novels — behind one self-hosted Go backend. Two separate Violentmonkey userscripts inject on-page UI (floating button + slide-in panel) and sync progress, so bookmarks unify across sites and devices:
- `manga-bookmark.user.js` — **asurascans.com** (current domain; asuracomic.net 301s here), **demonicscans.org**, **comix.to**, **kagane.to**.
- `manga-bookmark.user.js` — **asurascans.com** (asuracomic.net is dropped: its deep links 301 to the asurascans.com root, discarding the path), **demonicscans.org**, **comix.to**, **kagane.to**.
- `novel-bookmark.user.js` — **novelfull.com**, **lightnovelworld.net**.
One backend, one `bookmarks` table: a `kind` column (`manga`|`novel`) splits the libraries and the web UI switches between them. Rows are keyed `<site>:<series_id>`.
@@ -19,8 +19,9 @@ Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free
- Every site is its **own origin with its own `localStorage`** — a shared remote store is the only way to unify bookmarks. Cloud sync required, not optional.
- Userscript run in **isolated world**, so embedded API token safe from site's JS.
- Cloudflare's block on manga sites **IP-reputation-based, not universal — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — Cloudflare's bot scoring can flip previously-clean IP without notice. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test.
- **kagane.to and novelfull.com are the exception to the above** — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`) and skips them entirely when that's unset. The four other sites poll fine over plain TLS.
- **The CDP sidecar must look like a real browser, and stock headless images don't.** Measured 2026-08-08 against kagane.to, all from the same IP: `chromedp/headless-shell:stable` never cleared the challenge in 90s (`navigator.webdriver` true, empty plugin list, Chromium-branded client hints — suppressing `webdriver` alone changed nothing); `zenika/alpine-chrome` ships Chrome 124, refused outright; real Chrome with the default `--headless=new` UA never cleared, because the UA says `HeadlessChrome`; real Chrome with a stock UA **and** a non-UTC clock zone cleared in ~4s. Hence `chrome/` — a Debian image with `google-chrome-stable`, a version-derived UA, and `TZ`/`BROWSER_TZ`. Chrome reads the zone *name* through ICU from `/etc/localtime`'s symlink target, ignoring the file's contents, so mounting the host's `/etc/localtime` does **not** work; `/etc/timezone` is mounted instead.
- **kagane.to and novelfull.com are the exception to the above** — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`). When that's unset, kagane is skipped entirely (a plain fetch would only retrieve a challenge page) while novelfull pages are still attempted over plain TLS — its challenge is a live time-varying fact and its cover bytes never need the browser. The four other sites poll fine over plain TLS.
- **The CDP browser must look like a real browser, and stock headless images don't.** Measured 2026-08-08 against kagane.to, all from the same IP: `chromedp/headless-shell:stable` never cleared the challenge in 90s (`navigator.webdriver` true, empty plugin list, Chromium-branded client hints — suppressing `webdriver` alone changed nothing); `zenika/alpine-chrome` ships Chrome 124, refused outright; real Chrome with the default `--headless=new` UA never cleared, because the UA says `HeadlessChrome`; real Chrome with a stock UA **and** a non-UTC clock zone cleared in ~4s. Hence `chrome/` — a Debian image with `google-chrome-stable`, a version-derived UA, and `TZ`/`BROWSER_TZ`. Chrome reads the zone *name* through ICU from `/etc/localtime`'s symlink target, ignoring the file's contents, so mounting the host's `/etc/localtime` does **not** work; `/etc/timezone` is mounted instead.
- **The browser is not in the API stack and must not be put back.** It's its own compose unit (`chrome/docker-compose.yml`) on a second machine, reached over the tailnet — it held 471 MiB on a 1974 MiB swapless VPS, and a residential egress scores better with Cloudflare anyway (ADR-0006). Consequences that constrain code: `BROWSER_WS_URL` must be a tailnet **IP** (a MagicDNS name 500s at `/json/version`, same trap as the old Docker service name); the CDP port binds to the tailnet address only, since CDP authenticates nothing and that host has a real LAN; and the browser is on-demand (ADR-0005), so an unreachable or asleep one must degrade exactly as an unset `BROWSER_WS_URL` — plain-TLS libraries unaffected, kagane/novelfull logged and skipped, stored covers still served. Never add `chromedp.NoModifyURL`: discovery per fetch is what makes a restarted Chrome invisible.
- **UTC is the tell, not a country mismatch.** A UTC clock is the datacenter default, so Cloudflare scores it as one; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while `Asia/Jakarta` **and** `America/New_York` both cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (`Asia/Bangkok`) rather than the measured egress. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one.
- **A challenged page needs the tab kept open.** The interstitial takes seconds to solve and only then writes clearance into the browser's shared cookie jar. Navigate-read-close never clears anything; `BrowserFetcher.run` holds one tab and re-reads until the payload arrives.
@@ -29,9 +30,13 @@ Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free
```
Two Violentmonkey userscripts (isolated world, per-site adapters, localStorage cache)
-- fetch() HTTPS --> reverse proxy (TLS + CORS) --> Go net/http --> Postgres (volume)
|
| CDP over tailnet
v
on-demand Chrome, separate machine (chrome/)
```
Backend-specific architecture (packages, endpoints, poller, config env vars) lives in `backend/AGENTS.md`. Userscript-specific structure (adapters, retry queue, UI, live URL shapes) lives in `userscript/AGENTS.md`.
Two deployable units on two machines: the API stack (`docker-compose.yml` + `docker-compose.prod.yml`, on the VPS) and the browser (`chrome/docker-compose.yml`, on the home machine). They share nothing but `BROWSER_WS_URL` and update independently. Backend-specific architecture (packages, endpoints, poller, config env vars) lives in `backend/AGENTS.md`. Userscript-specific structure (adapters, retry queue, UI, live URL shapes) lives in `userscript/AGENTS.md`. Deploy order `DEPLOY.md` (§7 for the browser), redeploy `REDEPLOY.md` (§8 for the browser).
## Commands
@@ -40,10 +45,10 @@ Backend (`cd backend`):
- Single test: `go test -run TestName ./...`
- Build static binary: `CGO_ENABLED=0 go build`
Local stack: `docker compose up` (bookmark-api + postgres + headless-shell; `postgres-data` named volume, `restart: unless-stopped`). The `headless-shell` service keeps its name but now builds `chrome/` — real Google Chrome, for the reason in the hard constraints above.
Local stack: `docker compose up` (bookmark-api + postgres only; `postgres-data` named volume, `restart: unless-stopped`). No browser — without `BROWSER_WS_URL` the poller logs and skips kagane and novelfull. To run one: `cd chrome && BROWSER_BIND_ADDR=172.17.0.1 docker compose up -d --build`, then `BROWSER_WS_URL=ws://172.17.0.1:9222` in the root `.env` (bridge gateway, so the API container can name it by IP).
Live CDP proof (needs a sidecar and network, skipped otherwise):
`SMOKE_BROWSER_WS_URL=ws://<host>:<port> go test -run TestSmokeKagane ./internal/latest`
Live CDP proof (needs that browser and network, skipped otherwise):
`SMOKE_BROWSER_WS_URL=ws://<ip>:<port> go test -run TestSmokeKagane ./internal/latest`
— fetches a real kagane cover and chapter list. A red run means the challenge is
not clearing from this IP, which is a live fact to re-check, not necessarily a defect.
@@ -135,12 +140,6 @@ Style: one dense comment over function beats one per line inside. Tight, no work
Test: "competent reader get this from code in few sec?" Yes → skip. Needs detour through another file/spec/git-blame → write it.
## Relevant skills
`multi-stage-dockerfile` and `docker-compose-orchestration` for container work (referenced in plan).
`golang-code-style`, `golang-error-handling`, `golang-performance`, `golang-testing` for backend Go work.
## Agent skills
`AGENTS.md` is the single source of truth for agent guidance; every `CLAUDE.md` in this repo is a symlink to the `AGENTS.md` beside it. Edit `AGENTS.md`.
@@ -162,11 +161,7 @@ Single-context: one root `CONTEXT.md` plus `docs/adr/`, both created lazily. See
Project has knowledge graph at graphify-out/ with god nodes, community structure, cross-file relationships.
Rules:
- For codebase questions, first run `graphify query "<question>"` when graphify-out/graph.json exists. Use `graphify path "<A>" "<B>"` for relationships and `graphify explain "<concept>"` for focused concepts. Return scoped subgraph, usually much smaller than GRAPH_REPORT.md or raw grep output.
- For codebase questions and exploration, always first run `graphify query "<question>"` when graphify-out/graph.json exists. Use `graphify path "<A>" "<B>"` for relationships and `graphify explain "<concept>"` for focused concepts. Return scoped subgraph, usually much smaller than GRAPH_REPORT.md or raw grep output.
- If graphify-out/wiki/index.md exists, use for broad navigation instead of raw source browsing.
- Read graphify-out/GRAPH_REPORT.md only for broad architecture review or when query/path/explain don't surface enough context.
- After modifying code, run `graphify update .` to keep graph current (AST-only, no API cost).
## Notes
- Keep comms terse — drop articles, fluff, pleasantries. Code/commits/security written normally.
+34 -8
View File
@@ -7,17 +7,31 @@ so progress survives across sites and devices.
## Language
**Series**:
One ongoing work — a manga or a novel — as published by a Site. Identified by its
stable slug on that Site, never by its title. A Series exists once and is shared by
every Reader who bookmarks it; it owns the facts that are true regardless of who is
reading — title, cover, Latest Chapter. A Reader cannot change them; they describe the
Series, not anyone's relationship to it.
One ongoing work — a manga or a novel — as published by a Site. Identified by the canonical
slug the Site itself publishes for it, never by its title and never by a Chapter Slug. A
Series exists once and is shared by every Reader who bookmarks it; it owns the facts that
are true regardless of who is reading — title, cover, Latest Chapter. A Reader cannot
change them; they describe the Series, not anyone's relationship to it.
_Avoid_: manga, title, book, comic
**Site**:
One third-party source a Series is published on. A Series on two Sites is two Series.
_Avoid_: source, host, provider, domain
**Chapter Slug**:
A slug a Site builds its chapter addresses from. Not an identity: one Series may have
several, any of them may differ from the slug that identifies the Series, and none is
computable from another. Only the Site's own links say which ones a Series uses, so a
Chapter Slug is always discovered, never derived.
_Avoid_: series slug, url slug, permalink, chapter path
**Cover**:
The image that stands for a Series wherever it is listed. A fact about the Series like
its title — one Cover per Series, shared by every Reader, never per-Reader. Defined by
what a Reader's browser can display, not by where the Site keeps the picture: an address
no client can load is not a Cover, it is a missing one.
_Avoid_: thumbnail, poster, image URL, artwork
**Reader**:
A person with their own Progress. Exactly one per set of credentials, so there is no
separate "account" concept to model — the credential belongs to the Reader.
@@ -42,9 +56,12 @@ is real activity, so only Progress reorders the list.
_Avoid_: position, bookmark (the noun is taken), last read
**Latest Chapter**:
The newest chapter a Site has published for a Series, discovered without the reader
present. Distinct from Progress in every way that matters: it is a fact about the Site,
not about the reader, and it must never reorder the list.
The highest-numbered chapter a Site has published for a Series, discovered without the
reader present. The number is what ranks it, never a date and never the Site's own
"newest chapter" banner — where a Site disagrees with itself, its list of chapters is
the record and its summary of that list is not. Distinct from Progress in every way
that matters: it is a fact about the Site, not about the reader, and it must never
reorder the list.
_Avoid_: newest, current chapter, update
**Poll**:
@@ -53,6 +70,15 @@ Reader present. Performed once per Series no matter how many Readers bookmarked
a Poll is work done on behalf of the Series, never on behalf of a Reader.
_Avoid_: scrape, refresh, check, sync
**Acquisition**:
The single read of a Series page made the moment the Series first exists, giving it
both its Latest Chapter and its Cover without waiting out the Poll queue. Distinct
from a Poll in the two ways that matter: a Reader is present — it is triggered by
their first Bookmark of that Series — and it is the only read that establishes a
Cover rather than refreshing facts. It happens once in a Series's life; every later
read of the same page is a Poll.
_Avoid_: initial poll, first fetch, prefetch, warm-up
**New Chapter**:
The state where Latest Chapter is ahead of Progress. The single condition the ember
accent is permitted to signal.
+229 -10
View File
@@ -43,7 +43,7 @@ TOKEN_KEY=<paste output of: openssl rand -hex 32>
OWNER_DISCORD_ID=<discord user id>
# CORS allowlist — leave as-is unless a site changes hostname.
ALLOWED_ORIGINS=https://asuracomic.net,https://asurascans.com,https://demonicscans.org,https://comix.to,https://kagane.to
ALLOWED_ORIGINS=https://asurascans.com,https://demonicscans.org,https://comix.to,https://kagane.to,https://novelfull.com,https://lightnovelworld.net
# Required — password for the bundled Postgres container. Compose builds the
# backend's DATABASE_URL out of it and has no fallback for either.
@@ -53,6 +53,15 @@ POSTGRES_PASSWORD=<paste output of: openssl rand -hex 24>
# not run; it then replaces the URL built from POSTGRES_PASSWORD above.
# DATABASE_URL=postgres://user:pass@host:5432/bookmarks?sslmode=require
# Required path inside bookmark-api. Compose builds the image and mounts the
# named cover-data volume at this path.
COVER_DIR=/covers
# Required — the origin this deployment answers on, no trailing slash. Cover
# URLs on the wire are absolute, because the userscript renders them on a
# Site's own origin (ADR-0007). Same host as BOOKMARK_API_HOST below.
PUBLIC_BASE_URL=https://bookmark-api.violetcrown.my.id
# Required for the Traefik override. Both have no fallback — compose refuses
# to start without them. BOOKMARK_WEB_HOST is required even if the web UI
# were unused; see 1b.
@@ -157,15 +166,17 @@ This merges the base file (build/image/env/volume) with the prod override
(no host port, Traefik network + router labels). Always pass **both** `-f`
flags — the prod file is not standalone.
Three services come up: `bookmark-api` (the backend), `postgres` (its database,
`postgres:17-alpine`), and `headless-shell`, a CDP sidecar the poller uses to
fetch kagane (behind a Cloudflare JS challenge). Neither of the latter two
publishes a port: `postgres` sits alone with `bookmark-api` on an
`internal: true` network, and `headless-shell` is reachable only over
`BROWSER_WS_URL`. A missing headless-shell just makes the poller skip kagane and
log it. A missing Postgres stops everything — `bookmark-api` waits for
`pg_isready` to pass, then applies its embedded migrations, and only then
listens. The schema is created that way; there is nothing to import by hand.
Two services come up: `bookmark-api` (the backend) and `postgres` (its
database, `postgres:17-alpine`). Postgres publishes no port — it sits alone
with `bookmark-api` on an `internal: true` network — and stops everything if it
is missing: `bookmark-api` waits for `pg_isready` to pass, then applies its
embedded migrations, and only then listens. The schema is created that way;
there is nothing to import by hand.
There is deliberately no browser here. Kagane and novelfull need one, and it
runs on a **separate machine** over the tailnet — §7. Until you do that step,
`BROWSER_WS_URL` is unset, the poller logs and skips those two sites, and
everything else works normally.
Check it's up and healthy:
@@ -259,6 +270,204 @@ copy immediately — reinstall on all devices, or they silently stop syncing.
---
## 7. The browser, on the home machine
Kagane and novelfull sit behind a Cloudflare JavaScript challenge no TLS
fingerprint clears, so the poller reaches them through a real Chrome over CDP.
That browser does **not** run on the VPS: it held 471 MiB of a 1974 MiB box
with no swap, and it scores better from a residential IP anyway (ADR-0006). It
is its own compose unit, deployed and updated independently of everything
above.
Do this after §2, on the second machine. Both machines must already be on the
same tailnet.
First, on the VPS, record what you are reclaiming — this is the whole point of
the move and there is no way to measure it afterwards:
```bash
free -m | awk '/^Mem:/ {print "available before:", $NF, "MiB"}'
```
Take it again after §7 is finished and the old sidecar is gone. Expect roughly
the sidecar's former footprint back (measured at 471 MiB working set, 595 MiB
cgroup).
**On the home machine:**
```bash
git clone <this repo> ~/mangaBookmark && cd ~/mangaBookmark/chrome
tailscale ip -4 # -> 100.x.y.z, this machine's tailnet IP
cp .env.example .env
echo "BROWSER_BIND_ADDR=$(tailscale ip -4)" >> .env
docker compose up -d --build
```
The clone is only for `chrome/`; nothing else on this machine reads the rest of
the repo. The unit is its own compose project (`bookmark-browser`), so it shares
no volume, network or lifecycle with an API stack that happens to sit beside it.
`BROWSER_BIND_ADDR` has no default on purpose. CDP authenticates nothing —
whatever reaches port 9222 drives the browser and, through it, this host — so
the bind address *is* the access control, backed by Tailscale device identity.
On the VPS that job was done by Docker network membership; this machine has a
real LAN, so `0.0.0.0` would be a hole punched into your home network. Compose
refuses to start rather than guess.
**Narrow it to the one device that needs it.** The bind address keeps CDP off
your LAN; it still leaves port 9222 open to every device on the tailnet, and
CDP has no login — a compromised phone is enough to drive this host. A new
tailnet's policy is allow-all, so this is the step that makes "Tailscale
identity is the access control" true rather than aspirational.
Tailscale has no `deny`, so a restriction is expressed by removing the blanket
grant and enumerating what is left. That only works if the browser machine can
be *excluded* from a selector that still covers your own devices — which is
what tagging buys: a tagged device has no user, so `autogroup:member` and
`autogroup:self` stop matching it. Tagging is the mechanism, not decoration.
In the admin console, under **Access controls**, the shipped policy grants
`{"src": ["*"], "dst": ["*"], "ip": ["*"]}`. Replace it:
```jsonc
{
"tagOwners": {
// Empty list: implicitly owned by the tailnet Owner/Admins, which is you.
"tag:bookmark-api": [],
"tag:bookmark-browser": [],
},
"grants": [
// The only thing on the tailnet that may drive the browser.
{
"src": ["tag:bookmark-api"],
"dst": ["tag:bookmark-browser"],
"ip": ["tcp:9222"],
},
// Your own devices reach your own devices, and the VPS, in full.
{
"src": ["autogroup:member"],
"dst": ["autogroup:self", "tag:bookmark-api"],
"ip": ["*"],
},
// On the browser machine you get SSH and nothing else. Widen this to `*`
// and the restriction above is void; delete it and you are locked out.
{
"src": ["autogroup:member"],
"dst": ["tag:bookmark-browser"],
"ip": ["tcp:22"],
},
// Uncomment if you route traffic through an exit node — dropping the
// blanket grant takes exit-node access with it.
// {"src": ["autogroup:member"], "dst": ["autogroup:internet"], "ip": ["*"]},
],
// Tagged devices left `autogroup:self`, so Tailscale SSH needs them named.
// Irrelevant if you reach these boxes with ordinary sshd over the tailnet —
// that is the `tcp:22` grant above.
"ssh": [
{
"action": "check",
"src": ["autogroup:member"],
"dst": ["autogroup:self", "tag:bookmark-api", "tag:bookmark-browser"],
"users": ["autogroup:nonroot", "root"],
},
],
// Run on every save, so a later edit that reopens 9222 is rejected outright.
"tests": [
{ "src": "tag:bookmark-api", "accept": ["tag:bookmark-browser:9222"] },
{
"src": "you@example.com",
"accept": ["tag:bookmark-browser:22"],
"deny": ["tag:bookmark-browser:9222"],
},
],
}
```
Then apply the tags — on the VPS and the home machine respectively:
```bash
sudo tailscale up --advertise-tags=tag:bookmark-api
sudo tailscale up --advertise-tags=tag:bookmark-browser
```
Each re-authenticates in a browser and issues a new node key; the tailnet IP is
unchanged, so `BROWSER_WS_URL` and `BROWSER_BIND_ADDR` still hold. Key expiry is
disabled once a device is tagged, which is what you want for a server — an
expired key would otherwise take the poller down every few months.
**Tagging replaces the device's user identity**, so do this only to machines
that exist to run these services. If your "home machine" is also your daily
driver, tag it anyway and reach it through the `:22` rule above, or skip the
tag and accept that any device of yours can reach CDP.
Enforcement is by the destination's packet filter, so the check below is real,
not advisory.
Prove the bind is tight, from the home machine itself:
```bash
curl -s -m 3 http://$(tailscale ip -4):9222/json/version # -> JSON
curl -s -m 3 http://<this machine's LAN IP>:9222/json/version
# -> curl: (7) Failed to connect ... Connection refused
```
The first call is also what wakes Chrome: it is not running until something
connects, and it is reaped again after five idle minutes. A cold first response
takes a few seconds; that is the browser starting, not a fault.
That check proves the *bind*, not the ACL — traffic that starts on the node is
not filtered. Prove the ACL from somewhere else: on your laptop or phone the
same URL must now time out, and from the VPS it must answer.
```bash
# on any other device of yours -> hangs until timeout
curl -s -m 5 http://<home machine tailnet IP>:9222/json/version
# on the VPS -> JSON
curl -s -m 20 http://<home machine tailnet IP>:9222/json/version
```
**On the VPS:**
```bash
cd ~/mangaBookmark
echo 'BROWSER_WS_URL=ws://100.x.y.z:9222' >> .env # the home machine's tailnet IP
docker compose -f docker-compose.yml -f docker-compose.prod.yml up -d
```
It must be the tailnet **IP**. A MagicDNS hostname fails: Chrome's DevTools HTTP
handler answers `/json/version` with a 500 for any `Host` header that is not an
IP or `localhost`, and the failure looks like a broken site rather than a broken
hostname.
**Prove it end to end.** This is the only check that says the challenge actually
clears from that machine's egress — it fetches a real kagane cover and a real
chapter list:
```bash
cd backend
SMOKE_BROWSER_WS_URL=ws://100.x.y.z:9222 go test -run TestSmokeKagane ./internal/latest
```
A red run means "not clearing from this address right now", which is a live
fact to re-check before it is a defect — Cloudflare's scoring moves. Then, from
the web UI, open a bookmarked kagane series and confirm the cover renders. Once
a cover is stored it is served from Postgres forever after, so the browser being
asleep, unreachable, or mid-power-outage costs chapter freshness and nothing
visible.
Finally, take the VPS `free -m` reading again and compare it against the one
from the top of this section.
**Updating the browser** is independent of the API stack and has its own
runbook — `REDEPLOY.md` §8.
---
## Updating
Pull new code, then rebuild:
@@ -272,6 +481,10 @@ server predates the Postgres migration, the old SQLite volume `bookmarks-data`
is still on disk and deliberately undeclared in compose so `down -v` cannot take
it; see `REDEPLOY.md` §1 for when to remove it.)
The browser is a separate unit on a separate machine with its own update
command — §7. Nothing above touches it, and it needs no coordination: the API
picks up a restarted Chrome's new debugger UUID by itself.
---
## Troubleshooting
@@ -287,6 +500,12 @@ it; see `REDEPLOY.md` §1 for when to remove it.)
| `compose ... config` errors about `TOKEN_KEY`, `OWNER_DISCORD_ID` or `POSTGRES_PASSWORD` | Run compose from the dir with `.env`, or export the vars. All three are required and none has a fallback. |
| `bookmark-api` restarts in a loop, `password authentication failed for user "bookmarks"` | `POSTGRES_PASSWORD` was changed after first boot; Postgres only applies it to an empty `postgres-data`. Restore the old value, or reset the role (`REDEPLOY.md` troubleshooting). |
| `bookmark-api` never logs `listening on :8080` | It is blocked on `postgres` passing `pg_isready`, or a migration failed. `docker compose -f docker-compose.yml -f docker-compose.prod.yml logs postgres`. |
| kagane rows never get a `latest_chapter`; log says `browser fetcher disabled` or nothing at all | `BROWSER_WS_URL` unset. Expected before §7 is done. |
| kagane polls all fail; log shows a 500 from `/json/version` | `BROWSER_WS_URL` names a MagicDNS hostname (or any name). Chrome's DevTools handler only accepts an IP or `localhost` — use the tailnet IP. |
| kagane polls fail with a connection error | Home machine off, off the tailnet, or the unit is down. `tailscale ping <machine>`, then `docker compose ps` in its `chrome/`. Costs freshness only; stored covers keep serving. |
| kagane cover is a placeholder for a newly bookmarked series | Its cover has never been fetched and the browser is unreachable. It fills in on the next successful poll of that series (up to `LATEST_CHAPTER_POLL_BROWSER_COOLDOWN`, default 6h). |
| `compose` in `chrome/` errors `set BROWSER_BIND_ADDR to this machine's tailnet IP` | No `chrome/.env`, or the variable is empty. Deliberate — it has no default so an unset value cannot publish CDP to the LAN. |
| browser container restarts, or is OOM-killed | `docker inspect bookmark-browser --format '{{.RestartCount}} {{.State.OOMKilled}}'`. The 512 MiB cap is sized against a measured 645 MiB untuned peak; a real breach is a Chrome regression worth reading `docker logs` for, not a number to raise reflexively. |
Backend config reference and endpoint list: see `README.md`.
+51 -13
View File
@@ -1,6 +1,6 @@
# Manga Bookmark
Track manga read-progress on **asurascans.com** (a.k.a. asuracomic.net),
Track manga read-progress on **asurascans.com**,
**demonicscans.org**, **comix.to**, and **kagane.to** from a phone (Bromite /
mobile Chromium), synced to a self-hosted Go backend so bookmarks unify across
all four sites and all devices.
@@ -15,8 +15,21 @@ Two parts:
```
Bromite userscript (isolated world, Shadow DOM UI, localStorage cache)
-- fetch() HTTPS --> reverse proxy (TLS + CORS) --> Go net/http --> Postgres (volume)
|
| CDP over the tailnet
v
headless Chrome, on-demand,
on a separate machine
(chrome/, ADR-0006)
```
Kagane and novelfull sit behind a Cloudflare JavaScript challenge no TLS
fingerprint clears, so the poller reaches those two through a real Chrome over
CDP. That browser is **not** part of the API stack: it is its own compose unit
on a second machine, spawned on the first connection and reaped when idle. The
API needs it only to discover new chapters and to fetch a kagane cover once —
covers are stored, so the library renders in full with the browser switched off.
---
## 1. Backend
@@ -29,8 +42,9 @@ Bromite userscript (isolated world, Shadow DOM UI, localStorage cache)
| `OWNER_DISCORD_ID` | *(required)* | Discord user ID of the owner: seeded as the first Reader, owns every pre-registration bookmark, and is the only Reader who can revoke another's sessions. |
| `ALLOWED_ORIGINS` | Asura + Demonic + Comix + Kagane origins | Comma-separated CORS allowlist. |
| `DATABASE_URL` | *(required)* | Postgres connection URL, e.g. `postgres://bookmarks:…@postgres:5432/bookmarks?sslmode=disable`. Compose builds it from `POSTGRES_PASSWORD`. |
| `COVER_DIR` | *(required)* | Filesystem volume for immutable, content-addressed Cover bytes. Compose builds the image and mounts `cover-data` at this path; standalone runs may choose another writable durable path. |
| `PORT` | `8080` | Plain HTTP; TLS terminated by the proxy. |
| `BROWSER_WS_URL` | `ws://172.28.0.10:9222` | Headless-shell CDP endpoint used to poll Kagane past its JS challenge. Must be an IP or `localhost` — Chrome's DevTools handler 500s any other Host header. |
| `BROWSER_WS_URL` | empty | CDP endpoint of the remote browser (`ws://<tailnet IP>:9222`), used to poll Kagane/Novelfull past their JS challenge and to fetch uncached Kagane covers. Must be an IP or `localhost` — Chrome's DevTools handler 500s any other Host header, MagicDNS names included. Unset disables both; stored covers still serve. |
| `DISCORD_CLIENT_ID` | *(required)* | Discord application credentials for the browser sign-in (ADR-0002). |
| `DISCORD_CLIENT_SECRET` | *(required)* | As above. Never logged, never echoed in an error. |
| `DISCORD_GUILD_ID` | *(required)* | The one guild whose membership gates sign-in, checked at login only. Membership *is* registration: any member becomes a Reader on first login. |
@@ -40,8 +54,9 @@ Bromite userscript (isolated world, Shadow DOM UI, localStorage cache)
| `USERSCRIPT_PATH` | `/userscript/manga-bookmark.user.js` | Bindmounted file served at `/u/{token}/manga-bookmark.user.js`. |
| `NOVEL_USERSCRIPT_PATH` | `/userscript/novel-bookmark.user.js` | Same, for the novel library. |
| `LATEST_CHAPTER_POLL_ENABLED` | `1` | `0` turns the poller off entirely. |
| `LATEST_CHAPTER_POLL_COOLDOWN` | `1h` | Rest between checks of one series; floor `15m`. |
| `LATEST_CHAPTER_POLL_INTERVAL` | `10m` | How often the poller wakes. Cannot shorten a cooldown. |
| `LATEST_CHAPTER_POLL_COOLDOWN` | `1h` | Rest between checks of one plain-TLS series; floor `15m`. |
| `LATEST_CHAPTER_POLL_BROWSER_COOLDOWN` | `6h` | Rest between checks of one browser-backed series; floor `15m`. |
| `LATEST_CHAPTER_POLL_INTERVAL` | `10m` | How often the poller wakes. Cannot shorten either cooldown. |
| `LATEST_CHAPTER_POLL_BATCH` | `14` | Series per wake. Keep `BATCH × STAGGER` under `INTERVAL`. |
| `LATEST_CHAPTER_POLL_STAGGER` | `20s` | Delay between fetches in a batch — this is the outbound request rate. |
@@ -49,8 +64,11 @@ Compose reads a few more from the same `.env` that the backend never sees:
`POSTGRES_PASSWORD` (required — `DATABASE_URL` is built from it, and Postgres
only applies it while `postgres-data` is empty), `BOOKMARK_API_HOST` and
`BOOKMARK_WEB_HOST` (required by the prod override), and the optional
`PROXY_NETWORK` / `TRAEFIK_ENTRYPOINT` / `TRAEFIK_CERTRESOLVER`. Full commentary
is in `.env.example`; deployment order is `DEPLOY.md`.
`PROXY_NETWORK` / `TRAEFIK_ENTRYPOINT` / `TRAEFIK_CERTRESOLVER`. The browser
unit has its own `chrome/.env` on its own machine — `BROWSER_BIND_ADDR`
(required, the tailnet IP the CDP port is published on) and the optional
`BROWSER_TZ`. Full commentary is in `.env.example` and `chrome/.env.example`;
deployment order is `DEPLOY.md`.
### Endpoints
@@ -96,6 +114,23 @@ cp .env.example .env
docker compose up -d --build # binds 127.0.0.1:8080
```
That brings up two services — the API and Postgres. The browser is deliberately
not one of them; without `BROWSER_WS_URL` the poller logs and skips kagane and
novelfull, and everything else works. To run one locally, publish it on the
Docker bridge gateway so the API container can name it by IP:
```bash
cd chrome
echo 'BROWSER_BIND_ADDR=172.17.0.1' > .env
docker compose up -d --build
# then in the repo's own .env: BROWSER_WS_URL=ws://172.17.0.1:9222
```
Bind it to `127.0.0.1` instead if you only want to drive it from the host, e.g.
`SMOKE_BROWSER_WS_URL=ws://127.0.0.1:9222 go test -run TestSmokeKagane ./internal/latest`.
In production that address is the home machine's tailnet IP and nothing else —
see `DEPLOY.md` §7 and ADR-0006.
Smoke test:
```bash
@@ -226,8 +261,10 @@ an API.
## Adapter reference (verified live 2026-07-24)
The site adapters key everything off URL regex, with `title`/`cover` from
`og:title` / `og:image`. Confirmed against live pages via Playwright:
The site adapters key everything off URL regex, with `title` from `og:title`
(or the page's own heading where a site ships none). No adapter reads a cover:
the backend acquires, stores and serves every Cover from its own origin
(ADR-0007). Confirmed against live pages via Playwright:
| Site | Series URL | Chapter URL | `series_id` |
|------|-----------|-------------|-------------|
@@ -237,11 +274,12 @@ The site adapters key everything off URL regex, with `title`/`cover` from
| **Kagane** (`kagane.to`) | `/series/<uuid>` | `/series/<uuid>/reader/<bookUuid>` | `<uuid>` |
Notes:
- **`asuracomic.net` deep links are dead (re-checked 2026-07-25).** They 301 to
the `asurascans.com` **root**, discarding the path, at the edge — before the
userscript gets a document — so nothing client-side can rescue them. Reach
series through `asurascans.com`. The host stays matched in case the redirect
starts preserving paths again.
- **`asuracomic.net` is no longer matched (deep links dead, re-checked
2026-07-25).** They 301 to the `asurascans.com` **root**, discarding the path,
at the edge — before the userscript gets a document — so nothing client-side
can rescue them. The backend rejects stored addresses on that host too, since
the poller pins each Site to one hostname. Reach series through
`asurascans.com`.
- Asura `og:title` carries a `Chapter N - Read Online \| Asura Scans` suffix that
the adapter strips; Demonic chapter `og:title` is `<Title> Chapter N`.
- Demonic's `<slug>` is identical on `/manga/…` and the canonical `/title/…`
+63 -5
View File
@@ -59,9 +59,13 @@ network can reach it — so every command below goes in through the container:
```bash
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks -c '\dt'
# -> bookmarks, readers, schema_migrations, series, sessions
# -> bookmarks, covers, readers, schema_migrations, series, sessions
```
The `covers` table is metadata only after the filesystem cutover: bytes live in
the separate `cover-data` volume. Back that volume up with the database dump;
restoring only Postgres leaves stored Cover addresses without files.
Inside the container that connects over the local socket as the `bookmarks`
superuser, so no password is needed anywhere in this section. `-T` is not
optional: without it Compose allocates a TTY, which rewrites `\n` to `\r\n` and
@@ -99,10 +103,11 @@ you will act as though you have one:
docker run --rm -v "$BACKUP_DIR":/backup postgres:17-alpine \
pg_restore --list "/backup/bookmarks-$STAMP.dump" | grep 'TABLE DATA'
# -> 1234; 0 0 TABLE DATA public bookmarks bookmarks
# -> 1235; 0 0 TABLE DATA public readers bookmarks
# -> 1236; 0 0 TABLE DATA public schema_migrations bookmarks
# -> 1237; 0 0 TABLE DATA public series bookmarks
# -> 1238; 0 0 TABLE DATA public sessions bookmarks
# -> 1235; 0 0 TABLE DATA public covers bookmarks
# -> 1236; 0 0 TABLE DATA public readers bookmarks
# -> 1237; 0 0 TABLE DATA public schema_migrations bookmarks
# -> 1238; 0 0 TABLE DATA public series bookmarks
# -> 1239; 0 0 TABLE DATA public sessions bookmarks
# 2. Sanity-check the live row count you just captured.
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks \
@@ -387,6 +392,55 @@ panel works on the phone.
---
## 8. The browser unit (separate machine, separate cadence)
Everything above is the API stack on the VPS. The headless browser is its own
compose unit on the home machine (ADR-0006, `DEPLOY.md` §7) and is redeployed
on its own schedule — it holds no data you can lose, so there is nothing to
back up and no ordering constraint against the API.
```bash
cd ~/mangaBookmark/chrome
git pull --ff-only
docker compose up -d --build
```
Then confirm it answers, and that a stopped-and-restarted Chrome is invisible
to the API:
```bash
curl -s -m 15 http://$(tailscale ip -4):9222/json/version | head -c 120
# -> {"Browser":"Chrome/1xx...","webSocketDebuggerUrl":"ws://...<new uuid>"}
```
The first call takes a few seconds: Chrome is not running until something
connects, and it is reaped again after five idle minutes. The debugger UUID
changes on every start and the API does not care — chromedp re-runs
`/json/version` discovery per fetch, which is exactly why `chromedp.NoModifyURL`
must never be added to `browser.go`.
**Rebuild is the Chrome upgrade path.** The image installs
`google-chrome-stable` unpinned on purpose: a stale browser is what Cloudflare
turns away, and the pinned Chrome 124 in `zenika/alpine-chrome` is the worked
example. The `chrome-profile` volume survives `--build`, so clearance cookies
are reused rather than re-solved.
Two things worth a glance after several days, both from the acceptance criteria
of the move:
```bash
docker inspect bookmark-browser --format '{{.RestartCount}} {{.State.OOMKilled}}'
# -> 0 false
free -m # the Gitea runner should still have its headroom
```
Nothing here needs doing during an API redeploy. The API stack does not
`depends_on` the browser, and an unreachable one degrades exactly as an unset
`BROWSER_WS_URL`: plain-TLS libraries unaffected, kagane and novelfull logged
and skipped, stored covers still served.
---
## Troubleshooting
| Symptom | Cause / fix |
@@ -404,6 +458,10 @@ panel works on the phone.
| `pg_restore`: `cannot drop … other objects depend on it` / `being accessed by other users` | Live connections block `--clean`. `$COMPOSE stop bookmark-api` first (§6). If they persist: `$COMPOSE exec -T postgres psql -U bookmarks -d postgres -c "select pg_terminate_backend(pid) from pg_stat_activity where datname='bookmarks' and pid <> pg_backend_pid()"`. |
| Dump is 0 bytes, or `pg_restore`: `did not find magic string in file header` | You ran `exec` without `-T`. The allocated TTY rewrites newlines in the binary stream and corrupts the archive in flight (§1). |
| `git pull`: `could not read Username for 'https://…'` | The checkout's remote is the HTTPS clone URL and the server has no credential helper, so the pull prompts into a closed stdin. Switch it to SSH once — `git remote set-url origin ssh://git@gitea.violetcrown.my.id:2222/sulthan/mangaBookmark.git`. Gitea's SSH listens on **2222**, not 22; port 22 is the host's own sshd and answers `Permission denied (publickey)` no matter which key is registered. |
| kagane rows stopped updating after a redeploy | Check `BROWSER_WS_URL` survived the `.env` edit and still names the home machine's tailnet **IP**. A hostname 500s at `/json/version`; an empty value disables the browser silently. Plain-TLS sites keep working either way, which is why this is easy to miss. |
| kagane covers went blank in the web UI | Covers use the `cover-data` volume now. Restore/check that volume alongside Postgres; rows in `covers` are metadata only. If the database has rows but files are missing, the next browser-backed request refetches them; without a browser it remains a 404. |
| Browser unit will not start: `set BROWSER_BIND_ADDR to this machine's tailnet IP` | `chrome/.env` is missing or the variable is empty. It has no default on purpose — an unset value must fail the deploy rather than publish an unauthenticated CDP port to the LAN. |
| `bookmark-browser` shows `OOMKilled true` | The cap did its job. Read `docker logs bookmark-browser` before raising it — the sizing and what the cap protects are in ADR-0006. |
Full first-time setup: `DEPLOY.md`. The one-off SQLite→Postgres move:
`CUTOVER.md`. Config reference and endpoints: `README.md`.
+51 -18
View File
@@ -91,13 +91,37 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
Fetches use `bogdanfinn/tls-client` with Chrome profile as defence in depth
against fingerprint-based blocking; any failure log and skip. kagane and
novelfull sit behind Cloudflare JavaScript challenges the TLS client can't
clear, so they are browser-only: fetched over CDP via `BROWSER_WS_URL`, and
simply not polled when that's unset. See
clear, so they are fetched over CDP via `BROWSER_WS_URL`; kagane is simply
not polled when that's unset, while novelfull falls back to a plain-TLS
attempt — its challenge is a live time-varying fact, and its cover bytes
never need the browser. See
`docs/superpowers/specs/2026-07-26-server-latest-chapter-polling-design.md`.
The poller's series write is a single-column UPDATE
(`Store.SetLatestChapter`), not a read-modify-write of the whole bookmark:
it cannot revert read progress or move `updated_at`, so the old
stale-re-read race is gone with the Get+Upsert flow.
- **Covers are acquired at creation, then served from our own origin
(ADR-0007):** the first Bookmark of a Series fires `Store.OnSeriesCreated`,
which `latest.Acquirer` turns into one series-page fetch yielding both the
Latest Chapter and the cover URL; the bytes then go through
`latest.CoverBytesFetcher` into `Store.SetSeriesCover`. It runs in a
goroutine — the Reader's PUT must neither block on a Site nor fail with one
— and every failure is logged and dropped, leaving the Bookmark intact. The
wire's `cover` is the absolute `PUBLIC_BASE_URL + /covers/{sha256}` once
bytes exist and `""` before, never an address that 404s. `GET /covers/{addr}`
is public and uncredentialed: the userscript renders it on a Site's origin,
where no cookie or token of ours travels. A client-sent `cover` is decoded
and discarded, permanently (ADR-0004 compatibility).
Browser-backed Sites join the same pipeline (issue #62): kagane pages *and*
cover bytes go through the browser sidecar (nothing falls back to a plain
fetch, which would only retrieve a challenge page), while novelfull needs
the browser only for its HTML — the cover URL comes out of the
browser-fetched page and the bytes go over plain TLS. With no browser
configured, kagane Covers are simply absent; novelfull still gets one — at
creation and on the poll — when its page body happens to answer a plain
request (the challenge is a live time-varying fact). The old kagane-only
serving path (`/img/kagane/{id}`, template rewrite, `CoverFetcher`) is gone
(issue #63): the one public route serves every Site.
- **`updated_at` drives list order, so moves only on real reading progress:** server apply its timestamp when row new or `last_chapter_num` changes, else keep stored value — favouriting series or recording newly published chapter must not reorder list. `PUT` therefore returns row **as stored**, clients must adopt that response rather than own payload. See `plans/2026-07-25-bookmark-list-favorites-design.md` §4.
- **Lifecycle buckets:** `status` on each bookmark is `reading` | `archived` |
`finished`, orthogonal to `favorite`. Archived and finished appear only in
@@ -114,32 +138,41 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
the owner of every pre-registration bookmark; required),
`ALLOWED_ORIGINS` (comma list),
`DATABASE_URL` (Postgres connection URL, required — no default),
`COVER_DIR` (required filesystem volume for content-addressed Cover bytes),
`PUBLIC_BASE_URL` (required origin this deployment answers on, trailing
slash trimmed; every Cover URL on the wire is built from it, absolute
because the userscript renders on a Site's origin — ADR-0007),
`PORT` (default `8080`), `DISCORD_CLIENT_ID`/`_CLIENT_SECRET`/`_GUILD_ID`/
`_REDIRECT_URI` (required; Discord OAuth for the browser UI),
`DISCORD_REQUIRED_ROLE` (optional role gate, empty by default),
`DISCORD_API_BASE` (default `https://discord.com/api/v10`),
`LATEST_CHAPTER_POLL_ENABLED`/`_COOLDOWN`/`_INTERVAL`/`_BATCH`/`_STAGGER`
(background latest-chapter poller; defaults on, `1h`/`10m`/`14`/`20s`).
`LATEST_CHAPTER_POLL_ENABLED`/`_COOLDOWN`/`_BROWSER_COOLDOWN`/`_INTERVAL`/
`_BATCH`/`_STAGGER` (background latest-chapter poller; defaults on,
`1h` plain-TLS cooldown, `6h` browser cooldown, `10m`/`14`/`20s`; both
cooldowns have a `15m` floor).
`USERSCRIPT_PATH` and `NOVEL_USERSCRIPT_PATH` (files served at
`/u/{token}/manga-bookmark.user.js` and `/u/{token}/novel-bookmark.user.js`,
defaults `/userscript/manga-bookmark.user.js` and
`/userscript/novel-bookmark.user.js`, both supplied by bindmount; the
`__API_TOKEN__` placeholder inside them is substituted with the requesting
Reader's credential at serve time).
`BROWSER_WS_URL` (CDP endpoint of the `chrome/` sidecar, used by the poller
for kagane and novelfull *and* by the web UI's kagane cover proxy; unset
disables browser polling and serves 404 from the proxy, leaving those sites
to the userscript alone).
- **kagane covers are proxied, not hot-linked:** kagane serves cover images
behind the same challenge as its pages and with
`cross-origin-resource-policy: same-origin`, so no `<img>` on the web UI's
origin can load one — not even from a browser holding the clearance cookie
(verified 2026-08-08). `Bookmark.CoverURL` rewrites a stored kagane
`og:image` to `/img/kagane/{id}`, served by `internal/web/cover.go` through
`latest.BrowserFetcher.Image` and memoised in-process. The templates render
`.CoverURL`, never `.Cover`. The id is matched against a UUID regex before it
reaches the browser: the stored value is client-supplied, so an unchecked one
is an SSRF primitive pointed at the deployment's own network.
`BROWSER_WS_URL` (CDP endpoint of the browser, which runs on a **separate
machine** and is reached over the tailnet — ADR-0006, `chrome/docker-compose.yml`.
Used by the poller for kagane and novelfull page fetches and by the cover
pipeline for kagane's image bytes (the browser is the only route that clears
the challenge kagane serves its covers behind); unset — the default —
disables browser polling and leaves kagane Covers blank until stored bytes
exist. Must be a tailnet IP, never a hostname: Chrome's DevTools handler 500s
`/json/version` for any Host that isn't an IP or `localhost`).
- **No per-Site cover path (issue #63):** every Cover — all six Sites — is
served by the one public `GET /covers/{addr}` route from content-addressed
bytes. There is no proxy, no per-Site rewrite, no second place that decides
a Cover's renderable address: the wire `cover` is it. The only place a Site
name still appears in cover code is the extraction module (`latest`), where
kagane's image URLs are claimed by `browserOnlyCoverURL` — they answer a
plain fetch with a challenge and `cross-origin-resource-policy: same-origin`;
every other Site's CDN answers plain TLS. Templates render `.Cover` — the
wire value — never anything else.
- **Web UI also owns:** session-gated `GET /install/{manga,novel}-bookmark.user.js`
(renders the bindmounted script with the acting Reader's derived credential
substituted in — the credential never appears in page markup, the address
+6 -1
View File
@@ -2,6 +2,7 @@
# --- build stage: compile a static, CGO-free binary ---
FROM golang:1.26-alpine AS build
ARG COVER_DIR=/covers
WORKDIR /src
# Dependencies first for layer caching (changes rarely).
@@ -19,11 +20,15 @@ COPY internal/ ./internal/
# -trimpath + -ldflags strip paths and debug info for a smaller image.
RUN CGO_ENABLED=0 GOOS=linux go build -trimpath -ldflags="-s -w" -o /out/server .
# Create the source directory; runtime COPY sets ownership for the named volume.
RUN mkdir -p "$COVER_DIR"
# --- runtime stage: distroless static, non-root ---
FROM gcr.io/distroless/static:nonroot
ARG COVER_DIR=/covers
WORKDIR /
COPY --from=build --chown=65532:65532 ${COVER_DIR} ${COVER_DIR}
COPY --from=build /out/server /server
EXPOSE 8080
USER nonroot:nonroot
ENV PORT=8080
+14 -8
View File
@@ -26,10 +26,14 @@ const testTokenKey = "test-token-key"
// credential is a function of it.
const testDiscordID = "test-owner"
// testCoverBaseURL is the public origin cover URLs are built from, standing in
// for PUBLIC_BASE_URL.
const testCoverBaseURL = "https://bookmarks.test"
func testConfig() Config {
return Config{
TokenKey: testTokenKey,
AllowedOrigins: []string{"https://asuracomic.net", "https://demonicscans.org"},
AllowedOrigins: []string{"https://asurascans.com", "https://demonicscans.org"},
Port: "8080",
}
}
@@ -60,7 +64,7 @@ func newTestStoreURL(t *testing.T) (*store.Store, string) {
url := pgtest.URL(t)
s, err := store.Open(url, store.Owner{
DiscordID: testDiscordID, TokenHash: token.Hash(ownerCredential()),
})
}, t.TempDir(), testCoverBaseURL)
if err != nil {
t.Fatalf("store.Open: %v", err)
}
@@ -165,7 +169,7 @@ func TestAuthAccepted(t *testing.T) {
func TestCORSPreflight(t *testing.T) {
srv := newTestServer(t)
req := httptest.NewRequest(http.MethodOptions, "/bookmarks/asura:foo-1", nil)
req.Header.Set("Origin", "https://asuracomic.net")
req.Header.Set("Origin", "https://asurascans.com")
req.Header.Set("Access-Control-Request-Method", "PUT")
rr := httptest.NewRecorder()
srv.ServeHTTP(rr, req)
@@ -173,7 +177,7 @@ func TestCORSPreflight(t *testing.T) {
if rr.Code != http.StatusNoContent {
t.Fatalf("preflight status = %d, want 204", rr.Code)
}
if got := rr.Header().Get("Access-Control-Allow-Origin"); got != "https://asuracomic.net" {
if got := rr.Header().Get("Access-Control-Allow-Origin"); got != "https://asurascans.com" {
t.Fatalf("Allow-Origin = %q, want reflected origin", got)
}
if got := rr.Header().Get("Access-Control-Allow-Methods"); got == "" {
@@ -200,11 +204,11 @@ func TestBookmarkRoundTrip(t *testing.T) {
key := "asura:solo-leveling-123"
in := store.Bookmark{
Title: "Solo Leveling",
SeriesURL: "https://asuracomic.net/series/solo-leveling-123",
Cover: "https://asuracomic.net/cover.jpg",
SeriesURL: "https://asurascans.com/series/solo-leveling-123",
Cover: "https://asurascans.com/cover.jpg",
LastChapter: "Chapter 10",
LastChapterNum: 10,
LastChapterURL: "https://asuracomic.net/series/solo-leveling-123/chapter/10",
LastChapterURL: "https://asurascans.com/series/solo-leveling-123/chapter/10",
}
body, _ := json.Marshal(in)
@@ -329,12 +333,14 @@ func TestFlatWireFieldSet(t *testing.T) {
latestNum := floatPtr(8)
want := store.Bookmark{
Key: key, Site: "comix", SeriesID: "some-title",
Title: in.Title, SeriesURL: in.SeriesURL, Cover: in.Cover,
Title: in.Title, SeriesURL: in.SeriesURL,
LastChapter: in.LastChapter, LastChapterNum: in.LastChapterNum,
LastChapterURL: in.LastChapterURL, Favorite: true,
LatestChapter: in.LatestChapter, LatestChapterNum: latestNum,
Status: store.StatusArchived, Kind: store.KindManga,
}
// Cover is deliberately absent above: the client's cover is discarded, and
// this wiring acquires none, so the field is present and empty (ADR-0007).
if stored.Title != want.Title || stored.SeriesURL != want.SeriesURL || stored.Cover != want.Cover ||
stored.LastChapter != want.LastChapter || stored.LastChapterNum != want.LastChapterNum ||
stored.LastChapterURL != want.LastChapterURL || stored.Favorite != want.Favorite ||
+96 -117
View File
@@ -1,35 +1,15 @@
package main
import (
"context"
"errors"
"database/sql"
"net/http"
"net/http/httptest"
"sync/atomic"
"strings"
"testing"
"bookmarkmanager/backend/internal/store"
)
// fakeCovers stands in for the headless browser. It counts calls so the test
// can prove the cache spares the browser a second navigation.
type fakeCovers struct {
body []byte
contentType string
err error
calls atomic.Int32
lastID atomic.Value
}
func (f *fakeCovers) Image(_ context.Context, imageID string) ([]byte, string, error) {
f.calls.Add(1)
f.lastID.Store(imageID)
if f.err != nil {
return nil, "", f.err
}
return f.body, f.contentType, nil
}
const testCoverID = "019fe11a-84c3-7fc3-a84b-88787374b617"
func getCover(t *testing.T, srv http.Handler, path string, cookie *http.Cookie) *httptest.ResponseRecorder {
t.Helper()
req := httptest.NewRequest(http.MethodGet, path, nil)
@@ -41,115 +21,114 @@ func getCover(t *testing.T, srv http.Handler, path string, cookie *http.Cookie)
return rr
}
// kagane serves its covers behind a Cloudflare challenge and with
// cross-origin-resource-policy: same-origin, so the UI can only show one by
// re-serving the bytes from its own origin.
func TestKaganeCoverProxiesAndCaches(t *testing.T) {
cf := &fakeCovers{body: []byte("\x00webp-bytes"), contentType: "image/webp"}
cfg := testConfig()
cfg.Covers = cf
srv, st := newWebTestServer(t, cfg)
cookie := sessionCookie(t, st)
// The acquired Cover is served from this deployment's own origin, to any
// browser rendering a third-party page — no session, no credential (ADR-0007).
func TestPublicCoverServesStoredBytesUnauthenticated(t *testing.T) {
const sourceURL = "https://cdn.asurascans.com/covers/solo.webp"
srv, st := newWebTestServer(t, testConfig())
if err := st.PutCover(sourceURL, []byte("\x00webp-bytes"), "image/webp"); err != nil {
t.Fatalf("PutCover: %v", err)
}
for i := range 2 {
rr := getCover(t, srv, "/img/kagane/"+testCoverID, cookie)
if rr.Code != http.StatusOK {
t.Fatalf("request %d: status = %d, want 200", i, rr.Code)
}
if got := rr.Body.String(); got != string(cf.body) {
t.Fatalf("request %d: body = %q, want %q", i, got, cf.body)
}
if got := rr.Header().Get("Content-Type"); got != "image/webp" {
t.Fatalf("request %d: Content-Type = %q, want image/webp", i, got)
}
// The wire URL is what a client actually requests, so the path under test
// is taken from it rather than rebuilt by hand.
wire := st.CoverWireURL(store.CoverAddress(sourceURL))
path, ok := strings.CutPrefix(wire, testCoverBaseURL)
if !ok {
t.Fatalf("wire URL %q is not on the public origin %q", wire, testCoverBaseURL)
}
if got := cf.calls.Load(); got != 1 {
t.Fatalf("fetcher called %d times, want 1 — the second read must come from the cache", got)
rr := getCover(t, srv, path, nil)
if rr.Code != http.StatusOK {
t.Fatalf("status = %d, want 200 without any credential", rr.Code)
}
if got := cf.lastID.Load(); got != testCoverID {
t.Fatalf("fetched image id = %v, want %s", got, testCoverID)
if got := rr.Body.String(); got != "\x00webp-bytes" {
t.Fatalf("body = %q, want the stored bytes", got)
}
if got := rr.Header().Get("Content-Type"); got != "image/webp" {
t.Fatalf("Content-Type = %q, want the stored one", got)
}
// Content-addressed bytes never change, so a client that has them must
// never need to ask again.
if got := rr.Header().Get("Cache-Control"); !strings.Contains(got, "immutable") {
t.Fatalf("Cache-Control = %q, want an immutable cache directive", got)
}
}
// The proxy reaches a headless browser, so it is not open to the internet.
func TestKaganeCoverRequiresSession(t *testing.T) {
cf := &fakeCovers{body: []byte("x"), contentType: "image/webp"}
cfg := testConfig()
cfg.Covers = cf
srv, _ := newWebTestServer(t, cfg)
rr := getCover(t, srv, "/img/kagane/"+testCoverID, nil)
if rr.Code != http.StatusUnauthorized {
t.Fatalf("status = %d, want 401", rr.Code)
func TestPublicCoverRejectsUnknownAddress(t *testing.T) {
srv, _ := newWebTestServer(t, testConfig())
cases := map[string]string{
"unknown": "/covers/" + store.CoverAddress("https://cdn.example/never-stored.jpg"),
"malformed": "/covers/not-an-address",
"traversal": "/covers/../../etc/passwd",
"empty": "/covers/",
}
if got := cf.calls.Load(); got != 0 {
t.Fatalf("fetcher called %d times for an unauthenticated request, want 0", got)
}
}
func TestKaganeCoverRejectsBadInput(t *testing.T) {
cases := []struct {
name string
id string
fetch *fakeCovers
}{
{
"an id that is not a uuid never reaches the browser",
"solo-leveling",
&fakeCovers{body: []byte("x"), contentType: "image/webp"},
},
{
"a uuid-shaped id with a trailing segment is rejected whole",
testCoverID + "x",
&fakeCovers{body: []byte("x"), contentType: "image/webp"},
},
{
"a challenged fetch is a missing cover",
testCoverID,
&fakeCovers{err: errors.New("challenge held")},
},
{
"a content type outside the image set is not echoed back",
testCoverID,
&fakeCovers{body: []byte("<script>"), contentType: "text/html"},
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
cfg := testConfig()
cfg.Covers = tc.fetch
srv, st := newWebTestServer(t, cfg)
rr := getCover(t, srv, "/img/kagane/"+tc.id, sessionCookie(t, st))
if rr.Code != http.StatusNotFound {
t.Fatalf("status = %d, want 404", rr.Code)
for name, path := range cases {
t.Run(name, func(t *testing.T) {
if rr := getCover(t, srv, path, nil); rr.Code == http.StatusOK {
t.Fatalf("%s: status = 200, want anything but a served body", path)
}
})
}
}
// ServeMux path-cleans a traversal into a redirect before the handler runs, so
// the guarantee to pin down is that no request shaped like one ever gets bytes.
func TestKaganeCoverTraversalServesNothing(t *testing.T) {
cf := &fakeCovers{body: []byte("secret"), contentType: "image/webp"}
cfg := testConfig()
cfg.Covers = cf
srv, st := newWebTestServer(t, cfg)
rr := getCover(t, srv, "/img/kagane/../../etc/passwd", sessionCookie(t, st))
// A content type outside the image set is never echoed back. The old kagane
// proxy could fetch text/html from a challenged fetch and had to refuse it;
// the general route's only input is the store, and the store refuses to
// record anything that is not an image — but the guarantee is pinned at the
// serving boundary, not the write gate, so a poisoned row (migrated data, a
// writer that skips the gate) is also never served.
func TestPublicCoverNeverEchoesNonImage(t *testing.T) {
const sourceURL = "https://cdn.example/cover"
st, dsn := newTestStoreURL(t)
// The write gate refuses non-image content types outright.
if err := st.PutCover(sourceURL, []byte("<script>"), "text/html"); err == nil {
t.Fatal("PutCover accepted a non-image content type")
}
// A legitimate row, then the content type flipped behind the store's back:
// the bytes exist at the address, so only the type is hostile.
address := store.CoverAddress(sourceURL)
if err := st.SetSeriesCover("asura", "solo", sourceURL, []byte("<script>"), "image/png"); err != nil {
t.Fatalf("seed row: %v", err)
}
db, err := sql.Open("pgx", dsn)
if err != nil {
t.Fatalf("open %s: %v", dsn, err)
}
defer db.Close()
if _, err := db.Exec(`UPDATE covers SET content_type = 'text/html' WHERE address = $1`, address); err != nil {
t.Fatalf("poison row: %v", err)
}
rr := getCover(t, newRouter(st, testConfig()), "/covers/"+address, nil)
if rr.Code == http.StatusOK {
t.Fatalf("status = 200, want anything but a served body")
}
if got := cf.calls.Load(); got != 0 {
t.Fatalf("fetcher called %d times for a traversal, want 0", got)
t.Fatalf("status = 200, want a refusal for a non-image row (body %q)", rr.Body.String())
}
}
// Without BROWSER_WS_URL there is no fetcher, and the endpoint must answer
// rather than reach for a nil one.
func TestKaganeCoverWithoutFetcher(t *testing.T) {
// The whole point of acquiring bytes is that the UI shows them: the card's
// <img> must carry the public address, not a third-party URL and not a
// placeholder.
func TestListRendersAcquiredCover(t *testing.T) {
const sourceURL = "https://static.comix.to/039d/i/1/34/6a6742bf15736@280.jpg"
srv, st := newWebTestServer(t, testConfig())
rr := getCover(t, srv, "/img/kagane/"+testCoverID, sessionCookie(t, st))
if rr.Code != http.StatusNotFound {
t.Fatalf("status = %d, want 404", rr.Code)
if _, err := st.Upsert(st.OwnerID(), store.Bookmark{
Key: "comix:n8we", Site: "comix", SeriesID: "n8we", Title: "Dungeons and Crayons",
SeriesURL: "https://comix.to/title/n8we", UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
if err := st.SetSeriesCover("comix", "n8we", sourceURL, []byte("\xff\xd8jpeg"), "image/jpeg"); err != nil {
t.Fatalf("SetSeriesCover: %v", err)
}
req := httptest.NewRequest(http.MethodGet, "/ui/list", nil)
req.AddCookie(sessionCookie(t, st))
rr := httptest.NewRecorder()
srv.ServeHTTP(rr, req)
if rr.Code != http.StatusOK {
t.Fatalf("status = %d, want 200", rr.Code)
}
want := `src="` + testCoverBaseURL + "/covers/" + store.CoverAddress(sourceURL) + `"`
if !strings.Contains(rr.Body.String(), want) {
t.Fatalf("rendered list does not contain %s", want)
}
}
+40
View File
@@ -50,6 +50,13 @@ func (h *Handler) Put(w http.ResponseWriter, r *http.Request) {
http.Error(w, "invalid JSON body", http.StatusBadRequest)
return
}
// A body may carry a cover, and it is discarded here rather than
// rejected: an older installed userscript may still send one, and
// ADR-0004's compatibility argument depends on those scripts continuing
// to work. The Cover is acquired server-side (ADR-0007), so the field is
// permanently inert - not pending removal, and not a value any later code
// should start reading.
b.Cover = ""
// Path key is authoritative; derive site/series_id from it when the body
// omits them so the stored row is always self-consistent.
@@ -124,3 +131,36 @@ func Healthz(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(http.StatusOK)
_, _ = w.Write([]byte("ok"))
}
// Cover serves stored cover bytes. GET /covers/{address}
//
// Public on purpose: the userscript renders these on Sites the deployment
// does not control, where no credential of ours may be sent, and the address
// is the SHA-256 of a URL the Site already publishes (ADR-0007). An unknown
// address is a 404 rather than an error - "no Cover yet" is a normal state,
// and the clients fall back to their placeholder.
func (h *Handler) Cover(w http.ResponseWriter, r *http.Request) {
body, contentType, ok, err := h.Store.CoverByAddress(r.PathValue("address"))
if err != nil {
log.Printf("cover: %v", err)
http.Error(w, "internal error", http.StatusInternalServerError)
return
}
if !ok {
http.NotFound(w, r)
return
}
// Refuse anything the write gate would not have recorded: a poisoned row
// (migrated data, a writer that skips the gate) must never be echoed back
// as bytes of a type no Cover may have.
if _, ok := store.CoverContentType(contentType); !ok {
log.Printf("cover %s: refusing non-image content type %q", r.PathValue("address"), contentType)
http.NotFound(w, r)
return
}
w.Header().Set("Content-Type", contentType)
// Content-addressed, so the bytes at this URL can never change. Public
// rather than private: no credential gates the route.
w.Header().Set("Cache-Control", "public, max-age=604800, immutable")
_, _ = w.Write(body)
}
+146
View File
@@ -0,0 +1,146 @@
package latest
import (
"context"
"errors"
"log"
"sync"
"time"
"bookmarkmanager/backend/internal/store"
)
// acquireTimeout bounds one creation-time acquisition end to end: the series
// page plus the cover bytes. Nothing is waiting on it — the Reader's write has
// already returned — so this only stops a stalled Site from holding a
// goroutine and a connection open forever.
const acquireTimeout = 45 * time.Second
// Acquirer gives a Series its Latest Chapter and its Cover the moment the
// first Bookmark creates it, instead of leaving the Reader to wait out the
// poll queue — which is ordered by Reader count, so a Series with one Reader
// sits behind every popular one (ADR-0007).
//
// Both facts come from a single series-page fetch, which is also why no
// client-supplied cover hint is worth accepting: the page has to be fetched
// for the chapter signal regardless, so a hint would save no request while
// adding a client-controlled input to a server-side fetch.
//
// Every failure path is "log and move on". The Bookmark, its progress and its
// Latest Chapter are already committed; a Site that is down or a Cover that
// cannot be produced must not disturb any of them, and the Series is simply
// left blank until the poll's own cover pass (#61) fills it.
type Acquirer struct {
Store *store.Store
// Fetch retrieves the series page over plain TLS. Nil with a nil
// BrowserFetch disables acquisition entirely.
Fetch Fetcher
// BrowserFetch retrieves kagane and novelfull pages through the browser
// sidecar, the only thing that clears their Cloudflare challenge. The
// per-site fallback policy lives in fetcherFor. Nil leaves those Sites
// unacquired when no fallback applies.
BrowserFetch Fetcher
// Covers retrieves the cover bytes. Nil leaves the Cover blank and the
// chapter half working.
Covers CoverBytesFetcher
// BrowserCoverFetch retrieves browser-claimed cover bytes through the
// sidecar. Nil leaves those Covers blank; nothing falls back to a plain
// fetch, which would only ever retrieve a challenge page.
BrowserCoverFetch BrowserCoverFetcher
// Ctx cancels in-flight acquisitions at shutdown. A hook signature has
// nowhere to pass one, so it lives here; nil means context.Background.
Ctx context.Context
inflight sync.WaitGroup
}
// acquireSlots caps how many creation-time fetches run at once. A Reader whose
// userscript bulk-syncs creates many Series at once, and a burst of
// simultaneous requests from one server IP is the traffic shape most likely to
// move that IP's bot score — the same reason the poller staggers its batch.
var acquireSlots = make(chan struct{}, 2)
// Acquire starts one acquisition and returns immediately: a Reader's bookmark
// action may not block on a third-party Site's latency, nor fail with it. It
// is the store's OnSeriesCreated hook, so it only ever runs for a Series no
// Reader had bookmarked before.
func (a *Acquirer) Acquire(sr store.Series) {
a.inflight.Add(1)
go func() {
defer a.inflight.Done()
defer func() {
if r := recover(); r != nil {
log.Printf("acquire %q: recovered from panic: %v", sr.Key(), r)
}
}()
parent := a.Ctx
if parent == nil {
parent = context.Background()
}
select {
case acquireSlots <- struct{}{}:
defer func() { <-acquireSlots }()
case <-parent.Done():
return
}
ctx, cancel := context.WithTimeout(parent, acquireTimeout)
defer cancel()
a.acquire(ctx, sr)
}()
}
// Wait blocks until every started acquisition has finished. It exists for
// tests: an asynchronous side effect is otherwise unobservable without
// polling for it.
func (a *Acquirer) Wait() { a.inflight.Wait() }
func (a *Acquirer) acquire(ctx context.Context, sr store.Series) {
if a.Fetch == nil && a.BrowserFetch == nil {
return
}
facts, err := readSeriesPage(ctx, sr.Site, sr.SeriesURL, a.BrowserFetch, a.Fetch)
if err != nil {
switch {
case errors.Is(err, errNotFetchable):
// series_url arrives in a client-supplied PUT body, so without the
// gate a token-holder chooses what the server fetches from its own
// network position.
log.Printf("acquire %q: not fetchable: site=%q url=%q", sr.Key(), sr.Site, sr.SeriesURL)
case errors.Is(err, errNoFetcher):
log.Printf("acquire %q: no fetcher for site %q", sr.Key(), sr.Site)
default:
log.Printf("acquire %q: %v", sr.Key(), err)
}
return
}
// This page just served the same purpose a poll tick would have; without
// the stamp the row stays due and the poller refetches it immediately.
//
// Stamped after success — the reverse of the poller, which stamps before
// the fetch: the Reader is here, watching the Series they just created, so
// a failed acquisition must leave the row due for a fast retry rather than
// consuming the cooldown. The stamp happens even when the page read
// succeeded but produced no facts to persist.
if err := a.Store.MarkLatestChecked(sr.Site, sr.SeriesID, time.Now().UnixMilli()); err != nil {
log.Printf("acquire %q: mark checked: %v", sr.Key(), err)
}
if facts.HasLatest {
if err := a.Store.SetLatestChapter(sr.Site, sr.SeriesID, facts.Latest.Label, facts.Latest.Num); err != nil {
log.Printf("acquire %q: set latest chapter: %v", sr.Key(), err)
}
}
if !facts.HasCover {
return
}
bytes, contentType, err := fetchCoverBytes(ctx, facts.Cover, a.BrowserCoverFetch, a.Covers)
if err != nil {
log.Printf("acquire %q: fetch cover %s: %v", sr.Key(), facts.Cover, err)
return
}
if err := a.Store.SetSeriesCover(sr.Site, sr.SeriesID, facts.Cover, bytes, contentType); err != nil {
log.Printf("acquire %q: persist cover: %v", sr.Key(), err)
}
}
+430
View File
@@ -0,0 +1,430 @@
package latest
import (
"context"
"errors"
"testing"
"time"
"bookmarkmanager/backend/internal/store"
)
// The series page carries both facts, which is the whole argument for taking
// them from one fetch.
const asuraSeriesAndCoverFixture = asuraSeriesFixture + asuraCoverFixture
const (
acquireKey = "asura:chronicles-of-the-demon-faction-f886a8af"
acquireSeriesID = "chronicles-of-the-demon-faction-f886a8af"
acquireSeriesURL = "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af"
acquireCoverURL = "https://cdn.asurascans.com/asura-images/covers/chronicles-of-the-demon-faction.d4dcb8.webp"
)
const (
kaganeKey = "kagane:019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
kaganeSeriesID = "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
kaganeSeriesURL = "https://kagane.to/series/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
kaganeImageID = "019fe11a-84c3-7fc3-a84b-88787374b617"
kaganeCoverSrc = "https://kagane.to/api/v2/image/" + kaganeImageID + "/compressed"
)
// kagane's browser-fetched body is one JSON object carrying both the chapter
// list (series_books) and the cover image ids (series_covers), so the single
// acquisition fetch yields both facts.
const kaganeSeriesAndCoverFixture = `{"series_id":"019f84bc-9ba0-7ed9-86f5-8b905ec7c28b",` +
`"series_books":[{"book_id":"b","title":"Episode 41","chapter_no":"41","sort_no":41}],` +
`"series_covers":[{"cover_id":"019fe11a-84d1-714b-9cf4-2827f277f3c0","language":"en",` +
`"image_id":"019fe11a-84c3-7fc3-a84b-88787374b617"}]}`
const (
novelfullKey = "novelfull:reverend-insanity"
novelfullSeriesID = "reverend-insanity"
novelfullSeriesURI = "https://novelfull.com/reverend-insanity.html"
novelfullCoverURL = "https://novelfull.com/uploads/webp/novel/reverend-insanity-82661d911a.webp"
)
// newAcquirer wires an acquirer onto the store's creation hook, which is how
// main wires it: the write path is what starts an acquisition.
func newAcquirer(s *store.Store, page *fakeFetcher, covers *fakeBytesCoverFetcher) *Acquirer {
a := &Acquirer{Store: s, Fetch: page, Covers: covers}
s.OnSeriesCreated = a.Acquire
return a
}
func bookmarkNewSeries(t *testing.T, s *store.Store, seriesURL string) store.Bookmark {
t.Helper()
stored, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: acquireKey, Site: "asura", SeriesID: acquireSeriesID,
Title: "Chronicles of the Demon Faction", SeriesURL: seriesURL,
Cover: "https://evil.example/client-supplied.jpg", UpdatedAt: 1000,
})
if err != nil {
t.Fatalf("Upsert: %v", err)
}
return stored
}
func readBookmark(t *testing.T, s *store.Store, key string) store.Bookmark {
t.Helper()
b, ok, err := s.Get(s.OwnerID(), key)
if err != nil || !ok {
t.Fatalf("Get %q = %v, %v", key, ok, err)
}
return b
}
func bookmarkNewKaganeSeries(t *testing.T, s *store.Store) store.Bookmark {
t.Helper()
stored, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: kaganeKey, Site: "kagane", SeriesID: kaganeSeriesID,
Title: "Infinite Decryption", SeriesURL: kaganeSeriesURL, UpdatedAt: 1000,
})
if err != nil {
t.Fatalf("Upsert: %v", err)
}
return stored
}
func bookmarkNewNovelfullSeries(t *testing.T, s *store.Store) store.Bookmark {
t.Helper()
stored, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: novelfullKey, Site: "novelfull", SeriesID: novelfullSeriesID,
Title: "Reverend Insanity", SeriesURL: novelfullSeriesURI, UpdatedAt: 1000,
})
if err != nil {
t.Fatalf("Upsert: %v", err)
}
return stored
}
// The reported bug: a Reader bookmarks a Series nobody holds and expects the
// Cover, not a broken image. Both facts come from the one series-page fetch.
func TestAcquireFillsChapterAndCoverFromOneFetch(t *testing.T) {
s, _ := newTestStore(t)
page := &fakeFetcher{body: asuraSeriesAndCoverFixture, status: 200}
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"}
acq := newAcquirer(s, page, covers)
// The write itself must not carry the acquisition: it returns before the
// Cover exists, and the field is empty until the bytes land.
stored := bookmarkNewSeries(t, s, acquireSeriesURL)
if stored.Cover != "" {
t.Fatalf("Cover on the creating write = %q, want empty", stored.Cover)
}
acq.Wait()
if got := page.callCount(); got != 1 {
t.Fatalf("series page fetches = %d, want exactly 1", got)
}
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetches = %d, want 1", got)
}
got := readBookmark(t, s, acquireKey)
if got.LatestChapterNum == nil || *got.LatestChapterNum != 181 {
t.Fatalf("LatestChapterNum = %v, want 181", got.LatestChapterNum)
}
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(acquireCoverURL); got.Cover != want {
t.Fatalf("Cover = %q, want the absolute address %q", got.Cover, want)
}
body, contentType, ok, err := s.CoverByAddress(store.CoverAddress(acquireCoverURL))
if err != nil || !ok {
t.Fatalf("CoverByAddress = %v, %v", ok, err)
}
if string(body) != "cover-bytes" || contentType != "image/jpeg" {
t.Fatalf("stored cover = (%q, %q), want the fetched bytes", body, contentType)
}
}
// A Series that already exists is not re-acquired: no fetch, and the Cover it
// already has is left alone.
func TestAcquireSkipsAnExistingSeries(t *testing.T) {
s, _ := newTestStore(t)
page := &fakeFetcher{body: asuraSeriesAndCoverFixture, status: 200}
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"}
acq := newAcquirer(s, page, covers)
bookmarkNewSeries(t, s, acquireSeriesURL)
acq.Wait()
bookmarkNewSeries(t, s, acquireSeriesURL)
acq.Wait()
if got := page.callCount(); got != 1 {
t.Fatalf("series page fetches = %d, want 1 — an existing series is not re-acquired", got)
}
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetches = %d, want 1", got)
}
got := readBookmark(t, s, acquireKey)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(acquireCoverURL); got.Cover != want {
t.Fatalf("Cover = %q, want the acquired one %q", got.Cover, want)
}
}
// A Site that is down costs the Cover and nothing else.
func TestAcquireFailureLeavesTheBookmarkIntact(t *testing.T) {
cases := []struct {
name string
page *fakeFetcher
covers *fakeBytesCoverFetcher
// wantLatest is the chapter that still lands; 0 means none did.
wantLatest float64
}{
{
"the series page is unreachable",
&fakeFetcher{err: errors.New("connection reset")},
&fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"},
0,
},
{
"the series page answers with a challenge",
&fakeFetcher{body: challengeFixture, status: 200},
&fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"},
0,
},
{
"only the cover bytes fail",
&fakeFetcher{body: asuraSeriesAndCoverFixture, status: 200},
&fakeBytesCoverFetcher{err: errors.New("403")},
181,
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
s, _ := newTestStore(t)
acq := newAcquirer(s, tc.page, tc.covers)
stored := bookmarkNewSeries(t, s, acquireSeriesURL)
acq.Wait()
got := readBookmark(t, s, acquireKey)
if got.Cover != "" {
t.Fatalf("Cover = %q, want empty rather than an address that 404s", got.Cover)
}
if got.Title != stored.Title || got.UpdatedAt != stored.UpdatedAt {
t.Fatalf("bookmark = %+v, want it untouched by the failed acquisition", got)
}
if tc.wantLatest == 0 {
if got.LatestChapterNum != nil {
t.Fatalf("LatestChapterNum = %v, want none captured", *got.LatestChapterNum)
}
return
}
if got.LatestChapterNum == nil || *got.LatestChapterNum != tc.wantLatest {
t.Fatalf("LatestChapterNum = %v, want %v", got.LatestChapterNum, tc.wantLatest)
}
})
}
}
// series_url arrives in a client-supplied body, so the acquisition reuses the
// poller's gate rather than deriving a second one: a non-https scheme, a
// site the parsers do not know, or a host pinned to another site is refused
// before the server spends a request from its own network position.
func TestAcquireRefusesAnUnfetchableSeriesURL(t *testing.T) {
for _, seriesURL := range []string{
"http://asurascans.com/comics/x",
"file:///etc/passwd",
"",
} {
t.Run(seriesURL, func(t *testing.T) {
s, _ := newTestStore(t)
page := &fakeFetcher{body: asuraSeriesAndCoverFixture, status: 200}
acq := newAcquirer(s, page, &fakeBytesCoverFetcher{})
bookmarkNewSeries(t, s, seriesURL)
acq.Wait()
if got := page.callCount(); got != 0 {
t.Fatalf("fetches for %q = %d, want 0", seriesURL, got)
}
})
}
}
// blockingFetcher stands in for a Site that never answers, so a synchronous
// acquisition would be visible as a stalled write rather than a slow one.
type blockingFetcher struct {
release <-chan struct{}
body string
}
func (f *blockingFetcher) Get(ctx context.Context, _ string) (string, int, error) {
select {
case <-f.release:
return f.body, 200, nil
case <-ctx.Done():
return "", 0, ctx.Err()
}
}
// The Reader's write may not wait on a third-party Site: with the acquisition
// wedged on an unanswering page, the PUT still returns.
func TestAcquireDoesNotBlockTheWrite(t *testing.T) {
s, _ := newTestStore(t)
release := make(chan struct{})
acq := &Acquirer{Store: s, Fetch: &blockingFetcher{release: release, body: asuraSeriesAndCoverFixture}}
s.OnSeriesCreated = acq.Acquire
upserted := make(chan error, 1)
go func() {
_, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: acquireKey, Site: "asura", SeriesID: acquireSeriesID,
Title: "Chronicles of the Demon Faction", SeriesURL: acquireSeriesURL, UpdatedAt: 1000,
})
upserted <- err
}()
select {
case err := <-upserted:
if err != nil {
t.Fatalf("Upsert: %v", err)
}
case <-time.After(10 * time.Second):
t.Fatal("the creating write blocked on the acquisition")
}
close(release)
acq.Wait()
}
// The second symptom of #47: a kagane Series bookmarked from a chapter page
// gets its Cover at creation, with the bytes fetched through the browser
// sidecar — the only path that clears the challenge — into the
// content-addressed store.
func TestAcquireKaganeCoverThroughBrowser(t *testing.T) {
s, _ := newTestStore(t)
tlsPage := &fakeFetcher{body: "", status: 403}
browserPage := &fakeFetcher{body: kaganeSeriesAndCoverFixture, status: 200}
covers := &fakeCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
acq := &Acquirer{
Store: s, Fetch: tlsPage, BrowserFetch: browserPage,
BrowserCoverFetch: covers, Covers: &fakeBytesCoverFetcher{},
}
s.OnSeriesCreated = acq.Acquire
bookmarkNewKaganeSeries(t, s)
acq.Wait()
if got := tlsPage.callCount(); got != 0 {
t.Fatalf("plain-TLS page fetches = %d, want 0 — kagane pages are browser-only", got)
}
if got := browserPage.callCount(); got != 1 {
t.Fatalf("browser page fetches = %d, want 1", got)
}
if got := covers.callCount(); got != 1 {
t.Fatalf("browser cover fetches = %d, want 1", got)
}
if got := covers.calls[0]; got != kaganeCoverSrc {
t.Fatalf("browser cover fetched URL %q, want %q", got, kaganeCoverSrc)
}
got := readBookmark(t, s, kaganeKey)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(kaganeCoverSrc); got.Cover != want {
t.Fatalf("Cover = %q, want the content-addressed URL %q", got.Cover, want)
}
body, contentType, ok, err := s.CoverByAddress(store.CoverAddress(kaganeCoverSrc))
if err != nil || !ok {
t.Fatalf("CoverByAddress = %v, %v", ok, err)
}
if string(body) != "cover-bytes" || contentType != "image/webp" {
t.Fatalf("stored cover = (%q, %q), want the browser-fetched bytes", body, contentType)
}
}
// novelfull needs the browser only for its HTML: the cover URL comes out of
// the browser-fetched page, but the bytes go over plain TLS through the
// ordinary gated fetcher, never through the browser (issue #62).
func TestAcquireNovelfullCoverOverPlainTLS(t *testing.T) {
s, _ := newTestStore(t)
browserPage := &fakeFetcher{body: novelfullSeriesFixture + novelfullCoverFixture, status: 200}
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
acq := &Acquirer{
Store: s, Fetch: &fakeFetcher{body: "", status: 403},
BrowserFetch: browserPage, Covers: covers,
}
s.OnSeriesCreated = acq.Acquire
bookmarkNewNovelfullSeries(t, s)
acq.Wait()
if got := browserPage.callCount(); got != 1 {
t.Fatalf("browser page fetches = %d, want 1", got)
}
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetches = %d, want 1 — novelfull bytes never touch the browser", got)
}
if got := covers.calls[0]; got != novelfullCoverURL {
t.Fatalf("cover fetched from %q, want %q", got, novelfullCoverURL)
}
got := readBookmark(t, s, novelfullKey)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(novelfullCoverURL); got.Cover != want {
t.Fatalf("Cover = %q, want %q", got.Cover, want)
}
}
// With no browser sidecar configured, kagane is simply not acquired: no
// request is spent on a page that could only ever answer with a challenge,
// and nothing falls back to a plain fetch.
func TestAcquireKaganeSkippedWithoutBrowser(t *testing.T) {
s, _ := newTestStore(t)
tlsPage := &fakeFetcher{body: kaganeSeriesAndCoverFixture, status: 200}
acq := &Acquirer{
Store: s, Fetch: tlsPage,
Covers: &fakeBytesCoverFetcher{body: []byte("x"), contentType: "image/webp"},
}
s.OnSeriesCreated = acq.Acquire
bookmarkNewKaganeSeries(t, s)
acq.Wait()
if got := tlsPage.callCount(); got != 0 {
t.Fatalf("plain-TLS fetches for kagane = %d, want 0", got)
}
if got := readBookmark(t, s, kaganeKey); got.Cover != "" {
t.Fatalf("Cover = %q, want empty without a browser", got.Cover)
}
}
// The byte half of "nothing falls back to a plain fetch": with a browser for
// the page but none for the bytes, a kagane Cover stays absent and the TLS
// cover fetcher is never consulted.
func TestAcquireKaganeBytesNeverFallBackToPlainTLS(t *testing.T) {
s, _ := newTestStore(t)
browserPage := &fakeFetcher{body: kaganeSeriesAndCoverFixture, status: 200}
tlsCovers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
acq := &Acquirer{
Store: s, Fetch: &fakeFetcher{body: "", status: 403},
BrowserFetch: browserPage, Covers: tlsCovers,
}
s.OnSeriesCreated = acq.Acquire
bookmarkNewKaganeSeries(t, s)
acq.Wait()
if got := tlsCovers.callCount(); got != 0 {
t.Fatalf("plain-TLS cover fetches = %d, want 0 — kagane bytes are browser-only", got)
}
if got := readBookmark(t, s, kaganeKey); got.Cover != "" {
t.Fatalf("Cover = %q, want empty without a browser cover fetcher", got.Cover)
}
}
// novelfull's no-browser degradation differs from kagane's: only its HTML
// needs the sidecar, so when the page body is available — the challenge is a
// live time-varying fact that sometimes answers a plain request — the Cover
// still lands, bytes over plain TLS.
func TestAcquireNovelfullCoverWithoutBrowser(t *testing.T) {
s, _ := newTestStore(t)
tlsPage := &fakeFetcher{body: novelfullSeriesFixture + novelfullCoverFixture, status: 200}
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
acq := &Acquirer{Store: s, Fetch: tlsPage, Covers: covers}
s.OnSeriesCreated = acq.Acquire
bookmarkNewNovelfullSeries(t, s)
acq.Wait()
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetches = %d, want 1", got)
}
got := readBookmark(t, s, novelfullKey)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(novelfullCoverURL); got.Cover != want {
t.Fatalf("Cover = %q, want %q", got.Cover, want)
}
}
+115 -62
View File
@@ -24,12 +24,6 @@ const challengeTimeout = 45 * time.Second
var kaganeSeriesRe = regexp.MustCompile(`^/series/([0-9a-f-]{36})/?$`)
// kaganeImageIDRe pins the only path segment Image interpolates into an
// outbound URL. The id arrives from a stored cover URL, which a client
// supplied, so it is matched rather than trusted: a headless browser is a
// strong SSRF primitive.
var kaganeImageIDRe = regexp.MustCompile(`^[0-9a-f-]{36}$`)
// BrowserFetcher retrieves pages through a remote headless Chrome over the
// DevTools Protocol.
//
@@ -53,19 +47,22 @@ var kaganeImageIDRe = regexp.MustCompile(`^[0-9a-f-]{36}$`)
type BrowserFetcher struct {
allocCtx context.Context
cancel context.CancelFunc
// One page at a time: caps the sidecar's memory and keeps series from
// sharing page state.
// One page at a time: caps the browser's memory — it runs under a hard
// cgroup cap on a shared machine — and keeps series from sharing page state.
mu sync.Mutex
}
var _ Fetcher = (*BrowserFetcher)(nil)
// NewBrowserFetcher connects to a headless-shell over CDP. wsURL must name the
// sidecar by IP, e.g. ws://172.28.0.10:9222 — not by Docker DNS name. Chrome's
// DevTools HTTP handler 500s any /json/version request whose Host header
// isn't an IP or "localhost" (confirmed 2026-08-03 against
// chromedp/headless-shell:stable), so the compose network pins the sidecar's
// address for this to resolve at all.
// NewBrowserFetcher connects to a Chrome over CDP. The browser is not a
// sidecar: it runs on a separate machine and is reached over the tailnet
// (ADR-0006), so wsURL is that machine's tailnet address, e.g.
// ws://100.64.0.5:9222.
//
// It must be an IP, never a hostname — not MagicDNS, not a Docker service
// name. Chrome's DevTools HTTP handler 500s any /json/version request whose
// Host header isn't an IP or "localhost" (confirmed 2026-08-03), so a name
// fails at discovery and surfaces as a dead site rather than a bad URL.
//
// Do not add chromedp.NoModifyURL here: that option skips the /json/version
// discovery request entirely and dials wsURL as if it were already the full
@@ -73,8 +70,10 @@ var _ Fetcher = (*BrowserFetcher)(nil)
// /devtools/browser/<uuid>, a path chosen fresh at every Chrome start — dialing
// the bare host:port 404s. The default (discovery) path works precisely
// because Chrome's /json/version response echoes back the Host header of the
// discovery request in webSocketDebuggerUrl, so as long as wsURL is a
// container-reachable IP, the URL chromedp gets back already points at it.
// discovery request in webSocketDebuggerUrl, so as long as wsURL is an IP this
// process can reach, the URL chromedp gets back already points at it. That is
// also why a Chrome restarted behind a stable endpoint needs no reconnect
// here: the fresh UUID arrives with the next discovery.
func NewBrowserFetcher(wsURL string) (*BrowserFetcher, error) {
if wsURL == "" {
return nil, fmt.Errorf("empty browser websocket url")
@@ -87,54 +86,69 @@ func (f *BrowserFetcher) Close() {
f.cancel()
}
// Get navigates to seriesURL, lets any challenge resolve, then reads either the
// site's JSON API (kagane) from inside the page so the request carries the
// clearance cookie, or the served HTML itself (novelfull) — see
// novelfullSeriesURL for the latter case. The returned body is whatever the
// site's chapter list lives in, which is what latestChapterFrom's per-site
// switch expects.
// Get navigates to seriesURL, lets any challenge resolve, then reads the
// payload the Site's registry entry describes — kagane's chapter-list API from
// inside the page so the request carries the clearance cookie, novelfull's
// served HTML. The returned body is whatever the Site's chapter list lives in,
// which is what the entry's LatestChapter parse expects.
func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int, error) {
apiURL, isKagane := kaganeAPIURL(seriesURL)
if !isKagane && !novelfullSeriesURL(seriesURL) {
return "", 0, fmt.Errorf("not a fetchable browser series url: %q", seriesURL)
}
var body string
// kagane's chapter list is only in its JSON API, which must be called from
// inside the page so the request carries the clearance cookie. novelfull
// renders its chapters into the HTML, so the cleared DOM is the answer.
// chromedp.OuterHTML returns a QueryAction and chromedp.Evaluate an
// EvaluateAction, so the variable has to be the interface both implement.
var read chromedp.Action = chromedp.OuterHTML("html", &body, chromedp.ByQuery)
if isKagane {
read = chromedp.Evaluate(
`fetch(`+jsString(apiURL)+`).then(r => r.ok ? r.text() : "")`,
&body,
awaitPromise,
)
}
// novelfull's payload is the DOM itself, and the interstitial has a DOM
// too, so "we have an answer" has to exclude it explicitly. kagane's
// in-page fetch just fails while challenged, which is already the signal.
done := func() bool { return body != "" && (isKagane || !isInterstitial(body)) }
if err := f.run(ctx, seriesURL, read, done); err != nil {
// Challenge never cleared, or the API refused. Indistinguishable from
// here and handled identically by the caller.
if errors.Is(err, errChallengeHeld) {
return "", 403, nil
// Sorted order (browserBackedSites sorts) makes dispatch deterministic:
// entries' Read funcs are expected to refuse any address owned by another
// Site, and the loop must not depend on that staying true.
for _, name := range browserBackedSites() {
s := sites[name]
read, ok := s.Browser.Read(seriesURL, &body)
if !ok {
continue
}
return "", 0, fmt.Errorf("browser fetch %q: %w", seriesURL, err)
if err := f.run(ctx, seriesURL, read,
func() bool { return s.Browser.Done(body) }); err != nil {
// Challenge never cleared, or the payload was refused.
// Indistinguishable from here and handled identically by the caller.
if errors.Is(err, errChallengeHeld) {
return "", 403, nil
}
return "", 0, fmt.Errorf("browser fetch %q: %w", seriesURL, err)
}
return body, 200, nil
}
return body, 200, nil
return "", 0, fmt.Errorf("not a fetchable browser series url: %q", seriesURL)
}
// Image retrieves one kagane cover as raw bytes and its content type.
// kaganeRead builds the in-tab fetch of kagane's chapter-list API: the
// request must be made from inside the page so it carries the clearance
// cookie, and the API is the only place the list exists. Refusing any other
// address is the per-Site half of the SSRF gate, kept deliberately behind
// fetchableSeriesURL (see browserRead.Read).
func kaganeRead(seriesURL string, out *string) (chromedp.Action, bool) {
apiURL, ok := kaganeAPIURL(seriesURL)
if !ok {
return nil, false
}
return chromedp.Evaluate(
`fetch(`+jsString(apiURL)+`).then(r => r.ok ? r.text() : "")`,
out, awaitPromise), true
}
// novelfullRead reads the cleared DOM. novelfull renders its chapter list
// into the served HTML, so there is no API to call from inside the page — the
// challenge-cleared DOM is the payload.
func novelfullRead(seriesURL string, out *string) (chromedp.Action, bool) {
if !novelfullSeriesURL(seriesURL) {
return nil, false
}
return chromedp.OuterHTML("html", out, chromedp.ByQuery), true
}
// Image retrieves one cover's bytes through the browser sidecar, and its
// content type.
//
// It exists because kagane serves covers behind the same challenge as its
// pages *and* with `cross-origin-resource-policy: same-origin`, so an <img> on
// the web UI's origin cannot load one even from a browser that already holds
// the clearance cookie (verified 2026-08-08). Proxying is the only route.
// pages *and* with `cross-origin-resource-policy: same-origin`, so the bytes
// are only reachable from inside a browser that already holds the clearance
// cookie (verified 2026-08-08). Acquisition through the sidecar is the only
// route.
//
// The image URL is navigated to rather than fetched from some other kagane
// page: the challenge only runs on a top-level navigation, and once it clears
@@ -144,12 +158,14 @@ func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int
// The challenge is not solved by the first read: WaitReady("body") is satisfied
// by the interstitial too. run holds the tab open until the in-page fetch
// succeeds, which is what gives the challenge script the seconds it needs.
func (f *BrowserFetcher) Image(ctx context.Context, imageID string) ([]byte, string, error) {
if !kaganeImageIDRe.MatchString(imageID) {
return nil, "", fmt.Errorf("not a kagane image id: %q", imageID)
func (f *BrowserFetcher) Image(ctx context.Context, imageURL string) ([]byte, string, error) {
m := kaganeImageURLRe.FindStringSubmatch(imageURL)
if m == nil {
return nil, "", fmt.Errorf("not a browser-fetchable cover url: %q", imageURL)
}
imageID := m[1]
var dataURL string
err := f.run(ctx, "https://kagane.to/api/v2/image/"+imageID+"/compressed",
err := f.run(ctx, imageURL,
chromedp.Evaluate(`fetch(location.href).then(r => r.ok
? r.blob().then(b => new Promise(res => {
const fr = new FileReader();
@@ -178,6 +194,36 @@ func (f *BrowserFetcher) Image(ctx context.Context, imageID string) ([]byte, str
// the poller answers with a 403 and its ordinary cooldown.
var errChallengeHeld = errors.New("challenge held")
// errBrowserInterrupted distinguishes a remote Chrome restart from the
// caller's own deadline. chromedp reports both as context.Canceled.
var errBrowserInterrupted = errors.New("browser interrupted")
func classifyBrowserError(ctx context.Context, browserLost bool, err error) error {
if err == nil || ctx.Err() != nil {
return err
}
if !browserLost {
return err
}
if !errors.Is(err, context.Canceled) {
return err
}
return fmt.Errorf("%w: %w", errBrowserInterrupted, err)
}
func browserConnectionLost(ctx context.Context) bool {
c := chromedp.FromContext(ctx)
if c == nil || c.Browser == nil {
return true
}
select {
case <-c.Browser.LostConnection:
return true
default:
return false
}
}
// challengePollInterval paces re-reads while a challenge solves itself.
const challengePollInterval = 2 * time.Second
@@ -203,6 +249,7 @@ func (f *BrowserFetcher) run(ctx context.Context, target string, read chromedp.A
f.mu.Lock()
defer f.mu.Unlock()
callerCtx := ctx
ctx, cancel := context.WithTimeout(ctx, challengeTimeout)
defer cancel()
tabCtx, cancelTab := chromedp.NewContext(f.allocCtx)
@@ -219,20 +266,26 @@ func (f *BrowserFetcher) run(ctx context.Context, target string, read chromedp.A
chromedp.Navigate(target),
chromedp.WaitReady("body", chromedp.ByQuery),
); err != nil {
return err
return classifyBrowserError(callerCtx, browserConnectionLost(tabCtx), err)
}
var lastErr error
for {
// The challenge reloads the page when it passes, which tears down the
// execution context mid-read. That is a retry, not a failure.
if err := chromedp.Run(tabCtx, read); err != nil {
err = classifyBrowserError(callerCtx, browserConnectionLost(tabCtx), err)
if errors.Is(err, errBrowserInterrupted) {
return err
}
lastErr = err
} else if done() {
return nil
}
select {
case <-ctx.Done():
if err := callerCtx.Err(); err != nil {
return err
}
if lastErr != nil {
return fmt.Errorf("%w (last read: %v)", errChallengeHeld, lastErr)
}
+19 -1
View File
@@ -1,6 +1,10 @@
package latest
import "testing"
import (
"context"
"errors"
"testing"
)
func TestKaganeAPIURL(t *testing.T) {
const uuid = "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
@@ -57,3 +61,17 @@ func TestNovelfullSeriesURL(t *testing.T) {
})
}
}
func TestClassifyBrowserInterruption(t *testing.T) {
if err := classifyBrowserError(context.Background(), true, context.Canceled); !errors.Is(err, errBrowserInterrupted) {
t.Fatalf("classifyBrowserError(context.Canceled) = %v, want browser interruption", err)
}
if err := classifyBrowserError(context.Background(), false, context.Canceled); errors.Is(err, errBrowserInterrupted) {
t.Fatalf("ordinary cancellation misclassified as browser interruption: %v", err)
}
caller, cancel := context.WithCancel(context.Background())
cancel()
if err := classifyBrowserError(caller, true, context.Canceled); errors.Is(err, errBrowserInterrupted) {
t.Fatalf("caller cancellation misclassified as browser interruption: %v", err)
}
}
+204
View File
@@ -0,0 +1,204 @@
package latest
import (
"context"
"errors"
"fmt"
"io"
"mime"
"net"
"net/http"
"net/netip"
"net/url"
"strings"
"time"
"bookmarkmanager/backend/internal/store"
)
// CoverBytesFetcher retrieves one cover from its source URL. The caller owns
// persistence; this seam keeps network policy independent from the store.
type CoverBytesFetcher interface {
Fetch(ctx context.Context, sourceURL string) (body []byte, contentType string, err error)
}
// fetchCoverBytes routes a cover's byte retrieval by URL shape, not by Site
// name: the browser fetcher's module claims the addresses only it can fetch
// (kagane's image route answers a plain fetch with a challenge and
// `cross-origin-resource-policy: same-origin`), and everything else goes over
// plain TLS. Missing fetchers degrade to an error the caller logs, never a
// fallback onto a path that cannot succeed. One routing rule for the poll and
// the acquirer, so the two cannot drift apart.
func fetchCoverBytes(ctx context.Context, cover string, browser BrowserCoverFetcher, tls CoverBytesFetcher) ([]byte, string, error) {
if browserOnlyCoverURL(cover) {
if browser == nil {
return nil, "", errors.New("no cover fetcher")
}
return browser.Image(ctx, cover)
}
if tls == nil {
return nil, "", errors.New("no cover fetcher")
}
return tls.Fetch(ctx, cover)
}
// CoverResolver resolves a host before any connection is attempted. Tests
// inject it to exercise hostile DNS results without touching the live network.
type CoverResolver func(context.Context, string) ([]netip.Addr, error)
// TLSCoverFetcher retrieves image bytes with the standard HTTPS client. Unlike
// TLSFetcher, it does not need a browser fingerprint: cover hosts are public
// CDNs and the response is accepted only after the destination gate passes.
type TLSCoverFetcher struct {
client *http.Client
resolve CoverResolver
}
var _ CoverBytesFetcher = (*TLSCoverFetcher)(nil)
const coverRequestTimeout = 30 * time.Second
var carrierGradeNAT = netip.MustParsePrefix("100.64.0.0/10")
// NewCoverFetcher builds the production cover client with the real resolver.
func NewCoverFetcher() *TLSCoverFetcher {
return NewCoverFetcherWithResolver(nil)
}
// NewCoverFetcherWithResolver builds a cover client using resolve, or the real
// system resolver when resolve is nil.
func NewCoverFetcherWithResolver(resolve CoverResolver) *TLSCoverFetcher {
if resolve == nil {
resolve = defaultCoverResolver
}
return newCoverFetcher(newCoverHTTPClient(resolve), resolve)
}
func newCoverFetcher(client *http.Client, resolve CoverResolver) *TLSCoverFetcher {
f := &TLSCoverFetcher{client: client, resolve: resolve}
client.CheckRedirect = func(req *http.Request, _ []*http.Request) error {
if err := f.validateURL(req.Context(), req.URL); err != nil {
return fmt.Errorf("redirect destination: %w", err)
}
return nil
}
return f
}
func defaultCoverResolver(ctx context.Context, host string) ([]netip.Addr, error) {
return net.DefaultResolver.LookupNetIP(ctx, "ip", host)
}
func newCoverHTTPClient(resolve CoverResolver) *http.Client {
base, ok := http.DefaultTransport.(*http.Transport)
if !ok {
base = &http.Transport{}
}
transport := base.Clone()
// A proxy would make the dial target the proxy rather than the cover host,
// defeating destination classification. Cover fetching is direct by design.
transport.Proxy = nil
dialer := &net.Dialer{}
transport.DialContext = func(ctx context.Context, network, address string) (net.Conn, error) {
host, port, err := net.SplitHostPort(address)
if err != nil {
return nil, fmt.Errorf("split cover address %q: %w", address, err)
}
addrs, err := resolveCoverHost(ctx, host, resolve)
if err != nil {
return nil, err
}
for _, addr := range addrs {
if !publicCoverAddress(addr) {
return nil, fmt.Errorf("cover host resolves to refused address %s", addr)
}
conn, err := dialer.DialContext(ctx, network, net.JoinHostPort(addr.String(), port))
if err == nil {
return conn, nil
}
}
return nil, fmt.Errorf("cover host %q has no reachable address", host)
}
return &http.Client{Transport: transport, Timeout: coverRequestTimeout}
}
func (f *TLSCoverFetcher) Fetch(ctx context.Context, sourceURL string) ([]byte, string, error) {
u, err := url.Parse(sourceURL)
if err != nil {
return nil, "", fmt.Errorf("parse cover URL: %w", err)
}
if err := f.validateURL(ctx, u); err != nil {
return nil, "", err
}
req, err := http.NewRequestWithContext(ctx, http.MethodGet, u.String(), nil)
if err != nil {
return nil, "", fmt.Errorf("build cover request: %w", err)
}
resp, err := f.client.Do(req)
if err != nil {
return nil, "", fmt.Errorf("fetch cover: %w", err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return nil, "", fmt.Errorf("fetch cover: status %d", resp.StatusCode)
}
raw, _, err := mime.ParseMediaType(resp.Header.Get("Content-Type"))
if err != nil {
return nil, "", fmt.Errorf("fetch cover: unsupported content type %q", resp.Header.Get("Content-Type"))
}
contentType, ok := store.CoverContentType(raw)
if !ok {
return nil, "", fmt.Errorf("fetch cover: unsupported content type %q", raw)
}
if resp.ContentLength > maxBodyBytes {
return nil, "", fmt.Errorf("fetch cover: response exceeds %d bytes", maxBodyBytes)
}
body, err := io.ReadAll(io.LimitReader(resp.Body, maxBodyBytes+1))
if err != nil {
return nil, "", fmt.Errorf("read cover: %w", err)
}
if len(body) > maxBodyBytes {
return nil, "", fmt.Errorf("fetch cover: response exceeds %d bytes", maxBodyBytes)
}
return body, contentType, nil
}
// This gate deliberately differs from fetchableSeriesURL: cover hosts are
// site-independent CDNs, so a Site host allowlist would reject valid covers.
func (f *TLSCoverFetcher) validateURL(ctx context.Context, u *url.URL) error {
if u == nil || u.Scheme != "https" || u.Host == "" || u.User != nil {
return errors.New("cover URL must use HTTPS without credentials")
}
host := u.Hostname()
if host == "" {
return errors.New("cover URL has no host")
}
addrs, err := resolveCoverHost(ctx, host, f.resolve)
if err != nil {
return fmt.Errorf("resolve cover host %q: %w", host, err)
}
if len(addrs) == 0 {
return fmt.Errorf("resolve cover host %q: no addresses", host)
}
for _, addr := range addrs {
if !publicCoverAddress(addr) {
return fmt.Errorf("cover host %q resolves to refused address %s", host, addr)
}
}
return nil
}
func resolveCoverHost(ctx context.Context, host string, resolve CoverResolver) ([]netip.Addr, error) {
if literal, err := netip.ParseAddr(host); err == nil {
return []netip.Addr{literal.Unmap()}, nil
}
return resolve(ctx, strings.TrimSuffix(host, "."))
}
func publicCoverAddress(addr netip.Addr) bool {
addr = addr.Unmap()
return addr.IsValid() && addr.IsGlobalUnicast() &&
!addr.IsLoopback() && !addr.IsPrivate() && !addr.IsLinkLocalUnicast() &&
!carrierGradeNAT.Contains(addr)
}
+226
View File
@@ -0,0 +1,226 @@
package latest
import (
"bytes"
"context"
"crypto/tls"
"io"
"net"
"net/http"
"net/http/httptest"
"net/netip"
"testing"
)
func TestCoverFetcherFetchesPublicHTTPSImage(t *testing.T) {
server := httptest.NewTLSServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.TLS == nil {
t.Fatal("cover request was not made over TLS")
}
w.Header().Set("Content-Type", "image/jpeg")
io.WriteString(w, "cover-bytes")
}))
defer server.Close()
transport := server.Client().Transport.(*http.Transport).Clone()
transport.TLSClientConfig = &tls.Config{InsecureSkipVerify: true} // test server certificate
transport.DialContext = func(ctx context.Context, network, _ string) (net.Conn, error) {
return (&net.Dialer{}).DialContext(ctx, network, server.Listener.Addr().String())
}
client := &http.Client{Transport: transport}
fetcher := newCoverFetcher(client, func(context.Context, string) ([]netip.Addr, error) {
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
})
body, contentType, err := fetcher.Fetch(context.Background(), "https://cdn.example/cover.jpg")
if err != nil {
t.Fatalf("Fetch: %v", err)
}
if string(body) != "cover-bytes" || contentType != "image/jpeg" {
t.Fatalf("Fetch = (%q, %q), want (cover-bytes, image/jpeg)", body, contentType)
}
}
func TestNewCoverFetcherRechecksResolverBeforeConnection(t *testing.T) {
var requests int
server := httptest.NewTLSServer(http.HandlerFunc(func(http.ResponseWriter, *http.Request) {
requests++
}))
defer server.Close()
_, port, err := net.SplitHostPort(server.Listener.Addr().String())
if err != nil {
t.Fatalf("server address: %v", err)
}
resolves := 0
fetcher := NewCoverFetcherWithResolver(func(context.Context, string) ([]netip.Addr, error) {
resolves++
if resolves == 1 {
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
}
return []netip.Addr{netip.MustParseAddr("127.0.0.1")}, nil
})
_, _, err = fetcher.Fetch(context.Background(), "https://cdn.example:"+port+"/cover.jpg")
if err == nil {
t.Fatal("Fetch accepted a destination that became private")
}
if resolves != 2 {
t.Fatalf("resolver calls = %d, want preflight and dial checks", resolves)
}
if requests != 0 {
t.Fatalf("requests = %d, want 0", requests)
}
}
type roundTripFunc func(*http.Request) (*http.Response, error)
func (f roundTripFunc) RoundTrip(r *http.Request) (*http.Response, error) { return f(r) }
func coverResponse(status int, contentType, location string, body []byte) *http.Response {
header := make(http.Header)
if contentType != "" {
header.Set("Content-Type", contentType)
}
if location != "" {
header.Set("Location", location)
}
return &http.Response{
StatusCode: status,
Status: http.StatusText(status),
Header: header,
Body: io.NopCloser(bytes.NewReader(body)),
ContentLength: int64(len(body)),
}
}
func TestCoverFetcherRefusesUnsafeDestinationsBeforeRequest(t *testing.T) {
var calls int
client := &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) {
calls++
return coverResponse(http.StatusOK, "image/jpeg", "", []byte("must not reach network")), nil
})}
resolve := func(_ context.Context, host string) ([]netip.Addr, error) {
switch host {
case "loopback.example":
return []netip.Addr{netip.MustParseAddr("127.0.0.1")}, nil
case "private.example":
return []netip.Addr{netip.MustParseAddr("10.0.0.1")}, nil
case "linklocal.example":
return []netip.Addr{netip.MustParseAddr("169.254.1.1")}, nil
case "unique-local.example":
return []netip.Addr{netip.MustParseAddr("fc00::1")}, nil
case "cgnat.example":
return []netip.Addr{netip.MustParseAddr("100.64.0.1")}, nil
default:
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
}
}
fetcher := newCoverFetcher(client, resolve)
tests := []string{
"http://public.example/cover.jpg",
"https://127.0.0.1/cover.jpg",
"https://10.0.0.1/cover.jpg",
"https://169.254.1.1/cover.jpg",
"https://[fc00::1]/cover.jpg",
"https://100.64.0.1/cover.jpg",
"https://loopback.example/cover.jpg",
"https://private.example/cover.jpg",
"https://linklocal.example/cover.jpg",
"https://unique-local.example/cover.jpg",
"https://cgnat.example/cover.jpg",
}
for _, sourceURL := range tests {
t.Run(sourceURL, func(t *testing.T) {
calls = 0
if _, _, err := fetcher.Fetch(context.Background(), sourceURL); err == nil {
t.Fatal("Fetch accepted refused destination")
}
if calls != 0 {
t.Fatalf("network calls = %d, want 0", calls)
}
})
}
}
func TestCoverFetcherStopsRedirectIntoPrivateAddress(t *testing.T) {
var calls int
client := &http.Client{Transport: roundTripFunc(func(req *http.Request) (*http.Response, error) {
calls++
if req.URL.Hostname() != "cdn.example" {
t.Fatalf("redirect reached %s", req.URL)
}
return coverResponse(http.StatusFound, "", "https://internal.example/cover.jpg", nil), nil
})}
fetcher := newCoverFetcher(client, func(_ context.Context, host string) ([]netip.Addr, error) {
if host == "internal.example" {
return []netip.Addr{netip.MustParseAddr("192.168.1.1")}, nil
}
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
})
if _, _, err := fetcher.Fetch(context.Background(), "https://cdn.example/cover.jpg"); err == nil {
t.Fatal("Fetch followed redirect into private address")
}
if calls != 1 {
t.Fatalf("network calls = %d, want only public first hop", calls)
}
}
func TestCoverFetcherRejectsOversizedBody(t *testing.T) {
var calls int
client := &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) {
calls++
response := coverResponse(http.StatusOK, "image/webp", "", bytes.Repeat([]byte("x"), maxBodyBytes+1))
response.ContentLength = -1
return response, nil
})}
fetcher := newCoverFetcher(client, func(context.Context, string) ([]netip.Addr, error) {
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
})
if _, _, err := fetcher.Fetch(context.Background(), "https://cdn.example/large.webp"); err == nil {
t.Fatal("Fetch accepted oversized body")
}
if calls != 1 {
t.Fatalf("network calls = %d, want 1", calls)
}
}
func TestCoverFetcherRejectsNonImage(t *testing.T) {
var calls int
client := &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) {
calls++
return coverResponse(http.StatusOK, "text/html", "", []byte("challenge")), nil
})}
fetcher := newCoverFetcher(client, func(context.Context, string) ([]netip.Addr, error) {
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
})
if _, _, err := fetcher.Fetch(context.Background(), "https://cdn.example/challenge"); err == nil {
t.Fatal("Fetch accepted non-image response")
}
if calls != 1 {
t.Fatalf("network calls = %d, want 1", calls)
}
}
// comix labels its covers "image/jpg", which is not a registered type but is
// what the Site actually answers with; the bytes are stored under the real
// name so one image cannot land under two spellings.
func TestCoverFetcherCanonicalisesJpgAlias(t *testing.T) {
client := &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) {
return coverResponse(http.StatusOK, "image/jpg", "", []byte("cover-bytes")), nil
})}
fetcher := newCoverFetcher(client, func(context.Context, string) ([]netip.Addr, error) {
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
})
body, contentType, err := fetcher.Fetch(context.Background(), "https://static.comix.to/cover.jpg")
if err != nil {
t.Fatalf("Fetch: %v", err)
}
if string(body) != "cover-bytes" || contentType != "image/jpeg" {
t.Fatalf("Fetch = (%q, %q), want (cover-bytes, image/jpeg)", body, contentType)
}
}
+3 -2
View File
@@ -11,8 +11,9 @@ import (
)
// maxBodyBytes caps what a single series page can cost in memory. Real pages
// measured 100-400 KB on 2026-07-26, so this is roughly 10x headroom and mostly
// guards against a proxy handing back something enormous.
// measured 100-400 KB on 2026-07-26; lightnovelworld runs larger — 685 KB and
// 1.18 MB measured 2026-08-11 — so the headroom there is roughly 3.5x, and the
// cap mostly guards against a proxy handing back something enormous.
const maxBodyBytes = 4 << 20
// chromeUA matches the client profile below. A Chrome fingerprint paired with a
+152 -89
View File
@@ -2,6 +2,7 @@ package latest
import (
"context"
"errors"
"log"
"net/url"
"time"
@@ -15,6 +16,13 @@ type Fetcher interface {
Get(ctx context.Context, url string) (body string, status int, err error)
}
// BrowserCoverFetcher retrieves one cover's bytes through the browser-backed
// path — the only route that clears the challenge kagane's image URLs answer
// a plain fetch with. Satisfied by BrowserFetcher.
type BrowserCoverFetcher interface {
Image(ctx context.Context, imageURL string) (body []byte, contentType string, err error)
}
// Poller re-checks each bookmarked series' newest published chapter on a
// schedule, independent of the userscript's own in-browser checks. The two run
// in parallel and report the same observable fact, so whichever writes last wins
@@ -23,12 +31,12 @@ type Fetcher interface {
// Two clocks, deliberately independent:
//
// - Interval is how often this goroutine wakes up and looks.
// - Cooldown is how long one series rests since its own last check.
// - Cooldowns are how long a series rests since its own last check. Browser-
// backed sites use the longer BrowserCooldown.
//
// Only the cooldown is per series, and it is enforced by the WHERE clause in
// DueForLatestCheck rather than by any timer. Shortening Interval therefore
// cannot shorten anyone's cooldown; it only makes the poller wake up and find
// nothing due more often.
// Cooldowns are enforced by the WHERE clause in DueForLatestCheck rather than
// by any timer. Shortening Interval therefore cannot shorten anyone's cooldown;
// it only makes the poller wake up and find nothing due more often.
type Poller struct {
Store *store.Store
Fetch Fetcher
@@ -36,24 +44,103 @@ type Poller struct {
// cannot clear. Nil disables those sites entirely rather than falling back
// to Fetch, which would only ever retrieve a challenge page.
BrowserFetch Fetcher
Now func() time.Time // injected so tests can freeze it
Cooldown time.Duration
Interval time.Duration
Stagger time.Duration
Batch int
// CoverFetch is optional; failures are logged and never affect the chapter poll.
CoverFetch BrowserCoverFetcher
// CoverBytesFetch is optional; it handles plain-TLS sources through the
// same failure-isolated prefetch path.
CoverBytesFetch CoverBytesFetcher
Now func() time.Time // injected so tests can freeze it
Cooldown time.Duration
BrowserCooldown time.Duration
Interval time.Duration
Stagger time.Duration
Batch int
}
// fetcherFor returns the fetcher a site needs, or nil when the site cannot be
// fetched at all right now. kagane and novelfull both sit behind a Cloudflare
// JavaScript challenge that no TLS fingerprint clears — kagane verified
// 2026-08-03, novelfull verified 2026-08-05, both against the same Chrome_133
// profile TLSFetcher uses — so they are browser-only or nothing.
func (p *Poller) fetcherFor(site string) Fetcher {
switch site {
case "kagane", "novelfull":
return p.BrowserFetch
// fillBlankCover gives a Series its Cover when it has none. The blank state is
// what "no Cover yet" means on the wire (ADR-0007): permanently-blank rows
// created before acquisition existed, and rows whose creation-time fetch
// failed, both heal here. A non-blank CoverAddress is left alone — refetching
// would add a request per Series per cycle and change artwork under the Reader
// for no visible reason. A row that already carries a source URL is owned by
// prefetchCover instead; this path only records a Cover address already
// extracted from the series page.
//
// Failures are logged against the Series and never returned: the chapter poll
// must not notice. A failed fill is retried the next time this Series is due;
// there is no separate retry queue.
func (p *Poller) fillBlankCover(ctx context.Context, sr store.Series, cover string) {
if sr.CoverAddress != "" || sr.Cover != "" {
return
}
return p.Fetch
if cover == "" {
return
}
p.storeCover(ctx, sr, cover)
}
// prefetchCover heals Series that already carry a third-party source URL but
// no stored address — the state left by client-supplied covers before
// acquisition moved server-side. Every Site takes the same path; fetchCoverBytes
// routes by URL shape, so browser-claimed URLs still need the sidecar. New
// blanks have no source URL and go through fillBlankCover from the series page
// instead.
func (p *Poller) prefetchCover(ctx context.Context, sr store.Series) {
if sr.Cover == "" || sr.CoverAddress != "" {
return
}
body, contentType, found, err := p.Store.GetCover(sr.Cover)
if err != nil {
log.Printf("latest poll %q: read cover: %v", sr.Key(), err)
return
}
if found {
if err := p.Store.SetSeriesCover(sr.Site, sr.SeriesID, sr.Cover, body, contentType); err != nil {
log.Printf("latest poll %q: persist cover: %v", sr.Key(), err)
}
return
}
p.storeCover(ctx, sr, sr.Cover)
}
// storeCover fetches bytes for sourceURL and points the Series at them. Every
// failure is logged against the Series and swallowed so the chapter poll
// cannot see it.
func (p *Poller) storeCover(ctx context.Context, sr store.Series, sourceURL string) {
bytes, contentType, err := fetchCoverBytes(ctx, sourceURL, p.CoverFetch, p.CoverBytesFetch)
if err != nil {
log.Printf("latest poll %q: fetch cover %s: %v", sr.Key(), sourceURL, err)
return
}
if err := p.Store.SetSeriesCover(sr.Site, sr.SeriesID, sourceURL, bytes, contentType); err != nil {
log.Printf("latest poll %q: persist cover: %v", sr.Key(), err)
}
}
// fetcherFor returns the fetcher a site's page needs, or nil when the site
// cannot be fetched at all right now. A Site whose registry entry carries a
// Browser read — kagane and novelfull, both behind a Cloudflare JavaScript
// challenge no TLS fingerprint clears — prefers the browser; when it is
// absent, the entry's Fallback decides whether plain TLS may take over. One
// routing rule for the poll and the acquirer, so the two cannot drift apart.
func fetcherFor(site string, browser, tls Fetcher) Fetcher {
s, known := sites[site]
if !known {
// No registry entry means nothing to fetch or parse; fail closed even
// though the only caller gates first, so a future caller that skips
// the gate cannot hand an arbitrary https URL to the TLS fetcher.
return nil
}
if s.Browser == nil {
return tls
}
if browser != nil {
return browser
}
if s.Browser.Fallback {
return tls
}
return nil
}
// Run polls until ctx is cancelled.
@@ -63,8 +150,8 @@ func (p *Poller) fetcherFor(site string) Fetcher {
// failure mode for a misconfigured batch x stagger: a slower cadence, never
// concurrent fetch storms.
func (p *Poller) Run(ctx context.Context) {
log.Printf("latest-chapter poller: interval=%s cooldown=%s batch=%d stagger=%s",
p.Interval, p.Cooldown, p.Batch, p.Stagger)
log.Printf("latest-chapter poller: interval=%s cooldown=%s browser-cooldown=%s batch=%d stagger=%s",
p.Interval, p.Cooldown, p.BrowserCooldown, p.Batch, p.Stagger)
t := time.NewTicker(p.Interval)
defer t.Stop()
for {
@@ -80,8 +167,10 @@ func (p *Poller) Run(ctx context.Context) {
// runOnce processes one batch of due series.
func (p *Poller) runOnce(ctx context.Context) {
cutoff := p.Now().Add(-p.Cooldown).UnixMilli()
due, err := p.Store.DueForLatestCheck(cutoff, p.Batch)
now := p.Now()
cutoff := now.Add(-p.Cooldown).UnixMilli()
browserCutoff := now.Add(-p.BrowserCooldown).UnixMilli()
due, err := p.Store.DueForLatestCheck(cutoff, browserCutoff, browserBackedSites(), p.Batch)
if err != nil {
log.Printf("latest poll: due query: %v", err)
return
@@ -135,39 +224,35 @@ func (p *Poller) checkOne(ctx context.Context, sr store.Series) {
return
}
// series_url is client-supplied (PUT /bookmarks/{key} accepts any string),
// so this is not just an optimisation against burning a request on an
// unknown site: without it, the server would issue a GET from its own
// network position to whatever URL a token-holder writes, including
// link-local/internal addresses or non-https schemes. The cooldown above
// is already consumed, so a row that never passes this check is retried at
// cooldown pace rather than hot-looping.
if !fetchableSeriesURL(sr.Site, sr.SeriesURL) {
log.Printf("latest poll %q: not fetchable: site=%q url=%q", sr.Key(), sr.Site, sr.SeriesURL)
return
}
f := p.fetcherFor(sr.Site)
if f == nil {
log.Printf("latest poll %q: no fetcher for site %q", sr.Key(), sr.Site)
return
}
body, status, err := f.Get(ctx, sr.SeriesURL)
facts, err := readSeriesPage(ctx, sr.Site, sr.SeriesURL, p.BrowserFetch, p.Fetch)
if err != nil {
log.Printf("latest poll %q: fetch %s: %v", sr.Key(), sr.SeriesURL, err)
switch {
case errors.Is(err, errNotFetchable):
// The cooldown above is already consumed, so a row that never
// passes the gate is retried at cooldown pace rather than
// hot-looping.
log.Printf("latest poll %q: not fetchable: site=%q url=%q", sr.Key(), sr.Site, sr.SeriesURL)
return
case errors.Is(err, errNoFetcher):
log.Printf("latest poll %q: no fetcher for site %q", sr.Key(), sr.Site)
return
}
// A legacy cover heals independently of the page read: its source may
// answer — a CDN — while the origin does not, so a fetch failure does
// not skip the heal, matching the order the shared read replaced.
p.prefetchCover(ctx, sr)
log.Printf("latest poll %q: %v", sr.Key(), err)
return
}
if status != 200 {
log.Printf("latest poll %q: fetch %s: status %d", sr.Key(), sr.SeriesURL, status)
return
}
latest, ok := latestChapterFrom(sr.Site, sr.SeriesURL, body)
if !ok {
// A legacy cover source is healed independently of the page read.
p.prefetchCover(ctx, sr)
// Cover fill is independent of the chapter signal: a page that lost its
// chapter list may keep its og:image, and a blank Series heals either way.
p.fillBlankCover(ctx, sr, facts.Cover)
if !facts.HasLatest {
// Most likely a challenge page or a layout change. Either way the row is
// already stamped, so this waits out a cooldown instead of hot-looping.
log.Printf("latest poll %q: no chapter links in %d bytes", sr.Key(), len(body))
log.Printf("latest poll %q: no chapter links in %d bytes", sr.Key(), facts.BodyLen)
return
}
@@ -175,7 +260,7 @@ func (p *Poller) checkOne(ctx context.Context, sr store.Series) {
// chapter should correct the stored number downward. The comparison is
// against the due-query snapshot; a concurrent write in between only costs
// one redundant UPDATE of the same absolute value, never a wrong one.
if sr.LatestChapterNum != nil && *sr.LatestChapterNum == latest.Num {
if sr.LatestChapterNum != nil && *sr.LatestChapterNum == facts.Latest.Num {
return
}
@@ -183,52 +268,30 @@ func (p *Poller) checkOne(ctx context.Context, sr store.Series) {
// bookmark joining to it, and the bookmark's updated_at is never touched —
// a newly published chapter is not reading progress and must not reorder
// the list.
if err := p.Store.SetLatestChapter(sr.Site, sr.SeriesID, latest.Label, latest.Num); err != nil {
if err := p.Store.SetLatestChapter(sr.Site, sr.SeriesID, facts.Latest.Label, facts.Latest.Num); err != nil {
log.Printf("latest poll %q: set latest chapter: %v", sr.Key(), err)
return
}
log.Printf("latest poll %q: latest is now %s", sr.Key(), latest.Label)
log.Printf("latest poll %q: latest is now %s", sr.Key(), facts.Latest.Label)
}
// fetchableSeriesURL reports whether site is a site latestChapterFrom knows how
// to parse and seriesURL is safe to hand to a fetcher: an https URL with a
// non-empty host. series_url comes from client-supplied PUT bodies, so this is
// a defence against the poller being used to probe arbitrary hosts from the
// server's own network position, not just a check against wasted requests.
//
// Three sites are held to a stricter rule, each for a different reason:
//
// - kagane and novelfull are fetched by a headless browser, which executes
// JavaScript and carries cookies, and is therefore a far stronger SSRF
// primitive than an HTTP GET. Their hosts must match exactly, not merely
// be non-empty.
// - lightnovelworld's parser regex hardcodes its host, so a URL anywhere
// else could never yield a match — reject it here rather than burn the
// request.
// fetchableSeriesURL reports whether site is a Site the registry knows and
// seriesURL is safe to hand to a fetcher: an https URL whose host matches the
// Site's pinned hostname exactly. series_url comes from client-supplied PUT
// bodies, so this is a defence against the poller being used to probe
// arbitrary hosts from the server's own network position, not just a check
// against wasted requests. The pin guards different things per Site — a
// browser Site guards a control that executes JavaScript and carries cookies,
// a parser Site guards a wasted request — but the rule is one rule, from the
// registry.
func fetchableSeriesURL(site, seriesURL string) bool {
switch site {
case "asura", "demonic", "comix", "kagane", "novelfull", "lightnovelworld":
default:
s, known := sites[site]
if !known {
return false
}
u, err := url.Parse(seriesURL)
if err != nil {
return false
}
if u.Scheme != "https" || u.Host == "" {
return false
}
switch site {
case "kagane":
return u.Hostname() == "kagane.to"
case "novelfull":
// Fetched by a real browser, same as kagane, so the host is pinned
// rather than merely non-empty.
return u.Hostname() == "novelfull.com"
case "lightnovelworld":
// Its parser regex hardcodes this host, so a URL anywhere else could
// never yield a match — reject it here rather than burn the request.
return u.Hostname() == "lightnovelworld.net"
}
return true
return u.Scheme == "https" && u.Hostname() == s.Host
}
+668 -21
View File
@@ -3,7 +3,9 @@ package latest
import (
"context"
"crypto/sha256"
"database/sql"
"errors"
"log"
"os"
"strings"
"sync"
@@ -20,13 +22,16 @@ func TestMain(m *testing.M) { os.Exit(pgtest.Main(m)) }
// test needs one, is created by opening the same database as a second owner.
var testOwner = store.Owner{DiscordID: "test-owner", TokenHash: sha256.Sum256([]byte("owner-token-hash"))}
// testCoverBaseURL is the public origin every stored cover URL is built from.
const testCoverBaseURL = "https://bookmarks.test"
// newTestStore opens a store on a Postgres database of this test's own and
// returns the URL, for helpers that need a second connection to the same
// database (see TestRunOnceFetchesSharedSeriesOnce).
func newTestStore(t *testing.T) (*store.Store, string) {
t.Helper()
url := pgtest.URL(t)
s, err := store.Open(url, testOwner)
s, err := store.Open(url, testOwner, t.TempDir(), testCoverBaseURL)
if err != nil {
t.Fatalf("Open: %v", err)
}
@@ -103,18 +108,145 @@ func (f *fakeFetcher) callCount() int {
return len(f.calls)
}
type fakeCoverFetcher struct {
mu sync.Mutex
calls []string
body []byte
contentType string
err error
}
func (f *fakeCoverFetcher) Image(_ context.Context, imageURL string) ([]byte, string, error) {
f.mu.Lock()
f.calls = append(f.calls, imageURL)
f.mu.Unlock()
if f.err != nil {
return nil, "", f.err
}
return f.body, f.contentType, nil
}
func (f *fakeCoverFetcher) callCount() int {
f.mu.Lock()
defer f.mu.Unlock()
return len(f.calls)
}
type fakeBytesCoverFetcher struct {
mu sync.Mutex
calls []string
body []byte
contentType string
err error
}
func (f *fakeBytesCoverFetcher) Fetch(_ context.Context, sourceURL string) ([]byte, string, error) {
f.mu.Lock()
f.calls = append(f.calls, sourceURL)
f.mu.Unlock()
if f.err != nil {
return nil, "", f.err
}
return f.body, f.contentType, nil
}
func (f *fakeBytesCoverFetcher) callCount() int {
f.mu.Lock()
defer f.mu.Unlock()
return len(f.calls)
}
// newTestPoller wires a poller with a frozen clock and no stagger, so tests run
// instantly and deterministically.
func newTestPoller(t *testing.T, s *store.Store, f Fetcher, at time.Time) *Poller {
t.Helper()
return &Poller{
Store: s,
Fetch: f,
Now: func() time.Time { return at },
Cooldown: time.Hour,
Interval: 10 * time.Minute,
Stagger: 0,
Batch: 14,
Store: s,
Fetch: f,
Now: func() time.Time { return at },
Cooldown: time.Hour,
BrowserCooldown: 6 * time.Hour,
Interval: 10 * time.Minute,
Stagger: 0,
Batch: 14,
}
}
// seedCoverSource writes a Series' cover source address without any stored
// bytes. Nothing in production produces that state any more — a client cover
// is discarded and an acquired one arrives with its bytes — but rows created
// before covers moved server-side still carry one, and the prefetch is what
// heals them.
func seedCoverSource(t *testing.T, dbURL, site, seriesID, coverURL string) {
t.Helper()
db, err := sql.Open("pgx", dbURL)
if err != nil {
t.Fatalf("open %s: %v", dbURL, err)
}
defer db.Close()
if _, err := db.Exec(
`UPDATE series SET cover = $3 WHERE site = $1 AND series_id = $2`,
site, seriesID, coverURL); err != nil {
t.Fatalf("seed cover source %s:%s: %v", site, seriesID, err)
}
}
func TestRunOncePrefetchesPublicCover(t *testing.T) {
s, dbURL := newTestStore(t)
const (
key = "asura:chronicles-of-the-demon-faction-f886a8af"
seriesURL = "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af"
coverURL = "https://cdn.example/covers/chronicles.jpg"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "asura", SeriesID: "chronicles-of-the-demon-faction-f886a8af",
SeriesURL: seriesURL, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
seedCoverSource(t, dbURL, "asura", "chronicles-of-the-demon-faction-f886a8af", coverURL)
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"}
p := &Poller{
Store: s, Fetch: &fakeFetcher{body: asuraSeriesFixture, status: 200}, CoverBytesFetch: covers,
Now: func() time.Time { return time.UnixMilli(5_000_000) }, Cooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetch calls = %d, want 1", got)
}
body, contentType, found, err := s.GetCover(coverURL)
if err != nil || !found {
t.Fatalf("GetCover: %v found=%v", err, found)
}
if string(body) != "cover-bytes" || contentType != "image/jpeg" {
t.Fatalf("stored cover = (%q, %q), want (cover-bytes, image/jpeg)", body, contentType)
}
}
func TestRunOnceDoesNotStoreNonImagePublicCover(t *testing.T) {
s, dbURL := newTestStore(t)
const (
key = "asura:non-image-cover"
seriesURL = "https://asurascans.com/comics/non-image-cover"
coverURL = "https://cdn.example/covers/challenge"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "asura", SeriesID: "non-image-cover", SeriesURL: seriesURL,
UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
seedCoverSource(t, dbURL, "asura", "non-image-cover", coverURL)
p := &Poller{
Store: s, Fetch: &fakeFetcher{body: asuraSeriesFixture, status: 200},
CoverBytesFetch: &fakeBytesCoverFetcher{body: []byte("challenge"), contentType: "text/html"},
Now: func() time.Time { return time.UnixMilli(5_000_000) }, Cooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if _, _, found, err := s.GetCover(coverURL); err != nil || found {
t.Fatalf("non-image cover = found %v, err %v; want missing", found, err)
}
}
@@ -248,7 +380,7 @@ func TestRunOnceFetchesSharedSeriesOnce(t *testing.T) {
// A second reader tracks the same series. The seed is the only
// reader-creation path, so a second Open as a different owner is how a
// test gets a second reader on the same database.
other, err := store.Open(url, store.Owner{DiscordID: "second-reader", TokenHash: sha256.Sum256([]byte("second-token-hash"))})
other, err := store.Open(url, store.Owner{DiscordID: "second-reader", TokenHash: sha256.Sum256([]byte("second-token-hash"))}, t.TempDir(), testCoverBaseURL)
if err != nil {
t.Fatalf("Open second reader: %v", err)
}
@@ -304,6 +436,26 @@ func TestRunOnceOneBadSeriesDoesNotStallBatch(t *testing.T) {
}
}
func TestRunLogsCooldowns(t *testing.T) {
var logs strings.Builder
previous := log.Writer()
log.SetOutput(&logs)
t.Cleanup(func() { log.SetOutput(previous) })
ctx, cancel := context.WithCancel(context.Background())
cancel()
(&Poller{
Cooldown: time.Hour,
BrowserCooldown: 6 * time.Hour,
Interval: time.Hour,
}).Run(ctx)
if got := logs.String(); !strings.Contains(got, "cooldown=1h") ||
!strings.Contains(got, "browser-cooldown=6h") {
t.Fatalf("startup log = %q, want both cooldowns", got)
}
}
// The cooldown is enforced by the due query, so a second immediate pass must do
// nothing at all — this is what makes the tick interval independent of it.
func TestRunOnceHonoursCooldownAcrossPasses(t *testing.T) {
@@ -334,6 +486,44 @@ func TestRunOnceHonoursCooldownAcrossPasses(t *testing.T) {
}
}
func TestRunOnceUsesBrowserCooldown(t *testing.T) {
s, _ := newTestStore(t)
const browserKey = "kagane:019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
seedForCheck(t, s, "asura:plain", "https://asurascans.com/comics/plain", 0)
seedForCheck(t, s, browserKey, "https://kagane.to/series/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b", 0)
hour := time.Hour
now := time.Unix(2*int64(hour/time.Second), 0)
tls := &fakeFetcher{status: 200}
browser := &fakeFetcher{status: 200}
p := &Poller{
Store: s,
Fetch: tls,
BrowserFetch: browser,
Now: func() time.Time { return now },
Cooldown: hour,
BrowserCooldown: 6 * hour,
Batch: 10,
}
p.runOnce(context.Background())
if got := tls.callCount(); got != 1 {
t.Fatalf("plain-TLS fetches after 2h = %d, want 1", got)
}
if got := browser.callCount(); got != 0 {
t.Fatalf("browser fetches after 2h = %d, want 0", got)
}
now = time.Unix(7*int64(hour/time.Second), 0)
p.runOnce(context.Background())
if got := tls.callCount(); got != 2 {
t.Fatalf("plain-TLS fetches after 7h = %d, want 2", got)
}
if got := browser.callCount(); got != 1 {
t.Fatalf("browser fetches after 7h = %d, want 1", got)
}
}
// A site that retracts a chapter should correct the stored number downward,
// mirroring the userscript's equality check (L427) rather than a >.
func TestRunOnceCorrectsDownward(t *testing.T) {
@@ -433,6 +623,13 @@ func TestFetchableSeriesURL(t *testing.T) {
{"demonic https", "demonic", "https://demonicscans.org/manga/X", true},
{"comix https", "comix", "https://comix.to/title/n8we-dungeons-and-crayons", true},
{"kagane on its own host", "kagane", "https://kagane.to/series/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b", true},
// The three plain-TLS Sites are pinned too: a client-supplied
// series_url must not aim a fetcher at a lookalike host, even when
// the fetcher is only an HTTP GET.
{"asura on a foreign host", "asura", "https://asurascans.com.evil.example/comics/x", false},
{"asura on the dead old domain", "asura", "https://asuracomic.net/comics/x", false},
{"demonic on a lookalike host", "demonic", "https://demonicscans.org.evil.example/manga/X", false},
{"comix on a foreign host", "comix", "https://evil.example/title/x", false},
// The browser fetcher runs JavaScript and carries cookies, so a
// client-supplied series_url must not be able to aim it anywhere else.
{"kagane on a foreign host", "kagane", "https://evil.example/series/x", false},
@@ -465,12 +662,14 @@ func TestKaganeSkippedWhenNoBrowserFetcher(t *testing.T) {
}); err != nil {
t.Fatalf("seed: %v", err)
}
f := &fakeFetcher{body: kaganeAPIFixture, status: 200}
p := &Poller{
Store: s, Fetch: f,
Store: s,
Fetch: f,
Now: func() time.Time { return time.UnixMilli(5_000_000) },
Cooldown: time.Hour, Interval: time.Hour, Batch: 10,
Cooldown: time.Hour, BrowserCooldown: time.Hour,
Interval: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
@@ -479,6 +678,51 @@ func TestKaganeSkippedWhenNoBrowserFetcher(t *testing.T) {
}
}
// novelfull without a browser is not skipped outright: its challenge is a
// live time-varying fact, so the plain-TLS page fetch is attempted and — when
// the body answers — fills both the chapter and the Cover, exactly the
// client-scraped rows #62 wants healed.
func TestNovelfullUsesTLSWhenNoBrowserFetcher(t *testing.T) {
s, _ := newTestStore(t)
key := "novelfull:reverend-insanity"
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key,
Site: "novelfull",
SeriesID: "reverend-insanity",
SeriesURL: "https://novelfull.com/reverend-insanity.html",
UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
tlsF := &fakeFetcher{body: novelfullSeriesFixture + novelfullCoverFixture, status: 200}
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
p := &Poller{
Store: s, Fetch: tlsF, CoverBytesFetch: covers,
Now: func() time.Time { return time.UnixMilli(5_000_000) },
Cooldown: time.Hour, BrowserCooldown: time.Hour,
Interval: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if len(tlsF.calls) != 1 {
t.Fatalf("TLS fetcher calls = %d, want 1", len(tlsF.calls))
}
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetches = %d, want 1", got)
}
got, found, err := s.Get(s.OwnerID(), key)
if err != nil || !found {
t.Fatalf("Get: %v found=%v", err, found)
}
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(novelfullCoverURL); got.Cover != want {
t.Fatalf("Cover = %q, want %q", got.Cover, want)
}
if got.LatestChapterNum == nil || *got.LatestChapterNum != 2334 {
t.Fatalf("LatestChapterNum = %v, want 2334", got.LatestChapterNum)
}
}
// With a browser fetcher wired up, kagane goes to it and not to the TLS one.
func TestKaganeUsesBrowserFetcher(t *testing.T) {
s, _ := newTestStore(t)
@@ -498,7 +742,8 @@ func TestKaganeUsesBrowserFetcher(t *testing.T) {
p := &Poller{
Store: s, Fetch: tlsF, BrowserFetch: browserF,
Now: func() time.Time { return time.UnixMilli(5_000_000) },
Cooldown: time.Hour, Interval: time.Hour, Batch: 10,
Cooldown: time.Hour, BrowserCooldown: time.Hour,
Interval: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
@@ -517,23 +762,209 @@ func TestKaganeUsesBrowserFetcher(t *testing.T) {
}
}
func TestRunOncePrefetchesKaganeCover(t *testing.T) {
s, dbURL := newTestStore(t)
const (
key = "kagane:019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
coverURL = "https://kagane.to/api/v2/image/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b/compressed"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "kagane", SeriesID: "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b",
SeriesURL: "https://kagane.to/series/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b",
UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
seedCoverSource(t, dbURL, "kagane", "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b", coverURL)
covers := &fakeCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
p := &Poller{
Store: s, Fetch: &fakeFetcher{body: kaganeAPIFixture, status: 200},
BrowserFetch: &fakeFetcher{body: kaganeAPIFixture, status: 200}, CoverFetch: covers,
Now: func() time.Time { return time.UnixMilli(5_000_000) },
Cooldown: time.Hour, BrowserCooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
body, contentType, ok, err := s.CoverByAddress(store.CoverAddress(coverURL))
if err != nil || !ok {
t.Fatalf("CoverByAddress: %v found=%v", err, ok)
}
if string(body) != "cover-bytes" || contentType != "image/webp" {
t.Fatalf("stored cover = (%q, %q), want (cover-bytes, image/webp)", body, contentType)
}
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetch calls = %d, want 1", got)
}
if got := readBookmark(t, s, key); got.Cover != testCoverBaseURL+"/covers/"+store.CoverAddress(coverURL) {
t.Fatalf("wire Cover = %q, want content-addressed URL", got.Cover)
}
}
func TestRunOnceDoesNotRefetchKaganeCover(t *testing.T) {
s, dbURL := newTestStore(t)
const (
key = "kagane:019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
seriesID = "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
coverURL = "https://kagane.to/api/v2/image/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b/compressed"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "kagane", SeriesID: seriesID,
SeriesURL: "https://kagane.to/series/" + seriesID, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
seedCoverSource(t, dbURL, "kagane", seriesID, coverURL)
at := time.UnixMilli(5_000_000)
covers := &fakeCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
p := &Poller{
Store: s, BrowserFetch: &fakeFetcher{body: kaganeAPIFixture, status: 200}, CoverFetch: covers,
Now: func() time.Time { return at }, Cooldown: time.Hour, BrowserCooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
at = at.Add(2 * time.Hour)
p.runOnce(context.Background())
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetch calls = %d, want 1 after two due cycles", got)
}
}
func TestRunOnceCoverFailureDoesNotBlockChapter(t *testing.T) {
s, dbURL := newTestStore(t)
const (
key = "kagane:019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
seriesID = "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
coverURL = "https://kagane.to/api/v2/image/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b/compressed"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "kagane", SeriesID: seriesID,
SeriesURL: "https://kagane.to/series/" + seriesID, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
seedCoverSource(t, dbURL, "kagane", seriesID, coverURL)
now := time.UnixMilli(5_000_000)
p := &Poller{
Store: s, BrowserFetch: &fakeFetcher{body: kaganeAPIFixture, status: 200},
CoverFetch: &fakeCoverFetcher{err: errors.New("browser unavailable")},
Now: func() time.Time { return now }, Cooldown: time.Hour, BrowserCooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
got, found, err := s.Get(s.OwnerID(), key)
if err != nil || !found {
t.Fatalf("Get: %v found=%v", err, found)
}
if got.LatestChapterNum == nil || *got.LatestChapterNum != 41 {
t.Fatalf("LatestChapterNum = %v, want 41", got.LatestChapterNum)
}
if checked := readLatestCheckedAt(t, s, key); checked != now.UnixMilli() {
t.Fatalf("latest_checked_at = %d, want %d", checked, now.UnixMilli())
}
}
func TestRunOnceRejectsInvalidKaganeCover(t *testing.T) {
s, dbURL := newTestStore(t)
const (
key = "kagane:019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
seriesID = "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
coverURL = "https://kagane.to/api/v2/image/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b/compressed"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "kagane", SeriesID: seriesID,
SeriesURL: "https://kagane.to/series/" + seriesID, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
seedCoverSource(t, dbURL, "kagane", seriesID, coverURL)
p := &Poller{
Store: s, BrowserFetch: &fakeFetcher{body: kaganeAPIFixture, status: 200},
CoverFetch: &fakeCoverFetcher{body: []byte("not an image"), contentType: "text/html"},
Now: func() time.Time { return time.UnixMilli(5_000_000) }, Cooldown: time.Hour, BrowserCooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if _, _, found, err := s.CoverByAddress(store.CoverAddress(coverURL)); err != nil || found {
t.Fatalf("invalid cover persisted = %v, err %v; want missing", found, err)
}
}
func TestRunOnceWithoutCoverFetcherStillPollsKagane(t *testing.T) {
s, dbURL := newTestStore(t)
const (
key = "kagane:019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
seriesID = "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
coverURL = "https://kagane.to/api/v2/image/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b/compressed"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "kagane", SeriesID: seriesID,
SeriesURL: "https://kagane.to/series/" + seriesID, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
seedCoverSource(t, dbURL, "kagane", seriesID, coverURL)
p := &Poller{
Store: s, BrowserFetch: &fakeFetcher{body: kaganeAPIFixture, status: 200},
Now: func() time.Time { return time.UnixMilli(5_000_000) }, Cooldown: time.Hour, BrowserCooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if _, _, found, err := s.CoverByAddress(store.CoverAddress(coverURL)); err != nil || found {
t.Fatalf("cover after nil CoverFetch = found %v, err %v; want missing", found, err)
}
}
func TestRunOnceRoutesNonKaganeCoverToPublicFetcher(t *testing.T) {
s, dbURL := newTestStore(t)
const key = "asura:solo"
const coverURL = "https://asurascans.com/covers/solo.jpg"
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "asura", SeriesID: "solo", SeriesURL: "https://asurascans.com/comics/solo",
UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
seedCoverSource(t, dbURL, "asura", "solo", coverURL)
browserCovers := &fakeCoverFetcher{body: []byte("must not be fetched"), contentType: "image/webp"}
publicCovers := &fakeBytesCoverFetcher{body: []byte("public cover"), contentType: "image/webp"}
p := &Poller{
Store: s, Fetch: &fakeFetcher{body: asuraSeriesFixture, status: 200}, CoverFetch: browserCovers,
CoverBytesFetch: publicCovers,
Now: func() time.Time { return time.UnixMilli(5_000_000) }, Cooldown: time.Hour, BrowserCooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if got := publicCovers.callCount(); got != 1 {
t.Fatalf("public cover fetch calls = %d, want 1", got)
}
if got := browserCovers.callCount(); got != 0 {
t.Fatalf("browser cover fetch calls for asura = %d, want 0", got)
}
}
func TestFetcherForRoutesNovelSites(t *testing.T) {
tls := &fakeFetcher{}
browser := &fakeFetcher{}
p := &Poller{Fetch: tls, BrowserFetch: browser}
cases := []struct {
site string
want Fetcher
site string
browser Fetcher
tls Fetcher
want Fetcher
}{
{"asura", tls},
{"lightnovelworld", tls},
{"kagane", browser},
{"novelfull", browser},
{"asura", browser, tls, tls},
{"lightnovelworld", browser, tls, tls},
{"kagane", browser, tls, browser},
{"novelfull", browser, tls, browser},
// browser-less deployment: kagane is nothing, novelfull degrades to TLS
{"kagane", nil, tls, nil},
{"novelfull", nil, tls, tls},
// unknown site: fail closed — nothing to fetch or parse
{"mangadex", browser, tls, nil},
}
for _, tc := range cases {
t.Run(tc.site, func(t *testing.T) {
if got := p.fetcherFor(tc.site); got != tc.want {
if got := fetcherFor(tc.site, tc.browser, tc.tls); got != tc.want {
t.Fatalf("fetcherFor(%q) = %v, want %v", tc.site, got, tc.want)
}
})
@@ -562,3 +993,219 @@ func TestFetchableSeriesURLPinsNovelHosts(t *testing.T) {
})
}
}
// A Series that has been blank since creation has no source URL to refetch.
// The poll extracts the Cover from the same series page it already fetched
// for the chapter signal and stores the bytes — every Site, both Libraries.
func TestRunOnceFillsBlankCoverFromSeriesPage(t *testing.T) {
cases := []struct {
name string
key string
site string
seriesID string
seriesURL string
kind string
body string
wantCover string
browser bool
}{
{
name: "asura manga",
key: "asura:chronicles-of-the-demon-faction-f886a8af", site: "asura",
seriesID: "chronicles-of-the-demon-faction-f886a8af",
seriesURL: "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af",
kind: store.KindManga, body: asuraSeriesFixture + asuraCoverFixture,
wantCover: "https://cdn.asurascans.com/asura-images/covers/chronicles-of-the-demon-faction.d4dcb8.webp",
},
{
name: "lightnovelworld novel",
key: "lightnovelworld:all-jobs-and-classes-i-just-wanted-one-skill-not-them-all", site: "lightnovelworld",
seriesID: "all-jobs-and-classes-i-just-wanted-one-skill-not-them-all", seriesURL: "https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/",
kind: store.KindNovel, body: lnwSeriesFixture + lnwCoverFixture,
wantCover: "https://i1.wp.com/lightnovelworld.net/wp-content/uploads/2025/10/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all.jpg",
},
{
name: "kagane manga",
key: "kagane:019fe11a-8670-7cf3-8343-0b02057d3787", site: "kagane",
seriesID: "019fe11a-8670-7cf3-8343-0b02057d3787",
seriesURL: "https://kagane.to/series/019fe11a-8670-7cf3-8343-0b02057d3787",
kind: store.KindManga, body: kaganeAPIFixtureWithCover, browser: true,
wantCover: "https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
s, _ := newTestStore(t)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: tc.key, Site: tc.site, SeriesID: tc.seriesID,
SeriesURL: tc.seriesURL, Kind: tc.kind, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
page := &fakeFetcher{body: tc.body, status: 200}
public := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"}
browser := &fakeCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
p := &Poller{
Store: s, Fetch: page, BrowserFetch: page,
CoverBytesFetch: public, CoverFetch: browser,
Now: func() time.Time { return time.UnixMilli(5_000_000) },
Cooldown: time.Hour, BrowserCooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
got := readBookmark(t, s, tc.key)
wantWire := testCoverBaseURL + "/covers/" + store.CoverAddress(tc.wantCover)
if got.Cover != wantWire {
t.Fatalf("Cover = %q, want %q", got.Cover, wantWire)
}
if tc.browser {
if got := browser.callCount(); got != 1 {
t.Fatalf("browser cover fetches = %d, want 1", got)
}
if got := public.callCount(); got != 0 {
t.Fatalf("public cover fetches = %d, want 0", got)
}
} else {
if got := public.callCount(); got != 1 {
t.Fatalf("public cover fetches = %d, want 1", got)
}
if got := browser.callCount(); got != 0 {
t.Fatalf("browser cover fetches = %d, want 0", got)
}
}
})
}
}
// Once a Cover exists the poll must leave it alone: refetching every cycle is
// noise for the Reader and a request per Series against Sites that already
// bot-score the deployment's single IP.
func TestRunOnceDoesNotReplaceExistingCover(t *testing.T) {
s, _ := newTestStore(t)
const (
key = "asura:chronicles-of-the-demon-faction-f886a8af"
seriesID = "chronicles-of-the-demon-faction-f886a8af"
seriesURL = "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af"
first = "https://cdn.example/covers/first.jpg"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "asura", SeriesID: seriesID, SeriesURL: seriesURL, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
if err := s.SetSeriesCover("asura", seriesID, first, []byte("first"), "image/jpeg"); err != nil {
t.Fatalf("seed cover: %v", err)
}
public := &fakeBytesCoverFetcher{body: []byte("second"), contentType: "image/jpeg"}
at := time.UnixMilli(5_000_000)
p := &Poller{
Store: s, Fetch: &fakeFetcher{body: asuraSeriesFixture + asuraCoverFixture, status: 200},
CoverBytesFetch: public,
Now: func() time.Time { return at }, Cooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
at = at.Add(2 * time.Hour)
p.runOnce(context.Background())
if got := public.callCount(); got != 0 {
t.Fatalf("cover fetch calls = %d, want 0", got)
}
got := readBookmark(t, s, key)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(first); got.Cover != want {
t.Fatalf("Cover = %q, want the first one %q", got.Cover, want)
}
}
// A blank Cover whose byte fetch fails is retried the next time the Series is
// polled. There is no separate retry queue — the due cycle is the queue.
func TestRunOnceRetriesFailedBlankCoverOnNextPoll(t *testing.T) {
s, _ := newTestStore(t)
const (
key = "asura:chronicles-of-the-demon-faction-f886a8af"
seriesID = "chronicles-of-the-demon-faction-f886a8af"
seriesURL = "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af"
coverURL = "https://cdn.asurascans.com/asura-images/covers/chronicles-of-the-demon-faction.d4dcb8.webp"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "asura", SeriesID: seriesID, SeriesURL: seriesURL, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
public := &fakeBytesCoverFetcher{err: errors.New("cdn down")}
at := time.UnixMilli(5_000_000)
p := &Poller{
Store: s, Fetch: &fakeFetcher{body: asuraSeriesFixture + asuraCoverFixture, status: 200},
CoverBytesFetch: public,
Now: func() time.Time { return at }, Cooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if got := readBookmark(t, s, key); got.Cover != "" {
t.Fatalf("Cover after failed fetch = %q, want blank", got.Cover)
}
if got := public.callCount(); got != 1 {
t.Fatalf("cover fetch calls after fail = %d, want 1", got)
}
public.err = nil
public.body = []byte("cover-bytes")
public.contentType = "image/jpeg"
at = at.Add(2 * time.Hour)
p.runOnce(context.Background())
if got := public.callCount(); got != 2 {
t.Fatalf("cover fetch calls after retry = %d, want 2", got)
}
got := readBookmark(t, s, key)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(coverURL); got.Cover != want {
t.Fatalf("Cover after retry = %q, want %q", got.Cover, want)
}
}
// Cover work is cosmetic: a failed blank fill must leave the chapter poll's
// result intact for every Site, not only kagane.
func TestRunOnceBlankCoverFailureDoesNotBlockChapter(t *testing.T) {
s, _ := newTestStore(t)
const (
key = "asura:chronicles-of-the-demon-faction-f886a8af"
seriesID = "chronicles-of-the-demon-faction-f886a8af"
seriesURL = "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "asura", SeriesID: seriesID, SeriesURL: seriesURL, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
now := time.UnixMilli(5_000_000)
var logs strings.Builder
prev := log.Writer()
log.SetOutput(&logs)
t.Cleanup(func() { log.SetOutput(prev) })
p := &Poller{
Store: s, Fetch: &fakeFetcher{body: asuraSeriesFixture + asuraCoverFixture, status: 200},
CoverBytesFetch: &fakeBytesCoverFetcher{err: errors.New("cdn down")},
Now: func() time.Time { return now }, Cooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
got := readBookmark(t, s, key)
if got.LatestChapterNum == nil || *got.LatestChapterNum != 181 {
t.Fatalf("LatestChapterNum = %v, want 181", got.LatestChapterNum)
}
if got.Cover != "" {
t.Fatalf("Cover = %q, want blank after failed fetch", got.Cover)
}
if !strings.Contains(logs.String(), key) {
t.Fatalf("cover failure log missing series key %q; got %q", key, logs.String())
}
}
// kaganeAPIFixture carries chapter data only. The blank-fill path needs a
// cover image id in the same body the chapter poll already retrieved.
const kaganeAPIFixtureWithCover = `
{"series_id":"019fe11a-8670-7cf3-8343-0b02057d3787","title":"Infinite Decryption",
"series_covers":[{"cover_id":"019fe11a-84d1-714b-9cf4-2827f277f3c0","language":"en","volume_number":"1","chapter_number":null,"note":null,"image_id":"019fe11a-84c3-7fc3-a84b-88787374b617"}],
"series_books":[{"book_id":"a","title":"Episode 1","chapter_no":"1","sort_no":1},
{"book_id":"b","title":"Episode 41","chapter_no":"41","sort_no":41},
{"book_id":"c","title":"Episode 40.5","chapter_no":"40.5","sort_no":40}]}
`
+58
View File
@@ -0,0 +1,58 @@
package latest
import (
"context"
"errors"
"fmt"
)
// seriesRead carries the two facts the poll and the acquirer both extract
// from a series page. Persistence, stamps and scheduling stay with the
// callers, so the policies that keep the two flows distinct (stamp order,
// cooldowns) are not swallowed by the module.
type seriesRead struct {
Latest latestChapter
HasLatest bool
Cover string
HasCover bool
// BodyLen is the fetched body's length, surfaced because the no-chapter
// log uses it to tell a markup change from a body the size cap cut short.
BodyLen int
}
// errNotFetchable and errNoFetcher separate the gate and the route from fetch
// failures so each caller keeps its own distinct log line for all three.
var (
errNotFetchable = errors.New("series url not fetchable")
errNoFetcher = errors.New("no fetcher for site")
)
// readSeriesPage performs the series-page read the poll and the acquirer have
// in common: gate the address, choose the route, fetch the page, extract the
// Latest Chapter and the Cover address. It persists nothing and stamps
// nothing.
//
// series_url arrives in a client-supplied PUT body (PUT /bookmarks/{key}
// accepts any string), so the gate is not an optimisation against burning a
// request on an unknown site: without it, the server would issue a GET from
// its own network position to whatever URL a token-holder writes, including
// link-local/internal addresses or non-https schemes.
func readSeriesPage(ctx context.Context, site, seriesURL string, browser, tls Fetcher) (seriesRead, error) {
if !fetchableSeriesURL(site, seriesURL) {
return seriesRead{}, fmt.Errorf("%w: site=%q url=%q", errNotFetchable, site, seriesURL)
}
f := fetcherFor(site, browser, tls)
if f == nil {
return seriesRead{}, fmt.Errorf("%w: site %q", errNoFetcher, site)
}
body, status, err := f.Get(ctx, seriesURL)
if err != nil {
return seriesRead{}, fmt.Errorf("fetch %s: %w", seriesURL, err)
}
if status != 200 {
return seriesRead{}, fmt.Errorf("fetch %s: status %d", seriesURL, status)
}
latest, hasLatest := latestChapterFrom(site, seriesURL, body)
cover, hasCover := coverFrom(site, seriesURL, body)
return seriesRead{Latest: latest, HasLatest: hasLatest, Cover: cover, HasCover: hasCover, BodyLen: len(body)}, nil
}
+370 -81
View File
@@ -1,10 +1,16 @@
package latest
import (
"encoding/json"
"html"
"log"
"net/url"
"regexp"
"sort"
"strconv"
"strings"
"github.com/chromedp/chromedp"
)
// latestChapter is the newest chapter a series page advertises.
@@ -13,6 +19,40 @@ type latestChapter struct {
Label string
}
// site answers the fixed questions every series-page read asks of its Site
// (ADR-0009): the host its addresses must carry, how to find the Latest
// Chapter and the Cover address in a body, and — for a Site behind a
// JavaScript challenge — how to read its payload from a cleared tab. One
// entry describes everything about one Site, and nowhere else gets to compare
// the site string.
type site struct {
// Host is the exact hostname a series_url for this Site must carry.
Host string
// LatestChapter finds the newest chapter in a fetched body.
LatestChapter func(seriesURL, body string) (latestChapter, bool)
// Cover finds the Cover address in a fetched body.
Cover func(seriesURL, body string) (string, bool)
// Browser reads this Site's payload from a cleared browser tab; nil
// means the page is fetched over plain TLS.
Browser *browserRead
}
type browserRead struct {
// Read builds the tab read for seriesURL, refusing (false) an address
// this Site will not open in a browser — the per-Site half of the SSRF
// gate, kept deliberately behind fetchableSeriesURL: a headless browser
// executes JavaScript and carries cookies, and series_url is
// client-supplied.
Read func(seriesURL string, out *string) (chromedp.Action, bool)
// Done reports whether the payload arrived.
Done func(body string) bool
// Fallback allows the plain-TLS fetcher when no browser is configured.
// False skips the Site instead. kagane is false — a plain fetch would
// only ever retrieve a challenge page — and novelfull is true, because
// its challenge is a live time-varying fact (AGENTS.md).
Fallback bool
}
// asuraSlugRe pulls the series slug out of a stored series_url.
// Shape verified live 2026-07-26: https://asurascans.com/comics/<slug>, where
// the slug carries a trailing build-hash suffix (e.g. "-f886a8af") that
@@ -34,6 +74,18 @@ var demonicChapterRe = regexp.MustCompile(`chaptered\.php\?manga=\d+&(?:amp;)?ch
// Only the id prefix is stable; the slug tail follows the title.
var comixSlugRe = regexp.MustCompile(`/title/([^/?#]+)`)
func comixSeriesID(seriesURL string) (string, bool) {
m := comixSlugRe.FindStringSubmatch(seriesURL)
if m == nil {
return "", false
}
id := m[1]
if i := strings.Index(id, "-"); i != -1 {
id = id[:i]
}
return id, true
}
// kaganeChapterRe matches the chapter numbers in a kagane API response. This
// branch is fed by the browser fetcher, so the body is JSON rather than HTML —
// there are no anchors to scan.
@@ -44,91 +96,37 @@ var kaganeChapterRe = regexp.MustCompile(`"chapter_no":"([0-9.]+)"`)
// "/<slug>/chapter-<n>[-<title-slug>].html". Verified live 2026-08-05.
var novelfullSlugRe = regexp.MustCompile(`^/([^/?#]+)\.html$`)
// lnwSlugRe does the same for lightnovelworld, whose series pages live under
// /novel/<slug>/ while its chapter URLs are flat at the site root:
// "/<slug>-chapter-<n>/", absolute in the page's own anchors. Verified live
// 2026-08-05.
var lnwSlugRe = regexp.MustCompile(`^/novel/([^/?#]+)/?$`)
// lnwChapterRe matches any chapter-shaped address on lightnovelworld. Unlike
// asura, novelfull and comix — which scope to their stored series slug so a
// foreign chapter link cannot contribute — this Site's chapter addresses carry
// the Chapter Slug, which is not the Series identity: one Series may publish
// under several Chapter Slugs (measured 2026-08-11: a sampled novel serves
// 1-99 under one slug and 100-423 under another), so no stored-slug pattern can
// cover a Series' whole list. An unscoped match is safe because
// lnwLatestChapter truncates the body at the comment thread before scanning
// (lnwCommentMarker); without that, a visitor's comment could set the Latest
// Chapter on the shared Series row.
var lnwChapterRe = regexp.MustCompile(`lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/`)
// latestChapterFrom returns the highest chapter number body advertises for this
// series. ok is false when the body yields nothing usable — an unknown site, an
// empty body, a Cloudflare challenge page, and a site redesign all land here,
// and the caller treats all four identically.
//
// Ported from the userscript's latestChapterFromAnchors (asura L123-133,
// demonic L183-193), including its reason for taking a maximum rather than a
// first or last: neither site lists chapters in a dependable order.
// lnwCommentMarker is the boundary of lightnovelworld's server-rendered
// wpdiscuz comment thread. It occurs exactly once per page and follows every
// chapter anchor (measured 2026-08-11,
// docs/research/lightnovelworld-chapter-vs-series-slug.md §6), so cutting the
// body at its first occurrence keeps the whole chapter list while excluding a
// region any visitor can write to. Absent means the page shape changed: the
// body is skipped, never scanned whole.
const lnwCommentMarker = "wpd-threads"
// maxChapter returns the highest chapter number the regex finds in body. A
// maximum rather than a first or last, ported from the userscript's
// latestChapterFromAnchors (asura L123-133, demonic L183-193): neither site
// lists chapters in a dependable order.
//
// The userscript's asura rule additionally requires the anchor text to match
// /Chapter\s+[\d.]+/i. That check exists only to skip the "First Chapter"
// shortcut, which points at chapter/1 and therefore can never win a maximum, so
// it is redundant here. For asura, scoping the pattern to this series' own slug
// replaces it with a stronger guarantee: a chapter link belonging to some other
// series cannot contribute even if the page starts carrying them. demonic has no
// such guarantee — demonicChapterRe matches any chaptered.php?manga=<id> anchor
// with no per-series scoping, because the stored series_id for demonic is a
// slug, not the numeric id the URL carries, so it cannot easily be scoped.
func latestChapterFrom(site, seriesURL, body string) (latestChapter, bool) {
var re *regexp.Regexp
switch site {
case "asura":
m := asuraSlugRe.FindStringSubmatch(seriesURL)
if m == nil {
return latestChapter{}, false
}
// Stored URLs predating a redeploy may carry a stale build hash;
// chapter hrefs in the fetched body carry the current one. Strip to
// the stable ID and make the hash optional in the pattern, so scoping
// survives rotations.
slug := asuraBuildHash.ReplaceAllString(m[1], "")
// Compiled per call rather than cached: this runs once per fetch, which
// is at most a few times a minute, and the slug varies per series.
re = regexp.MustCompile(`/comics/` + regexp.QuoteMeta(slug) + `(?:-[0-9a-f]{8})?/chapter/([0-9.]+)`)
case "demonic":
re = demonicChapterRe
case "comix":
m := comixSlugRe.FindStringSubmatch(seriesURL)
if m == nil {
return latestChapter{}, false
}
// comix ships an SPA: the served HTML carries a JSON state blob instead
// of chapter anchors, and latestChapterUrl is the only place the newest
// chapter appears. Scoping to this series' id prefix keeps a
// "recommended" strip's entries from winning the maximum.
id := m[1]
if i := strings.Index(id, "-"); i != -1 {
id = id[:i]
}
re = regexp.MustCompile(`"latestChapterUrl":"/title/` + regexp.QuoteMeta(id) + `-[^"]*-chapter-([0-9.]+)"`)
case "kagane":
re = kaganeChapterRe
case "novelfull":
u, err := url.Parse(seriesURL)
if err != nil {
return latestChapter{}, false
}
m := novelfullSlugRe.FindStringSubmatch(u.Path)
if m == nil {
return latestChapter{}, false
}
// Scoped to this series' slug for the same reason asura is: page 1
// carries a "latest chapters" widget and a "you may also like" strip,
// and neither may contribute to the maximum.
re = regexp.MustCompile(`/` + regexp.QuoteMeta(m[1]) + `/chapter-([0-9.]+)`)
case "lightnovelworld":
u, err := url.Parse(seriesURL)
if err != nil {
return latestChapter{}, false
}
m := lnwSlugRe.FindStringSubmatch(u.Path)
if m == nil {
return latestChapter{}, false
}
re = regexp.MustCompile(`lightnovelworld\.net/` + regexp.QuoteMeta(m[1]) + `-chapter-([0-9.]+)/`)
default:
return latestChapter{}, false
}
// shortcut, which points at chapter/1 and therefore can never win a maximum,
// so it is redundant once a maximum is taken.
func maxChapter(re *regexp.Regexp, body string) (latestChapter, bool) {
var best latestChapter
found := false
for _, m := range re.FindAllStringSubmatch(body, -1) {
@@ -146,3 +144,294 @@ func latestChapterFrom(site, seriesURL, body string) (latestChapter, bool) {
}
return best, found
}
// asuraLatestChapter scopes chapter links to this series' own slug, which
// replaces the userscript's anchor-text check with a stronger guarantee: a
// chapter link belonging to some other series cannot contribute even if the
// page starts carrying them.
func asuraLatestChapter(seriesURL, body string) (latestChapter, bool) {
m := asuraSlugRe.FindStringSubmatch(seriesURL)
if m == nil {
return latestChapter{}, false
}
// Stored URLs predating a redeploy may carry a stale build hash; chapter
// hrefs in the fetched body carry the current one. Strip to the stable ID
// and make the hash optional in the pattern, so scoping survives
// rotations.
slug := asuraBuildHash.ReplaceAllString(m[1], "")
// Compiled per call rather than cached: this runs once per fetch, which is
// at most a few times a minute, and the slug varies per series.
re := regexp.MustCompile(`/comics/` + regexp.QuoteMeta(slug) + `(?:-[0-9a-f]{8})?/chapter/([0-9.]+)`)
return maxChapter(re, body)
}
// demonicLatestChapter is not scoped: demonicChapterRe matches any
// chaptered.php?manga=<id> anchor, because the stored series_id is a slug,
// not the numeric id the URL carries, so it cannot be scoped.
func demonicLatestChapter(_, body string) (latestChapter, bool) {
return maxChapter(demonicChapterRe, body)
}
// comixLatestChapter reads comix's SPA: the served HTML carries a JSON state
// blob instead of chapter anchors, and latestChapterUrl is the only place the
// newest chapter appears. Scoping to this series' id prefix keeps a
// "recommended" strip's entries from winning the maximum.
func comixLatestChapter(seriesURL, body string) (latestChapter, bool) {
id, ok := comixSeriesID(seriesURL)
if !ok {
return latestChapter{}, false
}
re := regexp.MustCompile(`"latestChapterUrl":"/title/` + regexp.QuoteMeta(id) + `-[^"]*-chapter-([0-9.]+)"`)
return maxChapter(re, body)
}
// kaganeLatestChapter scans the kagane series API JSON that the browser read
// fetched from inside the page; the match rides on the property name,
// regardless of the surrounding JSON shape.
func kaganeLatestChapter(_, body string) (latestChapter, bool) {
return maxChapter(kaganeChapterRe, body)
}
// novelfullLatestChapter is scoped to this series' slug for the same reason
// asura is: page 1 carries a "latest chapters" widget and a "you may also
// like" strip, and neither may contribute to the maximum.
func novelfullLatestChapter(seriesURL, body string) (latestChapter, bool) {
u, err := url.Parse(seriesURL)
if err != nil {
return latestChapter{}, false
}
m := novelfullSlugRe.FindStringSubmatch(u.Path)
if m == nil {
return latestChapter{}, false
}
re := regexp.MustCompile(`/` + regexp.QuoteMeta(m[1]) + `/chapter-([0-9.]+)`)
return maxChapter(re, body)
}
// lnwLatestChapter truncates the body at the comment thread before scanning:
// it is the one region of the page any visitor can write to (see lnwChapterRe).
// A body without the marker is skipped, never scanned whole — a redesign must
// degrade into staleness, not into a wrong shared value; the logged body length
// tells a markup change from a body the size cap cut short.
func lnwLatestChapter(seriesURL, body string) (latestChapter, bool) {
i := strings.Index(body, lnwCommentMarker)
if i < 0 {
log.Printf("latest poll %q: no %s marker in %d bytes", seriesURL, lnwCommentMarker, len(body))
return latestChapter{}, false
}
return maxChapter(lnwChapterRe, body[:i])
}
// latestChapterFrom returns the highest chapter number body advertises for this
// series, via the Site's registry entry. ok is false when the body yields
// nothing usable — an unknown site, an empty body, a Cloudflare challenge page,
// and a site redesign all land here, and the caller treats all four identically.
func latestChapterFrom(site, seriesURL, body string) (latestChapter, bool) {
if fn := sites[site].LatestChapter; fn != nil {
return fn(seriesURL, body)
}
return latestChapter{}, false
}
var metaTagRe = regexp.MustCompile(`(?is)<meta\b[^>]*>`)
var doubleQuotedMetaAttrRe = regexp.MustCompile(`(?is)([a-z][a-z0-9:_-]*)\s*=\s*"([^"]*)"`)
var singleQuotedMetaAttrRe = regexp.MustCompile(`(?is)([a-z][a-z0-9:_-]*)\s*=\s*'([^']*)'`)
// comix's server-rendered page embeds query data in this JSON script; parsing
// the target detail entry avoids matching posters from recommended results.
var comixInitialDataRe = regexp.MustCompile(`(?is)<script\b[^>]*\bid\s*=\s*["']initial-data["'][^>]*>(.*?)</script>`)
// kaganeImageURLRe matches the canonical compressed image route kagane's API
// publishes — the only cover URL form the extractor emits and the browser
// fetcher accepts. The URL is matched in full (scheme, host, id shape) rather
// than trusted: the value a fetcher is pointed at may have been client-
// supplied, and a headless browser is a strong SSRF primitive.
var kaganeImageURLRe = regexp.MustCompile(`^https://kagane\.to/api/v2/image/([0-9a-f-]{36})/compressed$`)
// browserOnlyCoverURL reports whether the browser sidecar is the only fetcher
// for cover bytes at imageURL. kagane's image route answers a plain fetch with
// a challenge and `cross-origin-resource-policy: same-origin`, so a TLS fetch
// would only ever retrieve a challenge page and must not be attempted
// (ADR-0007). This is the byte-fetch router's per-Site knowledge; it lives in
// the extraction module, which owns kagane's URL shapes.
func browserOnlyCoverURL(imageURL string) bool {
return kaganeImageURLRe.MatchString(imageURL)
}
// kagane's browser-fetched series response publishes cover image IDs under
// series_covers. The API's canonical compressed image route is the only URL
// form accepted by the store and browser fetcher; no rendition is guessed.
func kaganeCoverURL(body string) string {
var response struct {
SeriesCovers []struct {
ImageID string `json:"image_id"`
} `json:"series_covers"`
}
if err := json.Unmarshal([]byte(body), &response); err != nil {
return ""
}
for _, cover := range response.SeriesCovers {
// Validate the assembled URL against the same regex the browser
// fetcher enforces, so the extractor can never emit an address the
// fetch would refuse.
imageURL := "https://kagane.to/api/v2/image/" + cover.ImageID + "/compressed"
if kaganeImageURLRe.MatchString(imageURL) {
return imageURL
}
}
return ""
}
func comixCoverURL(seriesURL, body string) string {
id, ok := comixSeriesID(seriesURL)
if !ok {
return ""
}
data := comixInitialDataRe.FindStringSubmatch(body)
if data == nil {
return ""
}
var state struct {
Queries map[string]json.RawMessage `json:"queries"`
}
if err := json.Unmarshal([]byte(data[1]), &state); err != nil {
return ""
}
raw := state.Queries[`["manga","detail","`+id+`"]`]
if len(raw) == 0 {
return ""
}
var detail struct {
Poster struct {
Medium string `json:"medium"`
} `json:"poster"`
}
if err := json.Unmarshal(raw, &detail); err != nil {
return ""
}
return publishedCoverURL(detail.Poster.Medium)
}
// ogImageCover reads the og:image metadata shared by asura, demonic and
// lightnovelworld.
func ogImageCover(_, body string) (string, bool) {
cover := metaContent(body, "property", "og:image")
return cover, cover != ""
}
func novelfullCoverEntry(_, body string) (string, bool) {
cover := metaContent(body, "name", "image")
return cover, cover != ""
}
func comixCoverEntry(seriesURL, body string) (string, bool) {
cover := comixCoverURL(seriesURL, body)
return cover, cover != ""
}
func kaganeCoverEntry(_, body string) (string, bool) {
cover := kaganeCoverURL(body)
return cover, cover != ""
}
// coverFrom reports false for unknown sites, challenge bodies, and pages with
// no usable cover, via the Site's registry entry.
func coverFrom(site, seriesURL, body string) (string, bool) {
if fn := sites[site].Cover; fn != nil {
return fn(seriesURL, body)
}
return "", false
}
// metaContent returns the content of the first <meta> whose attrName is
// attrValue. It keeps scanning after an empty match so a later published cover
// is not hidden by an empty tag.
func metaContent(body, attrName, attrValue string) string {
for _, tag := range metaTagRe.FindAllString(body, -1) {
attrs := make(map[string]string)
for _, m := range doubleQuotedMetaAttrRe.FindAllStringSubmatch(tag, -1) {
attrs[strings.ToLower(m[1])] = m[2]
}
for _, m := range singleQuotedMetaAttrRe.FindAllStringSubmatch(tag, -1) {
attrs[strings.ToLower(m[1])] = m[2]
}
if strings.EqualFold(attrs[strings.ToLower(attrName)], attrValue) {
if cover := publishedCoverURL(attrs["content"]); cover != "" {
return cover
}
}
}
return ""
}
func publishedCoverURL(value string) string {
value = strings.TrimSpace(html.UnescapeString(value))
return strings.ReplaceAll(value, " ", "%20")
}
// sites is the registry: one entry per Site, keyed by the stored site string.
// Adding a Site means adding an entry here and nowhere else — the dispatch
// functions above and the poller's route list are lookups into this map. An
// unknown site string resolves to the zero entry, which fails the existing
// not-fetchable and no-fetcher paths unchanged.
var sites = map[string]site{
"asura": {
Host: "asurascans.com",
LatestChapter: asuraLatestChapter,
Cover: ogImageCover,
},
"demonic": {
Host: "demonicscans.org",
LatestChapter: demonicLatestChapter,
Cover: ogImageCover,
},
"comix": {
Host: "comix.to",
LatestChapter: comixLatestChapter,
Cover: comixCoverEntry,
},
"kagane": {
Host: "kagane.to",
LatestChapter: kaganeLatestChapter,
Cover: kaganeCoverEntry,
Browser: &browserRead{
Read: kaganeRead,
Done: func(body string) bool { return body != "" },
// Never falls back: a plain fetch of a kagane page or cover would
// only ever retrieve a challenge page (verified 2026-08-03).
Fallback: false,
},
},
"novelfull": {
Host: "novelfull.com",
LatestChapter: novelfullLatestChapter,
Cover: novelfullCoverEntry,
Browser: &browserRead{
Read: novelfullRead,
// The interstitial has a DOM too, so "the payload arrived" has to
// exclude it explicitly.
Done: func(body string) bool { return body != "" && !isInterstitial(body) },
Fallback: true,
},
},
"lightnovelworld": {
Host: "lightnovelworld.net",
LatestChapter: lnwLatestChapter,
Cover: ogImageCover,
},
}
// browserBackedSites is derived from the registry: the Sites whose pages are
// read through the browser sidecar, which are also the ones granted the longer
// cooldown. Sorted so callers that range it (the due query, the browser
// fetcher's dispatch) see a stable order instead of map-iteration noise.
func browserBackedSites() []string {
out := make([]string, 0, len(sites))
for name, s := range sites {
if s.Browser != nil {
out = append(out, name)
}
}
sort.Strings(out)
return out
}
+247 -16
View File
@@ -1,6 +1,9 @@
package latest
import "testing"
import (
"strings"
"testing"
)
// Trimmed from https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af
// fetched 2026-07-26. The first anchor is the "First Chapter" shortcut: it is a
@@ -69,16 +72,230 @@ const novelfullSeriesFixture = `
<a href="/release-that-witch/chapter-9999.html">Chapter 9999</a>
`
// Trimmed from https://lightnovelworld.net/novel/a-will-eternal/ fetched
// 2026-08-05. Its chapter anchors are absolute and flat — /<slug>-chapter-<n>/
// at the site root, not under /novel/. The last anchor is another series'.
// Trimmed from
// https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/
// fetched 2026-08-11, whole document (~307 KB decoded; the wire body is ~32 KB
// zstd-compressed). This
// novel publishes its chapters under two Chapter Slugs: 1–99 at
// …-not-chapter-<n>/ and 100–423 at …-not-them-all-chapter-<n>/, and the page
// lists them newest-first, so the …-not-them-all anchors precede the …-not
// anchors. Every anchor through the comment-thread marker is verbatim page
// text (the site renders this novel's chapter titles as "[ ... words ]"). The
// comment block after the marker is the real wpdiscuz comment #wpd-comm-358_0
// from https://lightnovelworld.net/novel/the-sword-illuminates-the-great-wilderness/
// — the pinned page serves zero comments — with its share/link/vote/reply
// boilerplate trimmed. The comment's body carried no link, so the bare <a
// href> to https://lightnovelworld.net/overgeared-chapter-2059/ inside
// wpd-comment-text is the one composed element; that URL is a real chapter of
// a real different novel (overgeared; fetched, HTTP 200).
const lnwSeriesFixture = `
<a href="https://lightnovelworld.net/a-will-eternal-chapter-1/">Chapter 1</a>
<a href="https://lightnovelworld.net/a-will-eternal-chapter-1317/">Chapter 1317</a>
<a href="https://lightnovelworld.net/a-will-eternal-chapter-1298/">Chapter 1298</a>
<a href="https://lightnovelworld.net/overgeared-chapter-9999/">Chapter 9999</a>
<li data-ID="102741">
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-404/">
<div class="epl-num">Vol. 1 Ch. 404</div>
<div class="epl-title">[ ... words ]</div>
<div class="epl-date">April 12, 2026</div>
</a>
</li>
<li data-ID="102780">
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-423/">
<div class="epl-num">Vol. 1 Ch. 423</div>
<div class="epl-title">[ ... words ]</div>
<div class="epl-date">April 7, 2026</div>
</a>
</li>
<li data-ID="102527">
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-300/">
<div class="epl-num">Vol. 1 Ch. 300</div>
<div class="epl-title">[ ... words ]</div>
<div class="epl-date">April 4, 2026</div>
</a>
</li>
<li data-ID="102325">
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-200/">
<div class="epl-num">Vol. 1 Ch. 200</div>
<div class="epl-title">[ ... words ]</div>
<div class="epl-date">March 29, 2026</div>
</a>
</li>
<li data-ID="102121">
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-100/">
<div class="epl-num">Vol. 1 Ch. 100</div>
<div class="epl-title">[ ... words ]</div>
<div class="epl-date">March 22, 2026</div>
</a>
</li>
<li data-ID="26014">
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-chapter-99/">
<div class="epl-num">Vol. 1 Ch. 99</div>
<div class="epl-title">Chapter 99</div>
<div class="epl-date">November 5, 2025</div>
</a>
</li>
<li data-ID="25916">
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-chapter-50/">
<div class="epl-num">Vol. 1 Ch. 50</div>
<div class="epl-title">Chapter 50</div>
<div class="epl-date">October 29, 2025</div>
</a>
</li>
<li class='tseplsfrst' data-ID="25818">
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-chapter-1/">
<div class="epl-num">Vol. 1 Ch. 1</div>
<div class="epl-title">Chapter 01</div>
<div class="epl-date">October 11, 2025</div>
</a>
</li>
<div id="wpd-threads" class="wpd-thread-wrapper">
<div class="wpd-thread-list">
<div id='wpd-comm-358_0' class='comment byuser comment-author-jimbear even thread-even depth-1 wpd-comment wpd_comment_level-1'><div class="wpd-comment-wrap wpd-blog-user wpd-blog-subscriber">
<div class="wpd-comment-left ">
<div class="wpd-avatar ">
<img alt='hasbi asy' src='https://secure.gravatar.com/avatar/3c1792cb31cab842f90e7c463f0948e98536cf938aebbd2f678f937ca60fb799?s=64&#038;d=mm&#038;r=g' srcset='https://secure.gravatar.com/avatar/3c1792cb31cab842f90e7c463f0948e98536cf938aebbd2f678f937ca60fb799?s=128&#038;d=mm&#038;r=g 2x' class='avatar avatar-64 photo' height='64' width='64' decoding='async'/>
</div>
<div class="wpd-comment-label" wpd-tooltip="Member" wpd-tooltip-position="right">
<span>Member</span>
</div>
</div>
<div id="comment-358" class="wpd-comment-right">
<div class="wpd-comment-header">
<div class="wpd-comment-author ">
hasbi asy
</div>
<div class="wpd-comment-date" title="July 9, 2026 1:44 am">
<i class='far fa-clock' aria-hidden='true'></i>
1 month ago
</div>
</div>
<div class="wpd-comment-text">
<p>where&#8217;s everyone</p>
<a href="https://lightnovelworld.net/overgeared-chapter-2059/">https://lightnovelworld.net/overgeared-chapter-2059/</a>
</div>
</div>
</div>
<div id='wpdiscuz_form_anchor-358_0'></div>
</div>
</div>
`
// Trimmed from https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af
// (redirected to ...-00dcbf97) on 2026-08-10.
const asuraCoverFixture = `<meta property="og:image" content="https://cdn.asurascans.com/asura-images/covers/chronicles-of-the-demon-faction.d4dcb8.webp">`
// Trimmed from https://demonicscans.org/manga/Catastrophic-Necromancer on 2026-08-10.
// The source publishes the raw space in this URL.
const demonicCoverFixture = `<meta property="og:image" content="https://readermc.org/images/thumbnails/Catastrophic Necromancer.webp">`
// Trimmed from https://comix.to/title/n8we-dungeons-and-crayons on 2026-08-10.
// The state includes a recommended poster before the target detail object and
// nested IDs inside that object; no og:image is present.
const comixCoverFixture = `<script type="application/json" id="initial-data">{"queries":{"[\"manga\",\"recommended\",\"n8we\",1]":{"poster":{"medium":"https://static.comix.to/recommended@280.jpg","large":"https://static.comix.to/recommended.jpg"}},"[\"manga\",\"detail\",\"n8we\"]":{"chapters":[{"hid":"nested"}],"poster":{"medium":"https://static.comix.to/039d/i/1/34/6a6742bf15736@280.jpg","large":"https://static.comix.to/039d/i/1/34/6a6742bf15736.jpg"}}}}</script>`
// Trimmed from GET https://kagane.to/api/v2/series/019fe11a-8670-7cf3-8343-0b02057d3787 on 2026-08-10.
const kaganeCoverFixture = `{"series_covers":[{"cover_id":"019fe11a-84d1-714b-9cf4-2827f277f3c0","language":"en","volume_number":"1","chapter_number":null,"note":null,"image_id":"019fe11a-84c3-7fc3-a84b-88787374b617"}]}`
// Trimmed from https://novelfull.com/reverend-insanity.html on 2026-08-10.
const novelfullCoverFixture = `<meta name="image" content="https://novelfull.com/uploads/webp/novel/reverend-insanity-82661d911a.webp">`
// Trimmed from
// https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/
// on 2026-08-11.
const lnwCoverFixture = `<meta property="og:image" content="https://i1.wp.com/lightnovelworld.net/wp-content/uploads/2025/10/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all.jpg" />`
func TestCoverFrom(t *testing.T) {
const comixURL = "https://comix.to/title/n8we-dungeons-and-crayons"
tests := []struct {
name string
site string
seriesURL string
body string
wantOK bool
wantCover string
}{
{
name: "asura uses published metadata URL",
site: "asura", body: asuraCoverFixture, wantOK: true,
wantCover: "https://cdn.asurascans.com/asura-images/covers/chronicles-of-the-demon-faction.d4dcb8.webp",
},
{
name: "demonic escapes raw spaces",
site: "demonic", body: demonicCoverFixture, wantOK: true,
wantCover: "https://readermc.org/images/thumbnails/Catastrophic%20Necromancer.webp",
},
{
name: "comix takes target medium poster",
site: "comix", seriesURL: comixURL, body: comixCoverFixture, wantOK: true,
wantCover: "https://static.comix.to/039d/i/1/34/6a6742bf15736@280.jpg",
},
{
name: "kagane reads API cover image ID",
site: "kagane", body: kaganeCoverFixture, wantOK: true,
wantCover: "https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
},
{
name: "novelfull reads image metadata",
site: "novelfull", body: novelfullCoverFixture, wantOK: true,
wantCover: "https://novelfull.com/uploads/webp/novel/reverend-insanity-82661d911a.webp",
},
{
name: "lightnovelworld reads og image",
site: "lightnovelworld", body: lnwCoverFixture, wantOK: true,
wantCover: "https://i1.wp.com/lightnovelworld.net/wp-content/uploads/2025/10/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all.jpg",
},
{
name: "later metadata cover survives empty match",
site: "asura",
body: `<meta property="og:image" content="">` + asuraCoverFixture,
wantOK: true,
wantCover: "https://cdn.asurascans.com/asura-images/covers/chronicles-of-the-demon-faction.d4dcb8.webp",
},
{
name: "page without cover is empty",
site: "asura", body: `<meta property="og:title" content="No Cover">`,
},
{
name: "unknown site is empty",
site: "unknown", body: asuraCoverFixture,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
got, ok := coverFrom(tt.site, tt.seriesURL, tt.body)
if ok != tt.wantOK {
t.Fatalf("ok = %v, want %v (got %q)", ok, tt.wantOK, got)
}
if got != tt.wantCover {
t.Errorf("cover = %q, want %q", got, tt.wantCover)
}
})
}
}
func TestCoverFromChallenge(t *testing.T) {
tests := []struct {
site string
seriesURL string
}{
{"asura", "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af"},
{"demonic", "https://demonicscans.org/manga/Catastrophic-Necromancer"},
{"comix", "https://comix.to/title/n8we-dungeons-and-crayons"},
{"kagane", "https://kagane.to/series/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"},
{"novelfull", "https://novelfull.com/reverend-insanity.html"},
{"lightnovelworld", "https://lightnovelworld.net/novel/a-will-eternal/"},
}
for _, tt := range tests {
t.Run(tt.site, func(t *testing.T) {
if got, ok := coverFrom(tt.site, tt.seriesURL, challengeFixture); ok || got != "" {
t.Fatalf("cover = %q, ok = %v, want empty", got, ok)
}
})
}
}
func TestLatestChapterFrom(t *testing.T) {
const asuraURL = "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af"
const demonicURL = "https://demonicscans.org/manga/Catastrophic-Necromancer"
@@ -194,24 +411,38 @@ func TestLatestChapterFrom(t *testing.T) {
body: novelfullSeriesFixture,
wantOK: false,
},
// Stored before the slug split, so the address carries the ...-not
// Chapter Slug; the 100-423 block under the other slug must still win.
// The body is the chapter-list portion of lnwSeriesFixture with the
// comment block omitted; the marker is kept, because a body without it
// is skipped, not scanned.
{
name: "lightnovelworld takes the max and ignores another series",
name: "lightnovelworld max spans both chapter slugs",
site: "lightnovelworld",
seriesURL: "https://lightnovelworld.net/novel/a-will-eternal/",
body: lnwSeriesFixture,
wantOK: true, wantNum: 1317, wantLabel: "Chapter 1317",
seriesURL: "https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not/",
body: strings.SplitN(lnwSeriesFixture, lnwCommentMarker, 2)[0] + lnwCommentMarker,
wantOK: true, wantNum: 423, wantLabel: "Chapter 423",
},
{
name: "lightnovelworld tolerates a series url with no trailing slash",
name: "lightnovelworld comment anchor cannot set the latest chapter",
site: "lightnovelworld",
seriesURL: "https://lightnovelworld.net/novel/a-will-eternal",
seriesURL: "https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/",
body: lnwSeriesFixture,
wantOK: true, wantNum: 1317, wantLabel: "Chapter 1317",
wantOK: true, wantNum: 423, wantLabel: "Chapter 423",
},
// Marker removed from the fixture, comment block still present: a
// redesign must degrade into a skip, never into the comment's number.
{
name: "lightnovelworld body without the comment marker is skipped",
site: "lightnovelworld",
seriesURL: "https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/",
body: strings.ReplaceAll(lnwSeriesFixture, lnwCommentMarker, ""),
wantOK: false,
},
{
name: "lightnovelworld yields nothing on a challenge page",
site: "lightnovelworld",
seriesURL: "https://lightnovelworld.net/novel/a-will-eternal/",
seriesURL: "https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/",
body: challengeFixture,
wantOK: false,
},
+78 -12
View File
@@ -6,25 +6,30 @@ import (
"os"
"testing"
"time"
"bookmarkmanager/backend/internal/store"
)
// TestSmokeKaganeImage is the live proof that the cover proxy's fetch actually
// clears Cloudflare and returns image bytes. It needs a real headless Chrome
// with outbound network, so it runs only when SMOKE_BROWSER_WS_URL is set:
// TestSmokeKaganeImage is the live proof that the acquisition path's browser
// fetch actually clears Cloudflare and returns image bytes. It needs the real
// browser unit with outbound network, so it runs only when SMOKE_BROWSER_WS_URL
// is set:
//
// docker run --rm --shm-size=1gb -p 19222:9222 chromedp/headless-shell:stable
// SMOKE_BROWSER_WS_URL=ws://127.0.0.1:19222 go test -run TestSmokeKaganeImage ./internal/latest
// cd chrome && BROWSER_BIND_ADDR=127.0.0.1 docker compose up -d --build
// SMOKE_BROWSER_WS_URL=ws://127.0.0.1:9222 go test -run TestSmokeKaganeImage ./internal/latest
//
// Not chromedp/headless-shell: its challenge never clears (see chrome/Dockerfile),
// so a red run there proves nothing about kagane.
func TestSmokeKaganeImage(t *testing.T) {
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
if ws == "" {
t.Skip("SMOKE_BROWSER_WS_URL unset")
}
const imageID = "019fe11a-84c3-7fc3-a84b-88787374b617" // SP Baby's cover
const imageURL = "https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed" // SP Baby's cover
// The same URL through a plain client is what the web UI's <img> gets.
// The same URL through a plain client is what any other fetcher would get.
// Asserting on it keeps the test honest about why the browser is needed.
req, err := http.NewRequest(http.MethodGet,
"https://kagane.to/api/v2/image/"+imageID+"/compressed", nil)
req, err := http.NewRequest(http.MethodGet, imageURL, nil)
if err != nil {
t.Fatal(err)
}
@@ -43,7 +48,7 @@ func TestSmokeKaganeImage(t *testing.T) {
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
defer cancel()
body, contentType, err := f.Image(ctx, imageID)
body, contentType, err := f.Image(ctx, imageURL)
if err != nil {
t.Fatalf("Image: %v", err)
}
@@ -59,8 +64,10 @@ func TestSmokeKaganeImage(t *testing.T) {
}
t.Logf("fetched %d bytes of %s", len(body), contentType)
if _, _, err := f.Image(ctx, "not-a-uuid"); err == nil {
t.Fatal("Image accepted a non-uuid id")
// The browser module claims only the cover URL shape it can clear a
// challenge for; anything else must be refused before any navigation.
if _, _, err := f.Image(ctx, "https://cdn.example/cover.jpg"); err == nil {
t.Fatal("Image accepted a cover URL the browser module does not claim")
}
}
@@ -89,3 +96,62 @@ func TestSmokeKaganeGet(t *testing.T) {
t.Fatalf("status = %d, want 200 — the sidecar is not clearing the challenge", status)
}
}
// TestSmokeAcquireKaganeCover proves the #62 acquisition path end to end
// against the real browser: a kagane Series bookmarked at creation gets its
// Cover, bytes fetched through the sidecar into the content-addressed store.
// Same SMOKE_BROWSER_WS_URL gate as the tests above; a red run means the
// challenge is not clearing from this IP (a live fact to re-check), not
// necessarily a defect in the pipeline.
func TestSmokeAcquireKaganeCover(t *testing.T) {
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
if ws == "" {
t.Skip("SMOKE_BROWSER_WS_URL unset")
}
const (
seriesID = "019fe11a-8670-7cf3-8343-0b02057d3787"
coverURL = "https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed"
)
s, _ := newTestStore(t)
bf, err := NewBrowserFetcher(ws)
if err != nil {
t.Fatalf("NewBrowserFetcher: %v", err)
}
defer bf.Close()
tlsF, err := NewTLSFetcher()
if err != nil {
t.Fatalf("NewTLSFetcher: %v", err)
}
acq := &Acquirer{
Store: s, Fetch: tlsF, BrowserFetch: bf,
BrowserCoverFetch: bf, Covers: NewCoverFetcher(),
}
s.OnSeriesCreated = acq.Acquire
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: "kagane:" + seriesID, Site: "kagane", SeriesID: seriesID,
Title: "smoke", SeriesURL: "https://kagane.to/series/" + seriesID, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("Upsert: %v", err)
}
acq.Wait()
got, found, err := s.Get(s.OwnerID(), "kagane:"+seriesID)
if err != nil || !found {
t.Fatalf("Get: %v found=%v", err, found)
}
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(coverURL); got.Cover != want {
t.Fatalf("Cover = %q, want %q — the acquire path did not store the browser-fetched bytes", got.Cover, want)
}
body, contentType, ok, err := s.CoverByAddress(store.CoverAddress(coverURL))
if err != nil || !ok {
t.Fatalf("CoverByAddress: %v found=%v", err, ok)
}
if len(body) < 1000 {
t.Fatalf("stored cover is %d bytes, want a real image", len(body))
}
if contentType != "image/webp" {
t.Fatalf("content type = %q, want image/webp", contentType)
}
t.Logf("stored %d bytes of %s", len(body), contentType)
}
+101
View File
@@ -0,0 +1,101 @@
package latest
import (
"context"
"fmt"
"net/http"
"os"
"strings"
"testing"
"time"
)
// lnwSeriesPageFloor is the smallest body that can still be a whole
// lightnovelworld Series page. Whole pages measured 685 KB..1.18 MB on
// 2026-08-11 and carry the marker at ~94% of the document, so a body under
// 100 KB is a challenge, a notice, or a truncated read — asserting on it
// would report the marker missing when it was never fetched.
const lnwSeriesPageFloor = 100 << 10
// TestSmokeLnwCommentBoundary is the live proof that the comment-thread marker
// the lightnovelworld chapter scan truncates at (lnwCommentMarker,
// "wpd-threads") still holds on the Site. The scan depends on it: when the
// marker vanishes every Series is skipped and logged — correct, but silent
// until a Reader notices their Latest Chapter has stopped moving. It runs only
// when SMOKE_LNW_SERIES_URL is set — the URL of the live Series page to check.
// The immortality-simulator page measured 2026-08-11
// (docs/research/lightnovelworld-chapter-vs-series-slug.md) is the default to
// point it at:
//
// SMOKE_LNW_SERIES_URL=https://lightnovelworld.net/novel/immortality-simulator/ go test -v -run TestSmokeLnwCommentBoundary ./internal/latest
//
// A red run means the Site's markup has moved — the marker is gone, occurs
// more than once, or no longer follows the last chapter anchor — and the scan
// in sites.go is now skipping this Site. Revisit sites.go before anything
// else; the test is not flaky. A Cloudflare challenge or a non-200 is
// distinguished from a marker failure by the "not a marker failure" messages
// below, which carry the observed status and body length.
func TestSmokeLnwCommentBoundary(t *testing.T) {
seriesURL := os.Getenv("SMOKE_LNW_SERIES_URL")
if seriesURL == "" {
t.Skip("SMOKE_LNW_SERIES_URL unset")
}
if !fetchableSeriesURL("lightnovelworld", seriesURL) {
t.Fatalf("%q is not a fetchable lightnovelworld series URL", seriesURL)
}
f, err := NewTLSFetcher()
if err != nil {
t.Fatalf("NewTLSFetcher: %v", err)
}
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
body, status, err := f.Get(ctx, seriesURL)
if err != nil {
t.Fatalf("Get: %v", err)
}
if status != http.StatusOK {
t.Fatalf("status = %d, body %d bytes — not a marker failure; the Site did not answer this IP with a Series page", status, len(body))
}
if len(body) < lnwSeriesPageFloor {
t.Fatalf("body %d bytes — not a whole Series page (measured 685 KB..1.18 MB); not a marker failure, likely a Cloudflare challenge or a non-Series response", len(body))
}
if !lnwChapterRe.MatchString(body) {
t.Fatalf("no chapter anchor in %d bytes — not a lightnovelworld Series page; not a marker failure, likely a Cloudflare challenge or a non-Series response", len(body))
}
markerIdx := strings.Index(body, lnwCommentMarker)
if failures := checkLnwCommentBoundary(body); len(failures) > 0 {
t.Fatalf("%s (body %d bytes)", strings.Join(failures, "; "), len(body))
}
t.Logf("ok: %q once at byte %d, body %d bytes", lnwCommentMarker, markerIdx, len(body))
}
// checkLnwCommentBoundary verifies the three marker assertions against a
// fetched Series body: the marker occurs exactly once, every chapter anchor
// precedes it, and at least one anchor precedes it at all. It returns one
// human-readable failure per broken assertion — with observed offsets and body
// length — and empty when the page is healthy.
func checkLnwCommentBoundary(body string) []string {
markerIdx := strings.Index(body, lnwCommentMarker)
switch n := strings.Count(body, lnwCommentMarker); {
case n == 0:
return []string{fmt.Sprintf("%q occurs 0 times in %d bytes, want exactly 1", lnwCommentMarker, len(body))}
case n != 1:
return []string{fmt.Sprintf("%q occurs %d times in %d bytes (first at byte %d), want exactly 1", lnwCommentMarker, n, len(body), markerIdx)}
}
lastAnchor, anchorsBefore := -1, 0
for _, m := range lnwChapterRe.FindAllStringIndex(body, -1) {
if m[0] < markerIdx {
anchorsBefore++
}
lastAnchor = m[0]
}
var failures []string
if lastAnchor >= markerIdx {
failures = append(failures, fmt.Sprintf("last chapter anchor at byte %d does not precede the marker at byte %d", lastAnchor, markerIdx))
}
if anchorsBefore == 0 {
failures = append(failures, fmt.Sprintf("no chapter anchor before the marker at byte %d — the truncated prefix the scan sees yields nothing", markerIdx))
}
return failures
}
@@ -0,0 +1,8 @@
-- Kagane cover bytes belong in their own table so image blobs never enter the
-- series queries that drive the latest-chapter poller.
CREATE TABLE covers (
image_id text PRIMARY KEY,
body bytea NOT NULL,
content_type text NOT NULL,
fetched_at timestamptz NOT NULL DEFAULT now()
);
@@ -0,0 +1,9 @@
-- Cover bytes move out of Postgres. Existing rows are intentionally dropped:
-- the old kagane path already refetches missing Covers on demand.
DROP TABLE covers;
CREATE TABLE covers (
address text PRIMARY KEY,
path text NOT NULL,
content_type text NOT NULL
);
@@ -0,0 +1,8 @@
-- The Cover splits into two facts. `cover` keeps the third-party address the
-- bytes come from, which is what the acquisition path refetches and dedupes
-- on; `cover_address` is the content address of the bytes once they are
-- actually stored, and is what the wire's absolute URL is built from.
--
-- Empty `cover_address` therefore means "no Cover yet" rather than "a Cover
-- that 404s", which is the distinction the API and the UI both depend on.
ALTER TABLE series ADD COLUMN cover_address text NOT NULL DEFAULT '';
+264 -51
View File
@@ -1,17 +1,22 @@
package store
import (
"crypto/sha256"
"database/sql"
"embed"
"encoding/hex"
"errors"
"fmt"
"io/fs"
"os"
"path"
"path/filepath"
"regexp"
"slices"
"strconv"
"strings"
"github.com/jackc/pgx/v5/pgtype"
_ "github.com/jackc/pgx/v5/stdlib"
)
@@ -25,12 +30,16 @@ import (
// between readers: progress, favourite, lifecycle bucket, updated_at. The wire
// format stays flat regardless — see ADR-0004.
type Bookmark struct {
Key string `json:"key"`
Site string `json:"site"`
SeriesID string `json:"series_id"`
Title string `json:"title"`
SeriesURL string `json:"series_url"`
Cover string `json:"cover"`
Key string `json:"key"`
Site string `json:"site"`
SeriesID string `json:"series_id"`
Title string `json:"title"`
SeriesURL string `json:"series_url"`
// Cover is the wire value: an absolute URL on this deployment's own
// origin once the bytes exist, and "" until they do — never a third-party
// address and never an address that 404s (ADR-0007). A client may still
// send this field and it is discarded on the way in; see Upsert.
Cover string `json:"cover"`
LastChapter string `json:"last_chapter"`
LastChapterNum float64 `json:"last_chapter_num"`
LastChapterURL string `json:"last_chapter_url"`
@@ -57,11 +66,16 @@ type Bookmark struct {
// bookmark's own fields. Never serialized: the wire format is the flat
// Bookmark (ADR-0004).
type Series struct {
Site string
SeriesID string
Title string
SeriesURL string
Site string
SeriesID string
Title string
SeriesURL string
// Cover is the third-party source address the bytes come from, and
// CoverAddress the content address they are stored under. A blank
// CoverAddress is what "no Cover yet" means: the poll fills it and never
// replaces a filled one (ADR-0007).
Cover string
CoverAddress string
Kind string
LatestChapter string
LatestChapterNum *float64 // nil until first captured
@@ -137,20 +151,19 @@ func (b Bookmark) Initial() string {
return "?"
}
// kaganeCoverRe matches the cover URL kagane's og:image carries, which is what
// the userscript stores for that site.
var kaganeCoverRe = regexp.MustCompile(`^https://kagane\.to/api/v2/image/([0-9a-f-]{36})/compressed$`)
// CoverURL is the src the web UI puts in an <img>. For every site but kagane
// that is Cover as stored. kagane serves its images behind a Cloudflare
// challenge *and* with `cross-origin-resource-policy: same-origin`, so no page
// on another origin can load one however it asks (verified 2026-08-08); those
// go through the backend's own proxy instead.
func (b Bookmark) CoverURL() string {
if m := kaganeCoverRe.FindStringSubmatch(b.Cover); m != nil {
return "/img/kagane/" + m[1]
// CoverContentType canonicalises a fetched response's media type and reports
// whether the bytes are safe to store and serve. comix answers "image/jpg",
// which no standard lists but browsers accept; it is stored as the real name
// rather than passed through, so one image never lands under two spellings.
func CoverContentType(contentType string) (string, bool) {
switch contentType {
case "image/jpg":
return "image/jpeg", true
case "image/webp", "image/jpeg", "image/png", "image/avif", "image/gif":
return contentType, true
default:
return "", false
}
return b.Cover
}
// Library buckets. A bookmark is in exactly one. This cannot be derived from
@@ -175,14 +188,14 @@ var migrations embed.FS
// compile-time constant; every request value is bound as a parameter. The
// series-owned fields are joined in from the series table, in scanBookmark
// order, so the flat Bookmark reads back whole despite the split (ADR-0004).
const bookmarkColumns = `b.site, b.series_id, s.title, s.series_url, s.cover,
const bookmarkColumns = `b.site, b.series_id, s.title, s.series_url, s.cover_address,
b.last_chapter, b.last_chapter_num, b.last_chapter_url,
b.favorite, s.latest_chapter, s.latest_chapter_num, b.updated_at, b.status, s.kind`
// seriesColumns is the series row in scanSeries order, used by the poller's
// due query. latest_checked_at lives only on series — see MarkLatestChecked
// for why it stays off every client-visible write.
const seriesColumns = `s.site, s.series_id, s.title, s.series_url, s.cover,
const seriesColumns = `s.site, s.series_id, s.title, s.series_url, s.cover, s.cover_address,
s.kind, s.latest_chapter, s.latest_chapter_num, s.latest_checked_at`
// Owner is the person running the service: the first Reader, seeded at startup
@@ -203,7 +216,18 @@ type Store struct {
// ownerID is the seeded owner Reader (issue #22) — the only Reader with
// administrative reach (revoking another Reader's sessions). Every store
// method takes a reader id explicitly, so ownership is never implicit.
ownerID int64
ownerID int64
coverDir string
// coverBaseURL is this deployment's public origin. Cover addresses are
// absolute because the userscript renders them on third-party origins,
// where a relative path would resolve against the Site (ADR-0007).
coverBaseURL string
// OnSeriesCreated fires once, after commit, for a Series no Reader had
// bookmarked before. It is how creation-time Cover and Latest Chapter
// acquisition is triggered without the write waiting on a third-party
// Site; nil disables it, which is what every test that does not care
// about acquisition leaves it as.
OnSeriesCreated func(Series)
}
// OwnerID returns the seeded owner Reader's id: the administrator, and the
@@ -338,8 +362,31 @@ const allMigrations = 0
// Open connects to Postgres at url — a libpq connection URL such as
// "postgres://user:pass@host:5432/bookmarks?sslmode=disable" — brings its
// schema up to date, and seeds the owner Reader.
func Open(url string, owner Owner) (*Store, error) {
// schema up to date, seeds the owner Reader, and prepares cover storage.
func Open(url string, owner Owner, coverDir, coverBaseURL string) (*Store, error) {
if strings.TrimSpace(coverDir) == "" {
return nil, errors.New("cover directory is required")
}
// Every wire Cover is this string with a path glued on, rendered by a
// userscript on a Site's own origin: anything but an absolute origin
// produces addresses no client can load, silently (ADR-0007).
base := strings.TrimRight(coverBaseURL, "/")
if host, ok := strings.CutPrefix(base, "https://"); !ok || host == "" {
if host, ok := strings.CutPrefix(base, "http://"); !ok || host == "" {
return nil, fmt.Errorf("cover base URL %q is not an absolute http(s) origin", coverBaseURL)
}
}
if err := os.MkdirAll(coverDir, 0o755); err != nil {
return nil, fmt.Errorf("create cover directory: %w", err)
}
info, err := os.Stat(coverDir)
if err != nil {
return nil, fmt.Errorf("stat cover directory: %w", err)
}
if !info.IsDir() {
return nil, fmt.Errorf("cover directory %q is not a directory", coverDir)
}
db, err := sql.Open("pgx", url)
if err != nil {
return nil, fmt.Errorf("open postgres: %w", err)
@@ -374,7 +421,9 @@ func Open(url string, owner Owner) (*Store, error) {
db.Close()
return nil, fmt.Errorf("resolve owner: %w", err)
}
return &Store{db: db, ownerID: ownerID}, nil
return &Store{
db: db, ownerID: ownerID, coverDir: coverDir, coverBaseURL: base,
}, nil
}
// seedOwner makes sure the configured owner exists as exactly one readers row.
@@ -477,18 +526,20 @@ func applyMigration(db *sql.DB, version int64, body string) error {
// scanBookmark reads one row in bookmarkColumns order. Every column is NOT
// NULL except latest_chapter_num, where NULL means "never captured" — a
// distinct state from chapter zero, and the reason for the pointer.
func scanBookmark(scan func(...any) error) (Bookmark, error) {
func (s *Store) scanBookmark(scan func(...any) error) (Bookmark, error) {
var (
b Bookmark
coverAddress string
latestChapterNum sql.NullFloat64
)
if err := scan(
&b.Site, &b.SeriesID, &b.Title, &b.SeriesURL, &b.Cover,
&b.Site, &b.SeriesID, &b.Title, &b.SeriesURL, &coverAddress,
&b.LastChapter, &b.LastChapterNum, &b.LastChapterURL,
&b.Favorite, &b.LatestChapter, &latestChapterNum, &b.UpdatedAt, &b.Status, &b.Kind,
); err != nil {
return Bookmark{}, err
}
b.Cover = s.CoverWireURL(coverAddress)
if latestChapterNum.Valid {
b.LatestChapterNum = &latestChapterNum.Float64
}
@@ -513,7 +564,7 @@ func scanSeries(scan func(...any) error) (Series, error) {
latestChapterNum sql.NullFloat64
)
if err := scan(
&sr.Site, &sr.SeriesID, &sr.Title, &sr.SeriesURL, &sr.Cover,
&sr.Site, &sr.SeriesID, &sr.Title, &sr.SeriesURL, &sr.Cover, &sr.CoverAddress,
&sr.Kind, &sr.LatestChapter, &latestChapterNum, &sr.LatestCheckedAt,
&sr.readerCount,
); err != nil {
@@ -528,6 +579,146 @@ func scanSeries(scan func(...any) error) (Series, error) {
// Close releases the underlying database handle.
func (s *Store) Close() error { return s.db.Close() }
func coverSourceAddress(sourceURL string) string {
sum := sha256.Sum256([]byte(sourceURL))
return hex.EncodeToString(sum[:])
}
func coverRelativePath(address string) string {
return address[:2] + "/" + address[2:4] + "/" + address
}
func (s *Store) getCover(sourceURL string) ([]byte, string, bool, error) {
return s.getCoverByAddress(coverSourceAddress(sourceURL))
}
func (s *Store) getCoverByAddress(address string) ([]byte, string, bool, error) {
var relativePath, contentType string
err := s.db.QueryRow(
`SELECT path, content_type FROM covers WHERE address = $1`, address,
).Scan(&relativePath, &contentType)
if errors.Is(err, sql.ErrNoRows) {
return nil, "", false, nil
}
if err != nil {
return nil, "", false, fmt.Errorf("get cover %q: %w", address, err)
}
expectedPath := coverRelativePath(address)
if relativePath != expectedPath {
return nil, "", false, fmt.Errorf("cover %q has unexpected path %q", address, relativePath)
}
body, err := os.ReadFile(filepath.Join(s.coverDir, filepath.FromSlash(relativePath)))
if errors.Is(err, fs.ErrNotExist) {
return nil, "", false, nil
}
if err != nil {
return nil, "", false, fmt.Errorf("read cover %q: %w", address, err)
}
return body, contentType, true, nil
}
func (s *Store) putCover(sourceURL string, body []byte, contentType string) error {
stored, ok := CoverContentType(contentType)
if !ok {
return fmt.Errorf("put cover %q: unsupported content type %q", sourceURL, contentType)
}
contentType = stored
address := coverSourceAddress(sourceURL)
relativePath := coverRelativePath(address)
coverPath := filepath.Join(s.coverDir, filepath.FromSlash(relativePath))
if err := os.MkdirAll(filepath.Dir(coverPath), 0o755); err != nil {
return fmt.Errorf("create cover shard: %w", err)
}
tmp, err := os.CreateTemp(filepath.Dir(coverPath), ".cover-*")
if err != nil {
return fmt.Errorf("create cover temp file: %w", err)
}
tmpName := tmp.Name()
defer os.Remove(tmpName)
if _, err := tmp.Write(body); err != nil {
tmp.Close()
return fmt.Errorf("write cover temp file: %w", err)
}
if err := tmp.Sync(); err != nil {
tmp.Close()
return fmt.Errorf("sync cover temp file: %w", err)
}
if err := tmp.Close(); err != nil {
return fmt.Errorf("close cover temp file: %w", err)
}
if err := os.Link(tmpName, coverPath); err != nil && !errors.Is(err, fs.ErrExist) {
return fmt.Errorf("install cover file: %w", err)
}
if _, err := s.db.Exec(`
INSERT INTO covers (address, path, content_type)
VALUES ($1, $2, $3)
ON CONFLICT (address) DO NOTHING`, address, relativePath, contentType); err != nil {
return fmt.Errorf("record cover %q: %w", address, err)
}
return nil
}
// GetCover returns the immutable object addressed by its source URL. Missing
// files are reported with ok=false so callers can retry acquisition later.
func (s *Store) GetCover(sourceURL string) ([]byte, string, bool, error) {
return s.getCover(sourceURL)
}
// PutCover persists bytes under the source URL's content address. A later
// write for the same URL cannot replace the immutable object.
func (s *Store) PutCover(sourceURL string, body []byte, contentType string) error {
return s.putCover(sourceURL, body, contentType)
}
// CoverAddress is the content address bytes fetched from sourceURL are stored
// under. It is a pure function of the URL, so the acquisition path can name a
// Cover before it has the bytes.
func CoverAddress(sourceURL string) string { return coverSourceAddress(sourceURL) }
// coverAddressRe is the shape of a stored address: the hex SHA-256 of a source
// URL. Request paths reach CoverByAddress, so the shape is checked before the
// value is ever turned into a filesystem path.
var coverAddressRe = regexp.MustCompile(`^[0-9a-f]{64}$`)
// CoverByAddress returns the immutable object at one content address. An
// address that is not a stored one - malformed, unknown, or recorded but with
// its file gone - is reported with ok=false rather than as an error.
func (s *Store) CoverByAddress(address string) ([]byte, string, bool, error) {
if !coverAddressRe.MatchString(address) {
return nil, "", false, nil
}
return s.getCoverByAddress(address)
}
// CoverWireURL is the absolute URL a client renders for a stored Cover, and ""
// for a Series that has none yet. A blank is a real state, not a placeholder
// address: it is what tells both clients to draw their own fallback instead of
// requesting bytes that do not exist (ADR-0007).
func (s *Store) CoverWireURL(address string) string {
if address == "" {
return ""
}
return s.coverBaseURL + "/covers/" + address
}
// SetSeriesCover stores the bytes and points the Series at them, but only
// while the Series has no Cover: acquisition at creation and the poll both
// call this, and whichever arrives second must not overwrite the first. The
// bytes themselves are content-addressed and immutable, so storing them twice
// is free.
func (s *Store) SetSeriesCover(site, seriesID, sourceURL string, body []byte, contentType string) error {
if err := s.putCover(sourceURL, body, contentType); err != nil {
return err
}
if _, err := s.db.Exec(`
UPDATE series SET cover = $3, cover_address = $4
WHERE site = $1 AND series_id = $2 AND cover_address = ''`,
site, seriesID, sourceURL, coverSourceAddress(sourceURL)); err != nil {
return fmt.Errorf("set cover for %q: %w", site+":"+seriesID, err)
}
return nil
}
// List returns every bookmark of one reader, newest activity first.
// Series-owned fields are joined in, so each Bookmark reads back whole and
// flat (ADR-0004).
@@ -544,7 +735,7 @@ func (s *Store) List(readerID int64) ([]Bookmark, error) {
out := []Bookmark{}
for rows.Next() {
b, err := scanBookmark(rows.Scan)
b, err := s.scanBookmark(rows.Scan)
if err != nil {
return nil, fmt.Errorf("scan bookmark: %w", err)
}
@@ -561,7 +752,7 @@ func (s *Store) Get(readerID int64, key string) (Bookmark, bool, error) {
if !ok {
return Bookmark{}, false, nil
}
b, err := scanBookmark(s.db.QueryRow(
b, err := s.scanBookmark(s.db.QueryRow(
`SELECT `+bookmarkColumns+` FROM bookmarks b
JOIN series s ON s.site = b.site AND s.series_id = b.series_id
WHERE b.reader_id = $1 AND b.site = $2 AND b.series_id = $3`,
@@ -618,18 +809,28 @@ func (s *Store) Upsert(readerID int64, b Bookmark) (Bookmark, error) {
// The ::text casts are load-bearing: inside COALESCE/NULLIF there is no
// target column to infer the parameter type from, and Postgres rejects the
// statement rather than guessing.
if _, err := tx.Exec(`
INSERT INTO series (site, series_id, title, series_url, cover, kind,
//
// The cover columns are absent on purpose: the Cover is acquired
// server-side (ADR-0007), so a client-supplied one is not written even
// when the row is brand new.
//
// xmax is zero only on a row this statement inserted, which is how a
// Series nobody had bookmarked before is told apart from one that already
// existed — DO UPDATE returns a row either way.
var created bool
if err := tx.QueryRow(`
INSERT INTO series (site, series_id, title, series_url, kind,
latest_chapter, latest_chapter_num)
VALUES ($1, $2, $3, $4, $5,
COALESCE(NULLIF($6::text, ''), (SELECT kind FROM series WHERE site = $1 AND series_id = $2), 'manga'),
$7, $8)
VALUES ($1, $2, $3, $4,
COALESCE(NULLIF($5::text, ''), (SELECT kind FROM series WHERE site = $1 AND series_id = $2), 'manga'),
$6, $7)
ON CONFLICT (site, series_id) DO UPDATE SET
kind=excluded.kind,
latest_chapter=excluded.latest_chapter,
latest_chapter_num=excluded.latest_chapter_num`,
b.Site, b.SeriesID, b.Title, b.SeriesURL, b.Cover, b.Kind,
b.LatestChapter, latestNum); err != nil {
latest_chapter_num=excluded.latest_chapter_num
RETURNING xmax = 0`,
b.Site, b.SeriesID, b.Title, b.SeriesURL, b.Kind,
b.LatestChapter, latestNum).Scan(&created); err != nil {
return Bookmark{}, fmt.Errorf("upsert series for %q: %w", b.Key, err)
}
@@ -659,7 +860,7 @@ func (s *Store) Upsert(readerID int64, b Bookmark) (Bookmark, error) {
return Bookmark{}, fmt.Errorf("upsert %q: %w", b.Key, err)
}
stored, err := scanBookmark(tx.QueryRow(
stored, err := s.scanBookmark(tx.QueryRow(
`SELECT `+bookmarkColumns+` FROM bookmarks b
JOIN series s ON s.site = b.site AND s.series_id = b.series_id
WHERE b.reader_id = $1 AND b.site = $2 AND b.series_id = $3`,
@@ -670,6 +871,14 @@ func (s *Store) Upsert(readerID int64, b Bookmark) (Bookmark, error) {
if err := tx.Commit(); err != nil {
return Bookmark{}, fmt.Errorf("commit %q: %w", b.Key, err)
}
// After commit, never inside the transaction: the hook reaches a
// third-party Site, and the Reader's write must not wait on it.
if created && s.OnSeriesCreated != nil {
s.OnSeriesCreated(Series{
Site: b.Site, SeriesID: b.SeriesID, Title: stored.Title,
SeriesURL: stored.SeriesURL, Kind: stored.Kind,
})
}
return stored, nil
}
@@ -689,16 +898,17 @@ func (s *Store) Delete(readerID int64, key string) error {
}
// DueForLatestCheck returns series whose server-side latest-chapter check has
// aged past cutoffMs, ordered by how many bookmarks reference them (descending)
// then least-recently-checked first, at most limit of them.
// aged past the appropriate cutoff, ordered by how many bookmarks reference
// them (descending) then least-recently-checked first, at most limit of them.
// Browser-backed sites use browserCutoffMs; every other site uses cutoffMs.
//
// The reader_count ordering is the point of the split (ADR-0003): a series
// shared by several readers is fetched once per due cycle, and the popular
// ones stay freshest while the long tail absorbs any shortfall. Within one
// reader count, oldest-first keeps the poll fair when the backlog outgrows its
// reader count, oldest-first keeps the poll fair when the backlog outgrows
// throughput: the most neglected series is always next, so a large collection
// refreshes uniformly slower rather than leaving a tail that never refreshes at
// all. The userscript sorts its own queue the same way (L453).
// refreshes uniformly slower rather than leaving a tail that never refreshes
// at all. The userscript sorts its own queue the same way (L453).
//
// Series with no series_url are skipped — there is nothing to fetch, which is
// the same filter the userscript applies at L452. Series whose only bookmarks
@@ -706,17 +916,20 @@ func (s *Store) Delete(readerID int64, key string) error {
// burns requests. Archived bookmarks still count — knowing what a shelved
// series is up to is the whole reason for archiving instead of deleting.
// A series with no bookmarks at all never appears: the join excludes it.
func (s *Store) DueForLatestCheck(cutoffMs int64, limit int) ([]Series, error) {
func (s *Store) DueForLatestCheck(cutoffMs, browserCutoffMs int64, browserSites []string, limit int) ([]Series, error) {
rows, err := s.db.Query(`SELECT `+seriesColumns+`, COUNT(*) AS reader_count
FROM series s
JOIN bookmarks b ON b.site = s.site AND b.series_id = s.series_id
WHERE s.series_url <> ''
AND s.latest_checked_at <= $1
AND s.latest_checked_at <= CASE
WHEN s.site = ANY($3::text[]) THEN $2::bigint
ELSE $1::bigint
END
GROUP BY s.site, s.series_id, s.title, s.series_url, s.cover,
s.kind, s.latest_chapter, s.latest_chapter_num, s.latest_checked_at
HAVING COUNT(*) FILTER (WHERE b.status <> 'finished') > 0
ORDER BY reader_count DESC, s.latest_checked_at ASC
LIMIT $2`, cutoffMs, limit)
LIMIT $4`, cutoffMs, browserCutoffMs, pgtype.FlatArray[string](browserSites), limit)
if err != nil {
return nil, fmt.Errorf("query due series: %w", err)
}
+298 -64
View File
@@ -4,7 +4,9 @@ import (
"bytes"
"crypto/sha256"
"database/sql"
"encoding/hex"
"os"
"path/filepath"
"strconv"
"strings"
"testing"
@@ -19,9 +21,12 @@ func TestMain(m *testing.M) { os.Exit(pgtest.Main(m)) }
// reader register one (see secondReader).
var testOwner = Owner{DiscordID: "test-owner", TokenHash: sha256.Sum256([]byte("owner-token-hash"))}
// testCoverBaseURL is the public origin every stored cover URL is built from.
const testCoverBaseURL = "https://bookmarks.test"
func newTestStore(t *testing.T) *Store {
t.Helper()
store, err := Open(pgtest.URL(t), testOwner)
store, err := Open(pgtest.URL(t), testOwner, t.TempDir(), testCoverBaseURL)
if err != nil {
t.Fatalf("Open: %v", err)
}
@@ -46,7 +51,8 @@ func secondReader(t *testing.T, s *Store) int64 {
// error, and must leave the rows alone.
func TestOpenIsIdempotent(t *testing.T) {
url := pgtest.URL(t)
first, err := Open(url, testOwner)
coverDir := t.TempDir()
first, err := Open(url, testOwner, coverDir, testCoverBaseURL)
if err != nil {
t.Fatalf("Open: %v", err)
}
@@ -57,7 +63,7 @@ func TestOpenIsIdempotent(t *testing.T) {
}
first.Close()
second, err := Open(url, testOwner)
second, err := Open(url, testOwner, coverDir, testCoverBaseURL)
if err != nil {
t.Fatalf("reopen: %v", err)
}
@@ -109,7 +115,8 @@ func TestReaderTokenInfo(t *testing.T) {
// epoch-0 hash only while the row has never been rotated.
func TestRotateTokenInvalidatesOldAndSurvivesRestart(t *testing.T) {
url := pgtest.URL(t)
store, err := Open(url, testOwner)
coverDir := t.TempDir()
store, err := Open(url, testOwner, coverDir, testCoverBaseURL)
if err != nil {
t.Fatalf("Open: %v", err)
}
@@ -141,7 +148,7 @@ func TestRotateTokenInvalidatesOldAndSurvivesRestart(t *testing.T) {
}
store.Close()
reopened, err := Open(url, testOwner)
reopened, err := Open(url, testOwner, coverDir, testCoverBaseURL)
if err != nil {
t.Fatalf("reopen: %v", err)
}
@@ -291,7 +298,7 @@ func TestDueForLatestCheck(t *testing.T) {
s := newTestStore(t)
seedForCheck(t, s, "asura:x", tt.seriesURL, tt.checkedAt)
due, err := s.DueForLatestCheck(now-hour, 10)
due, err := s.DueForLatestCheck(now-hour, now-hour, nil, 10)
if err != nil {
t.Fatalf("DueForLatestCheck: %v", err)
}
@@ -309,7 +316,7 @@ func TestDueForLatestCheckOldestFirstAndLimited(t *testing.T) {
seedForCheck(t, s, "asura:b", "https://asurascans.com/comics/b", 200)
seedForCheck(t, s, "asura:a", "https://asurascans.com/comics/a", 100)
due, err := s.DueForLatestCheck(1000, 2)
due, err := s.DueForLatestCheck(1000, 1000, nil, 2)
if err != nil {
t.Fatalf("DueForLatestCheck: %v", err)
}
@@ -489,7 +496,7 @@ func TestDueForLatestCheckSkipsFinishedKeepsArchived(t *testing.T) {
}
}
due, err := store.DueForLatestCheck(time.Now().UnixMilli(), 10)
due, err := store.DueForLatestCheck(time.Now().UnixMilli(), time.Now().UnixMilli(), nil, 10)
if err != nil {
t.Fatalf("DueForLatestCheck: %v", err)
}
@@ -543,38 +550,6 @@ func TestDisplayChapter(t *testing.T) {
}
}
func TestCoverURL(t *testing.T) {
cases := []struct {
name string
cover string
want string
}{
{
"kagane routes through the proxy",
"https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
"/img/kagane/019fe11a-84c3-7fc3-a84b-88787374b617",
},
{
"another site is served as stored",
"https://gg.asuracomic.net/storage/media/1/conversions/cover.webp",
"https://gg.asuracomic.net/storage/media/1/conversions/cover.webp",
},
{
"a lookalike host is not rewritten",
"https://evil.example/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
"https://evil.example/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
},
{"no cover stays empty", "", ""},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
if got := (Bookmark{Cover: tc.cover}).CoverURL(); got != tc.want {
t.Errorf("CoverURL() = %q, want %q", got, tc.want)
}
})
}
}
func TestUpsertKindDefaultsToManga(t *testing.T) {
store := newTestStore(t)
got, err := store.Upsert(store.OwnerID(), Bookmark{
@@ -668,7 +643,7 @@ func TestMigration0002BackfillsExistingBookmarks(t *testing.T) {
// Bring it current through the production path: Open runs the schema to
// 0003, seeds the owner, then applies 0004 which attaches this row. 0002
// must have backfilled the series row, not lost data.
st, err := Open(url, testOwner)
st, err := Open(url, testOwner, t.TempDir(), testCoverBaseURL)
if err != nil {
t.Fatalf("Open after migrate: %v", err)
}
@@ -697,6 +672,48 @@ func TestMigration0002BackfillsExistingBookmarks(t *testing.T) {
}
}
func TestMigration0008DropsLegacyCoverRows(t *testing.T) {
db, err := sql.Open("pgx", pgtest.URL(t))
if err != nil {
t.Fatalf("open: %v", err)
}
defer db.Close()
if err := migrate(db, readersMigration); err != nil {
t.Fatalf("migrate readers: %v", err)
}
if err := seedOwner(db, testOwner); err != nil {
t.Fatalf("seed owner: %v", err)
}
if err := migrate(db, 7); err != nil {
t.Fatalf("migrate legacy covers: %v", err)
}
if _, err := db.Exec(`
INSERT INTO covers (image_id, body, content_type)
VALUES ('legacy-image', 'legacy-bytes', 'image/jpeg')`); err != nil {
t.Fatalf("seed legacy cover: %v", err)
}
if err := migrate(db, 0); err != nil {
t.Fatalf("migrate filesystem covers: %v", err)
}
var count int
if err := db.QueryRow(`SELECT count(*) FROM covers`).Scan(&count); err != nil {
t.Fatalf("count covers: %v", err)
}
if count != 0 {
t.Fatalf("legacy covers = %d, want 0", count)
}
var bodyColumn int
if err := db.QueryRow(`SELECT count(*) FROM information_schema.columns
WHERE table_name = 'covers' AND column_name = 'body'`).Scan(&bodyColumn); err != nil {
t.Fatalf("cover columns: %v", err)
}
if bodyColumn != 0 {
t.Fatal("legacy covers table still has body column")
}
}
// readSeries reads the series row directly, for asserting on what Upsert
// actually stored rather than what the joined Bookmark reports.
func readSeries(t *testing.T, s *Store, site, seriesID string) Series {
@@ -710,39 +727,143 @@ func readSeries(t *testing.T, s *Store, site, seriesID string) Series {
return sr
}
// The first PUT for a series creates its row from the client's title, cover
// and URL — there is no other source for them (ADR-0003).
// The first PUT for a series creates its row from the client's title and URL —
// there is no other source for them (ADR-0003). The Cover is not among them:
// it is acquired server-side, so a client-supplied one is dropped even on a
// brand-new row (ADR-0007).
func TestUpsertCreatesSeriesFromClient(t *testing.T) {
store := newTestStore(t)
if _, err := store.Upsert(store.OwnerID(), Bookmark{
stored, err := store.Upsert(store.OwnerID(), Bookmark{
Key: "asura:solo", Site: "asura", SeriesID: "solo",
Title: "Solo Leveling", SeriesURL: "https://asurascans.com/comics/solo",
Cover: "https://asurascans.com/covers/solo.jpg", Kind: KindManga,
UpdatedAt: 1000,
}); err != nil {
})
if err != nil {
t.Fatalf("Upsert: %v", err)
}
if stored.Cover != "" {
t.Fatalf("Cover = %q, want empty — a client cover is never stored", stored.Cover)
}
sr := readSeries(t, store, "asura", "solo")
if sr.Title != "Solo Leveling" || sr.SeriesURL != "https://asurascans.com/comics/solo" ||
sr.Cover != "https://asurascans.com/covers/solo.jpg" {
t.Fatalf("series = %+v, want client title/url/cover stored", sr)
if sr.Title != "Solo Leveling" || sr.SeriesURL != "https://asurascans.com/comics/solo" {
t.Fatalf("series = %+v, want client title/url stored", sr)
}
if sr.Cover != "" {
t.Fatalf("series cover = %q, want empty", sr.Cover)
}
}
// A PUT naming an existing series must not overwrite its title, cover or URL:
// the row is shared, and those values are scraped page content (ADR-0003).
// The hook is what starts creation-time acquisition, so it must fire exactly
// once per Series — on the PUT that created it, and on no later one, whichever
// Reader sends it.
func TestOnSeriesCreatedFiresOnceForANewSeries(t *testing.T) {
store := newTestStore(t)
var created []Series
store.OnSeriesCreated = func(sr Series) { created = append(created, sr) }
b := Bookmark{
Key: "comix:solo", Site: "comix", SeriesID: "solo", Title: "Solo Leveling",
SeriesURL: "https://comix.to/series/solo", Kind: KindManga, UpdatedAt: 1000,
}
if _, err := store.Upsert(store.OwnerID(), b); err != nil {
t.Fatalf("Upsert: %v", err)
}
b.LastChapterNum = 12
b.UpdatedAt = 2000
if _, err := store.Upsert(store.OwnerID(), b); err != nil {
t.Fatalf("second Upsert: %v", err)
}
if _, err := store.Upsert(secondReader(t, store), b); err != nil {
t.Fatalf("second reader Upsert: %v", err)
}
if len(created) != 1 {
t.Fatalf("hook fired %d times, want 1: %+v", len(created), created)
}
if created[0].Site != "comix" || created[0].SeriesID != "solo" ||
created[0].SeriesURL != "https://comix.to/series/solo" {
t.Fatalf("hook got %+v, want the created series' identity and URL", created[0])
}
}
// Acquisition at creation and the poll both write covers, and whichever
// arrives second must leave the first one alone: a Cover is replaced by
// nothing short of the series row being rebuilt.
func TestSetSeriesCoverDoesNotOverwrite(t *testing.T) {
store := newTestStore(t)
if _, err := store.Upsert(store.OwnerID(), Bookmark{
Key: "asura:solo", Site: "asura", SeriesID: "solo", UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
first := "https://asurascans.com/covers/first.jpg"
if err := store.SetSeriesCover("asura", "solo", first, []byte("first"), "image/jpeg"); err != nil {
t.Fatalf("SetSeriesCover: %v", err)
}
if err := store.SetSeriesCover("asura", "solo", "https://asurascans.com/covers/second.jpg",
[]byte("second"), "image/jpeg"); err != nil {
t.Fatalf("second SetSeriesCover: %v", err)
}
got, ok, err := store.Get(store.OwnerID(), "asura:solo")
if err != nil || !ok {
t.Fatalf("Get = %v, %v", ok, err)
}
if want := "https://bookmarks.test/covers/" + CoverAddress(first); got.Cover != want {
t.Fatalf("Cover = %q, want the first one %q", got.Cover, want)
}
}
// The address comes straight off a public request path, so anything that is
// not a stored address must be a miss rather than a filesystem lookup.
func TestCoverByAddress(t *testing.T) {
store := newTestStore(t)
if _, err := store.Upsert(store.OwnerID(), Bookmark{
Key: "asura:solo", Site: "asura", SeriesID: "solo", UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
source := "https://asurascans.com/covers/solo.jpg"
if err := store.SetSeriesCover("asura", "solo", source, []byte("bytes"), "image/jpeg"); err != nil {
t.Fatalf("SetSeriesCover: %v", err)
}
body, contentType, ok, err := store.CoverByAddress(CoverAddress(source))
if err != nil || !ok {
t.Fatalf("CoverByAddress = %v, %v", ok, err)
}
if string(body) != "bytes" || contentType != "image/jpeg" {
t.Fatalf("CoverByAddress = %q, %q, want the stored bytes", body, contentType)
}
for _, address := range []string{"", "../../etc/passwd", "ZZ" + CoverAddress(source)[2:],
CoverAddress("never stored")} {
_, _, ok, err := store.CoverByAddress(address)
if err != nil || ok {
t.Fatalf("CoverByAddress(%q) = %v, %v, want a clean miss", address, ok, err)
}
}
}
// A PUT naming an existing series must not overwrite its title or URL: the row
// is shared, and those values are scraped page content (ADR-0003). An acquired
// Cover is likewise untouched by any client.
func TestUpsertExistingSeriesIgnoresClientTitleCoverURL(t *testing.T) {
store := newTestStore(t)
base := Bookmark{
Key: "asura:solo", Site: "asura", SeriesID: "solo",
Title: "Solo Leveling", SeriesURL: "https://asurascans.com/comics/solo",
Cover: "https://asurascans.com/covers/solo.jpg", LastChapterNum: 10,
UpdatedAt: 1000,
LastChapterNum: 10, UpdatedAt: 1000,
}
if _, err := store.Upsert(store.OwnerID(), base); err != nil {
t.Fatalf("seed: %v", err)
}
acquired := "https://asurascans.com/covers/solo.jpg"
if err := store.SetSeriesCover("asura", "solo", acquired, []byte("bytes"), "image/jpeg"); err != nil {
t.Fatalf("SetSeriesCover: %v", err)
}
// Same series, hostile/compromised values, real progress advance.
base.Title = "Scraped Rename"
@@ -753,8 +874,9 @@ func TestUpsertExistingSeriesIgnoresClientTitleCoverURL(t *testing.T) {
if err != nil {
t.Fatalf("Upsert: %v", err)
}
wantCover := "https://bookmarks.test/covers/" + CoverAddress(acquired)
if got.Title != "Solo Leveling" || got.SeriesURL != "https://asurascans.com/comics/solo" ||
got.Cover != "https://asurascans.com/covers/solo.jpg" {
got.Cover != wantCover {
t.Fatalf("stored = %+v, want original title/url/cover kept", got)
}
if got.LastChapterNum != 11 {
@@ -790,16 +912,19 @@ func TestUpsertExistingSeriesAcceptsKindAndLatest(t *testing.T) {
}
// Deleting the last bookmark must leave the series row behind, so a later
// re-bookmark shows title and cover immediately instead of waiting for a poll.
// re-bookmark shows title and cover immediately instead of re-acquiring them.
func TestDeleteKeepsSeriesRow(t *testing.T) {
store := newTestStore(t)
if _, err := store.Upsert(store.OwnerID(), Bookmark{
Key: "asura:solo", Site: "asura", SeriesID: "solo",
Title: "Solo Leveling", Cover: "https://asurascans.com/covers/solo.jpg",
UpdatedAt: 1000,
Title: "Solo Leveling", UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
acquired := "https://asurascans.com/covers/solo.jpg"
if err := store.SetSeriesCover("asura", "solo", acquired, []byte("bytes"), "image/jpeg"); err != nil {
t.Fatalf("SetSeriesCover: %v", err)
}
if err := store.Delete(store.OwnerID(), "asura:solo"); err != nil {
t.Fatalf("Delete: %v", err)
}
@@ -817,7 +942,8 @@ func TestDeleteKeepsSeriesRow(t *testing.T) {
if err != nil {
t.Fatalf("re-upsert: %v", err)
}
if stored.Title != "Solo Leveling" || stored.Cover != "https://asurascans.com/covers/solo.jpg" {
wantCover := "https://bookmarks.test/covers/" + CoverAddress(acquired)
if stored.Title != "Solo Leveling" || stored.Cover != wantCover {
t.Fatalf("re-bookmark = %+v, want title/cover from the surviving series row", stored)
}
}
@@ -844,7 +970,7 @@ func TestDueForLatestCheckOrdersByReaderCountThenAge(t *testing.T) {
seedSecondReader(t, s, "asura:pop:2", "asura", "pop", 1001)
seedForCheck(t, s, "asura:solo", "https://asurascans.com/comics/solo", 100)
due, err := s.DueForLatestCheck(1000, 10)
due, err := s.DueForLatestCheck(1000, 1000, nil, 10)
if err != nil {
t.Fatalf("DueForLatestCheck: %v", err)
}
@@ -870,7 +996,7 @@ func TestDueForLatestCheckExcludesOrphanSeries(t *testing.T) {
t.Fatalf("seed orphan series: %v", err)
}
due, err := s.DueForLatestCheck(1000, 10)
due, err := s.DueForLatestCheck(1000, 1000, nil, 10)
if err != nil {
t.Fatalf("DueForLatestCheck: %v", err)
}
@@ -893,14 +1019,15 @@ func TestDueForLatestCheckExcludesOrphanSeries(t *testing.T) {
// rotations.
func TestSeedOwnerIdempotentAndRefreshesTokenHash(t *testing.T) {
url := pgtest.URL(t)
first, err := Open(url, Owner{DiscordID: "owner", TokenHash: sha256.Sum256([]byte("hash-v1"))})
coverDir := t.TempDir()
first, err := Open(url, Owner{DiscordID: "owner", TokenHash: sha256.Sum256([]byte("hash-v1"))}, coverDir, testCoverBaseURL)
if err != nil {
t.Fatalf("Open: %v", err)
}
ownerID := first.OwnerID()
first.Close()
second, err := Open(url, Owner{DiscordID: "owner", TokenHash: sha256.Sum256([]byte("hash-v2"))})
second, err := Open(url, Owner{DiscordID: "owner", TokenHash: sha256.Sum256([]byte("hash-v2"))}, coverDir, testCoverBaseURL)
if err != nil {
t.Fatalf("reopen: %v", err)
}
@@ -957,7 +1084,7 @@ func TestMigration0004AttachesBookmarksToOwner(t *testing.T) {
t.Fatalf("migrate to 0002: %v", err)
}
st, err := Open(url, testOwner)
st, err := Open(url, testOwner, t.TempDir(), testCoverBaseURL)
if err != nil {
t.Fatalf("Open: %v", err)
}
@@ -1207,7 +1334,7 @@ func TestTwoReadersShareOneSeriesWithIndependentProgress(t *testing.T) {
t.Fatalf("series rows = %d, want 1 shared row for two bookmarks", series)
}
due, err := s.DueForLatestCheck(time.Now().UnixMilli(), 10)
due, err := s.DueForLatestCheck(time.Now().UnixMilli(), time.Now().UnixMilli(), nil, 10)
if err != nil {
t.Fatalf("DueForLatestCheck: %v", err)
}
@@ -1223,7 +1350,7 @@ func TestTwoReadersShareOneSeriesWithIndependentProgress(t *testing.T) {
if b, ok, err := s.Get(s.OwnerID(), "asura:solo"); err != nil || !ok || b.LastChapterNum != 200 {
t.Fatalf("owner's bookmark after the other's delete = %+v ok=%v err=%v, want it intact", b, ok, err)
}
due, err = s.DueForLatestCheck(time.Now().UnixMilli(), 10)
due, err = s.DueForLatestCheck(time.Now().UnixMilli(), time.Now().UnixMilli(), nil, 10)
if err != nil {
t.Fatalf("DueForLatestCheck after delete: %v", err)
}
@@ -1231,3 +1358,110 @@ func TestTwoReadersShareOneSeriesWithIndependentProgress(t *testing.T) {
t.Fatalf("due after one Reader left = %+v, want the series still polled", due)
}
}
func TestCoverPersistsAcrossReopen(t *testing.T) {
url := pgtest.URL(t)
coverDir := t.TempDir()
first, err := Open(url, testOwner, coverDir, testCoverBaseURL)
if err != nil {
t.Fatalf("Open: %v", err)
}
body := []byte("stored-cover")
const sourceURL = "https://cdn.example/covers/series.jpg"
if err := first.PutCover(sourceURL, body, "image/webp"); err != nil {
t.Fatalf("PutCover: %v", err)
}
if err := first.Close(); err != nil {
t.Fatalf("close first store: %v", err)
}
second, err := Open(url, testOwner, coverDir, testCoverBaseURL)
if err != nil {
t.Fatalf("reopen: %v", err)
}
defer second.Close()
got, contentType, ok, err := second.GetCover(sourceURL)
if err != nil {
t.Fatalf("GetCover: %v", err)
}
if !ok || !bytes.Equal(got, body) || contentType != "image/webp" {
t.Fatalf("stored cover = (%q, %q, %v), want (%q, image/webp, true)", got, contentType, ok, body)
}
}
func TestOpenRequiresCoverDirectory(t *testing.T) {
if _, err := Open(pgtest.URL(t), testOwner, "", testCoverBaseURL); err == nil || !strings.Contains(err.Error(), "cover directory is required") {
t.Fatalf("Open without cover directory = %v, want required-directory error", err)
}
}
// A base URL without a scheme reads like a hostname and starts cleanly, but
// every Cover it puts on the wire is an address no browser can resolve.
func TestOpenRequiresAbsoluteCoverBaseURL(t *testing.T) {
for _, base := range []string{"", "bookmarks.test", "https://", "ftp://bookmarks.test"} {
if _, err := Open(pgtest.URL(t), testOwner, t.TempDir(), base); err == nil ||
!strings.Contains(err.Error(), "absolute http(s) origin") {
t.Fatalf("Open with base %q = %v, want absolute-origin error", base, err)
}
}
}
func TestCoverIsContentAddressedOnFilesystem(t *testing.T) {
url := pgtest.URL(t)
coverDir := t.TempDir()
first, err := Open(url, testOwner, coverDir, testCoverBaseURL)
if err != nil {
t.Fatalf("Open: %v", err)
}
body := []byte("stored-cover")
const sourceURL = "https://cdn.example/covers/series.jpg"
if err := first.PutCover(sourceURL, body, "image/webp"); err != nil {
first.Close()
t.Fatalf("PutCover: %v", err)
}
defer first.Close()
addressBytes := sha256.Sum256([]byte(sourceURL))
address := hex.EncodeToString(addressBytes[:])
wantPath := filepath.Join(address[:2], address[2:4], address)
var gotPath, contentType string
if err := first.db.QueryRow(`SELECT path, content_type FROM covers WHERE address = $1`, address).Scan(&gotPath, &contentType); err != nil {
t.Fatalf("cover row: %v", err)
}
if gotPath != wantPath || contentType != "image/webp" {
t.Fatalf("cover row = (%q, %q), want (%q, image/webp)", gotPath, contentType, wantPath)
}
if got, err := os.ReadFile(filepath.Join(coverDir, gotPath)); err != nil || !bytes.Equal(got, body) {
t.Fatalf("cover file = (%q, %v), want (%q, nil)", got, err, body)
}
var bodyColumn int
if err := first.db.QueryRow(`SELECT count(*) FROM information_schema.columns WHERE table_name = 'covers' AND column_name = 'body'`).Scan(&bodyColumn); err != nil {
t.Fatalf("cover columns: %v", err)
}
if bodyColumn != 0 {
t.Fatalf("covers still has body column")
}
}
func TestCoverStoreAcceptsAnySourceURL(t *testing.T) {
s := newTestStore(t)
const sourceURL = "https://cdn.example/covers/series.jpg"
want := []byte("cover-bytes")
if err := s.PutCover(sourceURL, want, "image/jpeg"); err != nil {
t.Fatalf("PutCover: %v", err)
}
got, contentType, ok, err := s.GetCover(sourceURL)
if err != nil {
t.Fatalf("GetCover: %v", err)
}
if !ok || !bytes.Equal(got, want) || contentType != "image/jpeg" {
t.Fatalf("GetCover = (%q, %q, %v), want (%q, image/jpeg, true)", got, contentType, ok, want)
}
if err := s.PutCover("https://cdn.example/not-image", []byte("html"), "text/html"); err == nil {
t.Fatal("PutCover accepted a non-image")
}
if _, _, ok, err := s.GetCover("https://cdn.example/not-image"); err != nil || ok {
t.Fatalf("rejected cover = found %v, err %v; want missing", ok, err)
}
}
-129
View File
@@ -1,129 +0,0 @@
package web
import (
"context"
"log"
"net/http"
"regexp"
"sync"
"time"
)
// CoverFetcher retrieves one kagane cover by image id. Satisfied by
// latest.BrowserFetcher, and nil when BROWSER_WS_URL is unset — which leaves
// kagane covers exactly as unavailable as they were before this endpoint
// existed, rather than hanging a request on a fetcher that cannot run.
type CoverFetcher interface {
Image(ctx context.Context, imageID string) (body []byte, contentType string, err error)
}
// coverIDRe matches the request path segment that becomes part of an outbound
// URL. The proxy is session-gated, but the id still reaches a headless browser,
// so it is validated at the boundary rather than passed through.
var coverIDRe = regexp.MustCompile(`^[0-9a-f-]{36}$`)
// coverTypes is the set of content types the proxy will echo back. A response
// header sourced from a third party is not repeated verbatim: anything outside
// this set is treated as "not a cover".
var coverTypes = map[string]bool{
"image/webp": true,
"image/jpeg": true,
"image/png": true,
"image/avif": true,
"image/gif": true,
}
// coverTimeout bounds one proxied cover. Shorter than the fetcher's own
// challenge budget on purpose: a browser page is waiting on this, and a cover
// that has not arrived by now is better left as a broken slot than as a request
// holding a connection open.
const coverTimeout = 20 * time.Second
// coverCacheMax caps the in-memory cover cache. Covers are immutable per image
// id and a library holds tens of series, so this is a ceiling that is never
// reached in practice; reaching it clears the map rather than evicting by age.
//
// ponytail: flush-on-full, not LRU. Swap it for an LRU if a library ever grows
// past this and the flush starts costing refetches.
const coverCacheMax = 500
type cachedCover struct {
body []byte
contentType string
}
type coverCache struct {
mu sync.Mutex
m map[string]cachedCover
}
func (c *coverCache) get(id string) (cachedCover, bool) {
c.mu.Lock()
defer c.mu.Unlock()
v, ok := c.m[id]
return v, ok
}
func (c *coverCache) put(id string, v cachedCover) {
c.mu.Lock()
defer c.mu.Unlock()
if c.m == nil || len(c.m) >= coverCacheMax {
c.m = make(map[string]cachedCover, coverCacheMax)
}
c.m[id] = v
}
// kaganeCover serves a kagane cover from the backend's own origin.
//
// kagane answers image requests with a Cloudflare challenge and
// `cross-origin-resource-policy: same-origin`, so the web UI cannot render one
// directly under any combination of referrer policy or crossorigin attribute
// (verified 2026-08-08). Fetching it through the headless browser that already
// clears the challenge, and re-serving it here, is what puts the bytes on an
// origin the page may load from.
//
// ponytail: covers are fetched on first view, one browser navigation at a time
// behind the fetcher's mutex, so a first load of a large kagane library
// trickles in over a few seconds. The cache makes it a one-off. Prefetching
// during the poll cycle is the upgrade if that ever grates.
func (h *Handler) kaganeCover(w http.ResponseWriter, r *http.Request) {
id := r.PathValue("id")
if !coverIDRe.MatchString(id) {
http.NotFound(w, r)
return
}
if h.covers == nil {
http.NotFound(w, r)
return
}
if v, ok := h.coverCache.get(id); ok {
writeCover(w, v)
return
}
ctx, cancel := context.WithTimeout(r.Context(), coverTimeout)
defer cancel()
body, contentType, err := h.covers.Image(ctx, id)
if err != nil {
log.Printf("kagane cover %s: %v", id, err)
http.NotFound(w, r)
return
}
if !coverTypes[contentType] {
log.Printf("kagane cover %s: unexpected content type %q", id, contentType)
http.NotFound(w, r)
return
}
v := cachedCover{body: body, contentType: contentType}
h.coverCache.put(id, v)
writeCover(w, v)
}
// writeCover sends the bytes with a long cache life: an image id names one
// immutable rendering, so a client that has it never needs to ask again.
func writeCover(w http.ResponseWriter, v cachedCover) {
w.Header().Set("Content-Type", v.contentType)
w.Header().Set("Cache-Control", "private, max-age=604800, immutable")
w.Write(v.body)
}
+1 -1
View File
@@ -6,7 +6,7 @@
<div class="row">
<a class="cover" href="{{.ContinueURL}}" target="_blank" rel="noopener noreferrer"
tabindex="-1" aria-hidden="true">
{{if .CoverURL}}<img src="{{.CoverURL}}" alt="" loading="lazy">
{{if .Cover}}<img src="{{.Cover}}" alt="" loading="lazy">
{{/* aria-hidden on the cover link is not enough — Chromium still exposes
the letter because the link is programmatically focusable — so the
monogram carries its own, same as the recent strip's. */}}
+1 -1
View File
@@ -14,7 +14,7 @@
<a class="recent-card {{if .HasNewChapter}}is-new{{end}}" href="{{.ContinueURL}}"
target="_blank" rel="noopener noreferrer">
<span class="recent-cover">
{{if .CoverURL}}<img src="{{.CoverURL}}" alt="" loading="lazy">
{{if .Cover}}<img src="{{.Cover}}" alt="" loading="lazy">
{{else}}<span class="monogram" aria-hidden="true">{{.Initial}}</span>{{end}}
{{if .HasNewChapter}}<span class="foot-rule"></span>
{{else if .Favorite}}<span class="foot-rule brass"></span>{{end}}
+1 -10
View File
@@ -50,10 +50,6 @@ type Handler struct {
// httpClient is the plain stdlib client that talks to Discord. It is not
// an injected interface: tests point APIBase at a stub server instead.
httpClient *http.Client
// covers proxies kagane cover images, which no browser can load directly.
// Nil disables the endpoint — see CoverFetcher.
covers CoverFetcher
coverCache coverCache
}
// listView is what every list-rendering template receives.
@@ -115,7 +111,7 @@ type loginView struct {
// New parses every template up front so a broken one kills the process at
// startup rather than the first request that touches it.
func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, novelPath string, covers CoverFetcher) (*Handler, error) {
func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, novelPath string) (*Handler, error) {
tmpl, err := template.ParseFS(templateFS, "templates/*.html")
if err != nil {
return nil, err
@@ -130,7 +126,6 @@ func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, nove
states: newOAuthStates(),
limiter: session.NewLoginLimiter(),
httpClient: &http.Client{Timeout: discordTimeout},
covers: covers,
}, nil
}
@@ -147,10 +142,6 @@ func (h *Handler) Register(mux *http.ServeMux) {
mux.HandleFunc("POST /ui/bookmarks/{key}/chapter", h.requireSession(h.uiChapter))
mux.HandleFunc("DELETE /ui/bookmarks/{key}", h.requireSession(h.uiDelete))
// Session-gated like every other UI route: the deployment proxies kagane's
// images for its own Readers, not for the internet.
mux.HandleFunc("GET /img/kagane/{id}", h.requireSession(h.kaganeCover))
// Install endpoints render the script directly under the session: the
// credential travels inside the served bytes, never in the address bar or
// the page markup. Updates after install use the credential-bearing /u/
+105 -40
View File
@@ -30,7 +30,17 @@ type Config struct {
// DatabaseURL is the Postgres connection URL; required, no default,
// because a wrong guess would silently start on an empty database.
DatabaseURL string
Port string
// CoverDir is the filesystem volume for immutable cover bytes. Required:
// serving a stored address without durable bytes would be worse than a
// startup failure.
CoverDir string
// PublicBaseURL is the origin this deployment answers on, e.g.
// "https://bookmarks.example.com". Required: cover URLs go out absolute
// because the userscript renders them on third-party origins, where a
// relative path would resolve against the Site (ADR-0007), and there is
// no way to guess it from a request the poller never sees.
PublicBaseURL string
Port string
// OwnerDiscordID identifies the seeded owner Reader (issue #22). Required:
// bookmarks are scoped to a Reader, and a fresh deployment needs one
// before anybody logs in. The owner is also the only Reader who can revoke
@@ -47,10 +57,6 @@ type Config struct {
NovelUserscriptPath string
// LatestPoll configures the background latest-chapter fetcher.
LatestPoll LatestPoll
// Covers proxies kagane cover images for the web UI. Not from the
// environment: it is the shared headless browser, wired in main once it
// connects, and nil in every test router.
Covers web.CoverFetcher
}
// LatestPoll configures the background latest-chapter poller.
@@ -60,18 +66,20 @@ type Config struct {
// deployment. Past that nothing breaks; the effective cadence stretches to
// N x interval / batch and the oldest-checked-first ordering keeps it uniform.
type LatestPoll struct {
Enabled bool
Cooldown time.Duration
Interval time.Duration
Stagger time.Duration
Batch int
Enabled bool
Cooldown time.Duration
BrowserCooldown time.Duration
Interval time.Duration
Stagger time.Duration
Batch int
}
const (
defaultPollCooldown = time.Hour
defaultPollInterval = 10 * time.Minute
defaultPollStagger = 20 * time.Second
defaultPollBatch = 14
defaultPollCooldown = time.Hour
defaultBrowserPollCooldown = 6 * time.Hour
defaultPollInterval = 10 * time.Minute
defaultPollStagger = 20 * time.Second
defaultPollBatch = 14
// minPollCooldown keeps a typo from turning a polite background check into
// a hammer against sites that are already bot-scoring us.
minPollCooldown = 15 * time.Minute
@@ -129,20 +137,27 @@ func envInt(key string, def int) int {
return n
}
func clampPollCooldown(name string, d time.Duration) time.Duration {
if d < minPollCooldown {
log.Printf("config: %s %s is below the %s floor, clamping", name, d, minPollCooldown)
return minPollCooldown
}
return d
}
// loadLatestPoll reads the poller's settings, clamping anything that would make
// it antisocial.
func loadLatestPoll() LatestPoll {
p := LatestPoll{
Enabled: envBool("LATEST_CHAPTER_POLL_ENABLED", true),
Cooldown: envDuration("LATEST_CHAPTER_POLL_COOLDOWN", defaultPollCooldown),
Interval: envDuration("LATEST_CHAPTER_POLL_INTERVAL", defaultPollInterval),
Stagger: envDuration("LATEST_CHAPTER_POLL_STAGGER", defaultPollStagger),
Batch: envInt("LATEST_CHAPTER_POLL_BATCH", defaultPollBatch),
}
if p.Cooldown < minPollCooldown {
log.Printf("config: cooldown %s is below the %s floor, clamping", p.Cooldown, minPollCooldown)
p.Cooldown = minPollCooldown
Enabled: envBool("LATEST_CHAPTER_POLL_ENABLED", true),
Cooldown: envDuration("LATEST_CHAPTER_POLL_COOLDOWN", defaultPollCooldown),
BrowserCooldown: envDuration("LATEST_CHAPTER_POLL_BROWSER_COOLDOWN", defaultBrowserPollCooldown),
Interval: envDuration("LATEST_CHAPTER_POLL_INTERVAL", defaultPollInterval),
Stagger: envDuration("LATEST_CHAPTER_POLL_STAGGER", defaultPollStagger),
Batch: envInt("LATEST_CHAPTER_POLL_BATCH", defaultPollBatch),
}
p.Cooldown = clampPollCooldown("cooldown", p.Cooldown)
p.BrowserCooldown = clampPollCooldown("browser cooldown", p.BrowserCooldown)
// batch x stagger has to fit inside one tick or a batch is still running
// when the next one is due. Run() serialises them, so this degrades to a
// slower cadence rather than to overlapping fetches — worth a warning, not
@@ -158,6 +173,8 @@ func loadConfig() Config {
c := Config{
TokenKey: os.Getenv("TOKEN_KEY"),
DatabaseURL: os.Getenv("DATABASE_URL"),
CoverDir: os.Getenv("COVER_DIR"),
PublicBaseURL: os.Getenv("PUBLIC_BASE_URL"),
Port: envOr("PORT", "8080"),
OwnerDiscordID: os.Getenv("OWNER_DISCORD_ID"),
UserscriptPath: envOr("USERSCRIPT_PATH", "/userscript/manga-bookmark.user.js"),
@@ -185,7 +202,12 @@ func loadConfig() Config {
// /healthz is public.
func newRouter(s *store.Store, cfg Config) http.Handler {
mux := http.NewServeMux()
h := &api.Handler{Store: s}
mux.HandleFunc("GET /healthz", api.Healthz)
// Public: cover bytes are rendered by the userscript on origins that may
// not send our credentials, and the address is the hash of a URL the Site
// already publishes (ADR-0007).
mux.HandleFunc("GET /covers/{address}", h.Cover)
// Outside httpmw.Auth (the updater sends no Authorization header) and
// outside the web UI's Discord auth (the script must be installable
@@ -197,7 +219,6 @@ func newRouter(s *store.Store, cfg Config) http.Handler {
mux.HandleFunc("GET /u/{token}/novel-bookmark.user.js",
userscript.Handler(s, cfg.NovelUserscriptPath))
h := &api.Handler{Store: s}
protected := http.NewServeMux()
protected.HandleFunc("GET /bookmarks", h.List)
protected.HandleFunc("PUT /bookmarks/{key}", h.Put)
@@ -210,7 +231,7 @@ func newRouter(s *store.Store, cfg Config) http.Handler {
// The browser UI is always registered; signing in is Discord OAuth, so
// there is no password to forget and no gate to leave unset.
wh, err := web.New(s, cfg.Discord, []byte(cfg.TokenKey),
cfg.UserscriptPath, cfg.NovelUserscriptPath, cfg.Covers)
cfg.UserscriptPath, cfg.NovelUserscriptPath)
if err != nil {
log.Fatalf("web handler: %v", err)
}
@@ -245,6 +266,12 @@ func main() {
if cfg.DatabaseURL == "" {
log.Fatal("DATABASE_URL is required")
}
if cfg.CoverDir == "" {
log.Fatal("COVER_DIR is required")
}
if cfg.PublicBaseURL == "" {
log.Fatal("PUBLIC_BASE_URL is required")
}
// The web UI signs in through Discord, so a deployment without the OAuth
// application is misconfigured rather than passwordless.
for key, v := range map[string]string{
@@ -265,7 +292,7 @@ func main() {
TokenHash: token.Hash(token.Token([]byte(cfg.TokenKey), cfg.OwnerDiscordID, 0)),
}
s, err := store.Open(cfg.DatabaseURL, owner)
s, err := store.Open(cfg.DatabaseURL, owner, cfg.CoverDir, cfg.PublicBaseURL)
if err != nil {
log.Fatalf("open store: %v", err)
}
@@ -276,9 +303,9 @@ func main() {
// chapters on its own.
//
// One headless browser serves both consumers that need a Cloudflare
// challenge cleared: the poller's kagane/novelfull fetches and the web
// UI's kagane cover proxy. Optional — unset leaves both degraded to what
// they were before the sidecar existed.
// challenge cleared: the poller's kagane/novelfull page fetches and
// kagane's cover bytes. Optional — unset leaves kagane unpolled and its
// Covers blank until the bytes exist.
var browser latest.Fetcher
pollCtx, stopPoll := context.WithCancel(context.Background())
defer stopPoll()
@@ -288,11 +315,37 @@ func main() {
log.Printf("browser fetcher disabled: %v", err)
} else {
browser = bf
cfg.Covers = bf
context.AfterFunc(pollCtx, bf.Close)
log.Printf("browser fetcher at %s", ws)
}
}
// A Series nobody had bookmarked before gets its Latest Chapter and its
// Cover from one fetch, at creation, instead of waiting out a poll queue
// ordered by Reader count. Off the write path: the hook returns as soon
// as the goroutine is started.
var tlsFetch latest.Fetcher
if f, err := latest.NewTLSFetcher(); err != nil {
log.Printf("creation-time acquisition: plain-TLS Sites disabled, cannot build client: %v", err)
} else {
tlsFetch = f
}
var browserCover latest.BrowserCoverFetcher
if b, ok := browser.(latest.BrowserCoverFetcher); ok {
browserCover = b
}
// The Acquirer must survive a TLS client failure: kagane needs only the
// sidecar, and novelfull degrades to whatever is left.
if tlsFetch != nil || browser != nil {
acq := &latest.Acquirer{
Store: s,
Fetch: tlsFetch,
BrowserFetch: browser,
BrowserCoverFetch: browserCover,
Covers: latest.NewCoverFetcher(),
Ctx: pollCtx,
}
s.OnSeriesCreated = acq.Acquire
}
startLatestPoller(pollCtx, s, cfg.LatestPoll, browser)
srv := &http.Server{
@@ -324,6 +377,27 @@ func main() {
}
}
// newLatestPoller wires the configured cooldowns and fetchers into the poller.
func newLatestPoller(s *store.Store, cfg LatestPoll, fetch, browser latest.Fetcher) *latest.Poller {
var covers latest.BrowserCoverFetcher
if f, ok := browser.(latest.BrowserCoverFetcher); ok {
covers = f
}
return &latest.Poller{
Store: s,
Fetch: fetch,
BrowserFetch: browser,
CoverFetch: covers,
CoverBytesFetch: latest.NewCoverFetcher(),
Now: time.Now,
Cooldown: cfg.Cooldown,
BrowserCooldown: cfg.BrowserCooldown,
Interval: cfg.Interval,
Stagger: cfg.Stagger,
Batch: cfg.Batch,
}
}
// startLatestPoller launches the background poller unless it is disabled or its
// HTTP client cannot be built. Any problem here is logged and skipped: this
// feature going missing degrades the service to userscript-only latest-chapter
@@ -341,16 +415,7 @@ func startLatestPoller(ctx context.Context, s *store.Store, cfg LatestPoll, brow
// Nil browser: sites behind a JavaScript challenge are simply not polled,
// and their latest_chapter comes from the userscript alone — which is how
// the service behaved before the sidecar existed.
p := &latest.Poller{
Store: s,
Fetch: f,
BrowserFetch: browser,
Now: time.Now,
Cooldown: cfg.Cooldown,
Interval: cfg.Interval,
Stagger: cfg.Stagger,
Batch: cfg.Batch,
}
p := newLatestPoller(s, cfg, f, browser)
go p.Run(ctx)
}
+49 -7
View File
@@ -17,25 +17,33 @@ import (
func TestLoadLatestPollDefaults(t *testing.T) {
for _, k := range []string{
"LATEST_CHAPTER_POLL_ENABLED", "LATEST_CHAPTER_POLL_COOLDOWN",
"LATEST_CHAPTER_POLL_INTERVAL", "LATEST_CHAPTER_POLL_STAGGER",
"LATEST_CHAPTER_POLL_BATCH",
"LATEST_CHAPTER_POLL_BROWSER_COOLDOWN", "LATEST_CHAPTER_POLL_INTERVAL",
"LATEST_CHAPTER_POLL_STAGGER", "LATEST_CHAPTER_POLL_BATCH",
} {
t.Setenv(k, "")
}
got := loadLatestPoll()
want := LatestPoll{
Enabled: true,
Cooldown: time.Hour,
Interval: 10 * time.Minute,
Stagger: 20 * time.Second,
Batch: 14,
Enabled: true,
Cooldown: time.Hour,
BrowserCooldown: 6 * time.Hour,
Interval: 10 * time.Minute,
Stagger: 20 * time.Second,
Batch: 14,
}
if got != want {
t.Fatalf("loadLatestPoll() = %+v, want %+v", got, want)
}
}
func TestLoadConfigReadsCoverDirectory(t *testing.T) {
t.Setenv("COVER_DIR", "/covers")
if got := loadConfig().CoverDir; got != "/covers" {
t.Fatalf("CoverDir = %q, want /covers", got)
}
}
func TestLoadLatestPollEnabledParsing(t *testing.T) {
tests := []struct {
raw string
@@ -74,6 +82,30 @@ func TestLoadLatestPollClampsAndFallsBack(t *testing.T) {
wantFrom: func(p LatestPoll) any { return p.Cooldown },
want: 15 * time.Minute,
},
{
name: "browser cooldown below the floor is clamped up",
env: map[string]string{"LATEST_CHAPTER_POLL_BROWSER_COOLDOWN": "1m"},
wantFrom: func(p LatestPoll) any { return p.BrowserCooldown },
want: 15 * time.Minute,
},
{
name: "browser cooldown at the floor is kept",
env: map[string]string{"LATEST_CHAPTER_POLL_BROWSER_COOLDOWN": "15m"},
wantFrom: func(p LatestPoll) any { return p.BrowserCooldown },
want: 15 * time.Minute,
},
{
name: "browser cooldown override is honoured",
env: map[string]string{"LATEST_CHAPTER_POLL_BROWSER_COOLDOWN": "8h"},
wantFrom: func(p LatestPoll) any { return p.BrowserCooldown },
want: 8 * time.Hour,
},
{
name: "browser cooldown unparseable value falls back",
env: map[string]string{"LATEST_CHAPTER_POLL_BROWSER_COOLDOWN": "six hours"},
wantFrom: func(p LatestPoll) any { return p.BrowserCooldown },
want: 6 * time.Hour,
},
{
name: "a valid override is honoured",
env: map[string]string{"LATEST_CHAPTER_POLL_INTERVAL": "5m"},
@@ -123,6 +155,16 @@ func TestLoadLatestPollClampsAndFallsBack(t *testing.T) {
}
}
func TestNewLatestPollerWiresCooldowns(t *testing.T) {
p := newLatestPoller(nil, LatestPoll{
Cooldown: time.Hour,
BrowserCooldown: 6 * time.Hour,
}, nil, nil)
if p.Cooldown != time.Hour || p.BrowserCooldown != 6*time.Hour {
t.Fatalf("poller cooldowns = %s/%s, want 1h/6h", p.Cooldown, p.BrowserCooldown)
}
}
func TestPutStatusValidation(t *testing.T) {
cases := []struct {
name string
+28
View File
@@ -0,0 +1,28 @@
# Copy to chrome/.env on the home machine. Never commit the real .env.
#
# This file configures the browser unit only. It is separate from the API
# stack's ../.env on purpose: the two run on different machines.
# The address the CDP port is published on — required, no default.
#
# Use this machine's **tailnet IP**, e.g. 100.x.y.z (`tailscale ip -4`). Not
# 0.0.0.0, not the LAN address: CDP has no authentication of its own, so
# anything that can reach this port has full control of the browser and a
# foothold on this host. Tailscale device identity plus an ACL is the access
# control; the bind address is what enforces it.
#
# For a throwaway local test, 127.0.0.1 is fine — but then only this machine
# can reach it, so the API must run here too.
# Left commented so `cp .env.example .env && docker compose up` fails with the
# variable's own message telling you what to set, rather than Docker rejecting
# "100.x.y.z" as an invalid IP.
# BROWSER_BIND_ADDR=100.x.y.z
# Clock zone the browser reports. Any real zone works and it need not match
# the egress IP's country — but it must not be UTC, which is itself the bot
# signal that stops the challenge clearing. The measurement is in entrypoint.sh.
#
# Unset falls back to the host's /etc/timezone, which is a real zone whenever
# the host clock is set to local time. Set this when the host runs UTC — a UTC
# server is exactly the case that fails.
# BROWSER_TZ=Asia/Jakarta
+8 -2
View File
@@ -18,7 +18,7 @@ FROM debian:trixie-slim
# stale, and a stale browser is exactly what Cloudflare turns away — the 124 in
# alpine-chrome is the worked example. Rebuild is the upgrade path.
RUN apt-get update \
&& apt-get install -y --no-install-recommends ca-certificates wget gnupg \
&& apt-get install -y --no-install-recommends ca-certificates wget gnupg util-linux \
&& wget -qO- https://dl.google.com/linux/linux_signing_key.pub \
| gpg --dearmor -o /usr/share/keyrings/google-chrome.gpg \
&& echo "deb [arch=amd64 signed-by=/usr/share/keyrings/google-chrome.gpg] https://dl.google.com/linux/chrome/deb/ stable main" \
@@ -27,9 +27,15 @@ RUN apt-get update \
&& apt-get install -y --no-install-recommends google-chrome-stable socat \
&& rm -rf /var/lib/apt/lists/*
# Keep the profile path present so Docker initializes the named volume with
# the unprivileged user's ownership.
RUN useradd --create-home --shell /usr/sbin/nologin chrome \
&& mkdir -p /home/chrome/profile /home/chrome/state \
&& chown -R chrome:chrome /home/chrome
# Unprivileged: Chrome refuses to run as root, and the CDP endpoint is a shell
# on whatever user owns it.
RUN useradd --create-home --shell /usr/sbin/nologin chrome
USER chrome
WORKDIR /home/chrome
+62
View File
@@ -0,0 +1,62 @@
# The browser, as its own deployable unit.
#
# This does NOT run beside the API. It runs on the home machine, reached from
# the VPS over the tailnet, and is updated without touching the API stack:
#
# cd chrome && docker compose up -d --build
#
# Set BROWSER_BIND_ADDR in chrome/.env to this machine's tailnet IP. See
# ../DEPLOY.md §7 for the full first-time procedure and ../docs/adr/
# 0006-browser-on-the-home-machine.md for why the browser lives here at all.
name: bookmark-browser
services:
browser:
build: .
image: bookmarkmanager-chrome:latest
container_name: bookmark-browser
restart: unless-stopped
environment:
# Any real zone works, but a UTC clock is itself the bot signal and the
# challenge then never clears — measurement in entrypoint.sh. Unset falls
# back to the host's /etc/timezone below, which is a real zone whenever
# the host clock is local; set BROWSER_TZ when the host runs UTC.
TZ: ${BROWSER_TZ:-}
volumes:
# The zone *name*, which is what Chrome's ICU needs — see entrypoint.sh.
# Absent on a non-Debian host, which the entrypoint handles by falling back to UTC.
- /etc/timezone:/etc/timezone:ro
# Cloudflare clearance must survive Chrome reaping and image recreation.
- chrome-profile:/home/chrome/profile
# Bound to the tailnet address only, never 0.0.0.0. CDP authenticates
# nothing: whatever reaches this port drives the browser and, through it,
# this host. On the VPS the safety was Docker network membership; here the
# machine has a real LAN, so the bind address *is* the access control,
# backed by Tailscale device identity. No default — an unset variable must
# fail the deploy rather than silently publish CDP to the LAN.
ports:
- "${BROWSER_BIND_ADDR:?set BROWSER_BIND_ADDR to this machine's tailnet IP}:9222:9222"
# Reaps zombie renderer processes, which otherwise accumulate for the
# container's lifetime.
init: true
# Chrome allocates shared memory per tab and dies on Docker's 64MB default.
# 128MB against a measured 19MB peak: the old 1GB reservation was sized by
# superstition, and this box has 1.8GB total.
shm_size: '128mb'
# The browser is the newcomer on a machine where a Gitea runner already
# holds ~1.2GiB of 1.8GiB. Load-bearing, not decorative: untuned Chrome
# peaked at 645MiB cgroup, which is more than is free here.
#
# memswap_limit is memory+swap combined, so this allows 512MiB of swap —
# Chrome reclaims its own cold pages onto this box's 5.9GiB of SATA swap
# instead of taking resident memory from the runner.
mem_limit: 512m
memswap_limit: 1g
# If the box does run out, the kernel takes the browser and never CI.
oom_score_adj: 800
# A challenge solve yields to a running build. Cold start degrades to ~3s
# at half a CPU, immaterial against a 45-second challenge budget.
cpu_shares: 512
volumes:
chrome-profile:
+203 -39
View File
@@ -16,46 +16,210 @@ set -eu
[ -n "${TZ:-}" ] || TZ=$(cat /etc/timezone 2>/dev/null || echo UTC)
export TZ
# Chrome's own UA advertises "HeadlessChrome" under --headless=new, and that one
# token is the difference between kagane.to's challenge clearing in ~4s and
# never clearing at all (measured 2026-08-08, same host, same Chrome, only the
# UA changed). Overriding it does not touch the Sec-CH-UA client hints, which
# report the real version, so the version is read back out of the binary rather
# than hardcoded: a hardcoded one would drift out of step with the hints on the
# next Chrome update and become a fresh tell.
state=/home/chrome/state
profile=/home/chrome/profile
lock_file=$state/lock
pid_file=$state/chrome.pid
connections_dir=$state/connections
last_use_file=$state/last-use
idle_seconds=300
mkdir -p "$state" "$profile" "$connections_dir"
exec 9>>"$lock_file"
# Chrome's own UA advertises "HeadlessChrome" under --headless=new, and that
# one token is the difference between kagane.to's challenge clearing in ~4s and
# never clearing at all. Read the installed major version so client hints and
# the UA stay aligned after an image rebuild.
major=$(google-chrome-stable --version | sed -E 's/[^0-9]*([0-9]+)\..*/\1/')
ua="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/${major}.0.0.0 Safari/537.36"
# Chrome binds its DevTools port to loopback and silently ignores
# --remote-debugging-address (verified 2026-08-08: Chrome 151 with
# --remote-debugging-address=0.0.0.0 still listened on 127.0.0.1 only), so the
# caller — another container — cannot reach it directly. socat fronting the
# loopback port is how chromedp/headless-shell solved the same problem and is
# why this image is a drop-in for it.
#
# Nothing publishes 9222; reachability is the `browser` network in
# docker-compose.yml, and an exposed CDP endpoint is remote code execution.
socat TCP-LISTEN:9222,fork,reuseaddr TCP:127.0.0.1:9223 &
lock() {
flock 9
}
# Chrome stays in the foreground so that its death takes the container down and
# compose's restart policy applies; a backgrounded browser behind a live socat
# would leave the sidecar looking healthy while answering nothing.
#
# No --enable-automation: it sets navigator.webdriver, the first thing a bot
# check reads.
#
# --no-sandbox because Chrome's zygote wants user namespaces, which Docker's
# default profile does not hand out; the alternative is --cap-add=SYS_ADMIN,
# which gives the container strictly more than it takes away. Containment here
# is the unprivileged user, the isolated network, and the fact that this
# browser only ever navigates to kagane.to and novelfull.com.
exec google-chrome-stable \
--headless=new \
--no-sandbox \
--remote-debugging-port=9223 \
--user-agent="$ua" \
--user-data-dir=/home/chrome/profile \
--no-first-run \
--no-default-browser-check \
--disable-gpu \
about:blank
unlock() {
flock -u 9
}
browser_alive() {
[ -s "$pid_file" ] || return 1
pid=$(cat "$pid_file")
[ -n "$pid" ] && kill -0 "$pid" 2>/dev/null
}
has_connections() {
for marker in "$connections_dir"/*; do
[ -e "$marker" ] || continue
pid=${marker##*/}
if kill -0 "$pid" 2>/dev/null; then
return 0
fi
# A SIGKILLed helper cannot run its cleanup trap. Reconcile its marker
# here so one dead client cannot pin Chrome forever.
rm -f "$marker"
done
return 1
}
start_browser() {
# No --enable-automation: it sets navigator.webdriver, the first thing a
# bot check reads. setsid gives Chrome a process group so the reaper can
# terminate its renderer children with the browser.
# --no-sandbox avoids granting SYS_ADMIN solely for Docker's unavailable
# user namespaces; containment is the unprivileged user and private network.
setsid google-chrome-stable \
--headless=new \
--no-sandbox \
--remote-debugging-port=9223 \
--user-agent="$ua" \
--user-data-dir="$profile" \
--no-first-run \
--no-default-browser-check \
--disable-gpu \
about:blank >/dev/null &
printf '%s\n' "$!" >"$pid_file"
}
stop_browser() {
pid=$(cat "$pid_file")
kill -TERM -- "-$pid" 2>/dev/null || kill -TERM "$pid" 2>/dev/null || true
i=0
while kill -0 "$pid" 2>/dev/null && [ "$i" -lt 100 ]; do
i=$((i + 1))
sleep 0.1
done
if kill -0 "$pid" 2>/dev/null; then
kill -KILL -- "-$pid" 2>/dev/null || kill -KILL "$pid" 2>/dev/null || true
fi
rm -f "$pid_file"
}
wait_for_browser() {
i=0
while [ "$i" -lt 300 ]; do
if wget -qO /dev/null http://127.0.0.1:9223/json/version; then
return 0
fi
browser_alive || return 1
i=$((i + 1))
sleep 0.1
done
return 1
}
finish_connection() {
lock
rm -f "$connection_marker"
date +%s >"$last_use_file"
unlock
}
connection_signal() {
trap - INT TERM HUP
finish_connection
exit 143
}
connection() {
connection_marker=$connections_dir/$$
lock
: >"$connection_marker"
if ! browser_alive; then
rm -f "$pid_file"
start_browser
fi
date +%s >"$last_use_file"
unlock
trap connection_signal INT TERM HUP
if wait_for_browser; then
if socat STDIO TCP:127.0.0.1:9223; then
result=0
else
result=$?
fi
else
# The client only ever sees a bare connection reset here, so this is
# the sole record that the browser, not the network, was the problem.
echo "browser did not come up; dropping connection" >&2
result=1
fi
finish_connection
return "$result"
}
reaper() {
while :; do
sleep 10
lock
if ! has_connections && browser_alive; then
now=$(date +%s)
last=$(cat "$last_use_file" 2>/dev/null || printf '%s' "$now")
if [ $((now - last)) -ge "$idle_seconds" ]; then
stop_browser
fi
fi
unlock
done
}
if [ "${1:-}" = connection ]; then
connection
exit $?
fi
# The files are process state, not the Chrome profile. The profile is a named
# volume in Compose, so clearance survives both a reap and a container rebuild.
for marker in "$connections_dir"/*; do
[ -e "$marker" ] || continue
rm -f "$marker"
done
rm -f "$pid_file" "$last_use_file"
# Chrome's singleton lock names the hostname and pid that took it, and a
# container rebuild changes both — so a Chrome killed uncleanly (OOM, docker
# kill) leaves a lock the next container reads as "another computer holds this
# profile" and refuses to start behind, permanently, with the only symptom a
# bare connection reset at 9222. Clearing it here is safe precisely because
# container_name pins this volume to one container: nothing can be holding the
# profile at the moment this line runs. The lock is process state; the
# clearance cookies it sits beside are not, and are left alone.
rm -f "$profile"/Singleton*
# Chrome binds DevTools to loopback and silently ignores
# --remote-debugging-address. socat remains the network front-end, but each
# accepted connection now starts a browser on demand and is tracked by a
# per-helper marker. A connection held by Go's transport delays reap by its
# idle timeout; the 300-second threshold starts once the last connection closes.
reaper &
reaper_pid=$!
socat TCP-LISTEN:9222,reuseaddr,fork EXEC:'/entrypoint.sh connection',nofork &
front_pid=$!
stop_browser_gracefully() {
lock
if browser_alive; then
# Chrome is a separate session, so stop its process group explicitly;
# this gives its cookie batch time to flush before the container exits.
stop_browser
fi
unlock
}
shutdown() {
trap - INT TERM HUP
stop_browser_gracefully
kill "$front_pid" "$reaper_pid" 2>/dev/null || true
exit 143
}
trap shutdown INT TERM HUP
if wait "$front_pid"; then
status=0
else
status=$?
fi
stop_browser_gracefully
kill "$reaper_pid" 2>/dev/null || true
exit "$status"
+10 -20
View File
@@ -17,24 +17,12 @@ services:
bookmark-api:
# Traffic arrives over the Traefik network, not a published port.
ports: !reset []
environment:
# Must be an IP, not the DNS name — see the base file's comment on this
# same key: Chrome's DevTools HTTP handler 500s any Host header that
# isn't an IP or "localhost".
BROWSER_WS_URL: ${BROWSER_WS_URL:-ws://172.28.0.10:9222}
depends_on:
headless-shell:
condition: service_started
postgres:
condition: service_healthy
# `networks:` here replaces the base file's list entirely, so all three must
# be named: `proxy` for Traefik routing, and `browser` / `db` (defined in
# the base file) to keep reaching headless-shell and Postgres without
# putting either on `proxy`.
# Compose *merges* this list with the base file's, so the service ends up on
# `default`, `db` and `proxy` — only the addition is named here. Do not
# "tidy" the base file down to `db` on the strength of `proxy` being present:
# `db` is `internal: true`, and egress comes from `default`.
networks:
- proxy
- browser
- db
labels:
- "traefik.enable=true"
- "traefik.docker.network=${PROXY_NETWORK:-proxy}"
@@ -52,10 +40,12 @@ services:
- "traefik.http.routers.bmweb.tls.certresolver=${TRAEFIK_CERTRESOLVER:-le}"
- "traefik.http.routers.bmweb.service=bmapi"
# headless-shell is untouched here: it keeps its `browser` network membership
# from the base file and must never join `proxy` — that network is shared
# with whatever else sits behind Traefik on this host, and an exposed
# CDP endpoint on it would be remote code execution for any of them.
# No browser service here. It runs on the home machine as its own unit
# (chrome/docker-compose.yml) and is reached over the tailnet — see
# docs/adr/0006-browser-on-the-home-machine.md. It must never be given a
# service on this host: `proxy` is shared with whatever else sits behind
# Traefik, and an unauthenticated CDP endpoint on it is remote code
# execution for any of them.
networks:
proxy:
+37 -57
View File
@@ -5,10 +5,16 @@
# If your proxy runs in Docker on its own network, use the prod override which
# attaches to that network instead of publishing a port:
# docker compose -f docker-compose.yml -f docker-compose.prod.yml up -d
#
# The browser is not here. It is its own unit on the home machine —
# chrome/docker-compose.yml — reached over the tailnet via BROWSER_WS_URL.
services:
bookmark-api:
build: ./backend
build:
context: ./backend
args:
COVER_DIR: ${COVER_DIR:?set COVER_DIR in .env}
image: bookmarkmanager-backend:latest
container_name: bookmark-api
restart: unless-stopped
@@ -19,15 +25,22 @@ services:
# Owner's Discord user ID — required. Seeds the owner Reader (the
# administrator); every other Reader registers on their first login.
OWNER_DISCORD_ID: ${OWNER_DISCORD_ID:?set OWNER_DISCORD_ID in .env}
ALLOWED_ORIGINS: ${ALLOWED_ORIGINS:-https://asuracomic.net,https://asurascans.com,https://demonicscans.org,https://comix.to,https://kagane.to,https://novelfull.com,https://lightnovelworld.net}
ALLOWED_ORIGINS: ${ALLOWED_ORIGINS:-https://asurascans.com,https://demonicscans.org,https://comix.to,https://kagane.to,https://novelfull.com,https://lightnovelworld.net}
# The bookmarks database. Host is the compose service name; the password
# comes from .env so it is never committed.
DATABASE_URL: ${DATABASE_URL:-postgres://bookmarks:${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD in .env}@postgres:5432/bookmarks?sslmode=disable}
# Required path inside the API. The build seeds ownership at this path
# and the named volume below mounts there.
COVER_DIR: ${COVER_DIR:?set COVER_DIR in .env}
# Public origin of this deployment, no trailing slash. Required: the
# Cover URLs on the wire are absolute, since the userscript renders them
# on a Site's origin rather than ours (ADR-0007).
PUBLIC_BASE_URL: ${PUBLIC_BASE_URL:?set PUBLIC_BASE_URL in .env}
PORT: "8080"
# Log timestamps only. Go's `log` stamps lines in local time, and this
# service has no other use for a zone: bookmark timestamps are unix ms
# and the two real time columns are timestamptz, both absolute instants.
# Purely so these lines read on the same clock as the sidecar's. Named
# Purely so these lines read on the same clock as the browser's. Named
# API_TZ rather than TZ so an operator's exported shell TZ cannot leak
# in; distroless already carries tzdata, so the name just resolves.
TZ: ${API_TZ:-Asia/Jakarta}
@@ -49,21 +62,24 @@ services:
# switch; it only takes effect because these are listed here.
LATEST_CHAPTER_POLL_ENABLED: ${LATEST_CHAPTER_POLL_ENABLED:-1}
LATEST_CHAPTER_POLL_COOLDOWN: ${LATEST_CHAPTER_POLL_COOLDOWN:-1h}
LATEST_CHAPTER_POLL_BROWSER_COOLDOWN: ${LATEST_CHAPTER_POLL_BROWSER_COOLDOWN:-6h}
LATEST_CHAPTER_POLL_INTERVAL: ${LATEST_CHAPTER_POLL_INTERVAL:-10m}
LATEST_CHAPTER_POLL_BATCH: ${LATEST_CHAPTER_POLL_BATCH:-14}
LATEST_CHAPTER_POLL_STAGGER: ${LATEST_CHAPTER_POLL_STAGGER:-20s}
# CDP endpoint for sites behind a JavaScript challenge (kagane). Unset
# disables browser polling for those sites; the userscript still covers them.
# Must be an IP, not the "headless-shell" DNS name: Chrome's DevTools HTTP
# handler rejects the discovery request (GET /json/version) with a 500
# unless the Host header is an IP address or "localhost" — confirmed
# 2026-08-03 against chromedp/headless-shell:stable, independent of
# chromedp's own dial logic. The sidecar's static address below exists so
# this URL survives container recreation.
BROWSER_WS_URL: ${BROWSER_WS_URL:-ws://172.28.0.10:9222}
# CDP endpoint for sites behind a JavaScript challenge (kagane,
# novelfull). The browser is not part of this stack — it runs on the home
# machine as its own unit (chrome/docker-compose.yml) and is reached over
# the tailnet. Unset disables browser polling for those sites and serves
# 404 from the cover proxy for covers not already stored; the userscript
# still covers them. Set it in .env to ws://<home machine tailnet IP>:9222.
#
# Must be an IP, not a MagicDNS hostname: Chrome's DevTools HTTP handler
# rejects the discovery request (GET /json/version) with a 500 unless the
# Host header is an IP address or "localhost" — confirmed 2026-08-03,
# independent of chromedp's own dial logic. The same trap that used to
# force a pinned Docker IP now forbids the tailnet name.
BROWSER_WS_URL: ${BROWSER_WS_URL:-}
depends_on:
headless-shell:
condition: service_started
# The migration runner is the first thing the binary does, so a Postgres
# that is still initialising means a crash-loop until it is not.
postgres:
@@ -74,12 +90,17 @@ services:
# no rebuild, no restart. `git pull` restores the committed version, which
# is why a redeploy always ships the repo's script.
- ./userscript:/userscript:ro
# Content-addressed cover bytes survive API restarts and redeploys.
- cover-data:${COVER_DIR:?set COVER_DIR in .env}
# Bound to loopback only: the proxy (or curl during smoke test) reaches it,
# the public internet does not.
ports:
- "127.0.0.1:8080:8080"
# `default` is not decoration: `db` is `internal: true`, and a container on
# nothing but an internal network gets neither a published port nor egress
# — which would silently kill every poller fetch.
networks:
- browser
- default
- db
postgres:
@@ -101,56 +122,15 @@ services:
networks:
- db
headless-shell:
# Real Google Chrome, not chromedp/headless-shell — see chrome/Dockerfile.
# The service name is kept so existing overrides and BROWSER_WS_URL stay put.
build: ./chrome
image: bookmarkmanager-chrome:latest
restart: unless-stopped
environment:
# A UTC clock is itself the bot signal: Cloudflare treats it as the
# datacenter default, and kagane's challenge then never clears. Measured
# 2026-08-08, identical container, one Indonesian egress IP: UTC never
# cleared in 60s (twice); Asia/Jakarta and America/New_York both cleared
# in 4s. So any real zone works and it need not match the IP's country —
# only UTC fails. Unset falls back to the host's /etc/timezone below,
# which is a real zone whenever the host clock is set to local time; set
# BROWSER_TZ when the host runs UTC.
TZ: ${BROWSER_TZ:-}
volumes:
# The zone *name*, which is what Chrome's ICU needs — see chrome/entrypoint.sh.
# Absent on a non-Debian host, which the entrypoint handles by falling back to UTC.
- /etc/timezone:/etc/timezone:ro
# Chrome allocates shared memory per tab and dies on Docker's 64MB default.
shm_size: '1gb'
# Reaps zombie renderer processes, which otherwise accumulate for the
# container's lifetime.
init: true
# Deliberately no `ports:` — an exposed CDP endpoint is remote code
# execution. Only bookmark-api, via the `browser` network below, may reach it.
# No `command:` either: every flag this browser needs is in its entrypoint,
# and the UA override there is load-bearing for the challenge.
networks:
browser:
# Pinned so BROWSER_WS_URL can name an IP (required, see above) that
# survives `docker compose up` recreating this container.
ipv4_address: 172.28.0.10
volumes:
postgres-data:
cover-data:
# The pre-Postgres SQLite volume (bookmarks-data) is deliberately no longer
# declared here: undeclared means `docker compose down -v` cannot take it
# with the rest, so the old database survives the cutover until someone
# removes it by hand.
networks:
# Not `internal: true`: headless Chrome still needs outbound access to reach
# kagane.to. Isolation here comes from membership (only bookmark-api and
# headless-shell join it), not from cutting egress.
browser:
ipam:
config:
- subnet: 172.28.0.0/24
# Postgres needs no egress and nothing outside bookmark-api needs to reach
# it, so this one really can be cut off from the outside world.
db:
+51
View File
@@ -0,0 +1,51 @@
# ADR-0005: On-demand browser sidecar
Date: 2026-08-09
Status: accepted
Superseded in part by ADR-0006: the lifecycle below is unchanged, but the
service no longer lives in the API stack and the name `headless-shell` is gone.
## Decision
Keep the `headless-shell` service and its CDP port alive, but start Google Chrome
only when the first CDP connection arrives. The entrypoint supervises a `socat`
front-end, serializes browser start/reap state with `flock`, and tracks each
connection with a marker named for its helper PID. A reaper stops Chrome after
300 seconds with no live markers. Marker reconciliation covers a helper killed
before its cleanup trap runs.
Chrome runs in its own process group so reap sends the termination signal to
Chrome and its renderer children. The explicit `/home/chrome/profile` user-data
directory remains: Chrome remaps remote debugging to loopback on modern builds,
and Chrome ignores remote-debugging flags on a default profile. `socat` therefore
continues to front Chrome's loopback CDP port.
The profile is a named Compose volume. Clearance cookies survive both a reap and
`docker compose up --build`; the browser still starts with a fresh debugger UUID,
so chromedp must keep endpoint discovery enabled and must not use
`chromedp.NoModifyURL`.
The socat front-end and explicit profile are retained because Chromium remaps a
non-loopback debugging address to loopback since M113, while Chrome ignores the
remote-debugging flags on a default profile since Chrome 136. Flag tuning is
deliberately not adopted: its roughly 30% idle-footprint saving is irrelevant
to a browser that exists for seconds per wake and risks an untested fingerprint.
## Constraints
The 300-second floor is deliberate. Chromium batches cookie persistence on a
roughly 31-second timer, and Go's default HTTP transport can keep the discovery
connection parked for about 90 seconds after use. Reaping only with zero live
connections holds Chrome through both windows and through the poller's staggered
batch plus cover prefetch.
The anti-bot properties remain unchanged: a plausible non-UTC timezone, a
Chrome-version-derived User-Agent without `HeadlessChrome`, and no automation
flag. A remote browser restart can surface as `context.Canceled`, the same error
as a caller deadline, so the backend wraps cancellation observed with a closed
CDP connection as `browser interrupted`; the focused test asserts that
classification without killing a real browser.
The same process-group stop runs during supervisor shutdown, not only during
idle reap, so Chrome can flush its cookie batch before a container rebuild or
graceful stop.
@@ -0,0 +1,74 @@
# ADR-0006: The browser runs on the home machine, over the tailnet
Date: 2026-08-09
Status: accepted
## Decision
The headless browser is no longer part of the API stack. It is its own compose
unit (`chrome/docker-compose.yml`), deployed on the home machine, and the API on
the VPS reaches it over the existing tailnet through `BROWSER_WS_URL`. No
fallback sidecar remains on the VPS.
The backend needs no code change for this. The CDP endpoint was already a
configuration seam and the fetcher only ever holds the endpoint URL, so
relocation — and reversal — is one environment variable.
## Why
The sidecar held 471 MiB working set (645 MiB peak) on a 1974 MiB VPS with no
swap, which also hosts Traefik, Gitea and its Postgres. That is 24% of the host
and 86% of this project's memory, for a service that at the time answered zero
requests: the poller's due query joins bookmarks, production held four kagane
series and no bookmarks on any of them, and with no kagane bookmark the web UI
never rendered a kagane cover either.
The home machine has 5.9 GiB of swap and a residential egress, which Cloudflare
scores better than a datacenter IP. Both machines were already on the tailnet.
This move is only safe because covers are persisted (ADR-0005's sibling work,
issue #43/#45) and the browser is on-demand (ADR-0005). Without stored covers a
sleeping home machine would blank the library; without on-demand start the CI
runner that already holds ~1.2 GiB of that box's 1.8 GiB would be squeezed
around the clock.
## Constraints
**`BROWSER_WS_URL` must be the tailnet IP, never a MagicDNS hostname.** Chrome's
DevTools HTTP handler answers `/json/version` with a 500 for any `Host` header
that is not an IP or `localhost`. This is the same trap that previously forced a
pinned Docker IP; the pinned subnet is gone, the constraint is not.
**The CDP port binds to the tailnet address only, never `0.0.0.0`.** CDP
authenticates nothing: whatever reaches the port drives the browser and, through
it, the host. On the VPS the safety came from Docker network membership; the
home machine has a real LAN, so a `0.0.0.0` bind is a hole punched into it. The
bind address is the enforcement and Tailscale device identity plus a per-device
ACL is the policy. `BROWSER_BIND_ADDR` deliberately has no default, so an unset
value fails the deploy instead of publishing CDP to the LAN.
No bearer-token proxy is added in front of CDP. It would only defend against a
device already inside the tailnet, and it would be one more thing between the
poller and a browser that is already hard enough to keep clearing challenges.
**Resource limits are load-bearing, not decorative.** The browser is the
newcomer on that box, not the incumbent. A hard 512 MiB cap with 1 GiB
memory+swap makes Chrome reclaim its own cold pages onto the machine's SATA swap
instead of taking resident memory from the runner; untuned Chrome peaked at
645 MiB cgroup, which is more than is free there. `oom_score_adj` biases the
kernel to kill the browser first and never CI. Reduced CPU weight makes a
challenge solve yield to a running build — cold start degrades to about 3 s at
half a CPU, immaterial against a 45-second challenge budget. The shared-memory
reservation drops from 1 GiB to 128 MiB against a measured 19 MiB peak.
## Consequences
An unreachable browser degrades exactly as an unset `BROWSER_WS_URL` already
does: plain-TLS libraries are unaffected, kagane and novelfull log and skip, the
series waits out its cooldown, and stored covers keep serving. A power outage at
home costs chapter freshness on two sites, never the appearance of the library.
The two units are deployed and updated independently. `REDEPLOY.md` §8 covers
the browser; everything before it covers the API stack. A local `docker compose
up` now brings up two services, not three, and polls kagane only if
`BROWSER_WS_URL` is pointed somewhere.
+105
View File
@@ -0,0 +1,105 @@
# ADR-0007: The backend hosts every Site's Cover bytes
Date: 2026-08-09
Status: accepted
## Decision
A Cover is the image a Reader's browser can display for a Series. Where the Site
keeps the picture is the backend's problem, not the client's: the backend fetches
the bytes, stores them, and serves them from its own origin. No client ever
renders a third-party URL, and no client ever supplies one.
Concretely:
- **Acquisition is server-side.** The Cover is extracted from the same series-page
fetch that already yields Latest Chapter. It runs once at Series creation rather
than waiting for the poll queue, so a newly bookmarked Series has both facts in
seconds instead of up to a queue's depth. The poll fills a blank Cover and never
overwrites a non-blank one.
- **Bytes live on a filesystem volume**, content-addressed by the SHA-256 of the
source URL, sharded `${COVER_DIR}/ab/cd/<sha256>`. The database holds the path
and content type, not the bytes.
- **One public route** serves them. No session, no credential.
- **The wire carries an absolute URL** built from a configured public base, and
carries `""` until the bytes exist.
## Why a future reader will find this surprising
Four of the six Sites let anyone hot-link their covers — `static.comix.to` even
answers `access-control-allow-origin: *`. Hosting copies looks like work we were
not obliged to do.
We were obliged. kagane serves covers with `cross-origin-resource-policy:
same-origin` behind a JavaScript challenge (measured 2026-08-08), so no `<img>`
outside kagane.to can load one under any combination of referrer policy and
`crossorigin` attribute. The first fix for that was a kagane-only proxy applied in
the web templates — and it produced issue #47, because the JSON API kept emitting
the raw kagane URL and the userscript rendered it into a broken-image glyph. A
per-Site exception that only one of two clients knows about is not a fix; it is a
bug with a delay on it. Uniformity is the property being bought: every client
renders every Cover the same way, and a Site changing its CORP header or its CDN
cannot break a client again.
## Considered options
**Per-Site exceptions, proxying only what must be proxied.** Cheapest, and what we
had. Rejected: it is what produced #47, and it requires every current and future
client to know which Sites are special.
**A host allowlist for the outbound fetch**, mirroring `fetchableSeriesURL`.
Rejected in favour of destination-class control — see below.
**Cover bytes in Postgres `bytea`**, extending the existing `covers` table.
Rejected: covers are immutable blobs served straight to browsers, which is what a
filesystem is for. The cost is real and accepted — durability is now two things to
back up instead of one, against ADR-0001's grain.
**Per-Reader Cover overrides.** Rejected, consistent with ADR-0003's rejection of
per-Reader title overrides. A Cover is a fact about the Series.
## Two deliberate relaxations
**Destination control is deny-class, not an allowlist.** The outbound fetch
requires `https`, resolves DNS first and refuses loopback, private, link-local and
CGNAT addresses, re-checks on every redirect hop, and caps body size and content
type. It does *not* pin a host set, which is what `fetchableSeriesURL` does for
`series_url`. Cover hosts are CDNs that move: `demonicscans.org` serves its covers
from `readermc.org`, a host with no visible relationship to the Site. An allowlist
would silently stop producing Covers the day a Site switched CDN, and the failure
would look like this bug. The resolved-IP check is the load-bearing part; without
it, an attacker-controlled page need only publish a DNS name pointing at
`127.0.0.1`.
**The cover route is public, where the kagane proxy was session-gated.** An `<img>`
in the userscript panel cannot send a bearer token, and it cannot be given one: the
panel's shadow root is `mode: "open"`, so the host page's own JavaScript can read
any `src` we set. A credential in an image URL is a credential handed to a
third-party site. The route serves public artwork from public Sites and its path
reveals nothing about which Reader holds what. The residual cost is that we can be
hot-linked by others.
## Consequences
- The poll's cover prefetch, today guarded by `sr.Site != "kagane"`, applies to
every Site in both Libraries. Nothing about Covers is conditioned on `kind`.
- Only kagane still needs the CDP browser for its bytes. The other five Sites fetch
over plain TLS — including novelfull, whose HTML answers `cf-mitigated: challenge`
while its image paths answer 200 with `access-control-allow-origin: *`
(measured 2026-08-09).
- Client-side cover scraping is deleted from both userscripts. It could not help: a
scraped URL has no render path left, and it is absent exactly when a Series is
created — neither comix nor lightnovelworld exposes a cover on a chapter page,
which is where a Reader bookmarks mid-read.
- `PUT /bookmarks/{key}` still accepts a `cover` field and ignores it. This extends
ADR-0003's "ignored after creation" to "ignored always", and keeps the flat wire
contract ADR-0004 requires so installed scripts keep working. The field is
therefore permanently inert rather than pending removal, and says so at the
decode site.
- A Cover that fails to load falls back to the placeholder in both clients. The
broken-image glyph reported in #47 is not a state we render.
- The existing kagane `covers` rows are dropped rather than migrated; that path
re-fetches on demand already.
- Two Sites deserve a note for whoever writes the extractor: asura's `.webp` cover
URL answers `Content-Type: image/jpeg`, so trust the header; demonic's `og:image`
carries a raw unencoded space and must be percent-encoded before fetching.
@@ -0,0 +1,102 @@
# ADR-0008: A Series identity is discovered from the Site's links, never derived from an address
Date: 2026-08-11
Status: accepted
## Decision
On lightnovelworld, a Series is identified by the slug in its `/novel/<slug>/`
address, and the userscript obtains that address by reading the chapter page's
`a[aria-label="All Chapter"]` anchor. It no longer constructs the address by
string manipulation of the chapter path. When neither that anchor nor the
microdata breadcrumb is present, the page resolves to `type: "other"` and no
Bookmark is offered.
A Chapter Slug — the slug a chapter address is built from — is not an identity
and is not stored. The backend finds chapters by matching
`lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/` against the series page
body **truncated at the first `wpd-threads`**.
## Why
Measured on lightnovelworld between 2026-08-10 and 2026-08-11.
The chapter slug and the series slug are two independent facts. In a 41-novel
sample, 3 diverged (~7%). `/my-longevity-simulation-chapter-1/` returns 200
while `/novel/my-longevity-simulation/` returns 404 with no redirect, and that
novel's real address is `/novel/immortality-simulator/`. Divergence runs in both
directions: `the-sword-illuminates-the-great-wilderness` is served by chapter
slug `radiant-blade-of-the-wilderness`. Neither slug is computable from the
other, and the Site publishes no alternative-names field, so the mapping exists
only in the chapter page's own markup.
Storing the Chapter Slug beside the identity does not work, because a Series may
have more than one. `/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/`
serves chapters 1–99 under `…-one-skill-not` and 100–423 under
`…-one-skill-not-them-all`. Both resolve, and both chapter pages point back at
the same Series.
Deriving the identity from the chapter path also made one Series produce two
rows: bookmarking from the series page yielded `lightnovelworld:immortality-simulator`,
and from a chapter page `lightnovelworld:my-longevity-simulation`.
The pointer is reliable. Across 8 chapter pages — chapter 1, chapter 1200, the
latest chapter, both slugs of the split novel, two divergent novels, and a novel
with a number in its title — the `All Chapter` anchor and breadcrumb position 2
were both present and agreed every time, including on the old-slug pages. Three
narrower selectors were rejected on evidence: `a[href*="/novel/"]` matches the
header nav index first; matching the text "All Chapter" false-matches the novel
titled "…Not Them All Chapter 200"; and the JSON-LD breadcrumb's position 2 is
the chapter, not the Series.
The scan is truncated because a series page server-renders a wpdiscuz comment
thread below the chapter list, and comment bodies are HTML that can carry an
anchor. Verified on `/novel/the-sword-illuminates-the-great-wilderness/`:
comment `#wpd-comm-358_0` rendered in the initial HTML, corroborated by that
page's comment RSS feed. The scanner takes the maximum chapter number with no
upper bound, and a Series row is shared by every Reader (ADR-0003), so one
comment containing a link to a high-numbered chapter would pin that Series'
Latest Chapter for everyone. `wpd-threads` occurs exactly once per page and
follows every chapter anchor on all 4 series pages measured. `wpdcom` and
`wpdiscuz` are unusable: they occur 111 to 143 times per page, including in
`<head>` before the chapter list.
## Considered options
**Correct `series_url` only, leaving the Chapter Slug as the identity.** This
repairs the Poll with no migration, because the Poll reads the stored address
rather than the identity. Rejected: it keeps an identity that the Site does not
guarantee to be stable, and leaves the duplicate-row hazard in place.
**Scope the match to the chapter-list container.** Rejected on measurement. The
prior research named `ul.clstyle`; on all 4 pages sampled that is the hidden,
empty "Latest Reading" template, and the real list is a classless `<ul>` in
`div.eplister.eplisterfull`. The backend scans raw HTML with no parser, so
container extraction means a second regex against class names, which is more
fragile than the one-off truncation and protects nothing extra.
**Repair the stored rows with a SQL migration.** Rejected as impossible: the
database holds no source for the correct slug. The `series` table has no chapter
address, and no endpoint on the Site maps one slug to the other.
## Consequences
Existing Bookmarks on divergent novels stop matching their own chapter pages,
because `keyOf` changes. The userscript therefore migrates a row in place when
it sees the mismatch: it rewrites the row's key, identity and address in the
cache, the retry-queue entry, and the last-checked map, then syncs. This heals
only when the Reader next opens a chapter page of that novel.
A migrated row leaves its old `series` row behind. Nothing deletes it, but
`DueForLatestCheck` is an inner join against bookmarks, so a Series with no
Bookmarks is never polled again. The old row is permanently stored and
permanently inert.
A Bookmark whose Reader never opens a chapter page of that novel does not heal.
It continues to return 404 at every cooldown, as it does today.
If the truncation marker disappears, the scan is skipped and logged rather than
run against the whole page. An env-gated live test, `TestSmokeLnwCommentBoundary`,
asserts that `wpd-threads` still occurs exactly once and still follows the last
chapter anchor. It skips when its environment variable is unset, matching the
existing `TestSmokeKagane*` convention.
@@ -0,0 +1,93 @@
# ADR-0009: A Site answers fixed questions; an unusual Site owns its own fetch
Date: 2026-08-11
Status: accepted
## Decision
Per-Site knowledge lives in one registry in `backend/internal/latest/sites.go`,
keyed by the stored site string. An entry answers a fixed set of questions: the
hostname a `series_url` must carry, how to find the Latest Chapter in a body,
how to find the Cover address in a body, and — for a Site behind a JavaScript
challenge — how to read its payload from a cleared tab, how to tell that the
payload arrived, and whether the plain-TLS fetcher may take over when no
browser is configured (`Fallback`; kagane never falls back, novelfull does,
each on measured evidence).
The question set does not grow to accommodate one Site. When a Site needs
something the set cannot express, that Site gets an optional override and
performs its own fetch, leaving the other entries untouched. The override is a
per-Site escape hatch, not a stage every Site passes through, and it is added to
the registry type when the first Site needs it rather than in advance.
Two things stay outside the registry. Cover *bytes* are routed by address shape
in `fetchCoverBytes`, never by Site name, so the Poll and the Acquisition cannot
drift apart. And the host pin inside a browser entry's read is kept even though
`fetchableSeriesURL` has already pinned the same host: a headless browser is a
strong SSRF primitive and `series_url` arrives in a client-supplied PUT body, so
the second check is deliberate and must not be deduplicated.
## Why
Before the registry, the site string was compared in six places across three
files: the Latest Chapter switch (`sites.go:102`), the Cover switch
(`sites.go:263`), the browser-backed list and the fetcher choice
(`poller.go:60`, `poller.go:132`), the host pins (`poller.go:314-325`), and the
payload read (`browser.go:96-119`). Nothing tied them together, so adding a
seventh Site meant finding all six unaided, and a Site added to five of them
failed at the sixth in production rather than at compile time.
The Sites are not alike and the registry does not ask them to be. asura strips a
rotating build hash from its slug before scoping a regex; comix reads a JSON
blob embedded in server-rendered HTML; kagane's chapter list exists only in its
JSON API, which must be called from inside the page so the request carries the
clearance cookie; lightnovelworld must truncate the body at the comment thread
first. What they have in common is not behaviour, it is the questions they
answer. Arbitrary behaviour behind one entry is the point.
Making a browser Site contribute a read and a completion test, rather than
letting it drive the browser, was chosen because the tab lifecycle in
`BrowserFetcher.run` is load-bearing and shared. It holds one tab open across
re-reads, because a Cloudflare interstitial needs several seconds of live page
to solve itself and write clearance into the shared cookie jar; reading once and
closing the tab, which is what this did before 2026-08-08, never clears
anything. It also serialises the browser, binds the caller's deadline to the
tab, distinguishes a lost browser from a retryable read, and paces re-reads.
Spreading that across per-Site adapters would put one subtle, measured loop
behind six doors.
This costs the adapters little, because `chromedp.Run` takes an Action and
`chromedp.Tasks` is an Action. A Site that must click, wait on a selector, and
then evaluate expresses all of it as its read. Only a Site needing something
outside the per-tab loop — its own cadence, two tabs, a tab held between calls,
cookies set before navigation — falls outside, and that Site takes the override.
## Considered options
**Widen the shared interface whenever a Site needs something new.** Rejected:
one Site's requirement becomes a field on all seven entries, and the entries
that ignore it still have to be read and understood by anyone adding the eighth.
**Give every Site the whole fetch.** Rejected: it makes the browser lifecycle
above a per-Site concern, and pulls `chromedp` into adapters for five Sites that
never open a browser.
**A Go `interface` with a method set instead of a registry of records.**
Rejected: most Sites differ in one or two answers, and three share a single
Cover implementation, so a method set produces near-empty types. A missing
answer is a nil value caught at dispatch, which is where an unknown Site is
already handled.
## Consequences
Adding a Site is one registry entry. The existing dispatch functions —
`latestChapterFrom`, `coverFrom`, `fetchableSeriesURL` — become registry
lookups, so the table tests that drive them by site string are unchanged.
A future architecture review will see an override that only one Site uses and
read it as an inconsistency to collapse. It is not. Collapsing it means either
widening the question set for every Site or moving the shared tab lifecycle into
the adapters, and both were rejected here on the evidence above.
An unknown site string resolves to the zero entry and fails the existing
not-fetchable and no-fetcher paths, which log and skip. That is unchanged.
@@ -0,0 +1,572 @@
# lightnovelworld.net — chapter slug vs. series slug
Research note for Gitea issue #77. All pages fetched live on **2026-08-11** with
plain `curl` and a desktop UA. **No Cloudflare challenge was encountered on any
request** — every fetch below returned real HTML on the first try, so the
Playwright fallback was never needed.
Every claim carries the URL it came from. Nothing here is inferred from the
existing code; where a claim is an interpretation rather than an observation it
is marked `[INFERENCE]`.
---
## 1. Summary answer table
| Question | Answer | Evidence |
|---|---|---|
| Does the chapter slug equal the series slug? | **No, not reliably.** 3 of 41 sampled novels diverge at their latest chapter (~7%); a 4th diverges historically. | §4 |
| Best chapter-page → series-URL pointer | `a[aria-label='All Chapter']` and the microdata breadcrumb `a[itemprop=item]` at `position 2`. Both carry the full absolute series URL. | §3 |
| Is JSON-LD `BreadcrumbList` usable? | **No.** The Rank Math JSON-LD breadcrumb on a chapter page contains only *Home* and *the chapter itself* — the series is absent. | §3.1 |
| Is `<link rel=canonical>` usable? | **No.** It points at the chapter page itself. | §3.2 |
| Are `og:` tags usable? | **No.** `og:url` is the chapter page; no `og:` tag names the series. | §3.3 |
| Does the series page list chapters in the initial HTML? | **Yes — all of them.** 1232 unique chapters (1…1232) present in one document. | §5 |
| Is there a paginated chapter endpoint? | **No.** `/novel/<slug>/chapters/` and `/novel/<slug>/chapters/page-2` both **404**. | §5 |
| Does the series page carry *other* novels' chapter anchors? | **No.** Across 30 series pages, every `-chapter-N/` URL belonged to the page's own novel. | §6 |
| Unscoped regex `lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/` on a series page | **SAFE** | §6 |
| Does `/novel/<chapter-slug>/` 404? | **Yes**, 404, no redirect. | §7 |
| Is there a chapter-slug → series-slug endpoint that avoids parsing a chapter page? | **No endpoint found.** But `/<chapter-slug>/` (no `-chapter-N`) **302s to chapter 1**, which still requires parsing a chapter page. | §7 |
| Is there an "Alternative names" field explaining divergence? | **No such field exists** on lightnovelworld series pages. | §4.3 |
---
## 2. The two reference pages
| URL | Status |
|---|---|
| `https://lightnovelworld.net/my-longevity-simulation-chapter-1/` | **200**, 0 redirects |
| `https://lightnovelworld.net/novel/immortality-simulator/` | **200**, 0 redirects |
| `https://lightnovelworld.net/novel/my-longevity-simulation/` | **404**, 0 redirects |
Confirmed: chapter slug `my-longevity-simulation`, series slug
`immortality-simulator`, and the naive `/novel/<chapter-slug>/` construction is
a hard 404.
---
## 3. Chapter page → series URL: every in-page pointer, in priority order
All snippets below are verbatim from
`https://lightnovelworld.net/my-longevity-simulation-chapter-1/`.
### Priority 1 — `a[aria-label='All Chapter']` ✅ WORKS
Single occurrence in the document, inside the chapter navigation bar:
```html
<div class="nvs nvsc"><a href='https://lightnovelworld.net/novel/immortality-simulator/' aria-label='All Chapter'><i class="fas fa-list-ul"></i> All Chapter</a></div>
```
Note the **single quotes** on both attributes — a regex written for `href="` will
miss it. This is the most narrowly-targeted pointer: exactly one element on the
page has `aria-label='All Chapter'`.
### Priority 2 — microdata `BreadcrumbList` / `a[itemprop=item]` ✅ WORKS
At line 363–379 of the served HTML. `position 2` is the series:
```html
<div itemscope="" itemtype="http://schema.org/BreadcrumbList">
<span itemprop="itemListElement" itemscope="" itemtype="http://schema.org/ListItem">
<a itemprop="item" href="https://lightnovelworld.net/"><span itemprop="name">Home</span></a>
<meta itemprop="position" content="1">
</span>
›
<span itemprop="itemListElement" itemscope="" itemtype="http://schema.org/ListItem">
<a itemprop="item" href="https://lightnovelworld.net/novel/immortality-simulator/"><span itemprop="name">Immortality Simulator</span></a>
<meta itemprop="position" content="2">
</span>
›
<span itemprop="itemListElement" itemscope="" itemtype="http://schema.org/ListItem">
<a itemprop="item" href="https://lightnovelworld.net/my-longevity-simulation-chapter-1/"><span itemprop="name">My Longevity Simulation Chapter 1</span></a>
<meta itemprop="position" content="3">
</span>
</div>
```
Bonus: `<span itemprop="name">Immortality Simulator</span>` gives the clean
series **title** as well, which the userscript currently derives by stripping
`Chapter <n>` off `h1.entry-title`.
Caveat: `itemprop="item"` is **not** unique to the breadcrumb — the site header
nav also uses `itemprop="url"` on anchors, and unrelated `itemprop="item"` usage
could appear. Scope any selector to the enclosing
`[itemtype="http://schema.org/BreadcrumbList"]`.
### Priority 3 — corroborating `/novel/` anchors ⚠️ AMBIGUOUS
The chapter page contains exactly **8** distinct `/novel/…` hrefs. Two are
navigational (`/novel/` index), one is the series, and **five are a
"recommended" strip of unrelated novels**:
```
href="/novel/"
href="https://lightnovelworld.net/novel/"
href="https://lightnovelworld.net/novel/dreadful-radio-game/"
href="https://lightnovelworld.net/novel/evil-god-average/"
href="https://lightnovelworld.net/novel/immortality-simulator/"
href="https://lightnovelworld.net/novel/omniscient-first-persons-viewpoint/"
href="https://lightnovelworld.net/novel/reverend-insanity/"
href="https://lightnovelworld.net/novel/surviving-in-a-romance-fantasy-novel/"
```
So **"grab the first `/novel/<slug>/` link on a chapter page" is unsafe** — the
recommendation strip contaminates it. Use Priority 1 or 2.
### ❌ NOT usable — JSON-LD `BreadcrumbList` (§3.1)
There is exactly one `application/ld+json` block on the page (Rank Math SEO).
Verbatim:
```html
<script type="application/ld+json" class="rank-math-schema">{"@context":"https://schema.org","@graph":[{"@type":"BreadcrumbList","@id":"https://lightnovelworld.net/my-longevity-simulation-chapter-1/#breadcrumb","itemListElement":[{"@type":"ListItem","position":"1","item":{"@id":"https://lightnovelworld.net","name":"Home"}},{"@type":"ListItem","position":"2","item":{"@id":"https://lightnovelworld.net/my-longevity-simulation-chapter-1/","name":"My Longevity Simulation Chapter 1"}}]}]}</script>
```
Only two items — Home and the chapter. **The series does not appear.** The
JSON-LD breadcrumb and the visible microdata breadcrumb disagree; the microdata
one is the richer of the two. Do not use the JSON-LD.
### ❌ NOT usable — `<link rel=canonical>` (§3.2)
```html
<link rel="canonical" href="https://lightnovelworld.net/my-longevity-simulation-chapter-1/" />
```
Self-referential.
### ❌ NOT usable — `og:` meta tags (§3.3)
```html
<meta property="og:type" content="article" />
<meta property="og:title" content="Read My Longevity Simulation Chapter 1 Online Free - LightNovelworld" />
<meta property="og:url" content="https://lightnovelworld.net/my-longevity-simulation-chapter-1/" />
<meta property="og:site_name" content="Light Novel World" />
```
`og:url` is the chapter. `og:title` carries the **chapter-slug-derived** title
("My Longevity Simulation"), *not* the series title ("Immortality Simulator") —
so og tags actively reinforce the wrong name.
### Also present — `rel=next` / `rel=prev` chapter navigation
```html
<a aria-label="next" href="https://lightnovelworld.net/my-longevity-simulation-chapter-2/" rel="next">Next <i class="fa fa-angle-right" aria-hidden="true"></i></a>
```
Chapter 1 has no `rel=prev`. Useful for walking chapters, useless for finding
the series.
---
## 4. The reverse direction, and how common divergence is
### 4.1 Sample method
Two independent samples, deduplicated:
- 14 novels taken from the latest-updates listing on `https://lightnovelworld.net/`
(distinct chapter slugs), each chapter page fetched and its
`aria-label='All Chapter'` href read for the series slug.
- 30 series URLs taken from `https://lightnovelworld.net/novel/`, each series
page fetched and its chapter anchors' slug prefix extracted. Two of the 30
(`/novel/feed/` and `/novel/list-mode/`) are not novels and are excluded,
leaving 28.
One novel (`investing-in-my-crippled-wife-every-return-makes-me-stronger`)
appears in both samples. **Distinct novels sampled: 41.**
### 4.2 Divergence results
| Series slug (`/novel/<…>/`) | Chapter slug (`/<…>-chapter-N/`) | Result |
|---|---|---|
| `immortality-simulator` | `my-longevity-simulation` | **DIVERGE** |
| `the-sword-illuminates-the-great-wilderness` | `radiant-blade-of-the-wilderness` | **DIVERGE** |
| `a-villains-will-to-survive` | `the-villain-wants-to-live` | **DIVERGE** |
| `all-jobs-and-classes-i-just-wanted-one-skill-not-them-all` | `all-jobs-and-classes-i-just-wanted-one-skill-not` (ch. 1–99) **and** `all-jobs-and-classes-i-just-wanted-one-skill-not-them-all` (ch. 100–423) | **SPLIT** (latest matches) |
| `86-eighty-six` | same | match |
| `a-dungeon-beneath-my-house-let-me-gain-800-exp` | same | match |
| `a-journey-of-black-and-red` | same | match |
| `a-knight-who-eternally-regresses` | same | match |
| `a-regressors-tale-of-cultivation` | same | match |
| `a-will-eternal` | same | match |
| `absolute-resonance` | same | match |
| `absolute-sword-sense` | same | match |
| `advent-of-the-three-calamities` | same | match |
| `after-changing-to-the-ruthless-way-the-brothers-cried-and-begged-for-forgiveness` | same | match |
| `against-the-gods` | same | match |
| `apocalypse-i-built-the-infinite-train` | same | match |
| `arcane-exfil` | same | match |
| `as-a-mafia-boss-i-refuse-to-be-an-extra` | same | match |
| `ascendance-of-a-bookworm` | same | match |
| `avatar-conquering-the-elements` | same | match |
| `battle-world-ascending-without-limits` | same | match |
| `became-the-patron-of-villains` | same | match |
| `ending-maker` | same | match |
| `got-dropped-into-a-ghost-story-still-gotta-work` | same | match |
| `greed-all-for-what` | same | match |
| `investing-in-my-crippled-wife-every-return-makes-me-stronger` | same | match |
| `lord-of-mysteries-2-circle-of-inevitability` | same | match |
| `lord-of-the-mysteries` | same | match |
| `my-vampire-system` | same | match |
| `cleaver-of-sin` | same | match |
| `greatest-legacy-of-the-magus-universe` | same | match |
| `magus-infinite` | same | match |
| `path-of-the-extra` | same | match |
| `regnum-aetern-dual-rebirth` | same | match |
| `shadow-slave` | same | match |
| `slime-evolution` | same | match |
| `sss-awakening-i-can-class-change-at-will` | same | match |
| `the-dao-of-reincarnation-ascending-through-a-hundred-lives` | same | match |
| `the-gamers-pov` | same | match |
| `the-insane-regressor-throne-of-pride` | same | match |
| `the-villains-pov` | same | match |
**Counts (41 distinct novels):**
- **37 match** (90.2%)
- **3 diverge at the latest chapter** (7.3%) — these are the ones that break the poller *today*
- **1 splits mid-series** but its newest chapters happen to match the series slug (2.4%)
`[INFERENCE]` The true site-wide divergence rate is probably in the same
ballpark, but 41 novels out of a catalogue of thousands is a small sample and
the listing page is alphabetically front-loaded. Treat ~7% as an order of
magnitude, not a precise figure.
### 4.3 Is there a derivable rule? **No.**
The divergence is **not directional**, so you cannot compute one slug from the
other:
- `immortality-simulator` — the *series* carries the polished English title
(`<h1 class="entry-title" itemprop="name">` is absent from the chapter page's
own naming; `og:title` says "My Longevity Simulation"), while the *chapters*
carry the literal translation.
- `the-sword-illuminates-the-great-wilderness` — the reverse: the *series* has
the literal title
(`<h1 class="entry-title" itemprop="name">The Sword Illuminates the Great Wilderness</h1>`,
`<title>Read The Sword Illuminates the Great Wilderness English Online Free - LightNovelworld</title>`)
while the *chapters* carry the polished one (`radiant-blade-of-the-wilderness`).
- `a-villains-will-to-survive` —
`<h1 class="entry-title" itemprop="name">A Villain’s Will to Survive</h1>`,
chapters at `the-villain-wants-to-live-chapter-N/`.
`[INFERENCE]` The consistent explanation is that a novel is retitled after
publication; WordPress updates the series post's slug but leaves the already-published
chapter posts' slugs alone. The direction of the retitle varies per novel, which
is why no rule exists. This is consistent with the split case in §4.4, but the
site exposes no field that states it.
**There is no "Alternative names" / "Associated names" field.** Scanning
`https://lightnovelworld.net/novel/a-villains-will-to-survive/` and
`https://lightnovelworld.net/novel/immortality-simulator/` for such a label
returns nothing; the series info panel exposes only **Author**, **Released**,
**Type**, **Status**. (The only `names?` match in the document is the `>Name*<`
label on the comment form.) So the old title is **not** recoverable from the
series page — the mapping only exists in the chapter anchors themselves.
### 4.4 The split case — a slug can change *mid-series*
`https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/`
carries **two** chapter slugs in one page:
| Chapter slug prefix | Anchors | Chapter range |
|---|---|---|
| `all-jobs-and-classes-i-just-wanted-one-skill-not` | 100 | 1 – 99 |
| `all-jobs-and-classes-i-just-wanted-one-skill-not-them-all` | 325 | 100 – 423 |
Both resolve, and **both point back at the same series**:
```
200 https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-chapter-1/
<a href='https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/' aria-label='All Chapter'>
200 https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-404/
<a href='https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/' aria-label='All Chapter'>
```
**Consequence:** a single stored chapter slug is not a sufficient key even for
one novel. A fix that captures "the chapter slug at bookmark time" and pins the
poller regex to it would still miss chapters published under a later slug. This
is the strongest argument for the unscoped regex over a stored-slug regex.
---
## 5. Series page → chapter list
From `https://lightnovelworld.net/novel/immortality-simulator/` (200, 655 037 bytes):
- **1234** anchors matching `href="https://lightnovelworld.net/<slug>-chapter-<n>/"`.
- **1232 unique** chapter URLs, covering chapter numbers **1 through 1232 with no gaps**.
- The 2 extras are the First/Last shortcut widget, which duplicates two chapters:
```html
<h2>Read Immortality Simulator</h2></div>
<div class="lastend">
<div class="inepcx">
<a href="https://lightnovelworld.net/my-longevity-simulation-chapter-1/">
<span>First Chapter</span>
<span class="epcur epcurfirst">Vol. 1 Ch. 1</span>
</a>
</div>
<div class="inepcx">
<a href="https://lightnovelworld.net/my-longevity-simulation-chapter-1232/">
```
**The entire chapter list is in the initial HTML.** No JS hydration, no
pagination, no separate endpoint. Confirmed by probing the shapes the issue
speculated about:
| URL | Status |
|---|---|
| `https://lightnovelworld.net/novel/immortality-simulator/chapters/` | **404** |
| `https://lightnovelworld.net/novel/immortality-simulator/chapters/page-2` | **404** |
The only `page-numbers` / pagination markup in the document belongs to
`class="wpdiscuz-comment-pagination"` (the comment widget), not the chapter list.
This holds for large novels too: `greed-all-for-what` served 2666 chapter
anchors and `my-vampire-system` 2547, all inline in one response.
**All chapter hrefs are absolute** (`https://lightnovelworld.net/…`). A grep for
relative `href="/<slug>-chapter-N/"` on the series page returned **0** matches,
so a pattern anchored on the literal host is safe.
---
## 6. Is the unscoped chapter regex SAFE on a series page?
### Verdict: **SAFE**
Proposed pattern: `lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/`
**Proof.** For each of the 30 series pages fetched from
`https://lightnovelworld.net/novel/`, every occurrence of `<slug>-chapter-<n>/`
anywhere in the document (href or not) was reduced to its slug prefix and
deduplicated:
- **27 of 28 real series pages had exactly 1 distinct chapter-slug prefix**, equal
to that novel's own chapter slug.
- **1 page had 2 prefixes** — `all-jobs-and-classes-…`, and §4.4 proves *both
belong to that same novel*. Not foreign.
- **0 pages carried any other novel's chapter URL.**
The recommendation and sidebar widgets on a series page link to **series** URLs
only, never chapter URLs. On
`https://lightnovelworld.net/novel/immortality-simulator/` the recommendation
strip renders as:
```html
<a href="https://lightnovelworld.net/novel/reverend-insanity/" itemprop="url" title="Reverend Insanity" class="tip" rel="292">
```
— `/novel/<slug>/`, which the pattern cannot match.
The one widget that *would* carry foreign chapter links, "Latest Reading", is an
**empty client-side template**. Its container is `display:none` and its `<ul>` is
empty in the served HTML; the row markup lives in an inert
`<span id="series-history-tpl" style='display:none'>` with mustache placeholders
and a `href="#/{{number}}"` — no real chapter URL:
```html
<div class="bixbox bxcl" id="series-history" style="display:none;">
<div class="releases"><h2>Latest Reading</h2></div>
<div class="series-history-pool">
<ul class="clstyle" id="series-history-ul"></ul>
</div>
</div>
<span id="series-history-tpl" style='display:none'>
<li data-id="{{id}}" data-num="{{chapter}}" data-vol="{{volume}}" data-title="{{title}}">
<div class="chbox"><div class="eph-num">
<a onclick="…return series_history.redirect({{id}});" href="#/{{number}}">
```
It is populated from the visitor's own local history, so a **server-side fetch
never sees content there** — the poller is immune. A browser-rendered fetch with
a fresh profile is likewise immune (no history to render).
### Caveats to record with the verdict
1. **`[INFERENCE]` This is an empirical guarantee, not a structural one.** Unlike
the asura/novelfull cases where scoping to the slug makes foreign contamination
*impossible*, here we only know that lightnovelworld's series template does not
currently emit foreign chapter anchors. If the theme ever adds a "latest site
updates" strip rendered server-side, the unscoped pattern breaks silently and
in the worst direction (a foreign chapter number *higher* than the real one
wins the maximum and the bookmark shows a phantom update).
**Correction, 2026-08-11.** This caveat understated the risk. A series page
server-renders a wpdiscuz comment thread below the chapter list, and comment
bodies are HTML that can carry an `<a href>`. Verified on
`/novel/the-sword-illuminates-the-great-wilderness/`: `#wpdcom` present,
comment `#wpd-comm-358_0` by "hasbi asy" rendered as `<p>where's everyone</p>`
inside `.wpd-comment-text`, corroborated by that page's comment RSS feed;
`firstLoadWithAjax: 0` confirms the thread is in the initial HTML. The sampled
pages carried 0 or 1 comments and no links, which is why §6 read clean. So the
unscoped pattern does not merely depend on the *theme* staying unchanged — it
reads a region any visitor can write to, and the scanner takes the maximum with
no upper bound. The scan must stop before the comment thread.
2. ~~Consider scoping the match to the chapter-list container rather than the whole
document, which would restore the structural guarantee at low cost. The list
items sit under `ul.clstyle` (`class="bixbox bxcl"`).~~
**Refuted, 2026-08-11**, measured on 4 series pages
(`the-sword-illuminates-the-great-wilderness`, `all-jobs-…-them-all`,
`a-will-eternal`, `immortality-simulator`). The single `ul.clstyle` per page is
the hidden, empty "Latest Reading" template (`#series-history-ul`, see §5), not
the chapter list — scoping to it would match nothing. The real chapter list is a
classless `<ul>` inside `div.eplister.eplisterfull`, in
`div.bixbox.bxcl.epcheck`.
The workable boundary is `wpd-threads`: it occurs exactly once per page, and on
all 4 pages every chapter anchor precedes it and no chapter-shaped href follows
it. `wpdcom` and `wpdiscuz` are unusable as markers — they occur 111 to 143
times per page, including in `<head>` roughly 15 KB *before* the chapter list,
so cutting there would discard the list itself. `id='comments'` also occurs once
but is single-quoted; the canonical `id="comments"` never appears.
3. `[0-9.]+` is fine: no decimal chapter numbers appeared anywhere in the sampled
pages, but the existing `strings.Trim(m[1], ".")` guard already handles them.
4. The current test fixture `lnwSeriesFixture` in
`backend/internal/latest/sites_test.go` includes
`<a href="https://lightnovelworld.net/overgeared-chapter-9999/">` and asserts it
is ignored. **That anchor is not representative of a real series page** — no
sampled page contained a foreign chapter anchor. Relaxing the regex will make
that assertion fail, and the correct response is to fix the fixture, not to
keep the scoping.
---
## 7. Redirects and reverse-lookup endpoints
| URL | Status | Redirects | Final |
|---|---|---|---|
| `https://lightnovelworld.net/novel/my-longevity-simulation/` | **404** | 0 | — |
| `https://lightnovelworld.net/novel/radiant-blade-of-the-wilderness/` | **404** | 0 | — |
| `https://lightnovelworld.net/immortality-simulator-chapter-1/` | **404** | 0 | — |
| `https://lightnovelworld.net/my-longevity-simulation/` | **200** | **1** | `https://lightnovelworld.net/my-longevity-simulation-chapter-1/` |
Two findings:
1. **`/novel/<chapter-slug>/` genuinely 404s.** Confirmed for both
`my-longevity-simulation` and `radiant-blade-of-the-wilderness`. There is no
alias, no 301, no fallback. Any stored series URL built by that construction is
permanently dead.
2. **The chapter slug alone (no `-chapter-N`) 302s to chapter 1** of that novel.
This is the closest thing to a reverse-lookup endpoint, but it lands on a
*chapter page*, so recovering the series URL still requires parsing that page's
`aria-label='All Chapter'` anchor or breadcrumb. **No endpoint maps chapter slug
→ series slug directly.**
3. The inverse construction also fails: `/<series-slug>-chapter-1/` 404s for a
divergent novel, so you cannot probe your way from a series slug to a chapter
slug either. The chapter slug must be read off the series page's anchors.
Also present but not a lookup path: the site is WordPress and exposes
`https://lightnovelworld.net/wp-json/wp/v2/posts/177695` (advertised via
`<link rel="alternate" title="JSON" type="application/json" …>` on the chapter
page). Whether the REST API exposes a chapter→series relation was **not tested**
— out of scope for this note, and it would still require fetching the chapter
page to learn the post ID.
---
## 8. Implications for issue #77
> The issue text itself could not be read: `gh` is not installed in this
> environment, so `issue://77` failed to resolve. The proposed regex quoted below
> is the one supplied in the task brief.
### 8.1 The bug
`backend/internal/latest/sites.go:128-137` derives the chapter-anchor pattern
from the **series slug**:
```go
m := lnwSlugRe.FindStringSubmatch(u.Path) // /novel/<series-slug>/
re = regexp.MustCompile(`lightnovelworld\.net/` + regexp.QuoteMeta(m[1]) + `-chapter-([0-9.]+)/`)
```
For `immortality-simulator` this compiles to a pattern matching
`…net/immortality-simulator-chapter-N/`, which appears **zero times** on the
series page — the anchors all say `my-longevity-simulation`. `latestChapterFrom`
returns `ok == false`, indistinguishable from a Cloudflare challenge or a site
redesign, and the bookmark silently stops tracking updates. Same for
`the-sword-illuminates-the-great-wilderness` and `a-villains-will-to-survive`.
### 8.2 The proposed fix is sound
Dropping the scoping to
`lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/` is sound in shape and fixes
all three currently-broken novels plus the split case in §4.4 — which no
stored-slug approach can fix, since that novel legitimately has two chapter
slugs. It must not be applied to the whole document, however: truncate the body
at `wpd-threads` first, per the corrected caveat 2 in §6. The `ul.clstyle`
suggestion originally recorded here is refuted.
Also worth removing: the `overgeared-chapter-9999` line in `lnwSeriesFixture`
and the `"lightnovelworld takes the max and ignores another series"` test case,
which encode a contamination scenario that §6 shows does not occur on this site.
Replacing them with a fixture that mirrors §4.4's two-slugs-one-novel reality
would be a better regression test.
### 8.3 The userscript has the same bug, and it is worse
`userscript/novel-bookmark.user.js` builds the series URL by string-concatenating
the chapter slug (lines 154 and 169):
```js
seriesUrl: "https://lightnovelworld.net/novel/" + m[1] + "/",
```
For a divergent novel this **writes a permanently-404 series URL into the
database at bookmark time**. Fixing only the backend regex leaves those rows
broken, and `fetchableSeriesURL` will happily accept the dead URL because it only
checks the hostname (`sites.go:321-324`).
The userscript should read the series URL off the page instead of constructing
it. On a chapter page, prefer in this order (§3):
```js
document.querySelector("a[aria-label='All Chapter']")?.href
?? document.querySelector("[itemtype$='BreadcrumbList'] a[itemprop=item][href*='/novel/']")?.href
```
The breadcrumb's `<span itemprop="name">` also supplies the correct series title
directly, replacing the current `h1.entry-title` + strip-`Chapter <n>` heuristic
— which today yields "My Longevity Simulation" where the series is actually
titled "Immortality Simulator".
A backfill for already-stored rows is likely needed: any `lightnovelworld` row
whose `series_url` 404s can be repaired by fetching
`https://lightnovelworld.net/<series_id>-chapter-1/` (or `/<series_id>/`, which
302s there per §7) and reading its `All Chapter` anchor.
### 8.4 Note on `series_id`
Rows are keyed `lightnovelworld:<series_id>` where `series_id` is currently the
**chapter slug** (from the userscript's `m[1]`). If the fix changes `series_id` to
the real series slug, existing divergent bookmarks change key and need migrating.
Whether that is worth it depends on whether `series_id` is used to rebuild URLs
anywhere — `sites.go` uses the stored `series_url`, not `series_id`, for
lightnovelworld, so storing the correct `series_url` may be sufficient without a
re-key.
---
## Reproduction
```sh
UA='Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0 Safari/537.36'
# §2 status codes
curl -sS -o /dev/null -w '%{http_code} %{url_effective}\n' -L -A "$UA" \
https://lightnovelworld.net/my-longevity-simulation-chapter-1/ \
https://lightnovelworld.net/novel/immortality-simulator/ \
https://lightnovelworld.net/novel/my-longevity-simulation/
# §3 the two working pointers
curl -sS -A "$UA" https://lightnovelworld.net/my-longevity-simulation-chapter-1/ \
| grep -E "aria-label='All Chapter'|itemprop=\"item\""
# §6 SAFE proof — distinct chapter-slug prefixes per series page (expect 1)
curl -sS -A "$UA" https://lightnovelworld.net/novel/immortality-simulator/ \
| grep -oE '[a-z0-9-]+-chapter-[0-9.]+/' | sed -E 's/-chapter-[0-9.]+\///' | sort | uniq -c
```
+235
View File
@@ -0,0 +1,235 @@
{
"0": "HTMX Library Internals",
"1": "Cover Fetch Test Helpers",
"2": "Manga Userscript Adapters",
"3": "Novel Userscript Adapters",
"4": "Series Acquisition Tests",
"5": "Bookmarks API Tests",
"6": "Storage Choice ADR",
"7": "Cover & Acquire Internals",
"8": "System Architecture Concepts",
"9": "Session Middleware",
"10": "Go Test Helpers",
"11": "Store Tests",
"12": "Bookmarks API Handler",
"13": "Web UI Handlers",
"14": "Go Error Handling",
"15": "CDP Browser Client",
"17": "Go Code Style Guide",
"18": "Agent Skills",
"20": "I/O Performance Patterns",
"21": "CPU Optimization",
"22": "Caching Patterns",
"23": "Browser Entrypoint",
"24": "Memory Allocation & GC",
"25": "Cover Fetcher Tests",
"26": "readSeries",
"27": "Find Skills Guide",
"28": "Allocation Patterns",
"29": "Observability & Alerting",
"31": "Memory Layout",
"32": "Repo Hard Constraints",
"33": "Go Testing Guide",
"34": "Session Store",
"35": "Web UI Filter Logic",
"36": "Userscript Test Harness",
"37": "Product & Security Context",
"38": "novel-logic.test.js",
"39": "UI Critique 2026-07-26A",
"40": "UI Critique 2026-07-26B",
"43": "Issue Tracker & Triage",
"44": "Ticket Workflow",
"45": "Go Perf Alert Rules",
"46": "Userscript Display Logic",
"47": "Go Perf Skill Docs",
"48": "Login Page Art",
"49": "BookmarkManager Logo",
"50": "Skills CLI",
"51": "Skills Leaderboard",
"52": "Complex Condition Extraction",
"53": "Sentinel Errors",
"54": "errors.As Patterns",
"55": "errors.Is Patterns",
"56": "errors.Join Patterns",
"57": "Error Wrapping",
"58": "Single Error Handling",
"59": "SIMD Optimizations",
"60": "GOGC Tuning",
"61": "GOMEMLIMIT",
"62": "Bottleneck Decision Tree",
"63": "pprof Profiling",
"64": "Test Timeout Helper",
"65": "httptest Patterns",
"66": "testify Suite Pattern",
"67": "go:embed Fixtures",
"68": "clockwork Time Mocking",
"69": "testify Mocking",
"70": "t.ArtifactDir Helper",
"71": "Subtests Pitfall",
"72": "golang-benchmark Skill",
"73": "golang-concurrency Skill",
"74": "golang-ci Skill",
"75": "golang-database Skill",
"76": "golang-lint Skill",
"77": "testify Skill",
"78": "Build Tag Integration Tests",
"79": "Test Naming Convention",
"80": "UI Critique A Finding",
"81": "UI Critique B Finding",
"82": "P0 Overflow Bug",
"83": "P1 hx-indicator Gap",
"84": "golang-benchmark Skill (ext)",
"85": "golang-concurrency Skill (ext)",
"86": "golang-ci Skill (ext)",
"87": "golang-data-structures Skill (ext)",
"88": "golang-database Skill (ext)",
"89": "golang-design-patterns Skill (ext)",
"90": "golang-documentation Skill (ext)",
"91": "golang-gopls Skill (ext)",
"92": "golang-lint Skill (ext)",
"93": "golang-naming Skill (ext)",
"94": "golang-observability Skill (ext)",
"95": "golang-refactoring Skill (ext)",
"96": "golang-safety Skill (ext)",
"97": "golang-samber-oops Skill (ext)",
"98": "golang-samber-slog Skill (ext)",
"99": "golang-structs-interfaces Skill (ext)",
"100": "golang-troubleshooting Skill (ext)",
"101": "promql-cli Skill",
"102": "Backend Module",
"103": "bookmark-api Service",
"104": "AGENTS.md",
"105": "reviewer.md",
"106": "Redeploy runbook",
"107": "1. Backend",
"108": "Deployment",
"109": "Cinder — BookmarkManager design system",
"110": "Implement tickets",
"111": "SQLite → Postgres cutover runbook",
"112": "Testing the userscript",
"113": "ADR-0007: The backend hosts every Site's Cover bytes",
"114": "Issue tracker: Gitea (`tea` CLI)",
"115": "ADR-0006: The browser runs on the home machine, over the tailnet",
"116": "ADR-0008: A Series identity is discovered from the Site's links, never derived from an address",
"117": "Domain Docs",
"118": "ticket-implementer.md",
"119": "implementer.md",
"120": "Series is a shared entity, and only the Poll may update it",
"121": "Postgres replaces SQLite as the primary datastore",
"122": "Identity comes from Discord OAuth; we store no passwords and send no email",
"123": "The wire format stays flat and deliberately does not mirror the schema",
"124": "ADR-0005: On-demand browser sidecar",
"126": "Bookmark Manager",
"127": "triage-labels.md",
"128": "Cross-Ticket Contract",
"129": "Implement Tickets Skill",
"130": "Orchestrator Role",
"131": "resolving-merge-conflicts Skill",
"132": "tdd Skill",
"133": "Ticket Wave Batching",
"134": "Four-Object Browser Stub",
"135": "Module Export Hook",
"136": "logic.test.js Test Harness",
"137": "manga-bookmark.user.js",
"138": "stripBuildHash",
"139": "Testing the Userscript Skill",
"140": "cr-spec Agent",
"141": "cr-standards Agent",
"142": "Escalate Rather Than Guess",
"143": "Status Contract",
"144": "Ticket Implementer Agent",
"145": "Worktree Isolation",
"146": "Escalate Rather Than Guess (opencode)",
"147": "Implementer Subagent (opencode)",
"148": "Subagent-Driven Development",
"149": "Code Quality Review",
"150": "Reviewer Subagent (opencode)",
"151": "Finding Severity Rubric",
"152": "Spec Compliance Review",
"161": "Why Use samber/oops",
"162": "singleflight Cache Stampede Prevention",
"163": "Struct Field Alignment",
"164": "testing/synctest Deterministic Goroutine Testing",
"165": "API Package (Bookmark JSON Handlers)",
"166": "Cover Acquisition & Serving Pipeline",
"167": "Backend AGENTS.md Guidance",
"168": "HTTP Middleware (Auth/Gzip/CORS)",
"169": "Latest Package (Site Parsers & Poller)",
"170": "Latest-Chapter Poller",
"171": "Main Composition Root",
"172": "Migration-Owned Schema",
"173": "Session Package (Cookie Signing & Rate Limit)",
"174": "updated_at List-Order Rule",
"175": "Userscript Package (Serving Handler)",
"176": "AGENTS.md",
"177": "Backend CLAUDE.md Guidance",
"178": "Graphify Knowledge Graph (graphify-out/)",
"179": "CLAUDE.md (Symlink to AGENTS.md)",
"192": "ADR-0001 (Drop modernc.org/sqlite)",
"193": "ADR-0003 (Split Shared Series Facts)",
"194": "SQLite-to-Postgres Cutover Runbook",
"195": "Import SQL Generation Rules",
"196": "Throwaway Import Generator",
"206": "Real scaling limit is the poller outbound fetch budget",
"207": "PostgreSQL (jackc/pgx/v5)",
"208": "SQLite (modernc.org/sqlite)",
"209": "Postgres chosen for future supportability, not concurrency",
"210": "Per-Reader bearer token for userscripts",
"211": "Discord OAuth2 (authorization code grant)",
"212": "ADR-0002: Discord OAuth, no passwords, no email",
"213": "Discord snowflake is the sole identity (lock-in)",
"214": "Bookmark (per-Reader state: Progress, Favourite, Lifecycle)",
"215": "Deduplicate polling per Series (reader_count DESC queue)",
"216": "ADR-0003: Series is shared, only the Poll updates it",
"217": "Only the Poll writes Series fields (security boundary)",
"218": "Series (shared entity keyed site+series_id)",
"219": "ADR-0004: Wire format stays flat, does not mirror schema",
"220": "Flat wire shape is a contract, not an implementation detail",
"221": "Installed userscripts must keep working (14-day grace window)",
"222": "CDP (Chrome DevTools Protocol) endpoint",
"223": "headless-shell service (socat-fronted CDP)",
"224": "Start Chrome on first CDP connection, reap after 300s idle",
"225": "BROWSER_WS_URL configuration seam",
"226": "ADR-0006: Browser runs on the home machine over the tailnet",
"227": "Browser moved home: VPS memory pressure, no requests served",
"228": "Tailnet (Tailscale network)",
"229": "Content-addressed filesystem storage (SHA-256 of source URL)",
"230": "Cover (Series image bytes)",
"231": "Deny-class destination control for outbound fetch",
"232": "ADR-0007: Backend hosts every Site's Cover bytes",
"233": "kagane CORP same-origin cover restriction",
"234": "Backend acquires, stores, serves every Cover (uniformity)",
"235": "a[aria-label='All Chapter'] anchor pointer",
"236": "Series identity is discovered from the Site's links",
"237": "ADR-0008: Series identity discovered, never derived",
"238": "Chapter slug vs series slug divergence (~7% measured)",
"239": "Scan truncated at first wpd-threads marker",
"240": "Surface ADR conflicts explicitly rather than silently overriding",
"241": "Domain docs: single-context layout guidance",
"242": "/domain-modeling skill (lazy CONTEXT.md creation)",
"243": "CONTEXT.md glossary (ubiquitous language)",
"244": "Gitea (tea CLI, gitea.violetcrown.my.id)",
"245": "wayfinder map/ticket mechanism",
"246": "Triage labels: canonical roles to tracker labels",
"247": "Canonical triage role labels (needs-triage ... wontfix)",
"248": "Cinder (BookmarkManager Web UI design system)",
"249": "Heat is typographic: ember reserved for unread chapters",
"250": "Design tokens (dark + light branches, no hardcoded hex)",
"251": "Three type roles: display serif / mono small-caps / sans",
"252": "a[aria-label='All Chapter'] priority pointer",
"253": "Research: lightnovelworld chapter slug vs series slug",
"254": "Gitea issue #77 (chapter vs series slug)",
"255": "Slug divergence measurements (3/41 diverge, 1 split)",
"256": "Unscoped chapter regex is SAFE, truncated at wpd-threads",
"257": "BookmarkManager",
"258": "Bromite (Primary Device)",
"259": "Dark-First Design Constraint",
"260": "Discord Guild Membership",
"261": "Reader Isolation Invariant",
"268": "Browser Unit Redeploy",
"269": "pg_dump Hot Backup",
"270": "Redeploy Runbook",
"271": "Rollback Strategy",
"273": "AGENTS.md",
"279": "Userscript CLAUDE guidance"
}
+1
View File
@@ -0,0 +1 @@
.
+558
View File
@@ -0,0 +1,558 @@
# Graph Report - mangaBookmark (2026-08-12)
## Corpus Check
- 110 files · ~267,791 words
- Verdict: corpus is large enough that graph structure adds value.
## Summary
- 1621 nodes · 3242 edges · 233 communities (63 shown, 170 thin omitted)
- Extraction: 90% EXTRACTED · 10% INFERRED · 0% AMBIGUOUS · INFERRED: 312 edges (avg confidence: 0.77)
- Token cost: 0 input · 0 output
## Graph Freshness
- Built from commit: `8ae98816`
- Run `git rev-parse HEAD` and compare to check if the graph is stale.
- Run `graphify update .` after code changes (no API cost).
## Community Hubs (Navigation)
- [[_COMMUNITY_HTMX Library Internals|HTMX Library Internals]]
- [[_COMMUNITY_Cover Fetch Test Helpers|Cover Fetch Test Helpers]]
- [[_COMMUNITY_Manga Userscript Adapters|Manga Userscript Adapters]]
- [[_COMMUNITY_Novel Userscript Adapters|Novel Userscript Adapters]]
- [[_COMMUNITY_Series Acquisition Tests|Series Acquisition Tests]]
- [[_COMMUNITY_Bookmarks API Tests|Bookmarks API Tests]]
- [[_COMMUNITY_Storage Choice ADR|Storage Choice ADR]]
- [[_COMMUNITY_Cover & Acquire Internals|Cover & Acquire Internals]]
- [[_COMMUNITY_System Architecture Concepts|System Architecture Concepts]]
- [[_COMMUNITY_Session Middleware|Session Middleware]]
- [[_COMMUNITY_Go Test Helpers|Go Test Helpers]]
- [[_COMMUNITY_Store Tests|Store Tests]]
- [[_COMMUNITY_Bookmarks API Handler|Bookmarks API Handler]]
- [[_COMMUNITY_Web UI Handlers|Web UI Handlers]]
- [[_COMMUNITY_Go Error Handling|Go Error Handling]]
- [[_COMMUNITY_CDP Browser Client|CDP Browser Client]]
- [[_COMMUNITY_Go Code Style Guide|Go Code Style Guide]]
- [[_COMMUNITY_Agent Skills|Agent Skills]]
- [[_COMMUNITY_IO Performance Patterns|I/O Performance Patterns]]
- [[_COMMUNITY_CPU Optimization|CPU Optimization]]
- [[_COMMUNITY_Caching Patterns|Caching Patterns]]
- [[_COMMUNITY_Browser Entrypoint|Browser Entrypoint]]
- [[_COMMUNITY_Memory Allocation & GC|Memory Allocation & GC]]
- [[_COMMUNITY_Cover Fetcher Tests|Cover Fetcher Tests]]
- [[_COMMUNITY_readSeries|readSeries]]
- [[_COMMUNITY_Find Skills Guide|Find Skills Guide]]
- [[_COMMUNITY_Allocation Patterns|Allocation Patterns]]
- [[_COMMUNITY_Observability & Alerting|Observability & Alerting]]
- [[_COMMUNITY_Memory Layout|Memory Layout]]
- [[_COMMUNITY_Repo Hard Constraints|Repo Hard Constraints]]
- [[_COMMUNITY_Go Testing Guide|Go Testing Guide]]
- [[_COMMUNITY_Session Store|Session Store]]
- [[_COMMUNITY_Web UI Filter Logic|Web UI Filter Logic]]
- [[_COMMUNITY_Userscript Test Harness|Userscript Test Harness]]
- [[_COMMUNITY_Product & Security Context|Product & Security Context]]
- [[_COMMUNITY_novel-logic.test.js|novel-logic.test.js]]
- [[_COMMUNITY_UI Critique 2026-07-26A|UI Critique 2026-07-26A]]
- [[_COMMUNITY_UI Critique 2026-07-26B|UI Critique 2026-07-26B]]
- [[_COMMUNITY_Issue Tracker & Triage|Issue Tracker & Triage]]
- [[_COMMUNITY_Ticket Workflow|Ticket Workflow]]
- [[_COMMUNITY_Go Perf Alert Rules|Go Perf Alert Rules]]
- [[_COMMUNITY_Userscript Display Logic|Userscript Display Logic]]
- [[_COMMUNITY_Go Perf Skill Docs|Go Perf Skill Docs]]
- [[_COMMUNITY_Login Page Art|Login Page Art]]
- [[_COMMUNITY_BookmarkManager Logo|BookmarkManager Logo]]
- [[_COMMUNITY_Skills CLI|Skills CLI]]
- [[_COMMUNITY_Skills Leaderboard|Skills Leaderboard]]
- [[_COMMUNITY_Complex Condition Extraction|Complex Condition Extraction]]
- [[_COMMUNITY_Sentinel Errors|Sentinel Errors]]
- [[_COMMUNITY_errors.As Patterns|errors.As Patterns]]
- [[_COMMUNITY_errors.Is Patterns|errors.Is Patterns]]
- [[_COMMUNITY_errors.Join Patterns|errors.Join Patterns]]
- [[_COMMUNITY_Error Wrapping|Error Wrapping]]
- [[_COMMUNITY_Single Error Handling|Single Error Handling]]
- [[_COMMUNITY_SIMD Optimizations|SIMD Optimizations]]
- [[_COMMUNITY_GOGC Tuning|GOGC Tuning]]
- [[_COMMUNITY_GOMEMLIMIT|GOMEMLIMIT]]
- [[_COMMUNITY_Bottleneck Decision Tree|Bottleneck Decision Tree]]
- [[_COMMUNITY_pprof Profiling|pprof Profiling]]
- [[_COMMUNITY_Test Timeout Helper|Test Timeout Helper]]
- [[_COMMUNITY_httptest Patterns|httptest Patterns]]
- [[_COMMUNITY_testify Suite Pattern|testify Suite Pattern]]
- [[_COMMUNITY_goembed Fixtures|go:embed Fixtures]]
- [[_COMMUNITY_clockwork Time Mocking|clockwork Time Mocking]]
- [[_COMMUNITY_testify Mocking|testify Mocking]]
- [[_COMMUNITY_t.ArtifactDir Helper|t.ArtifactDir Helper]]
- [[_COMMUNITY_Subtests Pitfall|Subtests Pitfall]]
- [[_COMMUNITY_golang-benchmark Skill|golang-benchmark Skill]]
- [[_COMMUNITY_golang-concurrency Skill|golang-concurrency Skill]]
- [[_COMMUNITY_golang-ci Skill|golang-ci Skill]]
- [[_COMMUNITY_golang-database Skill|golang-database Skill]]
- [[_COMMUNITY_golang-lint Skill|golang-lint Skill]]
- [[_COMMUNITY_testify Skill|testify Skill]]
- [[_COMMUNITY_Build Tag Integration Tests|Build Tag Integration Tests]]
- [[_COMMUNITY_Test Naming Convention|Test Naming Convention]]
- [[_COMMUNITY_UI Critique A Finding|UI Critique A Finding]]
- [[_COMMUNITY_UI Critique B Finding|UI Critique B Finding]]
- [[_COMMUNITY_P0 Overflow Bug|P0 Overflow Bug]]
- [[_COMMUNITY_P1 hx-indicator Gap|P1 hx-indicator Gap]]
- [[_COMMUNITY_golang-benchmark Skill (ext)|golang-benchmark Skill (ext)]]
- [[_COMMUNITY_golang-concurrency Skill (ext)|golang-concurrency Skill (ext)]]
- [[_COMMUNITY_golang-ci Skill (ext)|golang-ci Skill (ext)]]
- [[_COMMUNITY_golang-data-structures Skill (ext)|golang-data-structures Skill (ext)]]
- [[_COMMUNITY_golang-database Skill (ext)|golang-database Skill (ext)]]
- [[_COMMUNITY_golang-design-patterns Skill (ext)|golang-design-patterns Skill (ext)]]
- [[_COMMUNITY_golang-documentation Skill (ext)|golang-documentation Skill (ext)]]
- [[_COMMUNITY_golang-gopls Skill (ext)|golang-gopls Skill (ext)]]
- [[_COMMUNITY_golang-lint Skill (ext)|golang-lint Skill (ext)]]
- [[_COMMUNITY_golang-naming Skill (ext)|golang-naming Skill (ext)]]
- [[_COMMUNITY_golang-observability Skill (ext)|golang-observability Skill (ext)]]
- [[_COMMUNITY_golang-refactoring Skill (ext)|golang-refactoring Skill (ext)]]
- [[_COMMUNITY_golang-safety Skill (ext)|golang-safety Skill (ext)]]
- [[_COMMUNITY_golang-samber-oops Skill (ext)|golang-samber-oops Skill (ext)]]
- [[_COMMUNITY_golang-samber-slog Skill (ext)|golang-samber-slog Skill (ext)]]
- [[_COMMUNITY_golang-structs-interfaces Skill (ext)|golang-structs-interfaces Skill (ext)]]
- [[_COMMUNITY_golang-troubleshooting Skill (ext)|golang-troubleshooting Skill (ext)]]
- [[_COMMUNITY_promql-cli Skill|promql-cli Skill]]
- [[_COMMUNITY_Backend Module|Backend Module]]
- [[_COMMUNITY_bookmark-api Service|bookmark-api Service]]
- [[_COMMUNITY_AGENTS|AGENTS.md]]
- [[_COMMUNITY_reviewer|reviewer.md]]
- [[_COMMUNITY_Redeploy runbook|Redeploy runbook]]
- [[_COMMUNITY_1. Backend|1. Backend]]
- [[_COMMUNITY_Deployment|Deployment]]
- [[_COMMUNITY_Cinder — BookmarkManager design system|Cinder — BookmarkManager design system]]
- [[_COMMUNITY_Implement tickets|Implement tickets]]
- [[_COMMUNITY_SQLite → Postgres cutover runbook|SQLite → Postgres cutover runbook]]
- [[_COMMUNITY_Testing the userscript|Testing the userscript]]
- [[_COMMUNITY_ADR-0007 The backend hosts every Site's Cover bytes|ADR-0007: The backend hosts every Site's Cover bytes]]
- [[_COMMUNITY_Issue tracker Gitea (`tea` CLI)|Issue tracker: Gitea (`tea` CLI)]]
- [[_COMMUNITY_ADR-0006 The browser runs on the home machine, over the tailnet|ADR-0006: The browser runs on the home machine, over the tailnet]]
- [[_COMMUNITY_ADR-0008 A Series identity is discovered from the Site's links, never derived from an address|ADR-0008: A Series identity is discovered from the Site's links, never derived from an address]]
- [[_COMMUNITY_Domain Docs|Domain Docs]]
- [[_COMMUNITY_ticket-implementer|ticket-implementer.md]]
- [[_COMMUNITY_implementer|implementer.md]]
- [[_COMMUNITY_Series is a shared entity, and only the Poll may update it|Series is a shared entity, and only the Poll may update it]]
- [[_COMMUNITY_Postgres replaces SQLite as the primary datastore|Postgres replaces SQLite as the primary datastore]]
- [[_COMMUNITY_Identity comes from Discord OAuth; we store no passwords and send no email|Identity comes from Discord OAuth; we store no passwords and send no email]]
- [[_COMMUNITY_The wire format stays flat and deliberately does not mirror the schema|The wire format stays flat and deliberately does not mirror the schema]]
- [[_COMMUNITY_ADR-0005 On-demand browser sidecar|ADR-0005: On-demand browser sidecar]]
- [[_COMMUNITY_Bookmark Manager|Bookmark Manager]]
- [[_COMMUNITY_triage-labels|triage-labels.md]]
- [[_COMMUNITY_Cross-Ticket Contract|Cross-Ticket Contract]]
- [[_COMMUNITY_Implement Tickets Skill|Implement Tickets Skill]]
- [[_COMMUNITY_Orchestrator Role|Orchestrator Role]]
- [[_COMMUNITY_resolving-merge-conflicts Skill|resolving-merge-conflicts Skill]]
- [[_COMMUNITY_tdd Skill|tdd Skill]]
- [[_COMMUNITY_Ticket Wave Batching|Ticket Wave Batching]]
- [[_COMMUNITY_Four-Object Browser Stub|Four-Object Browser Stub]]
- [[_COMMUNITY_Module Export Hook|Module Export Hook]]
- [[_COMMUNITY_logic.test.js Test Harness|logic.test.js Test Harness]]
- [[_COMMUNITY_manga-bookmark.user.js|manga-bookmark.user.js]]
- [[_COMMUNITY_stripBuildHash|stripBuildHash]]
- [[_COMMUNITY_Testing the Userscript Skill|Testing the Userscript Skill]]
- [[_COMMUNITY_cr-spec Agent|cr-spec Agent]]
- [[_COMMUNITY_cr-standards Agent|cr-standards Agent]]
- [[_COMMUNITY_Escalate Rather Than Guess|Escalate Rather Than Guess]]
- [[_COMMUNITY_Status Contract|Status Contract]]
- [[_COMMUNITY_Ticket Implementer Agent|Ticket Implementer Agent]]
- [[_COMMUNITY_Worktree Isolation|Worktree Isolation]]
- [[_COMMUNITY_Escalate Rather Than Guess (opencode)|Escalate Rather Than Guess (opencode)]]
- [[_COMMUNITY_Implementer Subagent (opencode)|Implementer Subagent (opencode)]]
- [[_COMMUNITY_Subagent-Driven Development|Subagent-Driven Development]]
- [[_COMMUNITY_Code Quality Review|Code Quality Review]]
- [[_COMMUNITY_Reviewer Subagent (opencode)|Reviewer Subagent (opencode)]]
- [[_COMMUNITY_Finding Severity Rubric|Finding Severity Rubric]]
- [[_COMMUNITY_Spec Compliance Review|Spec Compliance Review]]
- [[_COMMUNITY_Why Use samberoops|Why Use samber/oops]]
- [[_COMMUNITY_singleflight Cache Stampede Prevention|singleflight Cache Stampede Prevention]]
- [[_COMMUNITY_Struct Field Alignment|Struct Field Alignment]]
- [[_COMMUNITY_testingsynctest Deterministic Goroutine Testing|testing/synctest Deterministic Goroutine Testing]]
- [[_COMMUNITY_API Package (Bookmark JSON Handlers)|API Package (Bookmark JSON Handlers)]]
- [[_COMMUNITY_Cover Acquisition & Serving Pipeline|Cover Acquisition & Serving Pipeline]]
- [[_COMMUNITY_Backend AGENTS.md Guidance|Backend AGENTS.md Guidance]]
- [[_COMMUNITY_HTTP Middleware (AuthGzipCORS)|HTTP Middleware (Auth/Gzip/CORS)]]
- [[_COMMUNITY_Latest Package (Site Parsers & Poller)|Latest Package (Site Parsers & Poller)]]
- [[_COMMUNITY_Latest-Chapter Poller|Latest-Chapter Poller]]
- [[_COMMUNITY_Main Composition Root|Main Composition Root]]
- [[_COMMUNITY_Migration-Owned Schema|Migration-Owned Schema]]
- [[_COMMUNITY_Session Package (Cookie Signing & Rate Limit)|Session Package (Cookie Signing & Rate Limit)]]
- [[_COMMUNITY_updated_at List-Order Rule|updated_at List-Order Rule]]
- [[_COMMUNITY_Userscript Package (Serving Handler)|Userscript Package (Serving Handler)]]
- [[_COMMUNITY_Backend CLAUDE.md Guidance|Backend CLAUDE.md Guidance]]
- [[_COMMUNITY_Graphify Knowledge Graph (graphify-out)|Graphify Knowledge Graph (graphify-out/)]]
- [[_COMMUNITY_CLAUDE.md (Symlink to AGENTS.md)|CLAUDE.md (Symlink to AGENTS.md)]]
- [[_COMMUNITY_ADR-0001 (Drop modernc.orgsqlite)|ADR-0001 (Drop modernc.org/sqlite)]]
- [[_COMMUNITY_ADR-0003 (Split Shared Series Facts)|ADR-0003 (Split Shared Series Facts)]]
- [[_COMMUNITY_SQLite-to-Postgres Cutover Runbook|SQLite-to-Postgres Cutover Runbook]]
- [[_COMMUNITY_Import SQL Generation Rules|Import SQL Generation Rules]]
- [[_COMMUNITY_Throwaway Import Generator|Throwaway Import Generator]]
- [[_COMMUNITY_Real scaling limit is the poller outbound fetch budget|Real scaling limit is the poller outbound fetch budget]]
- [[_COMMUNITY_PostgreSQL (jackcpgxv5)|PostgreSQL (jackc/pgx/v5)]]
- [[_COMMUNITY_SQLite (modernc.orgsqlite)|SQLite (modernc.org/sqlite)]]
- [[_COMMUNITY_Postgres chosen for future supportability, not concurrency|Postgres chosen for future supportability, not concurrency]]
- [[_COMMUNITY_Per-Reader bearer token for userscripts|Per-Reader bearer token for userscripts]]
- [[_COMMUNITY_Discord OAuth2 (authorization code grant)|Discord OAuth2 (authorization code grant)]]
- [[_COMMUNITY_ADR-0002 Discord OAuth, no passwords, no email|ADR-0002: Discord OAuth, no passwords, no email]]
- [[_COMMUNITY_Discord snowflake is the sole identity (lock-in)|Discord snowflake is the sole identity (lock-in)]]
- [[_COMMUNITY_Bookmark (per-Reader state Progress, Favourite, Lifecycle)|Bookmark (per-Reader state: Progress, Favourite, Lifecycle)]]
- [[_COMMUNITY_Deduplicate polling per Series (reader_count DESC queue)|Deduplicate polling per Series (reader_count DESC queue)]]
- [[_COMMUNITY_ADR-0003 Series is shared, only the Poll updates it|ADR-0003: Series is shared, only the Poll updates it]]
- [[_COMMUNITY_Only the Poll writes Series fields (security boundary)|Only the Poll writes Series fields (security boundary)]]
- [[_COMMUNITY_Series (shared entity keyed site+series_id)|Series (shared entity keyed site+series_id)]]
- [[_COMMUNITY_ADR-0004 Wire format stays flat, does not mirror schema|ADR-0004: Wire format stays flat, does not mirror schema]]
- [[_COMMUNITY_Flat wire shape is a contract, not an implementation detail|Flat wire shape is a contract, not an implementation detail]]
- [[_COMMUNITY_Installed userscripts must keep working (14-day grace window)|Installed userscripts must keep working (14-day grace window)]]
- [[_COMMUNITY_CDP (Chrome DevTools Protocol) endpoint|CDP (Chrome DevTools Protocol) endpoint]]
- [[_COMMUNITY_headless-shell service (socat-fronted CDP)|headless-shell service (socat-fronted CDP)]]
- [[_COMMUNITY_Start Chrome on first CDP connection, reap after 300s idle|Start Chrome on first CDP connection, reap after 300s idle]]
- [[_COMMUNITY_BROWSER_WS_URL configuration seam|BROWSER_WS_URL configuration seam]]
- [[_COMMUNITY_ADR-0006 Browser runs on the home machine over the tailnet|ADR-0006: Browser runs on the home machine over the tailnet]]
- [[_COMMUNITY_Browser moved home VPS memory pressure, no requests served|Browser moved home: VPS memory pressure, no requests served]]
- [[_COMMUNITY_Tailnet (Tailscale network)|Tailnet (Tailscale network)]]
- [[_COMMUNITY_Content-addressed filesystem storage (SHA-256 of source URL)|Content-addressed filesystem storage (SHA-256 of source URL)]]
- [[_COMMUNITY_Cover (Series image bytes)|Cover (Series image bytes)]]
- [[_COMMUNITY_Deny-class destination control for outbound fetch|Deny-class destination control for outbound fetch]]
- [[_COMMUNITY_ADR-0007 Backend hosts every Site's Cover bytes|ADR-0007: Backend hosts every Site's Cover bytes]]
- [[_COMMUNITY_kagane CORP same-origin cover restriction|kagane CORP same-origin cover restriction]]
- [[_COMMUNITY_Backend acquires, stores, serves every Cover (uniformity)|Backend acquires, stores, serves every Cover (uniformity)]]
- [[_COMMUNITY_aaria-label='All Chapter' anchor pointer|a[aria-label='All Chapter'] anchor pointer]]
- [[_COMMUNITY_Series identity is discovered from the Site's links|Series identity is discovered from the Site's links]]
- [[_COMMUNITY_ADR-0008 Series identity discovered, never derived|ADR-0008: Series identity discovered, never derived]]
- [[_COMMUNITY_Chapter slug vs series slug divergence (~7% measured)|Chapter slug vs series slug divergence (~7% measured)]]
- [[_COMMUNITY_Scan truncated at first wpd-threads marker|Scan truncated at first wpd-threads marker]]
- [[_COMMUNITY_Surface ADR conflicts explicitly rather than silently overriding|Surface ADR conflicts explicitly rather than silently overriding]]
- [[_COMMUNITY_Domain docs single-context layout guidance|Domain docs: single-context layout guidance]]
- [[_COMMUNITY_domain-modeling skill (lazy CONTEXT.md creation)|/domain-modeling skill (lazy CONTEXT.md creation)]]
- [[_COMMUNITY_CONTEXT.md glossary (ubiquitous language)|CONTEXT.md glossary (ubiquitous language)]]
- [[_COMMUNITY_Gitea (tea CLI, gitea.violetcrown.my.id)|Gitea (tea CLI, gitea.violetcrown.my.id)]]
- [[_COMMUNITY_wayfinder mapticket mechanism|wayfinder map/ticket mechanism]]
- [[_COMMUNITY_Triage labels canonical roles to tracker labels|Triage labels: canonical roles to tracker labels]]
- [[_COMMUNITY_Canonical triage role labels (needs-triage ... wontfix)|Canonical triage role labels (needs-triage ... wontfix)]]
- [[_COMMUNITY_Cinder (BookmarkManager Web UI design system)|Cinder (BookmarkManager Web UI design system)]]
- [[_COMMUNITY_Heat is typographic ember reserved for unread chapters|Heat is typographic: ember reserved for unread chapters]]
- [[_COMMUNITY_Design tokens (dark + light branches, no hardcoded hex)|Design tokens (dark + light branches, no hardcoded hex)]]
- [[_COMMUNITY_Three type roles display serif mono small-caps sans|Three type roles: display serif / mono small-caps / sans]]
- [[_COMMUNITY_aaria-label='All Chapter' priority pointer|a[aria-label='All Chapter'] priority pointer]]
- [[_COMMUNITY_Research lightnovelworld chapter slug vs series slug|Research: lightnovelworld chapter slug vs series slug]]
- [[_COMMUNITY_Gitea issue 77 (chapter vs series slug)|Gitea issue #77 (chapter vs series slug)]]
- [[_COMMUNITY_Slug divergence measurements (341 diverge, 1 split)|Slug divergence measurements (3/41 diverge, 1 split)]]
- [[_COMMUNITY_Unscoped chapter regex is SAFE, truncated at wpd-threads|Unscoped chapter regex is SAFE, truncated at wpd-threads]]
- [[_COMMUNITY_BookmarkManager|BookmarkManager]]
- [[_COMMUNITY_Bromite (Primary Device)|Bromite (Primary Device)]]
- [[_COMMUNITY_Dark-First Design Constraint|Dark-First Design Constraint]]
- [[_COMMUNITY_Discord Guild Membership|Discord Guild Membership]]
- [[_COMMUNITY_Reader Isolation Invariant|Reader Isolation Invariant]]
- [[_COMMUNITY_Browser Unit Redeploy|Browser Unit Redeploy]]
- [[_COMMUNITY_pg_dump Hot Backup|pg_dump Hot Backup]]
- [[_COMMUNITY_Redeploy Runbook|Redeploy Runbook]]
- [[_COMMUNITY_Rollback Strategy|Rollback Strategy]]
- [[_COMMUNITY_AGENTS|AGENTS.md]]
- [[_COMMUNITY_Userscript CLAUDE guidance|Userscript CLAUDE guidance]]
## God Nodes (most connected - your core abstractions)
1. `testConfig()` - 53 edges
2. `newWebTestServer()` - 49 edges
3. `newTestStore()` - 42 edges
4. `newTestStore()` - 41 edges
5. `e()` - 33 edges
6. `Handler` - 29 edges
7. `ne()` - 28 edges
8. `De()` - 28 edges
9. `Open()` - 27 edges
10. `se()` - 27 edges
## Surprising Connections (you probably didn't know these)
- `Browser Sidecar Service` --semantically_similar_to--> `Browser Sidecar (BROWSER_WS_URL)` [INFERRED] [semantically similar]
chrome/docker-compose.yml → backend/AGENTS.md
- `el()` --indirect_call--> `c()` [INFERRED]
userscript/manga-bookmark.user.js → backend/internal/web/static/htmx.min.js
- `el()` --indirect_call--> `c()` [INFERRED]
userscript/novel-bookmark.user.js → backend/internal/web/static/htmx.min.js
- `latestChapterFromAnchors()` --indirect_call--> `re()` [INFERRED]
userscript/manga-bookmark.user.js → backend/internal/web/static/htmx.min.js
- `el()` --indirect_call--> `k()` [INFERRED]
userscript/manga-bookmark.user.js → backend/internal/web/static/htmx.min.js
## Import Cycles
- None detected.
## Hyperedges (group relationships)
- **Batch Ticket Implementation Pipeline** — _claude_skills_implement_tickets_skill_implement_tickets, _omp_agents_ticket_implementer_ticket_implementer, _omp_agents_ticket_implementer_cr_spec, _omp_agents_ticket_implementer_cr_standards [INFERRED 0.85]
- **Subagent-Driven Development Pipeline** — _opencode_agent_implementer_implementer, _opencode_agent_reviewer_reviewer, _opencode_agent_implementer_subagent_driven_development [INFERRED 0.85]
- **Go HTML Template Family** — backend_internal_web_templates_app_doc, backend_internal_web_templates_card_doc, backend_internal_web_templates_list_doc, backend_internal_web_templates_chrome_doc, backend_internal_web_templates_login_doc, backend_internal_web_templates_setup_doc, backend_internal_web_templates_readers_doc, backend_internal_web_templates_icons_doc [INFERRED 0.95]
- **htmx Fragment Swap Flow** — backend_internal_web_templates_app_doc, backend_internal_web_templates_card_doc, backend_internal_web_templates_chrome_doc, backend_internal_web_templates_setup_doc, backend_internal_web_templates_readers_doc [INFERRED 0.95]
- **Backend owns the truth (single-writer ownership of shared facts)** — docs_adr_0003_series_shared_and_poll_owned_poll_owned_writes, docs_adr_0004_wire_format_does_not_mirror_the_schema_flat_wire_contract, docs_adr_0007_backend_hosts_cover_bytes_server_side_covers [INFERRED 0.85]
- **Headless browser infrastructure (sidecar, on-demand, home deployment)** — docs_adr_0005_on_demand_browser_headless_shell, docs_adr_0005_on_demand_browser_cdp, docs_adr_0005_on_demand_browser_on_demand_start, docs_adr_0006_browser_on_the_home_machine_home_machine_rationale [INFERRED 0.85]
- **lightnovelworld series-identity investigation and fix** — docs_research_lightnovelworld_chapter_vs_series_slug_issue_77, docs_research_lightnovelworld_chapter_vs_series_slug_unscoped_regex, docs_adr_0008_series_identity_is_discovered_not_derived_discovered_identity [INFERRED 0.85]
## Communities (233 total, 170 thin omitted)
### Community 0 - "HTMX Library Internals"
Cohesion: 0.08
Nodes (101): A(), ae(), an(), at(), B(), be(), bn(), bt() (+93 more)
### Community 1 - "Cover Fetch Test Helpers"
Cohesion: 0.10
Nodes (84): floatPtr(), testConfig(), getCover(), Cookie, Handler, ResponseRecorder, T, TestListRendersAcquiredCover() (+76 more)
### Community 2 - "Manga Userscript Adapters"
Cohesion: 0.06
Nodes (76): adapterFor(), anchorsFromDocument(), anchorsFromHTML(), apiDelete(), apiGet(), apiPut(), applyFabPos(), applyLatestChapterIfChanged() (+68 more)
### Community 3 - "Novel Userscript Adapters"
Cohesion: 0.06
Nodes (79): adapterFor(), anchorsFromDocument(), anchorsFromHTML(), apiDelete(), apiGet(), apiPut(), applyFabPos(), applyLatestChapterIfChanged() (+71 more)
### Community 4 - "Series Acquisition Tests"
Cohesion: 0.10
Nodes (67): bookmarkNewKaganeSeries(), bookmarkNewNovelfullSeries(), bookmarkNewSeries(), Context, Store, T, newAcquirer(), readBookmark() (+59 more)
### Community 5 - "Bookmarks API Tests"
Cohesion: 0.08
Nodes (67): auth(), getBookmarks(), Handler, Request, Store, T, newTestServer(), newTestStore() (+59 more)
### Community 7 - "Cover & Acquire Internals"
Cohesion: 0.09
Nodes (31): Addr, Context, Store, defaultCoverResolver(), fetchCoverBytes(), Client, Context, NewCoverFetcher() (+23 more)
### Community 8 - "System Architecture Concepts"
Cohesion: 0.10
Nodes (26): Confirm-Gated Destructive Actions, Discord OAuth & Guild-Membership Gate, Lifecycle Buckets (reading/archived/finished), Reader-Owned Store, HMAC-Derived Reader Credentials, Web Package (Browser UI + Templates), app.html — App Shell Template, Manga/Novel Library Switch (+18 more)
### Community 9 - "Session Middleware"
Cohesion: 0.07
Nodes (35): ClearCookie(), ClientIP(), Duration, Mutex, Request, ResponseWriter, Time, isHTTPS() (+27 more)
### Community 10 - "Go Test Helpers"
Cohesion: 0.05
Nodes (39): Test Helpers, Test Timeout, Basic Handler Test, HTTP Handler Testing, Query Parameters and Headers, Docker Compose Fixture, Integration Testing, SQL Schema Fixture (+31 more)
### Community 11 - "Store Tests"
Cohesion: 0.07
Nodes (79): M, TestMain(), M, TestMain(), M, Main(), start(), URL() (+71 more)
### Community 12 - "Bookmarks API Handler"
Cohesion: 0.08
Nodes (33): Handler, Request, ResponseWriter, Store, Healthz(), writeJSON(), Auth(), compressible() (+25 more)
### Community 13 - "Web UI Handlers"
Cohesion: 0.06
Nodes (24): coverRelativePath(), coverSourceAddress(), displayChapter(), Store, currentLib(), currentTab(), filterBookmarks(), Client (+16 more)
### Community 14 - "Go Error Handling"
Cohesion: 0.06
Nodes (33): Creating Errors, Custom Error Types, Custom types that wrap other errors, Decision table: which error strategy to use, Error Creation, Error String Conventions, Errors as Values, `errors.New` — static error messages (+25 more)
### Community 15 - "CDP Browser Client"
Cohesion: 0.08
Nodes (31): awaitPromise(), browserConnectionLost(), classifyBrowserError(), Action, Context, Mutex, jsString(), kaganeAPIURL() (+23 more)
### Community 17 - "Go Code Style Guide"
Cohesion: 0.08
Nodes (23): Code Style Details, Extract Complex Conditions, Value vs Pointer Arguments, Code Organization Within Files, Complex Conditions & Init Scope, Composite Literals, Control Flow, Cross-References (+15 more)
### Community 20 - "I/O Performance Patterns"
Cohesion: 0.11
Nodes (18): Avoid io.ReadAll for large payloads, Batch Operations, Buffered I/O, Cgo Overhead, Channel: batch processing from a stream, Concurrent Multi-Stage Pipelines, Connection pooling, Database: batch inserts over row-by-row (+10 more)
### Community 21 - "CPU Optimization"
Cohesion: 0.13
Nodes (15): Cache Locality, Contiguous 2D allocation, CPU Optimization, False Sharing, Function Inlining, Handling CPU-specific instruction sets, Instruction-Level Parallelism, Monotonic Time (+7 more)
### Community 22 - "Caching Patterns"
Cohesion: 0.13
Nodes (14): Algorithmic Complexity, Avoid iterator chains, Caching Patterns, Compiled Pattern Caching, Early returns and short-circuit loops, LRU caches, Map lookups over slice scanning, Precomputed lookup tables (+6 more)
### Community 23 - "Browser Entrypoint"
Cohesion: 0.35
Nodes (14): browser_alive(), connection(), connection_signal(), finish_connection(), has_connections(), lock(), reaper(), entrypoint.sh script (+6 more)
### Community 24 - "Memory Allocation & GC"
Cohesion: 0.13
Nodes (15): Allocation Rate Reduction, Ballast pattern (pre-Go 1.19), Garbage Collector Tuning, GC pacing, GC Profiling and Diagnostics, GODEBUG=gctrace=1, GOGC (default: 100), GOMAXPROCS in Containers (+7 more)
### Community 25 - "Cover Fetcher Tests"
Cohesion: 0.33
Nodes (12): coverResponse(), Request, T, TestCoverFetcherCanonicalisesJpgAlias(), TestCoverFetcherFetchesPublicHTTPSImage(), TestCoverFetcherRefusesUnsafeDestinationsBeforeRequest(), TestCoverFetcherRejectsNonImage(), TestCoverFetcherRejectsOversizedBody() (+4 more)
### Community 26 - "readSeries"
Cohesion: 0.13
Nodes (28): asuraLatestChapter(), browserOnlyCoverURL(), comixCoverEntry(), comixCoverURL(), comixLatestChapter(), comixSeriesID(), coverFrom(), demonicLatestChapter() (+20 more)
### Community 27 - "Find Skills Guide"
Cohesion: 0.14
Nodes (13): Common Skill Categories, Find Skills, How to Help Users Find Skills, Step 1: Understand What They Need, Step 2: Check the Leaderboard First, Step 3: Search for Skills, Step 4: Verify Quality Before Recommending, Step 5: Present Options to the User (+5 more)
### Community 28 - "Allocation Patterns"
Cohesion: 0.14
Nodes (14): Allocation Patterns, Backing Array Leaks, Direct indexing vs append, Eliminate redundant map lookups, Interface boxing, Map never shrinks, Map size hints, Memory Optimization (+6 more)
### Community 29 - "Observability & Alerting"
Cohesion: 0.22
Nodes (9): Alerting rules (examples), CPU saturation, GC pressure, Goroutine leaks, Grafana Dashboards, Memory leaks, Prometheus Metrics for Go, PromQL Queries for Performance Diagnosis (+1 more)
### Community 31 - "Memory Layout"
Cohesion: 0.40
Nodes (5): Map of pointers for large, frequently updated structs, Memory Layout, Pointer receivers for large structs, Struct field alignment, Zero-size field at end of struct
### Community 33 - "Go Testing Guide"
Cohesion: 0.20
Nodes (10): CI Regression Detection, Common Mistakes, Core Philosophy, Cross-References, Decision Tree: Where Is Time Spent?, Deep Dives, Go Performance Optimization, Iterative Optimization Methodology (+2 more)
### Community 34 - "Session Store"
Cohesion: 0.28
Nodes (4): Duration, Store, Time, Session
### Community 35 - "Web UI Filter Logic"
Cohesion: 0.31
Nodes (5): closeCardPanels(), setActiveTab(), toggleChapterForm(), toggleConfirmRow(), togglePanel()
### Community 37 - "Product & Security Context"
Cohesion: 0.17
Nodes (11): Accessibility & Inclusion, Brand Commitments, Capabilities and Constraints, Evidence on Hand, Operating Context, Platform, Positioning, Product (+3 more)
### Community 38 - "novel-logic.test.js"
Cohesion: 0.33
Nodes (5): ADR-0009: A Site answers fixed questions; an unusual Site owns its own fetch, Consequences, Considered options, Decision, Why
### Community 39 - "UI Critique 2026-07-26A"
Cohesion: 0.29
Nodes (6): Design Health Score, Design Specificity Verdict, Minor Observations, Persona Red Flags, Priority Issues, Questions to Consider
### Community 40 - "UI Critique 2026-07-26B"
Cohesion: 0.29
Nodes (6): Design Health Score, Design Specificity Verdict, Minor Observations, Persona Red Flags, Priority Issues, Questions to Consider
### Community 45 - "Go Perf Alert Rules"
Cohesion: 0.50
Nodes (4): Prometheus Alerting Rules (Go Performance), GoroutineLeak Alert, HighGCPauseTime Alert, MemoryNearLimit Alert
### Community 46 - "Userscript Display Logic"
Cohesion: 0.07
Nodes (27): 1. Summary answer table, 2. The two reference pages, 3. Chapter page → series URL: every in-page pointer, in priority order, 4.1 Sample method, 4.2 Divergence results, 4.3 Is there a derivable rule? **No.**, 4.4 The split case — a slug can change *mid-series*, 4. The reverse direction, and how common divergence is (+19 more)
### Community 47 - "Go Perf Skill Docs"
Cohesion: 0.22
Nodes (5): Continuous Profiling, Production Observability for Performance, Pyroscope pull mode (via Grafana Alloy), Pyroscope push mode, Real-Time Visualization (Development)
### Community 48 - "Login Page Art"
Cohesion: 0.67
Nodes (3): Fantasy Sword, Fiery Volcanic Scene, Login Art: Sword in Volcanic Rock
### Community 49 - "BookmarkManager Logo"
Cohesion: 1.00
Nodes (3): Mirrored Double Bookmark Mark, Ember Flame Accent, BookmarkManager Logo
### Community 103 - "bookmark-api Service"
Cohesion: 0.24
Nodes (10): Backend Go Service (stdlib net/http), Browser Sidecar (BROWSER_WS_URL), Browser Sidecar Service, CDP Endpoint (Tailnet-Bound :9222), Persistent Chrome Profile Volume, chrome/docker-compose.yml — Browser Deployable Unit, bookmark-api Prod Override, CDP Never on Shared Proxy Network (+2 more)
### Community 104 - "AGENTS.md"
Cohesion: 0.12
Nodes (14): Agent skills, Architecture, Commands, Comments, Design system, Domain docs, Forge: Gitea, not GitHub, graphify (+6 more)
### Community 105 - "reviewer.md"
Cohesion: 0.12
Nodes (15): Assessment, Calibration, Critical (Must Fix), Do Not Trust the Report, Important (Should Fix), Inputs, Issues, Method (+7 more)
### Community 106 - "Redeploy runbook"
Cohesion: 0.12
Nodes (15): 0. Preflight, 1. Back up the database, 2. Pull the new code, 3. Rebuild and restart, 4. Verify the deploy, 5. Smoke-test the full loop, 6. Rollback, 7. The whole thing, as one block (+7 more)
### Community 107 - "1. Backend"
Cohesion: 0.13
Nodes (14): 1. Backend, 2. Userscript, Adapter reference (verified live 2026-07-24), Config (env), Deploy behind your reverse proxy, Desktop iteration (optional), Develop / test, Endpoints (+6 more)
### Community 108 - "Deployment"
Cohesion: 0.14
Nodes (13): 0. Prerequisites, 1. Configure `.env`, 1b. Web UI, 2. Build + start, 3. Verify over HTTPS, 4. Configure the userscript, 5. Install on Bromite, 6. Smoke-test the full loop (+5 more)
### Community 109 - "Cinder — BookmarkManager design system"
Cohesion: 0.20
Nodes (9): 1. The one idea, 2. Tokens, 3. Type, 4. Components (web UI), 5. Components (userscript panel), 6. Motion, 7. Accessibility floor (not negotiable), 8. Adding something new — checklist (+1 more)
### Community 110 - "Implement tickets"
Cohesion: 0.22
Nodes (8): 1. Collect the tickets, 2. Plan the batch, 3. Get the plan approved, 4. Run a wave, 5. Land the wave, 6. Close the batch, Implement tickets, Ticket #<n> — <title>
### Community 111 - "SQLite → Postgres cutover runbook"
Cohesion: 0.22
Nodes (8): 0. The generator is throwaway, 1. Stop the old API and take a fresh export, 2. Bring up Postgres with the schema and the owner Reader, 3. Generate the import SQL, 4. Review it by eye, 5. Apply it, 6. Afterwards, SQLite → Postgres cutover runbook
### Community 112 - "Testing the userscript"
Cohesion: 0.29
Nodes (6): Adding a test, Commands, Gotchas, How the harness works, Testing the userscript, What is NOT testable here
### Community 113 - "ADR-0007: The backend hosts every Site's Cover bytes"
Cohesion: 0.29
Nodes (6): ADR-0007: The backend hosts every Site's Cover bytes, Consequences, Considered options, Decision, Two deliberate relaxations, Why a future reader will find this surprising
### Community 114 - "Issue tracker: Gitea (`tea` CLI)"
Cohesion: 0.29
Nodes (6): Conventions, Issue tracker: Gitea (`tea` CLI), Pull requests as a triage surface, Wayfinding operations, When a skill says "fetch the relevant ticket", When a skill says "publish to the issue tracker"
### Community 115 - "ADR-0006: The browser runs on the home machine, over the tailnet"
Cohesion: 0.33
Nodes (5): ADR-0006: The browser runs on the home machine, over the tailnet, Consequences, Constraints, Decision, Why
### Community 116 - "ADR-0008: A Series identity is discovered from the Site's links, never derived from an address"
Cohesion: 0.33
Nodes (5): ADR-0008: A Series identity is discovered from the Site's links, never derived from an address, Consequences, Considered options, Decision, Why
### Community 117 - "Domain Docs"
Cohesion: 0.33
Nodes (5): Before exploring, read these, Domain Docs, File structure, Flag ADR conflicts, Use the glossary's vocabulary
### Community 118 - "ticket-implementer.md"
Cohesion: 0.33
Nodes (5): Escalate rather than guess, Order of work, Report, Review, The worktree is your whole world
### Community 119 - "implementer.md"
Cohesion: 0.33
Nodes (5): Before You Begin, Report Format, Self-Review Before Reporting, When You're in Over Your Head, Your Job
### Community 120 - "Series is a shared entity, and only the Poll may update it"
Cohesion: 0.40
Nodes (4): Consequences, Only the Poll writes Series fields, Series is a shared entity, and only the Poll may update it, Why
### Community 121 - "Postgres replaces SQLite as the primary datastore"
Cohesion: 0.50
Nodes (3): Consequences, Considered options, Postgres replaces SQLite as the primary datastore
### Community 122 - "Identity comes from Discord OAuth; we store no passwords and send no email"
Cohesion: 0.50
Nodes (3): Consequences, Considered options, Identity comes from Discord OAuth; we store no passwords and send no email
### Community 123 - "The wire format stays flat and deliberately does not mirror the schema"
Cohesion: 0.50
Nodes (3): Consequence, The wire format stays flat and deliberately does not mirror the schema, Why a future reader will find this surprising
### Community 124 - "ADR-0005: On-demand browser sidecar"
Cohesion: 0.50
Nodes (3): ADR-0005: On-demand browser sidecar, Constraints, Decision
### Community 273 - "AGENTS.md"
Cohesion: 0.50
Nodes (3): Live URL shapes (verified 2026-07-26, may drift — re-check against live pages before trust), Second script: `novel-bookmark.user.js`, Userscript structure (single IIFE, `manga-bookmark.user.js`)
## Knowledge Gaps
- **508 isolated node(s):** `bookmarkmanager/backend`, `ctxKey`, `loginView`, `ctxKey`, `test` (+503 more)
These have ≤1 connection - possible missing edges or undocumented components.
- **170 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
## Suggested Questions
_Questions this graph is uniquely positioned to answer:_
- **Why does `New()` connect `Series Acquisition Tests` to `Bookmarks API Tests`, `Cover & Acquire Internals`, `Session Middleware`, `Store Tests`, `Web UI Handlers`?**
_High betweenness centrality (0.052) - this node is a cross-community bridge._
- **Why does `Open()` connect `Store Tests` to `Cover Fetch Test Helpers`, `Web UI Handlers`, `Series Acquisition Tests`, `Bookmarks API Tests`?**
_High betweenness centrality (0.034) - this node is a cross-community bridge._
- **Why does `newRouter()` connect `Bookmarks API Tests` to `Cover Fetch Test Helpers`, `Bookmarks API Handler`, `Series Acquisition Tests`?**
_High betweenness centrality (0.027) - this node is a cross-community bridge._
- **Are the 47 inferred relationships involving `testConfig()` (e.g. with `TestListRendersAcquiredCover()` and `TestPublicCoverNeverEchoesNonImage()`) actually correct?**
_`testConfig()` has 47 INFERRED edges - model-reasoned connections that need verification._
- **Are the 8 inferred relationships involving `newWebTestServer()` (e.g. with `TestListRendersAcquiredCover()` and `TestPublicCoverRejectsUnknownAddress()`) actually correct?**
_`newWebTestServer()` has 8 INFERRED edges - model-reasoned connections that need verification._
- **Are the 6 inferred relationships involving `newTestStore()` (e.g. with `TestCreateAndGetSession()` and `TestDeleteSessionIsPerReader()`) actually correct?**
_`newTestStore()` has 6 INFERRED edges - model-reasoned connections that need verification._
- **What connects `bookmarkmanager/backend`, `ctxKey`, `loginView` to the rest of the system?**
_548 weakly-connected nodes found - possible documentation gaps or missing edges._
File diff suppressed because one or more lines are too long
File diff suppressed because it is too large Load Diff
+622
View File
@@ -0,0 +1,622 @@
{
".agents/skills/golang-code-style/evals/evals.json": {
"mtime": 1784884678.6627614,
"ast_hash": "bec0e12446e7af3cd05de9b6d42badd8",
"semantic_hash": "bec0e12446e7af3cd05de9b6d42badd8"
},
".agents/skills/golang-error-handling/evals/evals.json": {
"mtime": 1784884678.6655047,
"ast_hash": "275d710b774fba1e1d0bc098d3646c6d",
"semantic_hash": "275d710b774fba1e1d0bc098d3646c6d"
},
".agents/skills/golang-performance/evals/evals.json": {
"mtime": 1784884678.6688662,
"ast_hash": "4f06df87f90aa0e4f6deaa318b47689f",
"semantic_hash": "4f06df87f90aa0e4f6deaa318b47689f"
},
".agents/skills/golang-testing/evals/evals.json": {
"mtime": 1784884678.6721346,
"ast_hash": "60a821bbfd20c6fe8bba996b8b540dd4",
"semantic_hash": "60a821bbfd20c6fe8bba996b8b540dd4"
},
"backend/go.mod": {
"mtime": 1786216141.668644,
"ast_hash": "dac242903b0e98c3e4395159d609e08e",
"semantic_hash": "dac242903b0e98c3e4395159d609e08e"
},
"backend/main.go": {
"mtime": 1786363889.5678573,
"ast_hash": "6e98a3ae91aaa132df251e43c4dfca6d",
"semantic_hash": "6e98a3ae91aaa132df251e43c4dfca6d"
},
"skills-lock.json": {
"mtime": 1784884678.6842625,
"ast_hash": "4a94ac85bad6bce330d085bcc0ae3ffd",
"semantic_hash": "4a94ac85bad6bce330d085bcc0ae3ffd"
},
"userscript/manga-bookmark.user.js": {
"mtime": 1786488438.9080842,
"ast_hash": "1f8bcddd3632d709f058a8401af8f127",
"semantic_hash": ""
},
".agents/skills/find-skills/SKILL.md": {
"mtime": 1784884338.760326,
"ast_hash": "62b297abdee9aea84c577ab2e04e1974",
"semantic_hash": "62b297abdee9aea84c577ab2e04e1974"
},
".agents/skills/golang-code-style/SKILL.md": {
"mtime": 1784884678.6623824,
"ast_hash": "d6a01e6f64550a5c8d59dac2e948000e",
"semantic_hash": "d6a01e6f64550a5c8d59dac2e948000e"
},
".agents/skills/golang-code-style/references/details.md": {
"mtime": 1784884678.6627865,
"ast_hash": "19891e396a986f1b24bf34b7a2854fcb",
"semantic_hash": "19891e396a986f1b24bf34b7a2854fcb"
},
".agents/skills/golang-error-handling/SKILL.md": {
"mtime": 1784884678.665052,
"ast_hash": "8b7970f472adb240e5bc4bde863d43a6",
"semantic_hash": "8b7970f472adb240e5bc4bde863d43a6"
},
".agents/skills/golang-error-handling/references/error-creation.md": {
"mtime": 1784884678.665524,
"ast_hash": "248dbf75492c68faef2334b8d83bd080",
"semantic_hash": "248dbf75492c68faef2334b8d83bd080"
},
".agents/skills/golang-error-handling/references/error-handling.md": {
"mtime": 1784884678.6655412,
"ast_hash": "c2424ee3999b05963b199e0100727aa3",
"semantic_hash": "c2424ee3999b05963b199e0100727aa3"
},
".agents/skills/golang-error-handling/references/error-wrapping.md": {
"mtime": 1784884678.6655717,
"ast_hash": "4a15d9266a1ea0c8bb1951e6b9c0f586",
"semantic_hash": "4a15d9266a1ea0c8bb1951e6b9c0f586"
},
".agents/skills/golang-performance/SKILL.md": {
"mtime": 1784884678.667833,
"ast_hash": "35a15fd129c5bedaada3fd4df0b6ba8c",
"semantic_hash": "35a15fd129c5bedaada3fd4df0b6ba8c"
},
".agents/skills/golang-performance/assets/prometheus-alerts.yml": {
"mtime": 1784884678.668527,
"ast_hash": "fa9357ffa87c4f894fc21afa9707c4db",
"semantic_hash": "fa9357ffa87c4f894fc21afa9707c4db"
},
".agents/skills/golang-performance/references/caching.md": {
"mtime": 1784884678.6690567,
"ast_hash": "807c42a82994e5548dbf7eb6e30c71e0",
"semantic_hash": "807c42a82994e5548dbf7eb6e30c71e0"
},
".agents/skills/golang-performance/references/cpu.md": {
"mtime": 1784884678.669087,
"ast_hash": "6d52e532ef51a35cc4134694c944517d",
"semantic_hash": "6d52e532ef51a35cc4134694c944517d"
},
".agents/skills/golang-performance/references/io-networking.md": {
"mtime": 1784884678.6691036,
"ast_hash": "95c5dd51f728fd69c945a92ff021d766",
"semantic_hash": "95c5dd51f728fd69c945a92ff021d766"
},
".agents/skills/golang-performance/references/memory.md": {
"mtime": 1784884678.6691158,
"ast_hash": "3b2108df06b4cfb3980fa80bbd9ebcff",
"semantic_hash": "3b2108df06b4cfb3980fa80bbd9ebcff"
},
".agents/skills/golang-performance/references/observability.md": {
"mtime": 1784884678.6691446,
"ast_hash": "0aa8a498e8d55ccdd4990ad187bae828",
"semantic_hash": "0aa8a498e8d55ccdd4990ad187bae828"
},
".agents/skills/golang-performance/references/runtime.md": {
"mtime": 1784884678.669426,
"ast_hash": "26386c33b3a3794aaef0556713bf8f3a",
"semantic_hash": "26386c33b3a3794aaef0556713bf8f3a"
},
".agents/skills/golang-testing/SKILL.md": {
"mtime": 1784884678.6716368,
"ast_hash": "0a9b9793bba2a239db94e980272a393e",
"semantic_hash": "0a9b9793bba2a239db94e980272a393e"
},
".agents/skills/golang-testing/references/helpers.md": {
"mtime": 1784884678.6722167,
"ast_hash": "0aecb7cfbeb9b374bf61c52e324d9d3f",
"semantic_hash": "0aecb7cfbeb9b374bf61c52e324d9d3f"
},
".agents/skills/golang-testing/references/http-testing.md": {
"mtime": 1784884678.67225,
"ast_hash": "9111110c28a7fbbffc3537aad786b390",
"semantic_hash": "9111110c28a7fbbffc3537aad786b390"
},
".agents/skills/golang-testing/references/integration-testing.md": {
"mtime": 1784884678.67225,
"ast_hash": "fcf9861bc36ae56e3fb7a1c18aa0e77f",
"semantic_hash": "fcf9861bc36ae56e3fb7a1c18aa0e77f"
},
".agents/skills/golang-testing/references/mocking.md": {
"mtime": 1784884678.6722653,
"ast_hash": "3a08979e4603aae5c32a58d5b6c39765",
"semantic_hash": "3a08979e4603aae5c32a58d5b6c39765"
},
"CLAUDE.md": {
"mtime": 1786488493.5922732,
"ast_hash": "7362a25f37333a2a6a55a5aeccf9b0cb",
"semantic_hash": ""
},
"DEPLOY.md": {
"mtime": 1786488464.5532806,
"ast_hash": "2b7b537aa1c0954400c19acc4b634029",
"semantic_hash": ""
},
"README.md": {
"mtime": 1786488487.8305292,
"ast_hash": "9d6be8aa8a2946c23ad48d8f2864b5ca",
"semantic_hash": ""
},
"docker-compose.prod.yml": {
"mtime": 1786292465.8305523,
"ast_hash": "0751998a532297b8ac507a01ec48dc31",
"semantic_hash": "0751998a532297b8ac507a01ec48dc31"
},
"docker-compose.yml": {
"mtime": 1786488450.3508112,
"ast_hash": "124fd581bf0a662ff15012abfdb40a92",
"semantic_hash": ""
},
".claude/settings.json": {
"mtime": 1784951973.1869545,
"ast_hash": "e51077b6a7f1f67afc748f1a32a1557d",
"semantic_hash": "e51077b6a7f1f67afc748f1a32a1557d"
},
"backend/web_test.go": {
"mtime": 1786216141.700644,
"ast_hash": "8f1b093b59eb1ed81bc7fc0c22495c50",
"semantic_hash": "8f1b093b59eb1ed81bc7fc0c22495c50"
},
"backend/main_test.go": {
"mtime": 1786363889.5678573,
"ast_hash": "8a165955cf28ad47481fec5ea7afb3d6",
"semantic_hash": "8a165955cf28ad47481fec5ea7afb3d6"
},
".claude/settings.local.json": {
"mtime": 1785697645.350201,
"ast_hash": "9a1ac6369f968e8df4be9dcff0948f70",
"semantic_hash": "9a1ac6369f968e8df4be9dcff0948f70"
},
"PRODUCT.md": {
"mtime": 1786216141.660644,
"ast_hash": "c52072d1978286060087fa0686f9c7f9",
"semantic_hash": "c52072d1978286060087fa0686f9c7f9"
},
"backend/.impeccable/critique/2026-07-26T15-50-42Z__backend-templates-app-html.md": {
"mtime": 1785128576.2412457,
"ast_hash": "d08627c27f22d453125db6c1ab4b71ec",
"semantic_hash": "d08627c27f22d453125db6c1ab4b71ec"
},
"backend/.impeccable/critique/2026-07-26T17-08-41Z__backend-templates-app-html.md": {
"mtime": 1785128576.2452524,
"ast_hash": "e69a8340a371579ca3ea689660f7d7bd",
"semantic_hash": "e69a8340a371579ca3ea689660f7d7bd"
},
"AGENTS.md": {
"mtime": 1786488493.5922732,
"ast_hash": "7362a25f37333a2a6a55a5aeccf9b0cb",
"semantic_hash": ""
},
"userscript/test/logic.test.js": {
"mtime": 1786488563.581765,
"ast_hash": "80d512bfe6fa8b847f3ba6169c321a74",
"semantic_hash": ""
},
".claude/skills/testing-the-userscript/SKILL.md": {
"mtime": 1786363889.5489495,
"ast_hash": "8f3c0132eb4787a2c8736eb99f7689af",
"semantic_hash": "8f3c0132eb4787a2c8736eb99f7689af"
},
"REDEPLOY.md": {
"mtime": 1786363889.552731,
"ast_hash": "d0baf08b95e7b5986234a9f36759c12e",
"semantic_hash": "d0baf08b95e7b5986234a9f36759c12e"
},
"docs/design-system.md": {
"mtime": 1786022513.9623306,
"ast_hash": "421cd7e57f02d4b467f120ca6ddd7b6a",
"semantic_hash": "421cd7e57f02d4b467f120ca6ddd7b6a"
},
"backend/api_test.go": {
"mtime": 1786488480.637822,
"ast_hash": "8e4b9293bc2e45ee3f42027315594fd5",
"semantic_hash": ""
},
"backend/cover_test.go": {
"mtime": 1786363889.552731,
"ast_hash": "c7e313d6c92eb28e6d370e5e89035984",
"semantic_hash": "c7e313d6c92eb28e6d370e5e89035984"
},
"backend/internal/api/handlers.go": {
"mtime": 1786363889.552731,
"ast_hash": "59e6b8767ab19839bb8f82891a7e4616",
"semantic_hash": "59e6b8767ab19839bb8f82891a7e4616"
},
"backend/internal/httpmw/middleware.go": {
"mtime": 1786216141.672644,
"ast_hash": "385b36f58488b7e6d93eb6d6034e9ee3",
"semantic_hash": "385b36f58488b7e6d93eb6d6034e9ee3"
},
"backend/internal/latest/browser.go": {
"mtime": 1786469354.923288,
"ast_hash": "ed129a7f00601ea90c877ff29fa21220",
"semantic_hash": ""
},
"backend/internal/latest/browser_test.go": {
"mtime": 1786262323.6964688,
"ast_hash": "e900f92971486f47d7ef76e9a95217fe",
"semantic_hash": "e900f92971486f47d7ef76e9a95217fe"
},
"backend/internal/latest/fetch.go": {
"mtime": 1786446006.9219902,
"ast_hash": "3e20ad86aa46783e9aa95c2b746551ee",
"semantic_hash": ""
},
"backend/internal/latest/poller.go": {
"mtime": 1786469641.8342345,
"ast_hash": "44fef6074ac2eaffc8233f46aad5236b",
"semantic_hash": ""
},
"backend/internal/latest/poller_test.go": {
"mtime": 1786469644.9306462,
"ast_hash": "64bc838c822f1bf33bbf9e291215454b",
"semantic_hash": ""
},
"backend/internal/latest/sites.go": {
"mtime": 1786469348.5394242,
"ast_hash": "b744cc685363317a526cc3bebceea39e",
"semantic_hash": ""
},
"backend/internal/latest/sites_test.go": {
"mtime": 1786446006.9219902,
"ast_hash": "eabca9014a306e3c71d238b0ae499f61",
"semantic_hash": ""
},
"backend/internal/latest/smoke_image_test.go": {
"mtime": 1786363889.5602942,
"ast_hash": "db068cb59575f8c82669acbaf84bcaed",
"semantic_hash": "db068cb59575f8c82669acbaf84bcaed"
},
"backend/internal/pgtest/pgtest.go": {
"mtime": 1786216141.6766438,
"ast_hash": "f60372d41516e66f7aaeb272da227d6e",
"semantic_hash": "f60372d41516e66f7aaeb272da227d6e"
},
"backend/internal/session/session.go": {
"mtime": 1786216141.6766438,
"ast_hash": "9952474ffb22c825d6f075b866ee26f4",
"semantic_hash": "9952474ffb22c825d6f075b866ee26f4"
},
"backend/internal/session/session_test.go": {
"mtime": 1786216141.680644,
"ast_hash": "37ffd00964e7a67350c68ed50c6503c5",
"semantic_hash": "37ffd00964e7a67350c68ed50c6503c5"
},
"backend/internal/store/migrations/0001_bookmarks.sql": {
"mtime": 1786216141.680644,
"ast_hash": "f87ccfb2c25c43f93021177ced0bfae4",
"semantic_hash": "f87ccfb2c25c43f93021177ced0bfae4"
},
"backend/internal/store/migrations/0002_series.sql": {
"mtime": 1786216141.6820722,
"ast_hash": "5dc98771e0c0f6e8416b434c280b0efb",
"semantic_hash": "5dc98771e0c0f6e8416b434c280b0efb"
},
"backend/internal/store/migrations/0003_reader.sql": {
"mtime": 1786216141.6820722,
"ast_hash": "444a97799f38f6222f87d4e9fb2d6258",
"semantic_hash": "444a97799f38f6222f87d4e9fb2d6258"
},
"backend/internal/store/migrations/0004_owner_bookmarks.sql": {
"mtime": 1786216141.684644,
"ast_hash": "e4fa900cc223865d3ecd4c60c5707a65",
"semantic_hash": "e4fa900cc223865d3ecd4c60c5707a65"
},
"backend/internal/store/migrations/0005_sessions.sql": {
"mtime": 1786216141.684644,
"ast_hash": "5158887ebc57cf8c16a7b821b61cd760",
"semantic_hash": "5158887ebc57cf8c16a7b821b61cd760"
},
"backend/internal/store/migrations/0006_reader_token_epoch.sql": {
"mtime": 1786216141.684644,
"ast_hash": "3093cc1c3aae0cd9643d04105d045402",
"semantic_hash": "3093cc1c3aae0cd9643d04105d045402"
},
"backend/internal/store/sessions.go": {
"mtime": 1786216141.684644,
"ast_hash": "eee3510cc6172ef4b1da820474c26b01",
"semantic_hash": "eee3510cc6172ef4b1da820474c26b01"
},
"backend/internal/store/sessions_test.go": {
"mtime": 1786216141.684644,
"ast_hash": "0b6764a0ee20f5cb7748eecd31a1d220",
"semantic_hash": "0b6764a0ee20f5cb7748eecd31a1d220"
},
"backend/internal/store/store.go": {
"mtime": 1786363889.5602942,
"ast_hash": "54367a8ab043983e2491b2eb2650961c",
"semantic_hash": "54367a8ab043983e2491b2eb2650961c"
},
"backend/internal/store/store_test.go": {
"mtime": 1786363889.5602942,
"ast_hash": "dc823fd77bcce2268114e31d759b20a5",
"semantic_hash": "dc823fd77bcce2268114e31d759b20a5"
},
"backend/internal/userscript/userscript.go": {
"mtime": 1786216141.692644,
"ast_hash": "aa13a71b1c9eefe4930fd31f27722328",
"semantic_hash": "aa13a71b1c9eefe4930fd31f27722328"
},
"backend/internal/userscript/userscript_test.go": {
"mtime": 1786216141.692644,
"ast_hash": "6c050968d7b8b67956da1a3136d2c3c7",
"semantic_hash": "6c050968d7b8b67956da1a3136d2c3c7"
},
"backend/internal/web/discord.go": {
"mtime": 1786216141.692644,
"ast_hash": "69e4959c65fa67d7491aedf6a71bb575",
"semantic_hash": "69e4959c65fa67d7491aedf6a71bb575"
},
"backend/internal/web/oauth_test.go": {
"mtime": 1786216141.692644,
"ast_hash": "a3bddeb70dd8d7eb14139da808723ea8",
"semantic_hash": "a3bddeb70dd8d7eb14139da808723ea8"
},
"backend/internal/web/static/filter.js": {
"mtime": 1786022513.9473197,
"ast_hash": "b4ee3306201bfd88b148b96801972617",
"semantic_hash": "b4ee3306201bfd88b148b96801972617"
},
"backend/internal/web/static/htmx.min.js": {
"mtime": 1785873769.5512016,
"ast_hash": "19a573773be4ca22570ca2f8543120c5",
"semantic_hash": "19a573773be4ca22570ca2f8543120c5"
},
"backend/internal/web/web.go": {
"mtime": 1786363889.5678573,
"ast_hash": "8308949c658d3a08ce1c3cdbea6a907c",
"semantic_hash": "8308949c658d3a08ce1c3cdbea6a907c"
},
"backend/reader_credential_test.go": {
"mtime": 1786216141.700644,
"ast_hash": "2a751ea6d9c06635aa3179db0aef1b2b",
"semantic_hash": "2a751ea6d9c06635aa3179db0aef1b2b"
},
"chrome/entrypoint.sh": {
"mtime": 1786363889.5678573,
"ast_hash": "8008a187690764436540fab47ba0cfcc",
"semantic_hash": "8008a187690764436540fab47ba0cfcc"
},
"userscript/novel-bookmark.user.js": {
"mtime": 1786454823.798836,
"ast_hash": "834effb0821f8d6c9f57f6554a5db462",
"semantic_hash": ""
},
"userscript/test/novel-logic.test.js": {
"mtime": 1786454566.070801,
"ast_hash": "b25a7377af210dd0aff8284501fd0252",
"semantic_hash": ""
},
".opencode/agent/implementer.md": {
"mtime": 1785873769.5402634,
"ast_hash": "000de469c18d68352027e10c8ce8acfb",
"semantic_hash": "000de469c18d68352027e10c8ce8acfb"
},
".opencode/agent/reviewer.md": {
"mtime": 1785873769.5416775,
"ast_hash": "e44a2f6f624db044e19508bc5ab05592",
"semantic_hash": "e44a2f6f624db044e19508bc5ab05592"
},
"CONTEXT.md": {
"mtime": 1786466512.6719532,
"ast_hash": "544b7d93f1d5cb9d347cb0b92fc3a709",
"semantic_hash": ""
},
"CUTOVER.md": {
"mtime": 1786216141.660644,
"ast_hash": "6c6f3e4c4c2f57867894280bce728c50",
"semantic_hash": "6c6f3e4c4c2f57867894280bce728c50"
},
"backend/AGENTS.md": {
"mtime": 1786363889.552731,
"ast_hash": "6356ee58447f299e5fa7aa83486bbbe0",
"semantic_hash": "6356ee58447f299e5fa7aa83486bbbe0"
},
"backend/CLAUDE.md": {
"mtime": 1786363889.552731,
"ast_hash": "6356ee58447f299e5fa7aa83486bbbe0",
"semantic_hash": "6356ee58447f299e5fa7aa83486bbbe0"
},
"backend/internal/web/templates/app.html": {
"mtime": 1786216141.692644,
"ast_hash": "3965e20e204afb71ba2a3aa86cb7c61c",
"semantic_hash": "3965e20e204afb71ba2a3aa86cb7c61c"
},
"backend/internal/web/templates/card.html": {
"mtime": 1786363889.5678573,
"ast_hash": "ab83ae0dbb34fd40c146a7cc1263173e",
"semantic_hash": "ab83ae0dbb34fd40c146a7cc1263173e"
},
"backend/internal/web/templates/chrome.html": {
"mtime": 1786363889.5678573,
"ast_hash": "d80b27cf3bd9d131075c485dd169dfcc",
"semantic_hash": "d80b27cf3bd9d131075c485dd169dfcc"
},
"backend/internal/web/templates/icons.html": {
"mtime": 1785873769.553139,
"ast_hash": "8e10c507c32934a92463b4bca9e34fe6",
"semantic_hash": "8e10c507c32934a92463b4bca9e34fe6"
},
"backend/internal/web/templates/list.html": {
"mtime": 1786216141.696644,
"ast_hash": "365548aace8c06559a1f66db0ae47256",
"semantic_hash": "365548aace8c06559a1f66db0ae47256"
},
"backend/internal/web/templates/login.html": {
"mtime": 1786216141.696644,
"ast_hash": "bcc3101498a66cf8b79f9d97c6c9cd6b",
"semantic_hash": "bcc3101498a66cf8b79f9d97c6c9cd6b"
},
"backend/internal/web/templates/readers.html": {
"mtime": 1786216141.696644,
"ast_hash": "c5034e76bd20a705d2799cb5ecb328f0",
"semantic_hash": "c5034e76bd20a705d2799cb5ecb328f0"
},
"backend/internal/web/templates/setup.html": {
"mtime": 1786216141.696644,
"ast_hash": "72e93c0b827414063596f7338987d879",
"semantic_hash": "72e93c0b827414063596f7338987d879"
},
"docs/adr/0001-postgresql-over-sqlite.md": {
"mtime": 1786216141.704644,
"ast_hash": "abfb08754cee58be67311377904f8ca4",
"semantic_hash": "abfb08754cee58be67311377904f8ca4"
},
"docs/adr/0002-discord-oauth-no-passwords-no-email.md": {
"mtime": 1786216141.7071996,
"ast_hash": "852a04d86659385085da6ffc8b489933",
"semantic_hash": "852a04d86659385085da6ffc8b489933"
},
"docs/adr/0003-series-shared-and-poll-owned.md": {
"mtime": 1786216141.7071996,
"ast_hash": "58bb4d24e20f6f7adc1a8b9cf7967c3f",
"semantic_hash": "58bb4d24e20f6f7adc1a8b9cf7967c3f"
},
"docs/adr/0004-wire-format-does-not-mirror-the-schema.md": {
"mtime": 1786216141.7071996,
"ast_hash": "a6ea2770dec2156f78a65b35ba06902a",
"semantic_hash": "a6ea2770dec2156f78a65b35ba06902a"
},
"docs/agents/domain.md": {
"mtime": 1786216141.7071996,
"ast_hash": "6f99318ac6cb9825b613bfde55d76091",
"semantic_hash": "6f99318ac6cb9825b613bfde55d76091"
},
"docs/agents/issue-tracker.md": {
"mtime": 1786216141.7087462,
"ast_hash": "1342e66ccb84a84fd579fb6dc0b8243a",
"semantic_hash": "1342e66ccb84a84fd579fb6dc0b8243a"
},
"docs/agents/triage-labels.md": {
"mtime": 1786216141.7087462,
"ast_hash": "69114d07ed792d6bb1d13758ba5435e1",
"semantic_hash": "69114d07ed792d6bb1d13758ba5435e1"
},
"userscript/AGENTS.md": {
"mtime": 1786454605.6727684,
"ast_hash": "e276ffe9a6e7b55fd3235466a1995c22",
"semantic_hash": ""
},
"userscript/CLAUDE.md": {
"mtime": 1786454605.6727684,
"ast_hash": "e276ffe9a6e7b55fd3235466a1995c22",
"semantic_hash": ""
},
"backend/internal/web/static/login-art.png": {
"mtime": 1786022513.9585779,
"ast_hash": "05d7863cba344a946256719a0c9ef959",
"semantic_hash": "05d7863cba344a946256719a0c9ef959"
},
"backend/internal/web/static/logo.svg": {
"mtime": 1786022513.9585779,
"ast_hash": "d0d34d0f08a25b53176cc55989b7babe",
"semantic_hash": "d0d34d0f08a25b53176cc55989b7babe"
},
"backend/internal/store/migrations/0007_covers.sql": {
"mtime": 1786262323.6964688,
"ast_hash": "6033ce0701be1236ed175362363bd96c",
"semantic_hash": "6033ce0701be1236ed175362363bd96c"
},
"docs/adr/0005-on-demand-browser.md": {
"mtime": 1786292465.8305523,
"ast_hash": "8cf2fb8c0a66c2c7d82ae433b99ca42b",
"semantic_hash": "8cf2fb8c0a66c2c7d82ae433b99ca42b"
},
"chrome/docker-compose.yml": {
"mtime": 1786292465.8305523,
"ast_hash": "5605599395a3f085904e78a2bfec1e58",
"semantic_hash": "5605599395a3f085904e78a2bfec1e58"
},
"docs/adr/0006-browser-on-the-home-machine.md": {
"mtime": 1786292465.8305523,
"ast_hash": "dfd6bbc045d23315f2942ee8d98eb7db",
"semantic_hash": "dfd6bbc045d23315f2942ee8d98eb7db"
},
"docs/adr/0007-backend-hosts-cover-bytes.md": {
"mtime": 1786292465.8305523,
"ast_hash": "b19e38045b3dcda7dd59634ed9227a68",
"semantic_hash": "b19e38045b3dcda7dd59634ed9227a68"
},
"backend/internal/store/migrations/0008_filesystem_covers.sql": {
"mtime": 1786363889.5602942,
"ast_hash": "46cf7822d4f667e3cab36b547abe5e97",
"semantic_hash": "46cf7822d4f667e3cab36b547abe5e97"
},
"backend/internal/latest/cover.go": {
"mtime": 1786363889.5565126,
"ast_hash": "e6749cfe3cd7c2e71d4392dde84f55f9",
"semantic_hash": "e6749cfe3cd7c2e71d4392dde84f55f9"
},
"backend/internal/latest/cover_fetch_test.go": {
"mtime": 1786363889.5565126,
"ast_hash": "60d9eb7c59a3751baf4f31c7655217e7",
"semantic_hash": "60d9eb7c59a3751baf4f31c7655217e7"
},
"backend/internal/latest/acquire.go": {
"mtime": 1786468950.9432797,
"ast_hash": "6c1ad34bbe9f5b49d0fd1eae9093f55d",
"semantic_hash": ""
},
"backend/internal/latest/acquire_test.go": {
"mtime": 1786363889.5565126,
"ast_hash": "7bd9f41814f6bf59d8998dfb81cf990a",
"semantic_hash": "7bd9f41814f6bf59d8998dfb81cf990a"
},
"backend/internal/store/migrations/0009_series_cover_address.sql": {
"mtime": 1786363889.5602942,
"ast_hash": "4c1f6328b2e1a95828fad6d88d474c2d",
"semantic_hash": "4c1f6328b2e1a95828fad6d88d474c2d"
},
".claude/skills/implement-tickets/SKILL.md": {
"mtime": 1786417841.7117643,
"ast_hash": "3060e32e19cc571d91871a98f18afe53",
"semantic_hash": "3060e32e19cc571d91871a98f18afe53"
},
".omp/agents/ticket-implementer.md": {
"mtime": 1786417841.7163916,
"ast_hash": "0150a46c0d21c572b71b3d87d21ac925",
"semantic_hash": "0150a46c0d21c572b71b3d87d21ac925"
},
"docs/adr/0008-series-identity-is-discovered-not-derived.md": {
"mtime": 1786417841.7163916,
"ast_hash": "3ce6d64ef6a8a39f27c257b39065389e",
"semantic_hash": "3ce6d64ef6a8a39f27c257b39065389e"
},
"docs/research/lightnovelworld-chapter-vs-series-slug.md": {
"mtime": 1786417841.7163916,
"ast_hash": "74a4e538875f0a7e8ca3d5dc48c0bb53",
"semantic_hash": "74a4e538875f0a7e8ca3d5dc48c0bb53"
},
"backend/internal/latest/smoke_lnw_test.go": {
"mtime": 1786446006.9219902,
"ast_hash": "2d65da8a081759172918fdf159760f45",
"semantic_hash": ""
},
"docs/adr/0009-a-site-answers-questions-its-own-way.md": {
"mtime": 1786469367.6875648,
"ast_hash": "8039012a5b6de2359ff1a47079f51b66",
"semantic_hash": ""
},
"backend/internal/latest/read.go": {
"mtime": 1786469329.7445614,
"ast_hash": "3cf29046ddaef39fafb1df70b9f9ae8c",
"semantic_hash": ""
}
}
+21 -11
View File
@@ -2,7 +2,7 @@ Guidance for OpenCode (and Claude Code) working under `userscript/`. See root `A
### Userscript structure (single IIFE, `manga-bookmark.user.js`)
1. **Site adapters** — one per host, `detect(location, document)` return page `type` + IDs. Identify type/IDs from **URL regex** (most stable); pull `title`/`cover` from **`og:title`/`og:image` meta tags**, not CSS classes.
1. **Site adapters** — one per host, `detect(location, document)` return page `type` + IDs. Identify type/IDs from **URL regex** (most stable); pull `title` from **`og:title`** (or the page heading where a site ships no og: tags), not CSS classes. **No adapter reads a cover**: the backend acquires, stores and serves every Cover from its own origin (ADR-0007), the wire's `cover` is already an address on our origin, and `apiPut` strips any `cover` off an outgoing body.
2. **API client** — `apiGet/apiPut/apiDelete` with bearer header; `localStorage` key `bmgr:manga:cache` for instant render + offline fallback.
3. **Progress logic** — auto-upsert `last_chapter` only when `chapterNum >= stored last_chapter_num` (re-reading old chapters must not regress progress; unparseable -> set current). Manual panel override forces any value.
4. **Retry queue** — every write go through `pushBookmark`/`pushDelete`, so
@@ -52,27 +52,37 @@ Guidance for OpenCode (and Claude Code) working under `userscript/`. See root `A
"Comix — Read Comics online for free" and after an in-page hop it is the
*previous* series' name. `document.title` is the one thing client routing does
update, so titles come from there, with the chapter page's `" · Ch.<n>"` tail
stripped. Covers likewise: `og:image` is absent, so the cover is the `img`
whose `alt` matches the cleaned title — verified live 2026-08-08.
stripped. It publishes no `og:image` either, which is one of the reasons cover
acquisition moved to the backend.
- **kagane.to**: series `/series/<uuid>`, reader
`/series/<uuid>/reader/<bookUuid>`. Reader URLs carry no chapter number, so
the number comes out of `og:title`. Two shapes exist: `"<Series> - Chapter
<n>[ - Episode <n>]"` and, for volume-numbered series, `"<Series> - Volume <v>
Chapter <n>"` with no episode name — both must yield a bare series title, or
the volume tail lands in the bookmark's title. Its covers are challenge- and
CORP-protected, so the web UI proxies them; the userscript still stores the
raw `og:image`. Behind a Cloudflare JS challenge, so the backend polls it
the volume tail lands in the bookmark's title.
Its covers are challenge- and CORP-protected, so nothing outside kagane.to can
load one directly; the panel renders the backend's own cover address like every
other Site. Behind a Cloudflare JS challenge, so the backend polls it
through the headless browser.
- **novelfull.com** (novel script): series `/<slug>.html`, chapter
`/<slug>/chapter-<n>[-<title-slug>].html`. No `og:*` tags at all — title from
`h3.title` (series) or `a.truyen-title` (chapter), cover from
`meta[name="image"]`. Behind a Cloudflare JS challenge no TLS fingerprint
`h3.title` (series) or `a.truyen-title` (chapter); the script reads no cover.
Behind a Cloudflare JS challenge no TLS fingerprint
clears, so the backend polls it through the headless browser.
- **lightnovelworld.net** (novel script): series `/novel/<slug>/`, chapter
`/<slug>-chapter-<n>/` — flat, at the site root. `h1.entry-title` is the clean
title on a series page and `<Title> Chapter <n>` on a chapter page. Chapter
pages carry no `og:image`. Its series page lists every chapter with an
`/<slug>-chapter-<n>/` — flat, at the site root. The chapter path's slug is a
Chapter Slug, not an identity: the Series address is read off the page's
`a[aria-label='All Chapter']` (fallback: the BreadcrumbList's second crumb),
and a Series may publish under several Chapter Slugs. A chapter page with no
pointer resolves to `other`, so no Bookmark is offered. `h1.entry-title` is
the clean title on a series page and `<Title> Chapter <n>` on a chapter page.
Its series page lists every chapter with an
absolute href, so the backend polls it with the plain TLS client.
The client performs no latest-chapter scan for this Site: the Poll's
one-hour cooldown dominates the client's four-hour throttle, so a scan
would add no freshness, and the page's wpdiscuz thread is a public write
surface a scan would have to truncate at. `computeLatestChapter` yields
null here and `backgroundRefreshLatest` skips the Site before any fetch.
### Second script: `novel-bookmark.user.js`
+21 -32
View File
@@ -6,7 +6,6 @@
// @author you
// @downloadURL https://bookmark-api.violetcrown.my.id/u/__API_TOKEN__/manga-bookmark.user.js
// @updateURL https://bookmark-api.violetcrown.my.id/u/__API_TOKEN__/manga-bookmark.user.js
// @match https://asuracomic.net/*
// @match https://asurascans.com/*
// @match https://demonicscans.org/*
// @match https://comix.to/*
@@ -56,9 +55,10 @@
// ============================================================
// Site adapters
//
// Page type + IDs come from URL regex (most stable); title/cover come from
// og: meta tags. Verified live 2026-07-24 against asurascans.com and
// demonicscans.org — see README "Adapter reference".
// Page type + IDs come from URL regex (most stable); the title comes from
// og: meta tags. Covers are never read here: the backend acquires and serves
// them itself (ADR-0007). Verified live 2026-07-24 against asurascans.com
// and demonicscans.org — see README "Adapter reference".
// ============================================================
function meta(prop) {
@@ -111,12 +111,10 @@
const asura = {
site: "asura",
// asuracomic.net deep links 301 to the asurascans.com *root*, dropping the
// path, and that happens at the edge before this script gets a document —
// so those URLs cannot be handled here at all (checked 2026-07-25). It stays
// matched in case the redirect starts preserving paths again; until then,
// reach series through asurascans.com.
matches: (loc) => /(^|\.)asurascans\.com$|(^|\.)asuracomic\.net$/.test(loc.hostname),
// asuracomic.net is not matched: its deep links 301 to the asurascans.com
// *root* at the edge, discarding the path, so this script never sees a
// series document there (re-checked 2026-07-25).
matches: (loc) => /(^|\.)asurascans\.com$/.test(loc.hostname),
detect(loc) {
const path = loc.pathname;
// /comics/<slug-hash>/chapter/<n>
@@ -128,7 +126,6 @@
site: this.site,
seriesId: stripBuildHash(m[1]),
title: cleanTitle(meta("og:title")),
cover: meta("og:image") || "",
seriesUrl: loc.origin + "/comics/" + m[1],
chapterLabel: "Chapter " + m[2],
chapterNum: isNaN(num) ? null : num,
@@ -143,7 +140,6 @@
site: this.site,
seriesId: stripBuildHash(m[1]),
title: cleanTitle(meta("og:title")),
cover: meta("og:image") || "",
seriesUrl: loc.origin + "/comics/" + m[1],
chapterLabel: null,
chapterNum: null,
@@ -192,7 +188,6 @@
site: this.site,
seriesId: decodeURIComponent(m[1]),
title: cleanTitle(meta("og:title")),
cover: meta("og:image") || "",
seriesUrl: loc.origin + "/manga/" + m[1],
chapterLabel: "Chapter " + m[2],
chapterNum: isNaN(num) ? null : num,
@@ -207,7 +202,6 @@
site: this.site,
seriesId: decodeURIComponent(m[1]),
title: cleanTitle(meta("og:title")),
cover: meta("og:image") || "",
seriesUrl: loc.origin + "/manga/" + m[1],
chapterLabel: null,
chapterNum: null,
@@ -259,7 +253,6 @@
site: this.site,
seriesId: comixSeriesId(m[1]),
title: pageTitle,
cover: coverFromPage(pageTitle),
seriesUrl: loc.origin + "/title/" + m[1],
chapterLabel: "Chapter " + m[2],
chapterNum: isNaN(num) ? null : num,
@@ -274,7 +267,6 @@
site: this.site,
seriesId: comixSeriesId(m[1]),
title: pageTitle,
cover: coverFromPage(pageTitle),
seriesUrl: loc.origin + "/title/" + m[1],
chapterLabel: null,
chapterNum: null,
@@ -289,17 +281,6 @@
return t.replace(/\s*·\s*Ch\.[\d.]+\s*$/i, "").trim();
}
// comix serves no og:image, so this is the one adapter that has to read
// the DOM for a cover. Matching on alt rather than a class keeps it off
// the site's styling: the cover is the image whose alt is the title.
// Do not "simplify" this into meta("og:image") — that returns null.
function coverFromPage(title) {
if (!title || !document.querySelectorAll) return "";
for (const img of document.querySelectorAll("img[alt]")) {
if (img.getAttribute("alt") === title) return img.getAttribute("src") || "";
}
return "";
}
},
// Scoped to this series' own id prefix so a recommendation strip's links
// cannot win the maximum. seriesId is passed in because the anchors alone
@@ -342,7 +323,6 @@
site: this.site,
seriesId: m[1],
title: cleanTitle(meta("og:title")),
cover: meta("og:image") || "",
seriesUrl: loc.origin + "/series/" + m[1],
chapterLabel: num === null ? null : "Chapter " + num,
chapterNum: num,
@@ -357,7 +337,6 @@
site: this.site,
seriesId: m[1],
title: cleanTitle(meta("og:title")),
cover: meta("og:image") || "",
seriesUrl: loc.origin + "/series/" + m[1],
chapterLabel: null,
chapterNum: null,
@@ -498,6 +477,10 @@
async function apiPut(key, obj, { sendStatus = false } = {}) {
const body = Object.assign({}, obj);
if (!sendStatus) delete body.status;
// Covers belong to the backend, which acquires and serves them itself
// (ADR-0007) and ignores an incoming one; a third-party address must never
// go back on the wire.
delete body.cover;
const res = await fetch(API_BASE + "/bookmarks/" + encodeURIComponent(key), {
method: "PUT",
headers: authHeaders({ "Content-Type": "application/json" }),
@@ -855,7 +838,6 @@
series_id: p.seriesId,
title: p.title || (existing && existing.title) || p.seriesId,
series_url: p.seriesUrl || (existing && existing.series_url) || "",
cover: p.cover || (existing && existing.cover) || "",
last_chapter: p.chapterLabel || (existing && existing.last_chapter) || "",
last_chapter_num:
p.chapterNum != null ? p.chapterNum : existing ? existing.last_chapter_num : null,
@@ -878,7 +860,6 @@
series_id: p.seriesId,
title: existing.title || p.title || p.seriesId,
series_url: existing.series_url || p.seriesUrl || "",
cover: existing.cover || p.cover || "",
last_chapter: p.chapterLabel || existing.last_chapter || "",
last_chapter_num: p.chapterNum != null ? p.chapterNum : existing.last_chapter_num,
last_chapter_url: p.chapterUrl || "",
@@ -1394,7 +1375,15 @@
return el("div", { class: "item" + heat }, [
el("a", { class: "go", href: cont }, [
b.cover
? el("img", { class: "cover", src: b.cover, loading: "lazy", alt: "" })
? el("img", {
class: "cover",
src: b.cover,
loading: "lazy",
alt: "",
// A Cover that will not load shows the designed placeholder
// rather than the browser's broken-image glyph (#47).
onerror: (e) => e.target.replaceWith(el("div", { class: "cover ph" })),
})
: el("div", { class: "cover ph" }),
]),
el("div", { class: "meta" }, [
+214 -38
View File
@@ -25,6 +25,13 @@
// This script owns the novel library; the manga script is a separate install
// with its own prefix, so the two never share a cache, a queue or a panel.
const STORE_PREFIX = "bmgr:novel:";
const CACHE_KEY = STORE_PREFIX + "cache";
// Per-device record of when each series was last checked for new chapters.
// Deliberately not synced: each device does its own checking.
const LASTCHECKED_KEY = STORE_PREFIX + "lastchecked";
const LATEST_CHECK_THROTTLE_MS = 4 * 60 * 60 * 1000;
const LATEST_CHECK_BATCH = 1; // series fetched per navigation // series fetched per navigation
// Which library this script's rows belong to. The manga script is a separate
// install that declares "manga"; the backend keeps whichever it is told.
@@ -33,16 +40,11 @@
// ============================================================
// Site adapters
//
// Page type + IDs come from URL regex (most stable); title/cover come from
// og: meta tags (with the novelfull name= meta as the exception).
// Page type + IDs come from URL regex (most stable); the title comes from
// og: meta tags or the page's own heading. Covers are never read here: the
// backend acquires and serves them itself (ADR-0007).
// ============================================================
function meta(prop) {
const el = document.querySelector('meta[property="' + prop + '"]');
return el ? el.getAttribute("content") : null;
}
// Chapter lists are read from two places: the page we are standing on, and
// series pages fetched in the background. Both are reduced to {href, text}
// pairs so each adapter needs only one rule for picking the latest chapter.
@@ -66,12 +68,6 @@
return out;
}
// novelfull ships no og: tags at all — its cover lives on a name= meta.
function metaName(name) {
const el = document.querySelector('meta[name="' + name + '"]');
return el ? el.getAttribute("content") : null;
}
function escapeRe(s) {
return s.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
}
@@ -108,7 +104,6 @@
// h3.title on a chapter page is the *chapter's* title; the breadcrumb
// link back to the series page carries the series name.
title: back ? (back.textContent || "").trim() : "",
cover: metaName("image") || "",
seriesUrl: loc.origin + "/" + m[1] + ".html",
chapterLabel: "Chapter " + m[2],
chapterNum: isNaN(num) ? null : num,
@@ -124,7 +119,6 @@
site: this.site,
seriesId: m[1],
title: h3 ? (h3.textContent || "").trim() : "",
cover: metaName("image") || "",
seriesUrl: loc.origin + "/" + m[1] + ".html",
chapterLabel: null,
chapterNum: null,
@@ -138,9 +132,30 @@
},
};
// The host shape this adapter owns. One definition so the pointer validator
// and `matches` cannot drift apart (a leading subdomain is allowed).
const lnwHostRe = /(^|\.)lightnovelworld\.net$/;
// A chapter page's pointer is its own link back to its Series. The href is
// page markup, so validate before trusting: resolve it against the page
// address, require the host above and the /novel/<slug>/ Series path.
// Anything else is not a pointer.
function seriesIdFromLnwPointer(href, base) {
if (!href) return null;
let u;
try {
u = new URL(href, base);
} catch (e) {
return null;
}
if (!lnwHostRe.test(u.hostname)) return null;
const m = u.pathname.match(/^\/novel\/([^/]+)\/?$/);
return m ? m[1] : null;
}
const lightnovelworld = {
site: "lightnovelworld",
matches: (loc) => /(^|\.)lightnovelworld\.net$/.test(loc.hostname),
matches: (loc) => lnwHostRe.test(loc.hostname),
detect(loc) {
const path = loc.pathname;
// /<slug>-chapter-<n>/ — flat, at the site root, not under /novel/. The
@@ -148,19 +163,38 @@
// words still resolves to the right series.
let m = path.match(/^\/(.+)-chapter-([0-9]+(?:\.[0-9]+)?)\/?$/);
if (m) {
// The address is not the identity on this Site: the slug in the path
// is a Chapter Slug, which can differ from the Series slug and is
// never computable from it. The Series address is read from the
// page's own pointer — a silent fallback to derivation is the defect
// this replaced, not a safety net.
const chapterSlug = m[1];
const pointer = document.querySelector("a[aria-label='All Chapter']");
let seriesId = pointer ? seriesIdFromLnwPointer(pointer.getAttribute("href"), loc.href) : null;
if (!seriesId) {
// Fallback: the microdata breadcrumb's second crumb is the Series.
// Scoped to the BreadcrumbList because itemprop="item" is not
// unique to it (the header nav uses microdata too).
const crumb = document.querySelector(
'[itemtype="http://schema.org/BreadcrumbList"] a[itemprop="item"][href*="/novel/"]'
);
seriesId = crumb ? seriesIdFromLnwPointer(crumb.getAttribute("href"), loc.href) : null;
}
// Neither pointer present, nor either pointing at a /novel/<slug>/
// address on this host: not a page the script understands, so no
// Bookmark under an invented identity.
if (!seriesId) return { type: "other" };
const num = parseFloat(m[2]);
const h1 = document.querySelector("h1.entry-title");
const heading = h1 ? h1.textContent || "" : "";
return {
type: "chapter",
site: this.site,
seriesId: m[1],
seriesId: seriesId,
chapterSlug: chapterSlug,
// The heading is "<Series> Chapter <n>"; drop the suffix.
title: heading.replace(/\s*Chapter\s+[0-9.]+\s*$/i, "").trim(),
// Chapter pages carry no og:image. Empty is safe: every write merges
// against the cached row, which keeps the cover the series page gave.
cover: "",
seriesUrl: "https://lightnovelworld.net/novel/" + m[1] + "/",
seriesUrl: "https://lightnovelworld.net/novel/" + seriesId + "/",
chapterLabel: "Chapter " + m[2],
chapterNum: isNaN(num) ? null : num,
chapterUrl: loc.href,
@@ -174,8 +208,8 @@
type: "series",
site: this.site,
seriesId: m[1],
chapterSlug: null,
title: h1 ? (h1.textContent || "").trim() : "",
cover: meta("og:image") || "",
seriesUrl: "https://lightnovelworld.net/novel/" + m[1] + "/",
chapterLabel: null,
chapterNum: null,
@@ -184,12 +218,6 @@
}
return { type: "other" };
},
latestChapterFromAnchors(anchors, seriesId) {
return maxChapter(
anchors,
new RegExp("lightnovelworld\\.net/" + escapeRe(seriesId) + "-chapter-([0-9.]+)/")
);
},
};
const ADAPTERS = [novelfull, lightnovelworld];
@@ -210,12 +238,109 @@
return ADAPTERS.find((a) => a.site === site) || null;
}
// A stale lightnovelworld row is one keyed under the Chapter Slug its
// address was built from, while the page's pointer names a different
// Series. All three storage sites must move together: the cache row alone
// leaves a queued write replaying under a key whose row no longer exists
// (that write is silently lost), and a last-checked timestamp left behind
// re-fetches the repaired row on the next visit. A repair must never cost a
// Reader the chapter they were tracking, so when a row already holds the
// repaired key the two rows merge and the farther-ahead progress wins.
// Pure: no storage access, no module state, no clock.
function repairLnwStaleRow(list, queue, lastChecked, page) {
if (
!page ||
page.site !== "lightnovelworld" ||
!page.chapterSlug ||
!page.seriesId ||
page.chapterSlug === page.seriesId
) {
return { list, queue, lastChecked };
}
const oldKey = "lightnovelworld:" + page.chapterSlug;
const newKey = "lightnovelworld:" + page.seriesId;
const stale = list.find((b) => b.key === oldKey);
if (!stale) return { list, queue, lastChecked };
let listOut;
const existing = list.find((b) => b.key === newKey);
if (existing) {
// Duplicate case (spec user story 2): one row must survive, and it is
// the one already under the repaired key — with the stale row's
// progress carried across when it is ahead, favourite OR'd, and the
// stronger lifecycle bucket kept (finished > archived > reading, so a
// merge can never silently un-archive or un-finish a row).
const merged = Object.assign({}, existing);
if (
existing.last_chapter_num == null ||
(stale.last_chapter_num != null && stale.last_chapter_num > existing.last_chapter_num)
) {
merged.last_chapter = stale.last_chapter;
merged.last_chapter_num = stale.last_chapter_num;
merged.last_chapter_url = stale.last_chapter_url;
}
merged.favorite = !!(existing.favorite || stale.favorite);
const rank = (s) => ({ finished: 2, archived: 1, reading: 0 }[s || "reading"] || 0);
merged.status = rank(stale.status) > rank(existing.status) ? stale.status : existing.status;
merged.updated_at = Math.max(existing.updated_at || 0, stale.updated_at || 0);
listOut = list.filter((b) => b.key !== oldKey).map((b) => (b.key === newKey ? merged : b));
} else {
listOut = list.map((b) =>
b.key === oldKey
? Object.assign({}, b, {
key: newKey,
series_id: page.seriesId,
series_url: page.seriesUrl,
})
: b
);
}
let queueOut = queue;
const oldEntry = queue.find((e) => e.key === oldKey);
if (oldEntry) {
const survivor = queue.find((x) => x.key === newKey);
if (!survivor) {
queueOut = queue.map((e) => (e.key === oldKey ? Object.assign({}, e, { key: newKey }) : e));
} else {
// Both keys hold a marker and one row survives, so the two collapse
// into one entry under the repaired key (the queue's one-entry-per-key
// invariant). sendStatus is sticky — an archive intent from either
// marker survives, the queue's own rule — and the worse attempts
// count wins. The stale-key marker's op is dropped: the row it
// described is retired by the merge itself.
queueOut = queue
.filter((e) => e.key !== oldKey && e.key !== newKey)
.concat([
{
key: newKey,
op: survivor.op,
sendStatus: survivor.sendStatus || oldEntry.sendStatus,
attempts: Math.max(survivor.attempts || 0, oldEntry.attempts || 0),
},
]);
}
}
let lastCheckedOut = lastChecked;
if (oldKey in lastChecked) {
lastCheckedOut = Object.assign({}, lastChecked);
// Max, not overwrite: a timestamp already under the repaired key must
// not be rolled back to an older one.
lastCheckedOut[newKey] = Math.max(lastCheckedOut[newKey] || 0, lastCheckedOut[oldKey]);
delete lastCheckedOut[oldKey];
}
return { list: listOut, queue: queueOut, lastChecked: lastCheckedOut };
}
// Highest chapter the site lists, or null when the markup yields nothing.
// seriesId is only consulted by adapters whose pages carry other series'
// chapter links; the rest ignore it.
// chapter links. An adapter without a scanner (lightnovelworld — see the
// AGENTS.md entry) yields null, not an error.
function computeLatestChapter(site, anchors, seriesId) {
const a = adapterFor(site);
return a ? a.latestChapterFromAnchors(anchors, seriesId) : null;
return a && a.latestChapterFromAnchors ? a.latestChapterFromAnchors(anchors, seriesId) : null;
}
function currentSite() {
@@ -286,6 +411,10 @@
async function apiPut(key, obj, { sendStatus = false } = {}) {
const body = Object.assign({}, obj);
if (!sendStatus) delete body.status;
// Covers belong to the backend, which acquires and serves them itself
// (ADR-0007) and ignores an incoming one; a third-party address must never
// go back on the wire.
delete body.cover;
const res = await fetch(API_BASE + "/bookmarks/" + encodeURIComponent(key), {
method: "PUT",
headers: authHeaders({ "Content-Type": "application/json" }),
@@ -613,7 +742,6 @@
series_id: p.seriesId,
title: p.title || (existing && existing.title) || p.seriesId,
series_url: p.seriesUrl || (existing && existing.series_url) || "",
cover: p.cover || (existing && existing.cover) || "",
last_chapter: p.chapterLabel || (existing && existing.last_chapter) || "",
last_chapter_num:
p.chapterNum != null ? p.chapterNum : existing ? existing.last_chapter_num : null,
@@ -636,7 +764,6 @@
series_id: p.seriesId,
title: existing.title || p.title || p.seriesId,
series_url: existing.series_url || p.seriesUrl || "",
cover: existing.cover || p.cover || "",
last_chapter: p.chapterLabel || existing.last_chapter || "",
last_chapter_num: p.chapterNum != null ? p.chapterNum : existing.last_chapter_num,
last_chapter_url: p.chapterUrl || "",
@@ -729,6 +856,14 @@
const site = currentSite();
if (!site) return;
// A Site with no client scanner (lightnovelworld) is refreshed by the Poll
// on a cooldown shorter than the client throttle; fetching its pages here
// would be a megabyte-scale request whose result is discarded. Skipping
// before the due filter records no freshness timestamp and consumes no
// per-navigation batch slot.
const adapter = adapterFor(site);
if (!adapter || !adapter.latestChapterFromAnchors) return;
const checked = loadLastChecked();
const now = Date.now();
const due = state.list
@@ -739,7 +874,6 @@
.slice(0, LATEST_CHECK_BATCH);
if (due.length === 0) return;
const adapter = adapterFor(site);
for (const bm of due) {
// Recorded even when the fetch fails, so a broken series is retried on
// the next throttle window rather than on every single page load.
@@ -1147,7 +1281,15 @@
return el("div", { class: "item" + heat }, [
el("a", { class: "go", href: cont }, [
b.cover
? el("img", { class: "cover", src: b.cover, loading: "lazy", alt: "" })
? el("img", {
class: "cover",
src: b.cover,
loading: "lazy",
alt: "",
// A Cover that will not load shows the designed placeholder
// rather than the browser's broken-image glyph (#47).
onerror: (e) => e.target.replaceWith(el("div", { class: "cover ph" })),
})
: el("div", { class: "cover ph" }),
]),
el("div", { class: "meta" }, [
@@ -1199,7 +1341,8 @@
// Refresh + navigation
// ============================================================
async function refresh() {
async function refresh(awaitFirst) {
if (awaitFirst) await awaitFirst; // repair sync lands before we adopt the server's view of it
await drain(); // push what we owe before adopting the server's view of it
loading = true;
render();
@@ -1214,9 +1357,41 @@
}
}
// Runs the repair on every navigation, silently: rewrites a stale
// lightnovelworld row to the page's discovered identity, persists all three
// sites, and syncs through the queue-backed path so an offline repair parks
// and replays later. No toast — the Reader never asked for this. Returns
// the sync promise (or null when nothing changed) so init can hold the
// server view until the repair has landed. When the transform returns its
// inputs unchanged there is nothing to do.
function applyLnwStaleRowRepair() {
const lastChecked = loadLastChecked();
const out = repairLnwStaleRow(state.list, queue, lastChecked, state.page);
if (out.list === state.list && out.queue === queue && out.lastChecked === lastChecked) {
return null;
}
state.list = out.list;
reindex(); // byKey answers under the old key until rebuilt
saveCache(state.list);
queue.splice(0, queue.length, ...out.queue); // closures hold this array instance
saveQueue(queue);
saveLastChecked(out.lastChecked);
// Queue-backed sync, outcome swallowed: the repaired row is PUT (with its
// bucket when archived — the only status a userscript write may send) and
// the old server bookmark is deleted so no duplicate survives on the
// wire. The delete parks under the retired key while offline — the
// teardown of the old identity, not a new record of the Chapter Slug.
const key = keyOf(state.page);
return Promise.all([
pushBookmark(key, statusOf(state.byKey[key]) === "archived"),
pushDelete("lightnovelworld:" + state.page.chapterSlug),
]).catch(() => {});
}
let lastUrl = location.href;
function onNavigate() {
state.page = detect();
applyLnwStaleRowRepair();
render();
maybeAutoUpdate();
maybeCaptureLatestOnSeriesPage();
@@ -1329,13 +1504,14 @@
function init() {
buildUI();
state.page = detect();
const repair = applyLnwStaleRowRepair();
render();
installNavWatcher();
installLongPress();
window.addEventListener("online", drain); // signal returned while the page stayed open
// Sync first: both auto-record and the latest-chapter checks below need to
// know which series are bookmarked and how fresh they are.
refresh().then(() => {
refresh(repair).then(() => {
maybeAutoUpdate();
maybeCaptureLatestOnSeriesPage();
backgroundRefreshLatest();
@@ -1576,7 +1752,7 @@
// Exposes pure logic only — see userscript/test/novel-logic.test.js.
// ============================================================
if (typeof window === "undefined" && typeof module === "object" && module.exports) {
module.exports = { novelfull, lightnovelworld, anchorsFromHTML, statusOf, kindOf, maxChapter, escapeRe };
module.exports = { novelfull, lightnovelworld, anchorsFromHTML, statusOf, kindOf, maxChapter, escapeRe, computeLatestChapter, repairLnwStaleRow };
}
// ============================================================
+14 -44
View File
@@ -34,9 +34,6 @@ let metaTags = {};
// document.title. comix's SPA rewrites this on client routing but never
// og:title, so the comix adapter reads it instead. Reassigned per test.
let docTitle = "";
// img[alt] elements comix's coverFromPage() scans. Reassigned per test; each
// entry is {alt, src}.
let pageImages = [];
globalThis.document = {
querySelector(sel) {
const m = sel.match(/^meta\[property="([^"]+)"\]$/);
@@ -44,12 +41,6 @@ globalThis.document = {
const v = metaTags[m[1]];
return v == null ? null : { getAttribute: () => v };
},
querySelectorAll(sel) {
if (sel !== "img[alt]") return [];
return pageImages.map((img) => ({
getAttribute: (attr) => img[attr] ?? null,
}));
},
addEventListener() {},
get title() {
return docTitle;
@@ -100,7 +91,7 @@ test("stripBuildHash ignores suffixes that are not exactly 8 hex chars", () => {
// ============================================================
test("asura.detect reads a series page, stripping the hash from the id only", () => {
metaTags = { "og:title": "Solo Leveling | Asura Scans", "og:image": "https://cdn.example/x.jpg" };
metaTags = { "og:title": "Solo Leveling | Asura Scans" };
const p = asura.detect(loc("https://asurascans.com/comics/solo-leveling-059befe1"));
assert.equal(p.type, "series");
assert.equal(p.site, "asura");
@@ -108,12 +99,11 @@ test("asura.detect reads a series page, stripping the hash from the id only", ()
// seriesUrl keeps the hash: navigation needs the current one (stale ones 302).
assert.equal(p.seriesUrl, "https://asurascans.com/comics/solo-leveling-059befe1");
assert.equal(p.title, "Solo Leveling");
assert.equal(p.cover, "https://cdn.example/x.jpg");
assert.equal(p.chapterNum, null);
});
test("asura.detect reads a chapter page including a decimal number", () => {
metaTags = { "og:title": "Solo Leveling Chapter 12.5 - Read Online | Asura Scans", "og:image": "" };
metaTags = { "og:title": "Solo Leveling Chapter 12.5 - Read Online | Asura Scans" };
const url = "https://asurascans.com/comics/solo-leveling-059befe1/chapter/12.5";
const p = asura.detect(loc(url));
assert.equal(p.type, "chapter");
@@ -130,6 +120,14 @@ test("asura.detect returns other for non-series paths", () => {
assert.equal(asura.detect(loc("https://asurascans.com/bookmarks")).type, "other");
});
test("asura.matches accepts only asurascans.com", () => {
assert.equal(asura.matches({ hostname: "asurascans.com" }), true);
assert.equal(asura.matches({ hostname: "www.asurascans.com" }), true);
// Dead domain: deep links 301 to the asurascans.com root, discarding the path.
assert.equal(asura.matches({ hostname: "asuracomic.net" }), false);
assert.equal(asura.matches({ hostname: "asurascans.com.evil.example" }), false);
});
test("asura.latestChapterFromAnchors takes the highest and skips the First Chapter shortcut", () => {
const best = asura.latestChapterFromAnchors([
{ href: "/comics/solo-leveling-059befe1/chapter/1", text: "Chapter 1" },
@@ -150,7 +148,7 @@ test("asura.latestChapterFromAnchors returns null when nothing matches", () => {
// ============================================================
test("demonic.detect reads a series page", () => {
metaTags = { "og:title": "The World After The Fall", "og:image": "https://cdn.example/y.jpg" };
metaTags = { "og:title": "The World After The Fall" };
const p = demonic.detect(loc("https://demonicscans.org/manga/the-world-after-the-fall"));
assert.equal(p.type, "series");
assert.equal(p.site, "demonic");
@@ -159,7 +157,7 @@ test("demonic.detect reads a series page", () => {
});
test("demonic.detect reads a chapter page and strips the suffix from the title", () => {
metaTags = { "og:title": "The World After The Fall Chapter 3", "og:image": "" };
metaTags = { "og:title": "The World After The Fall Chapter 3" };
const p = demonic.detect(loc("https://demonicscans.org/title/the-world-after-the-fall/chapter/3/1"));
assert.equal(p.type, "chapter");
assert.equal(p.chapterNum, 3);
@@ -247,26 +245,6 @@ test("comix parses decimal chapter numbers", () => {
assert.equal(p.chapterNum, 80.5);
});
test("comix.detect reads the cover from an img whose alt matches the cleaned title", () => {
metaTags = { "og:title": COMIX_STALE_HOME };
docTitle = "Dungeons and Crayons";
pageImages = [
{ alt: "Some Other Series", src: "https://cdn.example/other.jpg" },
{ alt: "Dungeons and Crayons", src: "https://cdn.example/cover.jpg" },
];
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
assert.equal(p.cover, "https://cdn.example/cover.jpg");
});
test("comix.detect leaves cover empty when no img alt matches the title", () => {
metaTags = {};
docTitle = "Dungeons and Crayons";
pageImages = [{ alt: "Some Other Series", src: "https://cdn.example/other.jpg" }];
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
assert.equal(p.cover, "");
pageImages = [];
});
test("comix ignores unrelated paths", () => {
assert.equal(comix.detect(loc("https://comix.to/browse")).type, "other");
});
@@ -308,23 +286,18 @@ const KAGANE_SERIES = "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b";
const KAGANE_BOOK = "019fa2e0-6dbd-73ca-b40b-fe06ab75eb0e";
test("kagane detects a series page", () => {
metaTags = {
"og:title": "Infinite Decryption: The Strongest Level 0",
"og:image": "https://kagane.to/api/v2/image/abc/compressed",
};
metaTags = { "og:title": "Infinite Decryption: The Strongest Level 0" };
const p = kagane.detect(loc("https://kagane.to/series/" + KAGANE_SERIES));
assert.equal(p.type, "series");
assert.equal(p.site, "kagane");
assert.equal(p.seriesId, KAGANE_SERIES);
assert.equal(p.title, "Infinite Decryption: The Strongest Level 0");
assert.equal(p.cover, "https://kagane.to/api/v2/image/abc/compressed");
assert.equal(p.seriesUrl, "https://kagane.to/series/" + KAGANE_SERIES);
});
test("kagane reads the chapter number out of og:title", () => {
metaTags = {
"og:title": "Infinite Decryption: The Strongest Level 0 - Chapter 41 - Episode 41",
"og:image": "https://kagane.to/api/v2/image/abc/compressed",
};
const p = kagane.detect(
loc("https://kagane.to/series/" + KAGANE_SERIES + "/reader/" + KAGANE_BOOK)
@@ -341,10 +314,7 @@ test("kagane reads the chapter number out of og:title", () => {
// no episode name, because the book carries volume_no and an empty title.
// Captured live 2026-08-08 from SP Baby.
test("kagane reads through a Volume-numbered chapter suffix", () => {
metaTags = {
"og:title": "SP Baby - Volume 1 Chapter 1",
"og:image": "https://kagane.to/api/v2/image/abc/compressed",
};
metaTags = { "og:title": "SP Baby - Volume 1 Chapter 1" };
const p = kagane.detect(
loc("https://kagane.to/series/" + KAGANE_SERIES + "/reader/" + KAGANE_BOOK)
);
+439 -20
View File
@@ -4,8 +4,7 @@ const test = require("node:test");
const assert = require("node:assert");
// ============================================================
// Minimal browser stub. Same shape as logic.test.js, plus a meta[name=...]
// branch: novelfull ships no og: tags, so its cover comes from name="image".
// Minimal browser stub. Same shape as logic.test.js.
// document.body stays UNDEFINED so the boot block waits for a DOMContentLoaded
// that never fires and no network call is ever made.
// ============================================================
@@ -20,20 +19,19 @@ globalThis.localStorage = {
globalThis.location = { href: "about:blank", hostname: "", pathname: "/", origin: "" };
let metaTags = {};
let namedMetas = {};
let elements = {};
// Attribute selectors (the lightnovelworld Series pointer) answer with an
// element exposing getAttribute, like the meta branch below.
let attrEls = {};
globalThis.document = {
querySelector(sel) {
let m = sel.match(/^meta\[property="([^"]+)"\]$/);
const attr = attrEls[sel];
if (attr != null) return attr;
const m = sel.match(/^meta\[property="([^"]+)"\]$/);
if (m) {
const v = metaTags[m[1]];
return v == null ? null : { getAttribute: () => v };
}
m = sel.match(/^meta\[name="([^"]+)"\]$/);
if (m) {
const v = namedMetas[m[1]];
return v == null ? null : { getAttribute: () => v };
}
const text = elements[sel];
return text == null ? null : { textContent: text };
},
@@ -47,8 +45,10 @@ globalThis.document = {
const {
novelfull,
lightnovelworld,
computeLatestChapter,
kindOf,
maxChapter,
repairLnwStaleRow,
} = require("../novel-bookmark.user.js");
function loc(href) {
@@ -58,8 +58,8 @@ function loc(href) {
function reset() {
metaTags = {};
namedMetas = {};
elements = {};
attrEls = {};
}
// ============================================================
@@ -68,21 +68,18 @@ function reset() {
test("novelfull.detect reads a series page", () => {
reset();
namedMetas = { image: "https://novelfull.com/uploads/thumbs/ri.jpg" };
elements = { "h3.title": "Reverend Insanity" };
const p = novelfull.detect(loc("https://novelfull.com/reverend-insanity.html"));
assert.equal(p.type, "series");
assert.equal(p.site, "novelfull");
assert.equal(p.seriesId, "reverend-insanity");
assert.equal(p.title, "Reverend Insanity");
assert.equal(p.cover, "https://novelfull.com/uploads/thumbs/ri.jpg");
assert.equal(p.seriesUrl, "https://novelfull.com/reverend-insanity.html");
assert.equal(p.chapterNum, null);
});
test("novelfull.detect reads a chapter page and points seriesUrl at the series", () => {
reset();
namedMetas = { image: "https://novelfull.com/uploads/thumbs/ri.jpg" };
elements = { "a.truyen-title": "Reverend Insanity" };
const url = "https://novelfull.com/reverend-insanity/chapter-2334-fang-yuan.html";
const p = novelfull.detect(loc(url));
@@ -120,19 +117,24 @@ test("novelfull.latestChapterFromAnchors takes the max and ignores other series"
test("lightnovelworld.detect reads a series page", () => {
reset();
metaTags = { "og:image": "https://lightnovelworld.net/wp-content/uploads/awe.webp" };
elements = { "h1.entry-title": "A Will Eternal" };
const p = lightnovelworld.detect(loc("https://lightnovelworld.net/novel/a-will-eternal/"));
assert.equal(p.type, "series");
assert.equal(p.site, "lightnovelworld");
assert.equal(p.seriesId, "a-will-eternal");
assert.equal(p.title, "A Will Eternal");
assert.equal(p.cover, "https://lightnovelworld.net/wp-content/uploads/awe.webp");
// A series page's address *is* its identity; there is no Chapter Slug.
assert.equal(p.chapterSlug, null);
});
test("lightnovelworld.detect strips the chapter suffix off the heading", () => {
reset();
elements = { "h1.entry-title": "A Will Eternal Chapter 1298" };
attrEls = {
"a[aria-label='All Chapter']": {
getAttribute: () => "https://lightnovelworld.net/novel/a-will-eternal/",
},
};
const url = "https://lightnovelworld.net/a-will-eternal-chapter-1298/";
const p = lightnovelworld.detect(loc(url));
assert.equal(p.type, "chapter");
@@ -141,8 +143,170 @@ test("lightnovelworld.detect strips the chapter suffix off the heading", () => {
assert.equal(p.chapterLabel, "Chapter 1298");
assert.equal(p.title, "A Will Eternal");
assert.equal(p.seriesUrl, "https://lightnovelworld.net/novel/a-will-eternal/");
// Chapter pages have no cover; the merge in bookmarkCurrent keeps the stored one.
assert.equal(p.cover, "");
});
test("lightnovelworld.detect reads the Series identity from the page pointer on a chapter page", () => {
reset();
elements = { "h1.entry-title": "A Will Eternal Chapter 1298" };
attrEls = {
// Verbatim from the real page (research note §3): the All Chapter anchor
// carries the absolute Series address.
"a[aria-label='All Chapter']": {
getAttribute: () => "https://lightnovelworld.net/novel/a-will-eternal/",
},
};
const url = "https://lightnovelworld.net/a-will-eternal-chapter-1298/";
const p = lightnovelworld.detect(loc(url));
assert.equal(p.type, "chapter");
assert.equal(p.seriesId, "a-will-eternal");
assert.equal(p.seriesUrl, "https://lightnovelworld.net/novel/a-will-eternal/");
assert.equal(p.chapterSlug, "a-will-eternal");
assert.equal(p.chapterNum, 1298);
});
test("lightnovelworld.detect falls back to the breadcrumb when the pointer is absent", () => {
reset();
elements = { "h1.entry-title": "My Longevity Simulation Chapter 1" };
const url = "https://lightnovelworld.net/my-longevity-simulation-chapter-1/";
// Same divergent page, both pointers from the real markup (research note
// §3): breadcrumb position 2 must say what the All Chapter anchor says.
attrEls = {
"a[aria-label='All Chapter']": {
getAttribute: () => "https://lightnovelworld.net/novel/immortality-simulator/",
},
};
const viaPointer = lightnovelworld.detect(loc(url));
attrEls = {
'[itemtype="http://schema.org/BreadcrumbList"] a[itemprop="item"][href*="/novel/"]': {
getAttribute: () => "https://lightnovelworld.net/novel/immortality-simulator/",
},
};
const viaBreadcrumb = lightnovelworld.detect(loc(url));
assert.equal(viaBreadcrumb.type, "chapter");
assert.equal(viaBreadcrumb.seriesId, viaPointer.seriesId);
assert.equal(viaBreadcrumb.seriesUrl, viaPointer.seriesUrl);
assert.equal(viaBreadcrumb.seriesId, "immortality-simulator");
assert.equal(viaBreadcrumb.chapterSlug, "my-longevity-simulation");
});
test("lightnovelworld.detect resolves a relative pointer against the page address", () => {
// Defensive: every measured page ships an absolute pointer href, but a
// theme change could go relative — the pointer is still this page's own
// link back to its Series.
reset();
elements = { "h1.entry-title": "A Will Eternal Chapter 1298" };
attrEls = {
"a[aria-label='All Chapter']": { getAttribute: () => "/novel/a-will-eternal/" },
};
const p = lightnovelworld.detect(loc("https://lightnovelworld.net/a-will-eternal-chapter-1298/"));
assert.equal(p.type, "chapter");
assert.equal(p.seriesId, "a-will-eternal");
assert.equal(p.seriesUrl, "https://lightnovelworld.net/novel/a-will-eternal/");
});
test("lightnovelworld.detect rejects a pointer that is not a /novel/ address on its own host", () => {
// Counterfactual pointers, exercising the criterion that a pointer "present
// but not parseable as /novel/<slug>/ on lightnovelworld.net" must not
// become a seriesUrl the backend is later asked to Poll: off-host, and
// on-host but the wrong path shape.
reset();
elements = { "h1.entry-title": "My Longevity Simulation Chapter 1" };
const url = "https://lightnovelworld.net/my-longevity-simulation-chapter-1/";
attrEls = {
"a[aria-label='All Chapter']": {
getAttribute: () => "https://evil.example/novel/immortality-simulator/",
},
};
assert.equal(lightnovelworld.detect(loc(url)).type, "other");
attrEls = {
"a[aria-label='All Chapter']": {
getAttribute: () => "https://lightnovelworld.net/fiction/immortality-simulator/",
},
};
assert.equal(lightnovelworld.detect(loc(url)).type, "other");
});
test("lightnovelworld.detect resolves to other when the page carries no pointer", () => {
reset();
elements = { "h1.entry-title": "My Longevity Simulation Chapter 1" };
const p = lightnovelworld.detect(
loc("https://lightnovelworld.net/my-longevity-simulation-chapter-1/")
);
assert.equal(p.type, "other");
});
test("lightnovelworld pins the divergent novel: the pointer's slug wins over the address's", () => {
// Regression pin for the derivation defect (research note §2/§3): the
// chapter address is built from "my-longevity-simulation" but the Series
// is published as "immortality-simulator". If the adapter ever derives the
// identity from the address again, this test goes red.
reset();
elements = { "h1.entry-title": "My Longevity Simulation Chapter 1" };
attrEls = {
"a[aria-label='All Chapter']": {
getAttribute: () => "https://lightnovelworld.net/novel/immortality-simulator/",
},
};
const url = "https://lightnovelworld.net/my-longevity-simulation-chapter-1/";
const p = lightnovelworld.detect(loc(url));
assert.equal(p.type, "chapter");
assert.equal(p.seriesId, "immortality-simulator");
assert.equal(p.seriesUrl, "https://lightnovelworld.net/novel/immortality-simulator/");
assert.equal(p.chapterSlug, "my-longevity-simulation");
assert.notEqual(p.chapterSlug, p.seriesId);
});
test("lightnovelworld.detect resolves a novel whose heading ends in a chapter number", () => {
// The split novel from research §4.4, chapter 200 (published under the
// current slug). Its heading ends "…Not Them All Chapter 200", which would
// false-match a selector that looks for the text "All Chapter".
reset();
elements = {
"h1.entry-title": "All Jobs and Classes I Just Wanted One Skill Not Them All Chapter 200",
};
attrEls = {
"a[aria-label='All Chapter']": {
getAttribute: () =>
"https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/",
},
};
const url =
"https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-200/";
const p = lightnovelworld.detect(loc(url));
assert.equal(p.type, "chapter");
assert.equal(p.seriesId, "all-jobs-and-classes-i-just-wanted-one-skill-not-them-all");
assert.equal(p.title, "All Jobs and Classes I Just Wanted One Skill Not Them All");
assert.equal(p.chapterNum, 200);
assert.equal(p.chapterSlug, "all-jobs-and-classes-i-just-wanted-one-skill-not-them-all");
});
test("lightnovelworld.detect resolves both Chapter Slugs of the split novel to one Series", () => {
// Research note §4.4: this Series serves chapters 1-99 under one Chapter
// Slug and 100-423 under another; both chapter addresses are live and both
// point back at the same Series. An old deep link must not get a different
// identity than a current one.
reset();
elements = { "h1.entry-title": "All Jobs and Classes I Just Wanted One Skill Not Chapter 1" };
attrEls = {
"a[aria-label='All Chapter']": {
getAttribute: () =>
"https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/",
},
};
const oldSlug = lightnovelworld.detect(
loc("https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-chapter-1/")
);
const currentSlug = lightnovelworld.detect(
loc(
"https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-404/"
)
);
assert.equal(oldSlug.type, "chapter");
assert.equal(oldSlug.seriesId, "all-jobs-and-classes-i-just-wanted-one-skill-not-them-all");
assert.equal(oldSlug.chapterSlug, "all-jobs-and-classes-i-just-wanted-one-skill-not");
assert.equal(currentSlug.type, "chapter");
assert.equal(currentSlug.seriesId, "all-jobs-and-classes-i-just-wanted-one-skill-not-them-all");
assert.equal(currentSlug.chapterSlug, "all-jobs-and-classes-i-just-wanted-one-skill-not-them-all");
});
test("lightnovelworld.detect returns other for non-series paths", () => {
@@ -151,17 +315,40 @@ test("lightnovelworld.detect returns other for non-series paths", () => {
assert.equal(lightnovelworld.detect(loc("https://lightnovelworld.net/az-lists/")).type, "other");
});
test("lightnovelworld.latestChapterFromAnchors takes the max and ignores other series", () => {
const best = lightnovelworld.latestChapterFromAnchors(
test("computeLatestChapter yields nothing for lightnovelworld, comment anchors included", () => {
// A realistic Series-page anchor set: current-slug chapters, a second
// Chapter Slug's chapters, and a wpdiscuz comment pasting a high-numbered
// chapter of another novel (the poisoning vector). The client never scans
// this Site — the Poll's cooldown dominates the client throttle and the
// comment thread is a public write surface — so this must fail the moment
// a scan is reintroduced, scoped or not.
const latest = computeLatestChapter(
"lightnovelworld",
[
{ href: "https://lightnovelworld.net/a-will-eternal-chapter-1/", text: "Chapter 1" },
{ href: "https://lightnovelworld.net/a-will-eternal-chapter-1317/", text: "Chapter 1317" },
{ href: "https://lightnovelworld.net/a-will-eternal-chapter-1298/", text: "Chapter 1298" },
// a divergent novel's second Chapter Slug
{ href: "https://lightnovelworld.net/my-longevity-simulation-chapter-400/", text: "Chapter 400" },
// a wpdiscuz comment anchor
{ href: "https://lightnovelworld.net/overgeared-chapter-9999/", text: "Chapter 9999" },
],
"a-will-eternal"
);
assert.deepEqual(best, { num: 1317, label: "Chapter 1317" });
assert.equal(latest, null);
});
test("computeLatestChapter still scans novelfull and ignores other series", () => {
const best = computeLatestChapter(
"novelfull",
[
{ href: "/reverend-insanity/chapter-2334-fang-yuan.html", text: "Chapter 2334" },
{ href: "/reverend-insanity/chapter-1.html", text: "Chapter 1" },
{ href: "/release-that-witch/chapter-9999.html", text: "Chapter 9999" },
],
"reverend-insanity"
);
assert.deepEqual(best, { num: 2334, label: "Chapter 2334" });
});
test("latestChapterFromAnchors returns null when nothing matches", () => {
@@ -169,6 +356,213 @@ test("latestChapterFromAnchors returns null when nothing matches", () => {
assert.equal(maxChapter([], /chapter-([0-9.]+)/), null);
});
// ============================================================
// repairLnwStaleRow — the row migration (spec seam 3)
// ============================================================
// A row as bookmarkCurrent builds it, keyed under the invented Chapter Slug
// identity the old adapter derived from the chapter address.
function staleRow(over) {
return Object.assign(
{
key: "lightnovelworld:my-longevity-simulation",
site: "lightnovelworld",
kind: "novel",
series_id: "my-longevity-simulation",
title: "My Longevity Simulation",
series_url: "https://lightnovelworld.net/novel/my-longevity-simulation/",
last_chapter: "Chapter 7",
last_chapter_num: 7,
last_chapter_url: "https://lightnovelworld.net/my-longevity-simulation-chapter-7/",
updated_at: 1000,
},
over
);
}
// The divergent novel from research §3: the address's slug is a Chapter Slug,
// the page pointer names the real Series.
function lnwPage(over) {
return Object.assign(
{
type: "chapter",
site: "lightnovelworld",
seriesId: "immortality-simulator",
chapterSlug: "my-longevity-simulation",
seriesUrl: "https://lightnovelworld.net/novel/immortality-simulator/",
},
over
);
}
test("repairLnwStaleRow rewrites the row, the queue entry and the last-checked map together", () => {
const list = [staleRow()];
const queue = [
{ key: "lightnovelworld:my-longevity-simulation", op: "put", sendStatus: false, attempts: 2 },
];
const lastChecked = { "lightnovelworld:my-longevity-simulation": 12345 };
const out = repairLnwStaleRow(list, queue, lastChecked, lnwPage());
assert.equal(out.list.length, 1);
assert.equal(out.list[0].key, "lightnovelworld:immortality-simulator");
assert.equal(out.list[0].series_id, "immortality-simulator");
assert.equal(out.list[0].series_url, "https://lightnovelworld.net/novel/immortality-simulator/");
assert.deepEqual(out.queue, [
{ key: "lightnovelworld:immortality-simulator", op: "put", sendStatus: false, attempts: 2 },
]);
assert.deepEqual(out.lastChecked, { "lightnovelworld:immortality-simulator": 12345 });
});
test("repairLnwStaleRow keeps Progress, Favourite and Lifecycle bucket on the rewritten row", () => {
const stale = staleRow({
last_chapter: "Chapter 42",
last_chapter_num: 42,
last_chapter_url: "https://lightnovelworld.net/my-longevity-simulation-chapter-42/",
favorite: true,
status: "archived",
updated_at: 777,
});
const out = repairLnwStaleRow([stale], [], {}, lnwPage());
const b = out.list[0];
assert.equal(b.last_chapter, "Chapter 42");
assert.equal(b.last_chapter_num, 42);
assert.equal(
b.last_chapter_url,
"https://lightnovelworld.net/my-longevity-simulation-chapter-42/"
);
assert.equal(b.favorite, true);
assert.equal(b.status, "archived");
assert.equal(b.updated_at, 777);
assert.equal(b.title, "My Longevity Simulation");
});
test("repairLnwStaleRow changes nothing when the Chapter Slug and Series slug agree", () => {
const list = [
staleRow({ key: "lightnovelworld:a-will-eternal", series_id: "a-will-eternal" }),
];
const queue = [{ key: "lightnovelworld:a-will-eternal", op: "put", sendStatus: true, attempts: 1 }];
const lastChecked = { "lightnovelworld:a-will-eternal": 99 };
const out = repairLnwStaleRow(
list,
queue,
lastChecked,
lnwPage({
chapterSlug: "a-will-eternal",
seriesId: "a-will-eternal",
seriesUrl: "https://lightnovelworld.net/novel/a-will-eternal/",
})
);
assert.equal(out.list, list); // same references, nothing rewritten
assert.equal(out.queue, queue);
assert.equal(out.lastChecked, lastChecked);
});
test("repairLnwStaleRow changes nothing when no row sits under the old key", () => {
const list = [
staleRow({
key: "lightnovelworld:immortality-simulator",
series_id: "immortality-simulator",
series_url: "https://lightnovelworld.net/novel/immortality-simulator/",
}),
];
const lastChecked = { "lightnovelworld:immortality-simulator": 99 };
const out = repairLnwStaleRow(list, [], lastChecked, lnwPage());
assert.equal(out.list, list);
assert.equal(out.queue.length, 0);
assert.equal(out.lastChecked, lastChecked);
});
test("repairLnwStaleRow ignores a page that carries no Chapter Slug", () => {
const list = [
staleRow({
key: "lightnovelworld:immortality-simulator",
series_id: "immortality-simulator",
}),
];
const lastChecked = { "lightnovelworld:immortality-simulator": 99 };
// A series page (chapterSlug null) and another site both stay untouched.
const series = repairLnwStaleRow(list, [], lastChecked, lnwPage({ chapterSlug: null }));
const other = repairLnwStaleRow(list, [], lastChecked, {
type: "chapter",
site: "novelfull",
seriesId: "x",
});
assert.equal(series.list, list);
assert.equal(other.list, list);
});
test("repairLnwStaleRow still moves the row and the last-checked map when the queue has no entry", () => {
const out = repairLnwStaleRow(
[staleRow()],
[],
{ "lightnovelworld:my-longevity-simulation": 5 },
lnwPage()
);
assert.equal(out.list[0].key, "lightnovelworld:immortality-simulator");
assert.deepEqual(out.queue, []);
assert.deepEqual(out.lastChecked, { "lightnovelworld:immortality-simulator": 5 });
});
test("repairLnwStaleRow carries a queued change to the repaired identity with its marker intact", () => {
const queue = [
{ key: "lightnovelworld:my-longevity-simulation", op: "put", sendStatus: true, attempts: 3 },
];
const out = repairLnwStaleRow([staleRow()], queue, {}, lnwPage());
assert.deepEqual(out.queue, [
{ key: "lightnovelworld:immortality-simulator", op: "put", sendStatus: true, attempts: 3 },
]);
});
test("repairLnwStaleRow merges a duplicate under the repaired key, keeping the farther-ahead progress", () => {
const canonical = staleRow({
key: "lightnovelworld:immortality-simulator",
series_id: "immortality-simulator",
series_url: "https://lightnovelworld.net/novel/immortality-simulator/",
last_chapter: "Chapter 5",
last_chapter_num: 5,
last_chapter_url: "https://lightnovelworld.net/immortality-simulator-chapter-5/",
favorite: false,
});
const stale = staleRow({
last_chapter: "Chapter 42",
last_chapter_num: 42,
last_chapter_url: "https://lightnovelworld.net/my-longevity-simulation-chapter-42/",
favorite: true,
status: "archived",
});
const queue = [
{ key: "lightnovelworld:my-longevity-simulation", op: "put", sendStatus: true, attempts: 3 },
{ key: "lightnovelworld:immortality-simulator", op: "put", sendStatus: false, attempts: 0 },
];
const out = repairLnwStaleRow([canonical, stale], queue, {}, lnwPage());
assert.equal(out.list.length, 1);
const merged = out.list[0];
assert.equal(merged.key, "lightnovelworld:immortality-simulator");
assert.equal(merged.last_chapter_num, 42); // the stale row is ahead — progress must not be lost
assert.equal(merged.favorite, true); // favourite survives from either row
assert.equal(merged.status, "archived"); // the stronger bucket survives
// One marker under the repaired key: sendStatus is sticky (the archive
// intent from the stale-key marker survives) and the worse attempts count
// wins — the queue's own coalescing rules.
assert.deepEqual(out.queue, [
{ key: "lightnovelworld:immortality-simulator", op: "put", sendStatus: true, attempts: 3 },
]);
});
test("repairLnwStaleRow does not regress progress when the repaired-key row is ahead", () => {
const canonical = staleRow({
key: "lightnovelworld:immortality-simulator",
series_id: "immortality-simulator",
series_url: "https://lightnovelworld.net/novel/immortality-simulator/",
last_chapter: "Chapter 100",
last_chapter_num: 100,
last_chapter_url: "https://lightnovelworld.net/immortality-simulator-chapter-100/",
});
const stale = staleRow({ last_chapter: "Chapter 42", last_chapter_num: 42 });
const out = repairLnwStaleRow([canonical, stale], [], {}, lnwPage());
assert.equal(out.list.length, 1);
assert.equal(out.list[0].last_chapter_num, 100);
});
// ============================================================
// kindOf
// ============================================================
@@ -180,3 +574,28 @@ test("kindOf defaults a missing kind to manga", () => {
test("kindOf passes through novel", () => {
assert.equal(kindOf({ kind: "novel" }), "novel");
});
// ============================================================
// Source guard
//
// Issue #74: the novel script was split off the manga one and lost four
// module-scope constants. The reads sit inside try/catch or a fire-and-forget
// promise, so the ReferenceError never surfaced — nothing but a static check
// catches this class.
// ============================================================
test("every SCREAMING_CASE constant the script uses is declared in it", () => {
const fs = require("node:fs");
for (const f of ["novel-bookmark.user.js", "manga-bookmark.user.js"]) {
const src = fs.readFileSync(require.resolve("../" + f), "utf8")
// comments and strings carry prose and SVG path data in the same shape
.replace(/\/\/[^\n]*|\/\*[\s\S]*?\*\/|"[^"\n]*"|'[^'\n]*'|`[\s\S]*?`/g, " ");
const declared = new Set(
[...src.matchAll(/\b(?:const|let|var|function)\s+([A-Z][A-Z0-9_]{2,})\b/g)].map((m) => m[1]),
);
for (const name of new Set(src.match(/\b[A-Z][A-Z0-9_]{2,}\b/g) || [])) {
if (name.startsWith("GM_") || name in globalThis) continue;
assert.ok(declared.has(name), `${f} uses ${name} but never declares it`);
}
}
});