Compare commits

..

17 Commits

Author SHA1 Message Date
sulthan 1ee5eb67ea Clear the stale Chrome singleton lock at browser boot (#76)
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 19:11:24 +07:00
sulthan b22ae82897 Restore the novel script's lost module-scope constants (#74) (#75)
> *This was generated by AI during triage.*

Fixes #74.

`novel-bookmark.user.js` was split out of `manga-bookmark.user.js` and lost four module-scope constants. Every use of them is behind a `try/catch` or a fire-and-forget promise, so the `ReferenceError`s were swallowed rather than reported.

| constant | used at | effect while missing |
| --- | --- | --- |
| `LATEST_CHECK_THROTTLE_MS` | `:722` | `backgroundRefreshLatest()` throws before computing `due` — no background latest-check ever runs for novels (the symptom in #74) |
| `LATEST_CHECK_BATCH` | `:724` | same throw |
| `CACHE_KEY` | `:215`, `:224` | `loadCache()` always returns `[]`, `saveCache()` silently no-ops — the local cache never persists |
| `LASTCHECKED_KEY` | `:232`, `:241` | last-checked map never persists, so the throttle would not hold even once the first two are defined |

#74 named only the two throttle constants. The two cache keys are the same lost lines with the same root cause, so they are restored here too — fixing only the pair the issue named would leave `backgroundRefreshLatest()` re-fetching every series on every navigation, because `saveLastChecked()` would still be a no-op.

Values and comments copied verbatim from `manga-bookmark.user.js:37-43`; throttle 4h, batch 1.

## Verification

- `node --check userscript/novel-bookmark.user.js` — clean.
- `node --test userscript/test/logic.test.js userscript/test/novel-logic.test.js` — 47/47 pass.
- New test `every SCREAMING_CASE constant the script uses is declared in it` scans both scripts (comments and string literals stripped first, so prose and SVG path data do not trip it). Confirmed it fails — 1 failing test — when `LATEST_CHECK_BATCH` is deleted again, and passes when restored.

A behavioural test cannot reach this: the storage helpers and the background refresh are exactly the layers the harness does not cover (see the `testing-the-userscript` skill), and the errors are swallowed anyway. A static guard is the only instrument that sees this bug class.

On-device confirmation that the ember now lights for novels is still outstanding — that needs Violentmonkey against a live novelfull/lightnovelworld page.

Reviewed-on: #75
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 18:12:03 +07:00
sulthan 7c7d597019 Delete the kagane-specific cover path (#63) (#73)
Closes #63

Deletes the second way to reach a Cover. Since #62, every Site's cover bytes land in the content-addressed store at creation or on the poll, and the one public route serves them all — nothing needs the kagane proxy anymore.

## What went

- **Template-level rewrite:** `Bookmark.CoverURL()` and both templates' use of it. Cards and chrome now render `.Cover` — the wire value — and nothing else. `Bookmark.CoverSource` was dead once `CoverURL` went, so it and its `bookmarkColumns` entry are gone too.
- **Kagane-only cover route and its identifier validation:** `GET /img/kagane/{id}`, `web.CoverFetcher`, `coverIDRe`, and the whole `internal/web/cover.go`.
- **The proxy's persistence:** `store.KaganeImageID`, `GetKaganeCover`, `PutKaganeCover`, `kaganeCoverSourceURL`, `kaganeCoverRe`.
- **The kagane-shaped branch in the byte-fetch routing:** `fetchCoverBytes` no longer takes a `site` argument and no longer names a Site. The URL shape kagane's API publishes is claimed by the browser module itself — `kaganeImageURLRe` + `browserCoverURL` live in `latest/browser.go` with the rest of the per-Site knowledge — and `BrowserFetcher.Image` is now URL-driven (it validates the URL it will navigate to, same SSRF discipline as before). The no-plain-TLS-fallback rule for a claimed URL is preserved: a claimed address with no browser is an error, never a challenge-page fetch.

## What stayed (deliberately)

- `BrowserFetcher.Image` and the browser-backed acquisition path: kagane genuinely serves cover bytes behind the challenge + `cross-origin-resource-policy: same-origin`, so the sidecar remains the only fetcher for them — it just routes by URL claim now instead of by Site name.
- `fetcherFor`'s per-Site page routing (kagane/novelfull page fetches) — that is the page path, not a cover path.

## Acceptance criteria

- [x] Template-level kagane cover rewrite gone
- [x] Kagane-only cover route and its identifier validation gone
- [x] Tests removed/rewritten against the general route, guarantees kept: unstored + traversal-shaped addresses serve nothing (`TestPublicCoverRejectsUnknownAddress`), non-image content types never echoed (`TestPublicCoverNeverEchoesNonImage` — new; the store-side gate was already pinned by `TestCoverStoreAcceptsAnySourceURL`). Store reopen-persistence and filesystem content-addressing tests rewritten against `PutCover`/`GetCover`, no guarantee lost.
- [x] No Site name in a cover code path outside the acquisition module (`grep kagane backend`: store/web/templates/api are clean; remaining hits are `latest/browser.go` + `latest/sites.go`, tests, docs)
- [x] Web UI and panel render Covers for all six Sites (templates render the wire address; panel renders `b.cover` — untouched, it never had a kagane path)
- [x] `go test ./...` green

## Verification

- `go vet ./...` clean
- `go test ./...` — all packages pass (root 16.9s, latest 12.7s, store 12.7s, web 0.004s)
- `CGO_ENABLED=0 go build` produces the static binary
- Cover-path tests run verbosely: `TestPublicCoverServesStoredBytesUnauthenticated`, `TestPublicCoverRejectsUnknownAddress` (unknown/malformed/traversal/empty), `TestPublicCoverNeverEchoesNonImage`, `TestListRendersAcquiredCover`, `TestAcquireKaganeCoverThroughBrowser`, `TestRunOncePrefetchesKaganeCover`, `TestRunOnceRoutesNonKaganeCoverToPublicFetcher` all pass; the three `SMOKE_*` tests skip without the browser sidecar, as designed

Live browser verification of the "web UI and panel render Covers for all six Sites" criterion is being run separately with Playwright against real Site pages and a locally mocked backend.

Reviewed-on: #73
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 18:02:47 +07:00
sulthan 78234f3c19 Browser-backed Sites join the Cover pipeline (#62) (#72)
Fixes #62

Browser-backed Sites join the Cover pipeline: kagane and novelfull Series now get their Covers at creation, through the same acquisition path as every other Site, instead of waiting for a poll pass.

## What changed

`latest.Acquirer` (creation-time acquisition, fired by the first Bookmark of a Series) previously skipped kagane and novelfull entirely — their pages only yield a Cloudflare challenge to the TLS client, so the request was spent for nothing. It now routes them like the poller does, with the two Sites split exactly as the issue demands:

- **kagane** — page fetched through the browser sidecar, cover URL extracted from the API JSON, bytes fetched through the browser sidecar (the only path that clears the challenge) into the content-addressed store. With no `BROWSER_WS_URL` configured, acquisition is skipped entirely and nothing falls back to a plain fetch.
- **novelfull** — page fetched through the browser sidecar, cover URL extracted from the HTML, bytes fetched over plain TLS through the ordinary gated fetcher (its image paths answer 200 with `access-control-allow-origin: *`, measured 2026-08-09). With no browser configured, the page fetch falls back to the TLS client — novelfull's challenge is a live time-varying fact (AGENTS.md), so when the page body answers, the Cover still lands; when it is challenged, nothing happens.

The byte-routing rule (kagane → browser, every other Site → TLS) is now one shared function (`latest.fetchCoverBytes`) used by both the Poller and the Acquirer, so the two cannot drift apart.

## Acceptance criteria

- [x] kagane cover bytes are fetched through the browser sidecar and stored in the content-addressed store — `TestAcquireKaganeCoverThroughBrowser`
- [x] novelfull cover URLs are extracted from the browser-fetched HTML, and its bytes are fetched over plain TLS — `TestAcquireNovelfullCoverOverPlainTLS`
- [x] With no browser sidecar configured, kagane Covers are absent and nothing falls back to a plain fetch — `TestAcquireKaganeSkippedWithoutBrowser`
- [x] With no browser sidecar configured, novelfull Covers still work if its page body is available — `TestAcquireNovelfullCoverWithoutBrowser`
- [x] Manually verified on-device: a kagane Series shows its Cover in the panel, not a broken-image glyph — being run by a separate manual-verification agent against a mocked scenario (no prod data); not part of this PR
- [x] `go test ./...` is green, with live-network checks gated behind `SMOKE_BROWSER_WS_URL` like the existing kagane image smoke test — new `TestSmokeAcquireKaganeCover` proves the end-to-end acquire path against the real browser when the env var is set

## Verification

- `go test ./...` green across all packages
- New unit tests exercise every routing decision with fakes — no network in the default suite
- Smoke test gated behind `SMOKE_BROWSER_WS_URL`, skipped by default

## Post-review changes (a66491a)

- **One routing rule for pages too** — `fetcherFor` is now a shared function used by both the Poller and the Acquirer; novelfull falls back to the plain-TLS fetcher in *both* when no browser is configured, so pre-existing (client-scraped) novelfull rows get healed by the poll as well, not just Series created after this change (`TestNovelfullUsesTLSWhenNoBrowserFetcher`).
- **Byte-level no-fallback proof** — `TestAcquireKaganeBytesNeverFallBackToPlainTLS` pins that kagane cover bytes never route to the TLS fetcher even when the page came through a browser.
- **Acquirer wired independent of the TLS client** — if `NewTLSFetcher` fails, kagane/novelfull acquisition still works via the sidecar (`main.go`).
- AGENTS.md (root + backend) updated for the novelfull plain-TLS fallback.

Reviewed-on: #72
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 11:06:36 +07:00
sulthan b9220b3dfc The Poll fills blank Covers for every Site and both Libraries (#70)
Closes #61.

## Summary

Permanently-blank Series (the half of #47 that creation-time acquisition cannot reach) heal on the next due poll cycle. The cover path is no longer kagane-only: every Site and both Libraries fill a blank Cover from the series page the chapter poll already fetched, and never replace a Cover that already exists.

## What changed

### `backend/internal/latest/poller.go`

- **`fillBlankCover`** — when `Cover` and `CoverAddress` are both blank, extract a source URL via `coverFrom` from the series-page body and store bytes through `SetSeriesCover`. Skips any Series that already has a source URL (owned by prefetch) or a stored address (never overwrite).
- **`prefetchCover`** — source-URL healing path, now site-uniform. Kagane no longer special-cases into `PutKaganeCover` alone; every Site lands on `SetSeriesCover`, so the wire Cover becomes a content-addressed public URL. Reuses already-stored bytes when present.
- **`storeCover` / `fetchCoverBytes`** — shared fetch+persist. Only kagane routes image bytes through the browser fetcher; every other Site uses plain TLS `CoverBytesFetch`. Failures log with the Series key and never return to the chapter path.
- **`checkOne`** — after a successful series-page fetch, calls `fillBlankCover` once regardless of whether chapter extraction succeeded (cover fill is independent of the chapter signal).

### `backend/internal/latest/poller_test.go`

Extended the existing poller harness (real store, fake fetchers) rather than a new one:

- `TestRunOnceFillsBlankCoverFromSeriesPage` — asura manga, lightnovelworld novel, kagane manga; asserts wire Cover + correct fetcher routing.
- `TestRunOnceDoesNotReplaceExistingCover` — second poll does not refetch.
- `TestRunOnceRetriesFailedBlankCoverOnNextPoll` — failed fill stays blank, next due cycle retries (no separate queue).
- `TestRunOnceBlankCoverFailureDoesNotBlockChapter` — chapter still lands; failure log carries the Series key.
- Kagane prefetch test now also asserts the content-addressed wire Cover.

## Acceptance criteria (#61)

| Criterion | Status |
|---|---|
| Cover prefetch runs for every Site | done |
| Cover prefetch runs for both Libraries | done |
| Poll fills a blank Cover | done |
| Poll never replaces an existing Cover | done |
| Failed cover fetch does not fail/block chapter poll | done |
| Failed cover fetch retried next poll, no separate queue | done |
| Failures logged with the Series | done |
| Existing poller tests extended | done |
| `go test ./...` green | done |
| Manually verified: blank Series gets Cover after a poll cycle | **left for you** |

## Out of scope / not closed

- Does **not** close #47 or #55 (per ticket).
- No migration/backfill script — the Poll walks every Series already.
- No admin refetch (#54).

## Review notes addressed

- Removed the kagane-only `PutKaganeCover` branch from prefetch so source-URL healing also sets `CoverAddress` (wire Cover).
- Guard so `fillBlankCover` does not double-fetch after `prefetchCover` healed the same snapshot.
- Single `fillBlankCover` call site after the series-page fetch.

## Test plan

- [x] `go test ./...` (backend; needs Docker/Postgres via `pgtest`)
- [ ] After deploy: pick a Series that was blank, wait one poll cycle, confirm Cover in web UI and userscript panel

Reviewed-on: #70
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 10:14:27 +07:00
sulthan e2c054e7ce Covers render in the userscript panel, from a public route (#60) (#69)
Closes #60.

Spec: #55. Originating bug: #47. Architecture: `docs/adr/0007-backend-hosts-cover-bytes.md`. Neither #47 nor #55 is closed from here.

## What this branch does

The panel now renders Covers from the deployment's own origin, and both userscripts stop having an opinion about where a Cover lives.

**The public route was already in place.** `GET /covers/{address}` landed with #59 (`92eba07`) and is registered on the bare mux, outside `httpmw.Auth` and outside the web UI's Discord session — `backend/main.go:210-214`, handler `backend/internal/api/handlers.go:142-158`. It reads no cookie and no header, answers `404` for an address that was never stored (and for a row whose file has gone missing — recorded-but-gone is not-found, never a fabricated body), refuses anything that is not `^[0-9a-f]{64}$` *before* the value becomes a path, and sets `Cache-Control: public, max-age=604800, immutable`. Those four properties are asserted by `backend/cover_test.go:231-278`. This branch re-verified them rather than re-implementing them; the only backend line it touches is a comment.

**Both userscripts lose cover scraping entirely.** Every adapter's `cover:` field is gone, along with the two helpers that fed them: the manga script's `coverFromPage()` (the `img[alt]` DOM scan comix needed, because comix publishes no `og:image`) and the novel script's `metaName()` plus the now-callerless module-level `meta()`. Nothing under `userscript/` reads `og:image`, `meta[name=image]`, or `img[alt]` any more.

**Nothing sends a cover either.** `delete body.cover` sits in `apiPut` — `manga-bookmark.user.js:486`, `novel-bookmark.user.js:275` — which is the single chokepoint every write passes through (`pushBookmark`, the retry-queue flush, `toggleFavorite`, `toggleArchive`). It operates on the `Object.assign` copy, so the in-memory row keeps the cover it renders with. This matters beyond tidiness: a Reader upgrading from an older copy has `localStorage` rows carrying third-party scraped URLs, and without the strip those would ride back up on the next write. The handler discards the field regardless (`handlers.go:53-59`) — it is permanently inert, not pending removal.

**Failed loads get the designed empty state, not the broken-image glyph.** `onerror: (e) => e.target.replaceWith(el("div", { class: "cover ph" }))` on the cover `<img>` in both card renderers (`manga:1380-1390`, `novel:1134-1144`). The replacement is byte-identical to the existing no-cover branch on the very next line, so it picks up the `.cover.ph` styling already in the panel CSS — no new tokens, no new rule. `el()` routes any `on*` prop through `addEventListener`, so this is a listener, not an inline attribute string, and the swap is a `createElement` + DOM call with no markup parsing anywhere near it. This is the half of #47 that was visible on kagane.

**The deleted scraping's tests went with it**: the two comix cover cases, the `pageImages` and `namedMetas` fixtures, the `img[alt]` and `meta[name=...]` stub branches, the now-dead `querySelectorAll` stub member, and every stale `og:image` fixture and `p.cover` assertion across both suites. The export lists needed no change and that was checked, not assumed — `coverFromPage` and `metaName` were module-private on `origin/main` and no cover symbol ever appeared in `module.exports`.

Docs that described the deleted behaviour were corrected in the same breath, because leaving them would instruct the next agent to put the scraping back: `userscript/AGENTS.md` (adapter contract + the per-site notes for comix, kagane and novelfull), the README's adapter reference, and the userscript testing skill's stub table.

## Verification

- `go test -count=1 ./...` — green across all nine packages (`backend` 29.8s, `latest`, `store`, `session`, `token`, `userscript`, `web`).
- `node --check` clean on both userscripts; `node --test` on both logic suites — 46 tests, 46 pass.
- `gofmt -l` clean; `go build ./...` clean.
- The `onerror` swap is DOM behaviour and deliberately has no coverage in the Node harness — that harness stubs a browser precisely so it never needs a DOM, and #60 says not to invent coverage for it. It was instead exercised for real: the `el()` helper and the exact render expression were loaded into a headless Chromium with a deliberately unloadable `src`, and the resulting DOM was `<div class="cover ph"></div>`. Ad hoc, not committed.
- **Not done, needs you:** the on-device criterion — a comix Series bookmarked mid-chapter showing its Cover in the panel. That needs a real install against the deployment and is the one box left unticked on #60.

## Reviewed

Both `/code-review` axes ran against `cc0fa92`. Spec found no missed requirement and no scope creep; standards found the diff clean on the four areas it scrutinised (the `delete body.cover` placement, the `onerror` handler's DOM safety, comment quality, dead-code removal). Their combined findings — the dead `querySelectorAll` stub, the stale README and skill text, and the handler comment whose premise this change invalidates — are fixed in `8b58019`.

## Out of scope, deliberately

The kagane-specific cover proxy still exists and still carries its session gate (#63 deletes it). The poll's blank-Cover fill (#61) and browser-backed Sites joining the pipeline (#62) are untouched.

Reviewed-on: #69
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 08:52:57 +07:00
sulthan 92eba07da7 A newly bookmarked Series acquires its Cover at creation (#59) (#68)
Closes #59.

Part of spec #55, and the ticket that fixes the reported bug #47. Architecture: `docs/adr/0007-backend-hosts-cover-bytes.md`. Does not close #47 or #55.

## What changed

A Reader bookmarks a Series nobody holds yet — the exact case in #47 — and within seconds the list shows its artwork instead of a broken image. The first Bookmark to create a Series fires `Store.OnSeriesCreated` after commit, and the new `latest.Acquirer` turns that into **one** series-page fetch that yields both the Latest Chapter and the cover URL. The bytes go through the gated cover fetcher from #57 and are stored content-addressed through #56, so the wire carries an absolute URL on this deployment's own origin — never a third-party address, and never one that 404s.

### Store

- Migration `0009_series_cover_address.sql` adds `series.cover_address`. The two facts are now split: `series.cover` is the third-party source address the bytes came from (the acquisition path's dedupe key), `series.cover_address` is the SHA-256 they are stored under. An empty `cover_address` is precisely what "no Cover yet" means, which is the distinction both the API and the UI depend on.
- `SetSeriesCover` writes the address only after the bytes are on disk, so the wire can never name an object that is not there.
- `CoverWireURL` builds `PUBLIC_BASE_URL + /covers/<sha256>` for every scanned row, and returns `""` for a blank address.
- The cover columns are gone from `Upsert`'s `INSERT` and its `DO UPDATE`. A client-supplied cover cannot reach the shared Series row on any path, not just the creation path.
- `Open` now rejects a base URL that is not an absolute `http(s)` origin: `PUBLIC_BASE_URL=bookmarks.example.com` would otherwise start cleanly and emit addresses no browser can load.

### Acquisition

- `internal/latest/acquire.go`: one fetch, gated by the poller's own `fetchableSeriesURL` (a `series_url` arrives in a client-supplied PUT body, so without the gate a token-holder chooses what the server fetches from its own network position).
- Asynchronous and log-and-drop. The Bookmark, its progress and its Latest Chapter are already committed; a Site that is down or a cover that cannot be produced disturbs none of them.
- Bounded by a two-slot semaphore. A bulk sync creating N Series would otherwise fire N simultaneous requests from one IP — the traffic shape the poller's stagger exists to avoid.
- Cancelled at shutdown (shares the poller's context) and stamps `latest_checked_at`, so the poller does not refetch the same page a tick later.
- Browser-backed Sites (kagane, novelfull) are deliberately skipped: their pages only yield a Cloudflare challenge to the TLS client, so the request would be spent for nothing. They arrive in #62.

### Wire and route

- `GET /covers/{address}` serves the bytes publicly and uncredentialed with `Cache-Control: public, max-age=604800, immutable`. The address is gated by a `^[0-9a-f]{64}$` pattern and cross-checked against a pure function of itself before any filesystem read, so no request shaped like a traversal reaches disk.
- `PUT /bookmarks/{key}` still accepts a `cover` field and discards it, permanently. Rejecting it would break every installed userscript the moment this deploys, and ADR-0004's compatibility argument depends on those scripts continuing to work. The decode site says so in place of a TODO nobody intends to keep.
- `store.CoverContentType` canonicalises comix's non-standard `image/jpg` to `image/jpeg`, so one image cannot land under two spellings. This one was found by the live smoke test, not by reading.

### Config

`PUBLIC_BASE_URL` is new and required (cover URLs must go out absolute — the userscript renders them on third-party origins, where a relative path resolves against the Site). Documented in `.env.example`, `docker-compose.yml` (`:?` so compose fails too), `DEPLOY.md` and `backend/AGENTS.md`.

## Acceptance criteria

All twelve of #59's criteria are met; the checklist on the issue is ticked with the evidence.

## Verification

- `go test ./...` green (Docker-backed Postgres suite).
- Live smoke against a real backend + Postgres: bookmarking `comix:n8we-dungeons-and-crayons` produced `"cover": "http://127.0.0.1:8099/covers/8ce74d80…"` and `"latest_chapter": "Chapter 81"` within seconds of the PUT; `curl` on that address returned `200`, `Content-Type: image/jpeg`, `Cache-Control: public, max-age=604800, immutable`, and a 280x420 JPEG. That run is what surfaced the `image/jpg` content type.
- Mutation-checked the asynchrony test: removing the `go` from `Acquire` turns `TestAcquireDoesNotBlockTheWrite` red.

## Reviewed

Both axes of `/code-review` were run against this diff before commit. Their findings that were actionable here are folded in: the concurrency bound, the shutdown tie, the `PUBLIC_BASE_URL` validation, the missing `latest_checked_at` stamp, and a test that could not fail.

## Known sequencing

A kagane/novelfull Series created between this deploy and #62 has no cover source at all: the acquisition skips those Sites and `Upsert` no longer persists the userscript-scraped address. This is #59's stated boundary rather than a defect, but it is a user-visible gap on two Sites and should order #62 accordingly.

Reviewed-on: #68
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 04:07:53 +07:00
sulthan b6b88bde8a feat(latest): extract per-site covers (#58) (#67)
Closes #58

## Summary

- Add pure per-Site cover extraction beside latest-chapter parsing for all six Sites.
- Read Asura, Demonic, LightNovelWorld, and NovelFull metadata; read the Comix target detail state; read Kagane's browser-fetched `series_covers[].image_id` JSON.
- Preserve published cover URLs, percent-encode Demonic raw spaces, select Comix's smaller published `medium`, and avoid thumbnail rendition URL synthesis.
- Add live-source fixtures plus no-cover and Cloudflare challenge coverage for every Site.

## Correctness

- Scope Comix extraction to the requested series detail key, avoiding recommended posters.
- Parse Kagane's current live API shape and emit its canonical compressed image route from the published image ID; unrelated JSON fields are ignored.
- Validate Kagane image IDs against the existing UUID-shaped route constraint.
- Keep extraction pure; storage, polling, and wire integration remain outside issue #58.

## Acceptance criteria

- [x] Cover extraction exists for all six Sites in the existing latest parser module.
- [x] Each Site has a live-source fixture with source URL and date.
- [x] Comix reads the state blob, not metadata.
- [x] Demonic raw spaces are percent-encoded.
- [x] Comix returns the smaller published rendition.
- [x] No-cover pages return empty.
- [x] Cloudflare challenge pages return empty.
- [x] No thumbnail URL is synthesized by editing a published URL.
- [x] `go test ./...` passes.

## Verification

- `go test ./...`
- `go vet ./...`
- `git diff --check`

Parent issues #47 and #55 remain open as requested.

Reviewed-on: #67
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 01:43:57 +07:00
sulthan 9d6d3bde72 Add gated cover byte fetcher (#66)
## Summary

Adds a plain-TLS cover byte fetcher with a destination-class SSRF gate and wires public cover sources through the content-addressed filesystem store.

## Changes

- Resolve hostnames before connecting; refuse non-HTTPS, loopback, private, link-local, unique-local, CGNAT, credentials, and mixed public/private DNS answers.
- Re-check every redirect and resolve/classify again at dial time to close DNS rebinding.
- Reuse `maxBodyBytes`; reject oversized responses and non-image content types before persistence.
- Add generic `Store.GetCover`/`PutCover` source-URL storage while preserving the browser-backed kagane path.
- Keep cover prefetch failures isolated from chapter polling.
- Add observable tests for TLS, no-connection refusals, all refused address classes, redirect blocking, streaming body caps, non-image rejection, content-addressed persistence, DNS rebinding, and poller routing.

## Verification

- `go test -count=1 ./...`
- `go vet ./...`

Both pass. No test touches the live network.

Closes #57

Reviewed-on: #66
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 00:45:55 +07:00
sulthan e8d1cba6c5 Move cover bytes to content-addressed filesystem storage (#65)
Refs #56

## Summary

Moves Kagane cover bytes out of Postgres bytea storage into an immutable, content-addressed filesystem store. Reader-visible behavior remains unchanged: the existing session-gated route serves stored bytes, missing bytes use the existing browser fetch path, and no browser still returns a missing cover.

## Changes

- Added migration 0008, which drops the legacy `covers` table and recreates it with only `address`, `path`, and `content_type`. Existing byte rows are intentionally dropped.
- Added SHA-256 source-URL addressing with two-level sharding (`ab/cd/<sha256>`). Writes use a temp file plus atomic link; reads validate the stored relative path before opening it.
- Made `COVER_DIR` required in runtime config and Compose. Compose passes it as a Docker build argument and volume target, so custom durable paths keep image ownership, runtime config, and the named `cover-data` volume aligned.
- Updated every `store.Open` caller and documented configuration, deployment, backup, and troubleshooting behavior.
- Added filesystem, restart, migration-drop, no-browser, and content-addressing coverage.

## Verification

- `go test ./...`
- `CGO_ENABLED=0 go build ./...`
- `docker build --build-arg COVER_DIR=/data/covers -t manga-bookmark-cover-check-custom ./backend`
- `docker compose config --format json` confirms custom `COVER_DIR` is the volume target
- `git diff --check origin/main`
- LSP diagnostics clean for touched Go files

Parents #47 and #55 remain open as required by #56.

Reviewed-on: #65
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 00:11:20 +07:00
sulthan 30c57bd39c Define Cover and record hosting its bytes (#47) (#64)
Defines **Cover** in the glossary and records ADR-0007, the decision behind #47's fix.

## Why these two files, and why now

`CONTEXT.md` named Cover inside the **Series** entry — "facts true regardless of who is reading — title, cover, Latest Chapter" — but never said *what* one is. That gap is the bug. Nothing in the model distinguished "an address on a Site" from "an image a Reader's browser can display", so both clients were left to work it out independently, and one of them got it wrong. kagane serves covers with `cross-origin-resource-policy: same-origin`, the web UI rewrote them to a proxy in its templates, the JSON API did not, and the panel rendered a broken-image glyph. The new entry closes the ambiguity: *an address no client can load is not a Cover, it is a missing one.*

ADR-0007 records what follows from that — the backend fetches, stores and serves every Site's cover bytes — plus the alternatives that were rejected and, more importantly, the two places this deliberately departs from existing precedent:

- **Destination-class control instead of a host allowlist.** `fetchableSeriesURL` sets the allowlist precedent for `series_url`, and covers do not follow it. Cover hosts are CDNs that move independently of their Site — demonicscans serves its covers from `readermc.org` — so an allowlist would stop producing Covers the day a Site switched CDN, and that failure would look exactly like #47. The resolve-then-classify step is what actually stops the SSRF.
- **A public cover route where the kagane proxy is session-gated.** An `<img>` cannot send a bearer token, and it cannot be given one either: the panel's shadow root is `mode: "open"`, so the host page's JavaScript can read any `src` the script sets.

Both are security-adjacent departures, which is precisely why they are written down rather than left in a commit message.

## Scope

Documentation only — no code, no schema, no behaviour. The implementation is #56–#63.

## Why this should merge promptly rather than sit

All eight implementation tickets cite `docs/adr/0007-backend-hosts-cover-bytes.md` as the authority for decisions they must not relitigate, and they are written in the vocabulary this glossary entry defines. An agent picking up #56 reads both from `main`. Until this lands they get a 404 and either invent a rationale or stall — so this PR gates the tickets, not the other way round.

Related: #47 (bug), #55 (spec), #54 (deferred admin refetch).
Reviewed-on: #64
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 23:19:37 +07:00
sulthan 8081a0a5d8 Give the Tailscale ACL step a working policy file (#53)
Follow-up to #52, which merged before this landed. Docs only — no code, no compose changes.

`DEPLOY.md` §7 told the operator to "tag the two machines" and showed a bare `acls` fragment. Following it literally does not work and is actively harmful:

- the fragment references `tag:bookmark-api` / `tag:bookmark-browser` without a `tagOwners` section, so the policy is rejected on save;
- it never says how a tag gets onto a device (`tailscale up --advertise-tags=...`, which re-authenticates);
- replacing the tailnet's default allow-all with only that one rule **removes the operator's own SSH access to the browser machine**.

Replaced with a complete, saveable policy file: `tagOwners`, the CDP rule, a second rule preserving own-device access including `:22`, and a `tests` block so a later edit that widens 9222 is rejected rather than silently applied.

Also records two things that were assumed rather than stated:

- **Why tagging is load-bearing.** Tailscale has no `deny`, so restricting 9222 means removing the blanket accept and enumerating what remains. That is only expressible if the browser machine falls outside a selector that still covers your own devices — which is exactly what a tag does, since a tagged device has no user and stops matching `autogroup:member` / `autogroup:self`. Without that, the whole step reads as arbitrary ceremony.
- **Tagging replaces a device's user identity**, so it suits a dedicated box and disrupts a daily driver. Both paths are now written down.

Finally, separates two checks the old text conflated: the existing `curl` runs on the home machine and proves only the **bind**, because node-local traffic is not filtered. Proving the **ACL** needs a third device, so that check is now its own step.

Verified: `tailscale.com/docs/reference/syntax/policy-file` and `/docs/features/tags` (validated Apr 2026 / Dec 2025) for `tagOwners`, `autogroup:self` semantics vs tagged devices, `--advertise-tags` re-auth and key-expiry behaviour. Markdown fences balanced.
Reviewed-on: #53
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 16:12:03 +07:00
sulthan 2a3bb6922d Move the browser off the VPS to its own unit (#46) (#52)
Closes #46 once deployed.

The headless browser leaves the API stack and becomes its own compose unit
(`chrome/docker-compose.yml`) intended for the home machine, reached over the
tailnet. No fallback sidecar is left on the VPS.

The backend needs no code change — `BROWSER_WS_URL` was already the only
coupling. Its default is now empty rather than a pinned Docker IP, so an
unconfigured or unreachable browser degrades exactly as it always has: plain-TLS
libraries unaffected, kagane/novelfull logged and skipped, stored covers still
served.

### What shipped

- `chrome/docker-compose.yml` + `chrome/.env.example` — the browser unit, with
  the CDP port bound to `${BROWSER_BIND_ADDR}` (no default) and the resource
  limits from the epic: 512 MiB / 1 GiB memory+swap, `oom_score_adj 800`,
  halved CPU weight, shm 1 GiB -> 128 MiB.
- API stack drops the service, its `depends_on` and the `browser` network.
- `bookmark-api` gains the `default` network. Dropping `browser` had left it on
  `db` alone, which is `internal: true` — no published port and, worse, no
  egress for the poller at all. Caught by actually bringing the stack up.
- ADR-0006 for the topology; `DEPLOY.md` §7 for first-time setup of the browser
  machine; `REDEPLOY.md` §8 for its independent update cadence; architecture
  diagrams, config tables and troubleshooting rows across README/AGENTS/env.

### Verified locally

- Browser unit builds and runs: Chrome 151, UA carries no `HeadlessChrome`,
  all limits applied as declared.
- **Live smoke passes through the new unit**: `TestSmokeKaganeImage` fetched
  56710 bytes of `image/webp`, `TestSmokeKaganeGet` got a 200 with a real
  chapter list. The challenge cleared under the reduced 128 MiB shm.
- Bind isolation proven: refused on the host's non-loopback address, accepted
  on the configured one.
- 321 MiB peak of the 512 MiB cap after a full solve; 0 restarts, no OOM kill.
- API stack comes up clean, `/healthz` 200; egress confirmed present on
  `default` and absent on `db`.
- `go test ./...`, `go vet`, `gofmt` clean.

### Left to the operator

Provisioning the home machine, the Tailscale ACL, setting `BROWSER_WS_URL` in
production, and observing acceptance criteria 5-7 (covers with the machine off,
several days of zero OOM/restarts, VPS memory improvement). `DEPLOY.md` §7 now
carries the before/after `free -m` reading those need.

Reviewed-on: #52
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 15:28:21 +07:00
sulthan d1800d0707 Prefetch Kagane covers during latest polling (#51)
## Summary
- Add an optional browser-backed cover fetcher to the latest-chapter poller.
- Prefetch missing Kagane covers during the existing due-series cycle and persist them before a Reader opens the web UI.
- Keep chapter polling, cooldown stamping, and on-first-view fallback independent from cover failures.

## Behavior and safety
- Stored Kagane covers are detected before browser work, so later poll cycles do not refetch them.
- Nil cover fetchers and non-Kagane series retain the existing behavior.
- Shared Kagane image-id and content-type validation prevents challenge or non-image responses from poisoning persistent cover storage.
- The browser is wired into both the chapter and cover poller paths from the composition root.

## Verification
- `go test ./...`
- Focused latest, store, and web package tests
- Deterministic tests cover missing covers, cached covers, failed fetches, invalid content types, nil fetchers, and non-Kagane series.

Closes #45

Reviewed-on: #51
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 14:58:59 +07:00
sulthan 84cfd1b2c1 Make browser sidecar on-demand (#44) (#50)
Closes #44. Chrome now starts on first CDP connection, tracks concurrent helpers, reaps after 300 seconds idle, preserves the named profile, and classifies reap interruptions. Shutdown stops Chrome's process group so cookie batches flush. ADR-0005 records the measured constraints and decisions. Verification: docker build, live CDP wake, graceful stop cleanup, sh -n, and go test ./... (7 packages, 3 no tests).

Reviewed-on: #50
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 07:39:07 +07:00
sulthan bfae84c5c3 Persist kagane covers in Postgres (#49)
Closes #43

Persist kagane cover bytes in a dedicated Postgres covers table keyed by image ID. The web handler reads storage before the browser, writes validated fetches through, and no longer keeps an in-process cover cache. Added migration, store persistence tests including reopen, handler coverage for stored/miss/rejected paths, and corrected repository guidance.

Verification:
- go test ./...
- CGO_ENABLED=0 go build ./...

Reviewed-on: #49
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 06:52:02 +07:00
sulthan cd3a7e3d01 feat(latest): split browser poll cooldown (#48)
## Summary

Split latest-chapter polling cooldowns by fetch cost. Browser-backed kagane and novelfull series now rest longer without changing the cadence of plain-TLS sites.

## Behavior

- Plain-TLS series keep the 1h default cooldown.
- Browser-backed series use `LATEST_CHAPTER_POLL_BROWSER_COOLDOWN`, defaulting to 6h.
- Both cooldowns share the existing 15m minimum floor; invalid values retain the existing fallback behavior.
- The poller still selects both classes in one due query per cycle.
- Existing ordering and exclusions remain unchanged: reader-count precedence, least-recently-checked ordering, finished exclusion, archived polling, and orphan exclusion.

## Implementation

- Added the browser cooldown to backend configuration and passed it through production poller construction.
- Added the browser-site list as the single routing source used for both due-query cutoff selection and fetcher choice.
- Kept all query values parameterized; the site list is passed as a bound PostgreSQL array parameter.
- Updated startup logging to report interval, plain cooldown, browser cooldown, batch, and stagger.
- Documented the variable, default, and floor in `README.md`, `.env.example`, `backend/AGENTS.md`, and `docker-compose.yml`.

## Review findings addressed

The first review found that configuration parsing was correct but `startLatestPoller` did not pass `BrowserCooldown` into `latest.Poller`; every browser-backed row would therefore have been due immediately. Production construction now goes through `newLatestPoller`, with a regression test covering both cooldown fields.

The review also identified duplicated browser-site knowledge in fetch routing. `slices.Contains(browserBackedSites, site)` now reuses the same list already supplied to the store query.

## Verification

- Focused backend tests pass: `go test ./internal/latest ./internal/store .`.
- Full suite passes: `go test ./...`.
- `graphify update .` completed.
- Issue #42 was updated and closed.

Reviewed-on: #48
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 06:25:48 +07:00
48 changed files with 4194 additions and 825 deletions
@@ -14,7 +14,7 @@ parsers, helpers. UI, network, and storage behaviour are verified on-device.
```bash
node --check userscript/manga-bookmark.user.js # parse check, silent on success
node --test userscript/test/logic.test.js # 14 tests as of 2026-07-28
node --test userscript/test/logic.test.js # 35 tests as of 2026-08-10
```
Run both before every commit that touches the userscript.
@@ -31,7 +31,7 @@ The test file installs four globals **before** requiring the userscript:
|---|---|---|
| `localStorage` | `Map`-backed stub | `loadCache`, `loadQueue`, and the key-migration IIFE touch it at module scope |
| `location` | `{href, hostname, pathname, origin}` | read during boot |
| `document` | `querySelector` for `meta[property="…"]` only, plus a no-op `addEventListener` | adapters read `og:title`/`og:image` |
| `document` | `querySelector` for `meta[property="…"]` only, plus a no-op `addEventListener` | adapters read `og:title` (covers are the backend's, never scraped) |
| `document.body` | **left `undefined`** | this is the whole trick |
`document.body === undefined` sends the userscript's boot block down its `else`
+37 -28
View File
@@ -25,6 +25,15 @@ POSTGRES_PASSWORD=changeme-generate-a-long-random-password
# Override only to point the backend at a Postgres compose does not run.
# DATABASE_URL=postgres://user:pass@host:5432/bookmarks?sslmode=require
# Directory inside bookmark-api for immutable, content-addressed Cover bytes.
# Compose builds the image and mounts its named volume at this path.
COVER_DIR=/covers
# Public origin this deployment answers on, no trailing slash. Required: Cover
# URLs go out absolute, because the userscript renders them on a Site's own
# origin where a relative path would resolve against the Site (ADR-0007).
PUBLIC_BASE_URL=https://bookmark-api.example.com
# --- Prod override (Traefik) only ---
# Subdomain Traefik routes to this service (required by the prod override).
# BOOKMARK_API_HOST=bookmark-api.example.com
@@ -67,13 +76,15 @@ DISCORD_REDIRECT_URI=
# Set to 0 to turn it off entirely.
# LATEST_CHAPTER_POLL_ENABLED=1
#
# Two independent clocks. COOLDOWN is how long one series rests between checks;
# INTERVAL is how often the poller wakes up and looks for series past that
# cooldown. Shortening INTERVAL cannot shorten a COOLDOWN.
# LATEST_CHAPTER_POLL_COOLDOWN=1h # per series, floor 15m
# LATEST_CHAPTER_POLL_INTERVAL=10m # how often to wake
# LATEST_CHAPTER_POLL_BATCH=14 # series per wake
# LATEST_CHAPTER_POLL_STAGGER=20s # delay between fetches in a batch
# Two independent clocks. COOLDOWN is how long a plain-TLS series rests between
# checks; BROWSER_COOLDOWN is the longer rest for kagane and novelfull. INTERVAL
# is how often the poller wakes up and looks for series past their cooldowns.
# Shortening INTERVAL cannot shorten either cooldown.
LATEST_CHAPTER_POLL_COOLDOWN=1h # plain-TLS per series, floor 15m
LATEST_CHAPTER_POLL_BROWSER_COOLDOWN=6h # browser-backed per series, floor 15m
LATEST_CHAPTER_POLL_INTERVAL=10m # how often to wake
LATEST_CHAPTER_POLL_BATCH=14 # series per wake
LATEST_CHAPTER_POLL_STAGGER=20s # delay between fetches in a batch
#
# Uses a ticker, not an immediate first run: the first poll happens one
# INTERVAL after startup, not at startup. A container restarting more often
@@ -83,29 +94,27 @@ DISCORD_REDIRECT_URI=
# defaults. Beyond that the cadence stretches uniformly rather than breaking;
# raise BATCH or lower INTERVAL. Keep BATCH x STAGGER under INTERVAL.
# Headless-shell CDP endpoint for sites behind a JavaScript challenge (kagane).
# Unset disables browser polling; those sites then rely on the userscript alone.
# Leave commented — the compose files' own default (ws://172.28.0.10:9222) is
# correct. Do NOT set this to the "headless-shell" DNS name: Chrome's DevTools
# HTTP handler 500s any /json/version request whose Host header isn't an IP or
# "localhost", which silently breaks every kagane poll.
# BROWSER_WS_URL=ws://172.28.0.10:9222
# Clock zone the headless browser reports. A UTC clock is itself the bot
# signal — Cloudflare treats it as the datacenter default — and kagane's
# challenge then never clears. Measured 2026-08-08, identical container, one
# Indonesian egress IP: UTC never cleared in 60s (twice); Asia/Jakarta and
# America/New_York both cleared in 4s. So any real zone works; it does not
# have to match the IP's country, it just must not be UTC.
# CDP endpoint of the browser, used for the two sites behind a Cloudflare
# JavaScript challenge (kagane, novelfull) and by the web UI's kagane cover
# proxy. Unset disables browser polling and serves 404 for covers not already
# stored; those sites then rely on the userscript alone. That is also exactly
# how an unreachable browser degrades, so a home machine that is off costs
# chapter freshness and nothing else.
#
# Unset falls back to the host's /etc/timezone, which is a real zone whenever
# the host clock is set to local time. Set this when the host runs UTC — a UTC
# server is exactly the case that fails. Only the browser sidecar reads it —
# the backend's own zone is API_TZ below, and is cosmetic.
# BROWSER_TZ=Asia/Jakarta
# The browser does NOT run in this stack. It is its own compose unit on the
# home machine (chrome/docker-compose.yml, chrome/.env.example) and is reached
# over the tailnet, so set this to that machine's tailnet address:
#
# BROWSER_WS_URL=ws://100.x.y.z:9222
#
# It must be the tailnet **IP**, never a MagicDNS hostname and never the old
# Docker service name: Chrome's DevTools HTTP handler 500s any /json/version
# request whose Host header isn't an IP or "localhost", which silently breaks
# every kagane poll. Left unset here on purpose — a wrong default would poll a
# stranger's address, and "no browser" is a safe, self-announcing state.
# BROWSER_WS_URL=ws://100.x.y.z:9222
# Zone the backend stamps its log lines in. Cosmetic only — it exists so the
# API's logs read on the same clock as the browser sidecar's. Nothing else in
# Zone the backend stamps its log lines in. Cosmetic only. Nothing else in
# the service has a zone: bookmark timestamps are unix ms, and the two real
# time columns are timestamptz. Defaults to Asia/Jakarta; set to UTC for the
# conventional server default.
+11 -6
View File
@@ -19,8 +19,9 @@ Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free
- Every site is its **own origin with its own `localStorage`** — a shared remote store is the only way to unify bookmarks. Cloud sync required, not optional.
- Userscript run in **isolated world**, so embedded API token safe from site's JS.
- Cloudflare's block on manga sites **IP-reputation-based, not universal — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — Cloudflare's bot scoring can flip previously-clean IP without notice. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test.
- **kagane.to and novelfull.com are the exception to the above** — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`) and skips them entirely when that's unset. The four other sites poll fine over plain TLS.
- **The CDP sidecar must look like a real browser, and stock headless images don't.** Measured 2026-08-08 against kagane.to, all from the same IP: `chromedp/headless-shell:stable` never cleared the challenge in 90s (`navigator.webdriver` true, empty plugin list, Chromium-branded client hints — suppressing `webdriver` alone changed nothing); `zenika/alpine-chrome` ships Chrome 124, refused outright; real Chrome with the default `--headless=new` UA never cleared, because the UA says `HeadlessChrome`; real Chrome with a stock UA **and** a non-UTC clock zone cleared in ~4s. Hence `chrome/` — a Debian image with `google-chrome-stable`, a version-derived UA, and `TZ`/`BROWSER_TZ`. Chrome reads the zone *name* through ICU from `/etc/localtime`'s symlink target, ignoring the file's contents, so mounting the host's `/etc/localtime` does **not** work; `/etc/timezone` is mounted instead.
- **kagane.to and novelfull.com are the exception to the above** — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`). When that's unset, kagane is skipped entirely (a plain fetch would only retrieve a challenge page) while novelfull pages are still attempted over plain TLS — its challenge is a live time-varying fact and its cover bytes never need the browser. The four other sites poll fine over plain TLS.
- **The CDP browser must look like a real browser, and stock headless images don't.** Measured 2026-08-08 against kagane.to, all from the same IP: `chromedp/headless-shell:stable` never cleared the challenge in 90s (`navigator.webdriver` true, empty plugin list, Chromium-branded client hints — suppressing `webdriver` alone changed nothing); `zenika/alpine-chrome` ships Chrome 124, refused outright; real Chrome with the default `--headless=new` UA never cleared, because the UA says `HeadlessChrome`; real Chrome with a stock UA **and** a non-UTC clock zone cleared in ~4s. Hence `chrome/` — a Debian image with `google-chrome-stable`, a version-derived UA, and `TZ`/`BROWSER_TZ`. Chrome reads the zone *name* through ICU from `/etc/localtime`'s symlink target, ignoring the file's contents, so mounting the host's `/etc/localtime` does **not** work; `/etc/timezone` is mounted instead.
- **The browser is not in the API stack and must not be put back.** It's its own compose unit (`chrome/docker-compose.yml`) on a second machine, reached over the tailnet — it held 471 MiB on a 1974 MiB swapless VPS, and a residential egress scores better with Cloudflare anyway (ADR-0006). Consequences that constrain code: `BROWSER_WS_URL` must be a tailnet **IP** (a MagicDNS name 500s at `/json/version`, same trap as the old Docker service name); the CDP port binds to the tailnet address only, since CDP authenticates nothing and that host has a real LAN; and the browser is on-demand (ADR-0005), so an unreachable or asleep one must degrade exactly as an unset `BROWSER_WS_URL` — plain-TLS libraries unaffected, kagane/novelfull logged and skipped, stored covers still served. Never add `chromedp.NoModifyURL`: discovery per fetch is what makes a restarted Chrome invisible.
- **UTC is the tell, not a country mismatch.** A UTC clock is the datacenter default, so Cloudflare scores it as one; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while `Asia/Jakarta` **and** `America/New_York` both cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (`Asia/Bangkok`) rather than the measured egress. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one.
- **A challenged page needs the tab kept open.** The interstitial takes seconds to solve and only then writes clearance into the browser's shared cookie jar. Navigate-read-close never clears anything; `BrowserFetcher.run` holds one tab and re-reads until the payload arrives.
@@ -29,9 +30,13 @@ Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free
```
Two Violentmonkey userscripts (isolated world, per-site adapters, localStorage cache)
-- fetch() HTTPS --> reverse proxy (TLS + CORS) --> Go net/http --> Postgres (volume)
|
| CDP over tailnet
v
on-demand Chrome, separate machine (chrome/)
```
Backend-specific architecture (packages, endpoints, poller, config env vars) lives in `backend/AGENTS.md`. Userscript-specific structure (adapters, retry queue, UI, live URL shapes) lives in `userscript/AGENTS.md`.
Two deployable units on two machines: the API stack (`docker-compose.yml` + `docker-compose.prod.yml`, on the VPS) and the browser (`chrome/docker-compose.yml`, on the home machine). They share nothing but `BROWSER_WS_URL` and update independently. Backend-specific architecture (packages, endpoints, poller, config env vars) lives in `backend/AGENTS.md`. Userscript-specific structure (adapters, retry queue, UI, live URL shapes) lives in `userscript/AGENTS.md`. Deploy order `DEPLOY.md` (§7 for the browser), redeploy `REDEPLOY.md` (§8 for the browser).
## Commands
@@ -40,10 +45,10 @@ Backend (`cd backend`):
- Single test: `go test -run TestName ./...`
- Build static binary: `CGO_ENABLED=0 go build`
Local stack: `docker compose up` (bookmark-api + postgres + headless-shell; `postgres-data` named volume, `restart: unless-stopped`). The `headless-shell` service keeps its name but now builds `chrome/` — real Google Chrome, for the reason in the hard constraints above.
Local stack: `docker compose up` (bookmark-api + postgres only; `postgres-data` named volume, `restart: unless-stopped`). No browser — without `BROWSER_WS_URL` the poller logs and skips kagane and novelfull. To run one: `cd chrome && BROWSER_BIND_ADDR=172.17.0.1 docker compose up -d --build`, then `BROWSER_WS_URL=ws://172.17.0.1:9222` in the root `.env` (bridge gateway, so the API container can name it by IP).
Live CDP proof (needs a sidecar and network, skipped otherwise):
`SMOKE_BROWSER_WS_URL=ws://<host>:<port> go test -run TestSmokeKagane ./internal/latest`
Live CDP proof (needs that browser and network, skipped otherwise):
`SMOKE_BROWSER_WS_URL=ws://<ip>:<port> go test -run TestSmokeKagane ./internal/latest`
— fetches a real kagane cover and chapter list. A red run means the challenge is
not clearing from this IP, which is a live fact to re-check, not necessarily a defect.
+7
View File
@@ -18,6 +18,13 @@ _Avoid_: manga, title, book, comic
One third-party source a Series is published on. A Series on two Sites is two Series.
_Avoid_: source, host, provider, domain
**Cover**:
The image that stands for a Series wherever it is listed. A fact about the Series like
its title — one Cover per Series, shared by every Reader, never per-Reader. Defined by
what a Reader's browser can display, not by where the Site keeps the picture: an address
no client can load is not a Cover, it is a missing one.
_Avoid_: thumbnail, poster, image URL, artwork
**Reader**:
A person with their own Progress. Exactly one per set of credentials, so there is no
separate "account" concept to model — the credential belongs to the Reader.
+228 -9
View File
@@ -53,6 +53,15 @@ POSTGRES_PASSWORD=<paste output of: openssl rand -hex 24>
# not run; it then replaces the URL built from POSTGRES_PASSWORD above.
# DATABASE_URL=postgres://user:pass@host:5432/bookmarks?sslmode=require
# Required path inside bookmark-api. Compose builds the image and mounts the
# named cover-data volume at this path.
COVER_DIR=/covers
# Required — the origin this deployment answers on, no trailing slash. Cover
# URLs on the wire are absolute, because the userscript renders them on a
# Site's own origin (ADR-0007). Same host as BOOKMARK_API_HOST below.
PUBLIC_BASE_URL=https://bookmark-api.violetcrown.my.id
# Required for the Traefik override. Both have no fallback — compose refuses
# to start without them. BOOKMARK_WEB_HOST is required even if the web UI
# were unused; see 1b.
@@ -157,15 +166,17 @@ This merges the base file (build/image/env/volume) with the prod override
(no host port, Traefik network + router labels). Always pass **both** `-f`
flags — the prod file is not standalone.
Three services come up: `bookmark-api` (the backend), `postgres` (its database,
`postgres:17-alpine`), and `headless-shell`, a CDP sidecar the poller uses to
fetch kagane (behind a Cloudflare JS challenge). Neither of the latter two
publishes a port: `postgres` sits alone with `bookmark-api` on an
`internal: true` network, and `headless-shell` is reachable only over
`BROWSER_WS_URL`. A missing headless-shell just makes the poller skip kagane and
log it. A missing Postgres stops everything — `bookmark-api` waits for
`pg_isready` to pass, then applies its embedded migrations, and only then
listens. The schema is created that way; there is nothing to import by hand.
Two services come up: `bookmark-api` (the backend) and `postgres` (its
database, `postgres:17-alpine`). Postgres publishes no port — it sits alone
with `bookmark-api` on an `internal: true` network — and stops everything if it
is missing: `bookmark-api` waits for `pg_isready` to pass, then applies its
embedded migrations, and only then listens. The schema is created that way;
there is nothing to import by hand.
There is deliberately no browser here. Kagane and novelfull need one, and it
runs on a **separate machine** over the tailnet — §7. Until you do that step,
`BROWSER_WS_URL` is unset, the poller logs and skips those two sites, and
everything else works normally.
Check it's up and healthy:
@@ -259,6 +270,204 @@ copy immediately — reinstall on all devices, or they silently stop syncing.
---
## 7. The browser, on the home machine
Kagane and novelfull sit behind a Cloudflare JavaScript challenge no TLS
fingerprint clears, so the poller reaches them through a real Chrome over CDP.
That browser does **not** run on the VPS: it held 471 MiB of a 1974 MiB box
with no swap, and it scores better from a residential IP anyway (ADR-0006). It
is its own compose unit, deployed and updated independently of everything
above.
Do this after §2, on the second machine. Both machines must already be on the
same tailnet.
First, on the VPS, record what you are reclaiming — this is the whole point of
the move and there is no way to measure it afterwards:
```bash
free -m | awk '/^Mem:/ {print "available before:", $NF, "MiB"}'
```
Take it again after §7 is finished and the old sidecar is gone. Expect roughly
the sidecar's former footprint back (measured at 471 MiB working set, 595 MiB
cgroup).
**On the home machine:**
```bash
git clone <this repo> ~/mangaBookmark && cd ~/mangaBookmark/chrome
tailscale ip -4 # -> 100.x.y.z, this machine's tailnet IP
cp .env.example .env
echo "BROWSER_BIND_ADDR=$(tailscale ip -4)" >> .env
docker compose up -d --build
```
The clone is only for `chrome/`; nothing else on this machine reads the rest of
the repo. The unit is its own compose project (`bookmark-browser`), so it shares
no volume, network or lifecycle with an API stack that happens to sit beside it.
`BROWSER_BIND_ADDR` has no default on purpose. CDP authenticates nothing —
whatever reaches port 9222 drives the browser and, through it, this host — so
the bind address *is* the access control, backed by Tailscale device identity.
On the VPS that job was done by Docker network membership; this machine has a
real LAN, so `0.0.0.0` would be a hole punched into your home network. Compose
refuses to start rather than guess.
**Narrow it to the one device that needs it.** The bind address keeps CDP off
your LAN; it still leaves port 9222 open to every device on the tailnet, and
CDP has no login — a compromised phone is enough to drive this host. A new
tailnet's policy is allow-all, so this is the step that makes "Tailscale
identity is the access control" true rather than aspirational.
Tailscale has no `deny`, so a restriction is expressed by removing the blanket
grant and enumerating what is left. That only works if the browser machine can
be *excluded* from a selector that still covers your own devices — which is
what tagging buys: a tagged device has no user, so `autogroup:member` and
`autogroup:self` stop matching it. Tagging is the mechanism, not decoration.
In the admin console, under **Access controls**, the shipped policy grants
`{"src": ["*"], "dst": ["*"], "ip": ["*"]}`. Replace it:
```jsonc
{
"tagOwners": {
// Empty list: implicitly owned by the tailnet Owner/Admins, which is you.
"tag:bookmark-api": [],
"tag:bookmark-browser": [],
},
"grants": [
// The only thing on the tailnet that may drive the browser.
{
"src": ["tag:bookmark-api"],
"dst": ["tag:bookmark-browser"],
"ip": ["tcp:9222"],
},
// Your own devices reach your own devices, and the VPS, in full.
{
"src": ["autogroup:member"],
"dst": ["autogroup:self", "tag:bookmark-api"],
"ip": ["*"],
},
// On the browser machine you get SSH and nothing else. Widen this to `*`
// and the restriction above is void; delete it and you are locked out.
{
"src": ["autogroup:member"],
"dst": ["tag:bookmark-browser"],
"ip": ["tcp:22"],
},
// Uncomment if you route traffic through an exit node — dropping the
// blanket grant takes exit-node access with it.
// {"src": ["autogroup:member"], "dst": ["autogroup:internet"], "ip": ["*"]},
],
// Tagged devices left `autogroup:self`, so Tailscale SSH needs them named.
// Irrelevant if you reach these boxes with ordinary sshd over the tailnet —
// that is the `tcp:22` grant above.
"ssh": [
{
"action": "check",
"src": ["autogroup:member"],
"dst": ["autogroup:self", "tag:bookmark-api", "tag:bookmark-browser"],
"users": ["autogroup:nonroot", "root"],
},
],
// Run on every save, so a later edit that reopens 9222 is rejected outright.
"tests": [
{ "src": "tag:bookmark-api", "accept": ["tag:bookmark-browser:9222"] },
{
"src": "you@example.com",
"accept": ["tag:bookmark-browser:22"],
"deny": ["tag:bookmark-browser:9222"],
},
],
}
```
Then apply the tags — on the VPS and the home machine respectively:
```bash
sudo tailscale up --advertise-tags=tag:bookmark-api
sudo tailscale up --advertise-tags=tag:bookmark-browser
```
Each re-authenticates in a browser and issues a new node key; the tailnet IP is
unchanged, so `BROWSER_WS_URL` and `BROWSER_BIND_ADDR` still hold. Key expiry is
disabled once a device is tagged, which is what you want for a server — an
expired key would otherwise take the poller down every few months.
**Tagging replaces the device's user identity**, so do this only to machines
that exist to run these services. If your "home machine" is also your daily
driver, tag it anyway and reach it through the `:22` rule above, or skip the
tag and accept that any device of yours can reach CDP.
Enforcement is by the destination's packet filter, so the check below is real,
not advisory.
Prove the bind is tight, from the home machine itself:
```bash
curl -s -m 3 http://$(tailscale ip -4):9222/json/version # -> JSON
curl -s -m 3 http://<this machine's LAN IP>:9222/json/version
# -> curl: (7) Failed to connect ... Connection refused
```
The first call is also what wakes Chrome: it is not running until something
connects, and it is reaped again after five idle minutes. A cold first response
takes a few seconds; that is the browser starting, not a fault.
That check proves the *bind*, not the ACL — traffic that starts on the node is
not filtered. Prove the ACL from somewhere else: on your laptop or phone the
same URL must now time out, and from the VPS it must answer.
```bash
# on any other device of yours -> hangs until timeout
curl -s -m 5 http://<home machine tailnet IP>:9222/json/version
# on the VPS -> JSON
curl -s -m 20 http://<home machine tailnet IP>:9222/json/version
```
**On the VPS:**
```bash
cd ~/mangaBookmark
echo 'BROWSER_WS_URL=ws://100.x.y.z:9222' >> .env # the home machine's tailnet IP
docker compose -f docker-compose.yml -f docker-compose.prod.yml up -d
```
It must be the tailnet **IP**. A MagicDNS hostname fails: Chrome's DevTools HTTP
handler answers `/json/version` with a 500 for any `Host` header that is not an
IP or `localhost`, and the failure looks like a broken site rather than a broken
hostname.
**Prove it end to end.** This is the only check that says the challenge actually
clears from that machine's egress — it fetches a real kagane cover and a real
chapter list:
```bash
cd backend
SMOKE_BROWSER_WS_URL=ws://100.x.y.z:9222 go test -run TestSmokeKagane ./internal/latest
```
A red run means "not clearing from this address right now", which is a live
fact to re-check before it is a defect — Cloudflare's scoring moves. Then, from
the web UI, open a bookmarked kagane series and confirm the cover renders. Once
a cover is stored it is served from Postgres forever after, so the browser being
asleep, unreachable, or mid-power-outage costs chapter freshness and nothing
visible.
Finally, take the VPS `free -m` reading again and compare it against the one
from the top of this section.
**Updating the browser** is independent of the API stack and has its own
runbook — `REDEPLOY.md` §8.
---
## Updating
Pull new code, then rebuild:
@@ -272,6 +481,10 @@ server predates the Postgres migration, the old SQLite volume `bookmarks-data`
is still on disk and deliberately undeclared in compose so `down -v` cannot take
it; see `REDEPLOY.md` §1 for when to remove it.)
The browser is a separate unit on a separate machine with its own update
command — §7. Nothing above touches it, and it needs no coordination: the API
picks up a restarted Chrome's new debugger UUID by itself.
---
## Troubleshooting
@@ -287,6 +500,12 @@ it; see `REDEPLOY.md` §1 for when to remove it.)
| `compose ... config` errors about `TOKEN_KEY`, `OWNER_DISCORD_ID` or `POSTGRES_PASSWORD` | Run compose from the dir with `.env`, or export the vars. All three are required and none has a fallback. |
| `bookmark-api` restarts in a loop, `password authentication failed for user "bookmarks"` | `POSTGRES_PASSWORD` was changed after first boot; Postgres only applies it to an empty `postgres-data`. Restore the old value, or reset the role (`REDEPLOY.md` troubleshooting). |
| `bookmark-api` never logs `listening on :8080` | It is blocked on `postgres` passing `pg_isready`, or a migration failed. `docker compose -f docker-compose.yml -f docker-compose.prod.yml logs postgres`. |
| kagane rows never get a `latest_chapter`; log says `browser fetcher disabled` or nothing at all | `BROWSER_WS_URL` unset. Expected before §7 is done. |
| kagane polls all fail; log shows a 500 from `/json/version` | `BROWSER_WS_URL` names a MagicDNS hostname (or any name). Chrome's DevTools handler only accepts an IP or `localhost` — use the tailnet IP. |
| kagane polls fail with a connection error | Home machine off, off the tailnet, or the unit is down. `tailscale ping <machine>`, then `docker compose ps` in its `chrome/`. Costs freshness only; stored covers keep serving. |
| kagane cover is a placeholder for a newly bookmarked series | Its cover has never been fetched and the browser is unreachable. It fills in on the next successful poll of that series (up to `LATEST_CHAPTER_POLL_BROWSER_COOLDOWN`, default 6h). |
| `compose` in `chrome/` errors `set BROWSER_BIND_ADDR to this machine's tailnet IP` | No `chrome/.env`, or the variable is empty. Deliberate — it has no default so an unset value cannot publish CDP to the LAN. |
| browser container restarts, or is OOM-killed | `docker inspect bookmark-browser --format '{{.RestartCount}} {{.State.OOMKilled}}'`. The 512 MiB cap is sized against a measured 645 MiB untuned peak; a real breach is a Chrome regression worth reading `docker logs` for, not a number to raise reflexively. |
Backend config reference and endpoint list: see `README.md`.
+44 -7
View File
@@ -15,8 +15,21 @@ Two parts:
```
Bromite userscript (isolated world, Shadow DOM UI, localStorage cache)
-- fetch() HTTPS --> reverse proxy (TLS + CORS) --> Go net/http --> Postgres (volume)
|
| CDP over the tailnet
v
headless Chrome, on-demand,
on a separate machine
(chrome/, ADR-0006)
```
Kagane and novelfull sit behind a Cloudflare JavaScript challenge no TLS
fingerprint clears, so the poller reaches those two through a real Chrome over
CDP. That browser is **not** part of the API stack: it is its own compose unit
on a second machine, spawned on the first connection and reaped when idle. The
API needs it only to discover new chapters and to fetch a kagane cover once —
covers are stored, so the library renders in full with the browser switched off.
---
## 1. Backend
@@ -29,8 +42,9 @@ Bromite userscript (isolated world, Shadow DOM UI, localStorage cache)
| `OWNER_DISCORD_ID` | *(required)* | Discord user ID of the owner: seeded as the first Reader, owns every pre-registration bookmark, and is the only Reader who can revoke another's sessions. |
| `ALLOWED_ORIGINS` | Asura + Demonic + Comix + Kagane origins | Comma-separated CORS allowlist. |
| `DATABASE_URL` | *(required)* | Postgres connection URL, e.g. `postgres://bookmarks:…@postgres:5432/bookmarks?sslmode=disable`. Compose builds it from `POSTGRES_PASSWORD`. |
| `COVER_DIR` | *(required)* | Filesystem volume for immutable, content-addressed Cover bytes. Compose builds the image and mounts `cover-data` at this path; standalone runs may choose another writable durable path. |
| `PORT` | `8080` | Plain HTTP; TLS terminated by the proxy. |
| `BROWSER_WS_URL` | `ws://172.28.0.10:9222` | Headless-shell CDP endpoint used to poll Kagane past its JS challenge. Must be an IP or `localhost` — Chrome's DevTools handler 500s any other Host header. |
| `BROWSER_WS_URL` | empty | CDP endpoint of the remote browser (`ws://<tailnet IP>:9222`), used to poll Kagane/Novelfull past their JS challenge and to fetch uncached Kagane covers. Must be an IP or `localhost` — Chrome's DevTools handler 500s any other Host header, MagicDNS names included. Unset disables both; stored covers still serve. |
| `DISCORD_CLIENT_ID` | *(required)* | Discord application credentials for the browser sign-in (ADR-0002). |
| `DISCORD_CLIENT_SECRET` | *(required)* | As above. Never logged, never echoed in an error. |
| `DISCORD_GUILD_ID` | *(required)* | The one guild whose membership gates sign-in, checked at login only. Membership *is* registration: any member becomes a Reader on first login. |
@@ -40,8 +54,9 @@ Bromite userscript (isolated world, Shadow DOM UI, localStorage cache)
| `USERSCRIPT_PATH` | `/userscript/manga-bookmark.user.js` | Bindmounted file served at `/u/{token}/manga-bookmark.user.js`. |
| `NOVEL_USERSCRIPT_PATH` | `/userscript/novel-bookmark.user.js` | Same, for the novel library. |
| `LATEST_CHAPTER_POLL_ENABLED` | `1` | `0` turns the poller off entirely. |
| `LATEST_CHAPTER_POLL_COOLDOWN` | `1h` | Rest between checks of one series; floor `15m`. |
| `LATEST_CHAPTER_POLL_INTERVAL` | `10m` | How often the poller wakes. Cannot shorten a cooldown. |
| `LATEST_CHAPTER_POLL_COOLDOWN` | `1h` | Rest between checks of one plain-TLS series; floor `15m`. |
| `LATEST_CHAPTER_POLL_BROWSER_COOLDOWN` | `6h` | Rest between checks of one browser-backed series; floor `15m`. |
| `LATEST_CHAPTER_POLL_INTERVAL` | `10m` | How often the poller wakes. Cannot shorten either cooldown. |
| `LATEST_CHAPTER_POLL_BATCH` | `14` | Series per wake. Keep `BATCH × STAGGER` under `INTERVAL`. |
| `LATEST_CHAPTER_POLL_STAGGER` | `20s` | Delay between fetches in a batch — this is the outbound request rate. |
@@ -49,8 +64,11 @@ Compose reads a few more from the same `.env` that the backend never sees:
`POSTGRES_PASSWORD` (required — `DATABASE_URL` is built from it, and Postgres
only applies it while `postgres-data` is empty), `BOOKMARK_API_HOST` and
`BOOKMARK_WEB_HOST` (required by the prod override), and the optional
`PROXY_NETWORK` / `TRAEFIK_ENTRYPOINT` / `TRAEFIK_CERTRESOLVER`. Full commentary
is in `.env.example`; deployment order is `DEPLOY.md`.
`PROXY_NETWORK` / `TRAEFIK_ENTRYPOINT` / `TRAEFIK_CERTRESOLVER`. The browser
unit has its own `chrome/.env` on its own machine — `BROWSER_BIND_ADDR`
(required, the tailnet IP the CDP port is published on) and the optional
`BROWSER_TZ`. Full commentary is in `.env.example` and `chrome/.env.example`;
deployment order is `DEPLOY.md`.
### Endpoints
@@ -96,6 +114,23 @@ cp .env.example .env
docker compose up -d --build # binds 127.0.0.1:8080
```
That brings up two services — the API and Postgres. The browser is deliberately
not one of them; without `BROWSER_WS_URL` the poller logs and skips kagane and
novelfull, and everything else works. To run one locally, publish it on the
Docker bridge gateway so the API container can name it by IP:
```bash
cd chrome
echo 'BROWSER_BIND_ADDR=172.17.0.1' > .env
docker compose up -d --build
# then in the repo's own .env: BROWSER_WS_URL=ws://172.17.0.1:9222
```
Bind it to `127.0.0.1` instead if you only want to drive it from the host, e.g.
`SMOKE_BROWSER_WS_URL=ws://127.0.0.1:9222 go test -run TestSmokeKagane ./internal/latest`.
In production that address is the home machine's tailnet IP and nothing else —
see `DEPLOY.md` §7 and ADR-0006.
Smoke test:
```bash
@@ -226,8 +261,10 @@ an API.
## Adapter reference (verified live 2026-07-24)
The site adapters key everything off URL regex, with `title`/`cover` from
`og:title` / `og:image`. Confirmed against live pages via Playwright:
The site adapters key everything off URL regex, with `title` from `og:title`
(or the page's own heading where a site ships none). No adapter reads a cover:
the backend acquires, stores and serves every Cover from its own origin
(ADR-0007). Confirmed against live pages via Playwright:
| Site | Series URL | Chapter URL | `series_id` |
|------|-----------|-------------|-------------|
+63 -5
View File
@@ -59,9 +59,13 @@ network can reach it — so every command below goes in through the container:
```bash
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks -c '\dt'
# -> bookmarks, readers, schema_migrations, series, sessions
# -> bookmarks, covers, readers, schema_migrations, series, sessions
```
The `covers` table is metadata only after the filesystem cutover: bytes live in
the separate `cover-data` volume. Back that volume up with the database dump;
restoring only Postgres leaves stored Cover addresses without files.
Inside the container that connects over the local socket as the `bookmarks`
superuser, so no password is needed anywhere in this section. `-T` is not
optional: without it Compose allocates a TTY, which rewrites `\n` to `\r\n` and
@@ -99,10 +103,11 @@ you will act as though you have one:
docker run --rm -v "$BACKUP_DIR":/backup postgres:17-alpine \
pg_restore --list "/backup/bookmarks-$STAMP.dump" | grep 'TABLE DATA'
# -> 1234; 0 0 TABLE DATA public bookmarks bookmarks
# -> 1235; 0 0 TABLE DATA public readers bookmarks
# -> 1236; 0 0 TABLE DATA public schema_migrations bookmarks
# -> 1237; 0 0 TABLE DATA public series bookmarks
# -> 1238; 0 0 TABLE DATA public sessions bookmarks
# -> 1235; 0 0 TABLE DATA public covers bookmarks
# -> 1236; 0 0 TABLE DATA public readers bookmarks
# -> 1237; 0 0 TABLE DATA public schema_migrations bookmarks
# -> 1238; 0 0 TABLE DATA public series bookmarks
# -> 1239; 0 0 TABLE DATA public sessions bookmarks
# 2. Sanity-check the live row count you just captured.
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks \
@@ -387,6 +392,55 @@ panel works on the phone.
---
## 8. The browser unit (separate machine, separate cadence)
Everything above is the API stack on the VPS. The headless browser is its own
compose unit on the home machine (ADR-0006, `DEPLOY.md` §7) and is redeployed
on its own schedule — it holds no data you can lose, so there is nothing to
back up and no ordering constraint against the API.
```bash
cd ~/mangaBookmark/chrome
git pull --ff-only
docker compose up -d --build
```
Then confirm it answers, and that a stopped-and-restarted Chrome is invisible
to the API:
```bash
curl -s -m 15 http://$(tailscale ip -4):9222/json/version | head -c 120
# -> {"Browser":"Chrome/1xx...","webSocketDebuggerUrl":"ws://...<new uuid>"}
```
The first call takes a few seconds: Chrome is not running until something
connects, and it is reaped again after five idle minutes. The debugger UUID
changes on every start and the API does not care — chromedp re-runs
`/json/version` discovery per fetch, which is exactly why `chromedp.NoModifyURL`
must never be added to `browser.go`.
**Rebuild is the Chrome upgrade path.** The image installs
`google-chrome-stable` unpinned on purpose: a stale browser is what Cloudflare
turns away, and the pinned Chrome 124 in `zenika/alpine-chrome` is the worked
example. The `chrome-profile` volume survives `--build`, so clearance cookies
are reused rather than re-solved.
Two things worth a glance after several days, both from the acceptance criteria
of the move:
```bash
docker inspect bookmark-browser --format '{{.RestartCount}} {{.State.OOMKilled}}'
# -> 0 false
free -m # the Gitea runner should still have its headroom
```
Nothing here needs doing during an API redeploy. The API stack does not
`depends_on` the browser, and an unreachable one degrades exactly as an unset
`BROWSER_WS_URL`: plain-TLS libraries unaffected, kagane and novelfull logged
and skipped, stored covers still served.
---
## Troubleshooting
| Symptom | Cause / fix |
@@ -404,6 +458,10 @@ panel works on the phone.
| `pg_restore`: `cannot drop … other objects depend on it` / `being accessed by other users` | Live connections block `--clean`. `$COMPOSE stop bookmark-api` first (§6). If they persist: `$COMPOSE exec -T postgres psql -U bookmarks -d postgres -c "select pg_terminate_backend(pid) from pg_stat_activity where datname='bookmarks' and pid <> pg_backend_pid()"`. |
| Dump is 0 bytes, or `pg_restore`: `did not find magic string in file header` | You ran `exec` without `-T`. The allocated TTY rewrites newlines in the binary stream and corrupts the archive in flight (§1). |
| `git pull`: `could not read Username for 'https://…'` | The checkout's remote is the HTTPS clone URL and the server has no credential helper, so the pull prompts into a closed stdin. Switch it to SSH once — `git remote set-url origin ssh://git@gitea.violetcrown.my.id:2222/sulthan/mangaBookmark.git`. Gitea's SSH listens on **2222**, not 22; port 22 is the host's own sshd and answers `Permission denied (publickey)` no matter which key is registered. |
| kagane rows stopped updating after a redeploy | Check `BROWSER_WS_URL` survived the `.env` edit and still names the home machine's tailnet **IP**. A hostname 500s at `/json/version`; an empty value disables the browser silently. Plain-TLS sites keep working either way, which is why this is easy to miss. |
| kagane covers went blank in the web UI | Covers use the `cover-data` volume now. Restore/check that volume alongside Postgres; rows in `covers` are metadata only. If the database has rows but files are missing, the next browser-backed request refetches them; without a browser it remains a 404. |
| Browser unit will not start: `set BROWSER_BIND_ADDR to this machine's tailnet IP` | `chrome/.env` is missing or the variable is empty. It has no default on purpose — an unset value must fail the deploy rather than publish an unauthenticated CDP port to the LAN. |
| `bookmark-browser` shows `OOMKilled true` | The cap did its job. Read `docker logs bookmark-browser` before raising it — the sizing and what the cap protects are in ADR-0006. |
Full first-time setup: `DEPLOY.md`. The one-off SQLite→Postgres move:
`CUTOVER.md`. Config reference and endpoints: `README.md`.
+51 -18
View File
@@ -91,13 +91,37 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
Fetches use `bogdanfinn/tls-client` with Chrome profile as defence in depth
against fingerprint-based blocking; any failure log and skip. kagane and
novelfull sit behind Cloudflare JavaScript challenges the TLS client can't
clear, so they are browser-only: fetched over CDP via `BROWSER_WS_URL`, and
simply not polled when that's unset. See
clear, so they are fetched over CDP via `BROWSER_WS_URL`; kagane is simply
not polled when that's unset, while novelfull falls back to a plain-TLS
attempt — its challenge is a live time-varying fact, and its cover bytes
never need the browser. See
`docs/superpowers/specs/2026-07-26-server-latest-chapter-polling-design.md`.
The poller's series write is a single-column UPDATE
(`Store.SetLatestChapter`), not a read-modify-write of the whole bookmark:
it cannot revert read progress or move `updated_at`, so the old
stale-re-read race is gone with the Get+Upsert flow.
- **Covers are acquired at creation, then served from our own origin
(ADR-0007):** the first Bookmark of a Series fires `Store.OnSeriesCreated`,
which `latest.Acquirer` turns into one series-page fetch yielding both the
Latest Chapter and the cover URL; the bytes then go through
`latest.CoverBytesFetcher` into `Store.SetSeriesCover`. It runs in a
goroutine — the Reader's PUT must neither block on a Site nor fail with one
— and every failure is logged and dropped, leaving the Bookmark intact. The
wire's `cover` is the absolute `PUBLIC_BASE_URL + /covers/{sha256}` once
bytes exist and `""` before, never an address that 404s. `GET /covers/{addr}`
is public and uncredentialed: the userscript renders it on a Site's origin,
where no cookie or token of ours travels. A client-sent `cover` is decoded
and discarded, permanently (ADR-0004 compatibility).
Browser-backed Sites join the same pipeline (issue #62): kagane pages *and*
cover bytes go through the browser sidecar (nothing falls back to a plain
fetch, which would only retrieve a challenge page), while novelfull needs
the browser only for its HTML — the cover URL comes out of the
browser-fetched page and the bytes go over plain TLS. With no browser
configured, kagane Covers are simply absent; novelfull still gets one — at
creation and on the poll — when its page body happens to answer a plain
request (the challenge is a live time-varying fact). The old kagane-only
serving path (`/img/kagane/{id}`, template rewrite, `CoverFetcher`) is gone
(issue #63): the one public route serves every Site.
- **`updated_at` drives list order, so moves only on real reading progress:** server apply its timestamp when row new or `last_chapter_num` changes, else keep stored value — favouriting series or recording newly published chapter must not reorder list. `PUT` therefore returns row **as stored**, clients must adopt that response rather than own payload. See `plans/2026-07-25-bookmark-list-favorites-design.md` §4.
- **Lifecycle buckets:** `status` on each bookmark is `reading` | `archived` |
`finished`, orthogonal to `favorite`. Archived and finished appear only in
@@ -114,32 +138,41 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
the owner of every pre-registration bookmark; required),
`ALLOWED_ORIGINS` (comma list),
`DATABASE_URL` (Postgres connection URL, required — no default),
`COVER_DIR` (required filesystem volume for content-addressed Cover bytes),
`PUBLIC_BASE_URL` (required origin this deployment answers on, trailing
slash trimmed; every Cover URL on the wire is built from it, absolute
because the userscript renders on a Site's origin — ADR-0007),
`PORT` (default `8080`), `DISCORD_CLIENT_ID`/`_CLIENT_SECRET`/`_GUILD_ID`/
`_REDIRECT_URI` (required; Discord OAuth for the browser UI),
`DISCORD_REQUIRED_ROLE` (optional role gate, empty by default),
`DISCORD_API_BASE` (default `https://discord.com/api/v10`),
`LATEST_CHAPTER_POLL_ENABLED`/`_COOLDOWN`/`_INTERVAL`/`_BATCH`/`_STAGGER`
(background latest-chapter poller; defaults on, `1h`/`10m`/`14`/`20s`).
`LATEST_CHAPTER_POLL_ENABLED`/`_COOLDOWN`/`_BROWSER_COOLDOWN`/`_INTERVAL`/
`_BATCH`/`_STAGGER` (background latest-chapter poller; defaults on,
`1h` plain-TLS cooldown, `6h` browser cooldown, `10m`/`14`/`20s`; both
cooldowns have a `15m` floor).
`USERSCRIPT_PATH` and `NOVEL_USERSCRIPT_PATH` (files served at
`/u/{token}/manga-bookmark.user.js` and `/u/{token}/novel-bookmark.user.js`,
defaults `/userscript/manga-bookmark.user.js` and
`/userscript/novel-bookmark.user.js`, both supplied by bindmount; the
`__API_TOKEN__` placeholder inside them is substituted with the requesting
Reader's credential at serve time).
`BROWSER_WS_URL` (CDP endpoint of the `chrome/` sidecar, used by the poller
for kagane and novelfull *and* by the web UI's kagane cover proxy; unset
disables browser polling and serves 404 from the proxy, leaving those sites
to the userscript alone).
- **kagane covers are proxied, not hot-linked:** kagane serves cover images
behind the same challenge as its pages and with
`cross-origin-resource-policy: same-origin`, so no `<img>` on the web UI's
origin can load one — not even from a browser holding the clearance cookie
(verified 2026-08-08). `Bookmark.CoverURL` rewrites a stored kagane
`og:image` to `/img/kagane/{id}`, served by `internal/web/cover.go` through
`latest.BrowserFetcher.Image` and memoised in-process. The templates render
`.CoverURL`, never `.Cover`. The id is matched against a UUID regex before it
reaches the browser: the stored value is client-supplied, so an unchecked one
is an SSRF primitive pointed at the deployment's own network.
`BROWSER_WS_URL` (CDP endpoint of the browser, which runs on a **separate
machine** and is reached over the tailnet — ADR-0006, `chrome/docker-compose.yml`.
Used by the poller for kagane and novelfull page fetches and by the cover
pipeline for kagane's image bytes (the browser is the only route that clears
the challenge kagane serves its covers behind); unset — the default —
disables browser polling and leaves kagane Covers blank until stored bytes
exist. Must be a tailnet IP, never a hostname: Chrome's DevTools handler 500s
`/json/version` for any Host that isn't an IP or `localhost`).
- **No per-Site cover path (issue #63):** every Cover — all six Sites — is
served by the one public `GET /covers/{addr}` route from content-addressed
bytes. There is no proxy, no per-Site rewrite, no second place that decides
a Cover's renderable address: the wire `cover` is it. The only place a Site
name still appears in cover code is the extraction module (`latest`), where
kagane's image URLs are claimed by `browserOnlyCoverURL` — they answer a
plain fetch with a challenge and `cross-origin-resource-policy: same-origin`;
every other Site's CDN answers plain TLS. Templates render `.Cover` — the
wire value — never anything else.
- **Web UI also owns:** session-gated `GET /install/{manga,novel}-bookmark.user.js`
(renders the bindmounted script with the acting Reader's derived credential
substituted in — the credential never appears in page markup, the address
+6 -1
View File
@@ -2,6 +2,7 @@
# --- build stage: compile a static, CGO-free binary ---
FROM golang:1.26-alpine AS build
ARG COVER_DIR=/covers
WORKDIR /src
# Dependencies first for layer caching (changes rarely).
@@ -19,11 +20,15 @@ COPY internal/ ./internal/
# -trimpath + -ldflags strip paths and debug info for a smaller image.
RUN CGO_ENABLED=0 GOOS=linux go build -trimpath -ldflags="-s -w" -o /out/server .
# Create the source directory; runtime COPY sets ownership for the named volume.
RUN mkdir -p "$COVER_DIR"
# --- runtime stage: distroless static, non-root ---
FROM gcr.io/distroless/static:nonroot
ARG COVER_DIR=/covers
WORKDIR /
COPY --from=build --chown=65532:65532 ${COVER_DIR} ${COVER_DIR}
COPY --from=build /out/server /server
EXPOSE 8080
USER nonroot:nonroot
ENV PORT=8080
+8 -2
View File
@@ -26,6 +26,10 @@ const testTokenKey = "test-token-key"
// credential is a function of it.
const testDiscordID = "test-owner"
// testCoverBaseURL is the public origin cover URLs are built from, standing in
// for PUBLIC_BASE_URL.
const testCoverBaseURL = "https://bookmarks.test"
func testConfig() Config {
return Config{
TokenKey: testTokenKey,
@@ -60,7 +64,7 @@ func newTestStoreURL(t *testing.T) (*store.Store, string) {
url := pgtest.URL(t)
s, err := store.Open(url, store.Owner{
DiscordID: testDiscordID, TokenHash: token.Hash(ownerCredential()),
})
}, t.TempDir(), testCoverBaseURL)
if err != nil {
t.Fatalf("store.Open: %v", err)
}
@@ -329,12 +333,14 @@ func TestFlatWireFieldSet(t *testing.T) {
latestNum := floatPtr(8)
want := store.Bookmark{
Key: key, Site: "comix", SeriesID: "some-title",
Title: in.Title, SeriesURL: in.SeriesURL, Cover: in.Cover,
Title: in.Title, SeriesURL: in.SeriesURL,
LastChapter: in.LastChapter, LastChapterNum: in.LastChapterNum,
LastChapterURL: in.LastChapterURL, Favorite: true,
LatestChapter: in.LatestChapter, LatestChapterNum: latestNum,
Status: store.StatusArchived, Kind: store.KindManga,
}
// Cover is deliberately absent above: the client's cover is discarded, and
// this wiring acquires none, so the field is present and empty (ADR-0007).
if stored.Title != want.Title || stored.SeriesURL != want.SeriesURL || stored.Cover != want.Cover ||
stored.LastChapter != want.LastChapter || stored.LastChapterNum != want.LastChapterNum ||
stored.LastChapterURL != want.LastChapterURL || stored.Favorite != want.Favorite ||
+96 -117
View File
@@ -1,35 +1,15 @@
package main
import (
"context"
"errors"
"database/sql"
"net/http"
"net/http/httptest"
"sync/atomic"
"strings"
"testing"
"bookmarkmanager/backend/internal/store"
)
// fakeCovers stands in for the headless browser. It counts calls so the test
// can prove the cache spares the browser a second navigation.
type fakeCovers struct {
body []byte
contentType string
err error
calls atomic.Int32
lastID atomic.Value
}
func (f *fakeCovers) Image(_ context.Context, imageID string) ([]byte, string, error) {
f.calls.Add(1)
f.lastID.Store(imageID)
if f.err != nil {
return nil, "", f.err
}
return f.body, f.contentType, nil
}
const testCoverID = "019fe11a-84c3-7fc3-a84b-88787374b617"
func getCover(t *testing.T, srv http.Handler, path string, cookie *http.Cookie) *httptest.ResponseRecorder {
t.Helper()
req := httptest.NewRequest(http.MethodGet, path, nil)
@@ -41,115 +21,114 @@ func getCover(t *testing.T, srv http.Handler, path string, cookie *http.Cookie)
return rr
}
// kagane serves its covers behind a Cloudflare challenge and with
// cross-origin-resource-policy: same-origin, so the UI can only show one by
// re-serving the bytes from its own origin.
func TestKaganeCoverProxiesAndCaches(t *testing.T) {
cf := &fakeCovers{body: []byte("\x00webp-bytes"), contentType: "image/webp"}
cfg := testConfig()
cfg.Covers = cf
srv, st := newWebTestServer(t, cfg)
cookie := sessionCookie(t, st)
// The acquired Cover is served from this deployment's own origin, to any
// browser rendering a third-party page — no session, no credential (ADR-0007).
func TestPublicCoverServesStoredBytesUnauthenticated(t *testing.T) {
const sourceURL = "https://cdn.asurascans.com/covers/solo.webp"
srv, st := newWebTestServer(t, testConfig())
if err := st.PutCover(sourceURL, []byte("\x00webp-bytes"), "image/webp"); err != nil {
t.Fatalf("PutCover: %v", err)
}
for i := range 2 {
rr := getCover(t, srv, "/img/kagane/"+testCoverID, cookie)
if rr.Code != http.StatusOK {
t.Fatalf("request %d: status = %d, want 200", i, rr.Code)
}
if got := rr.Body.String(); got != string(cf.body) {
t.Fatalf("request %d: body = %q, want %q", i, got, cf.body)
}
if got := rr.Header().Get("Content-Type"); got != "image/webp" {
t.Fatalf("request %d: Content-Type = %q, want image/webp", i, got)
}
// The wire URL is what a client actually requests, so the path under test
// is taken from it rather than rebuilt by hand.
wire := st.CoverWireURL(store.CoverAddress(sourceURL))
path, ok := strings.CutPrefix(wire, testCoverBaseURL)
if !ok {
t.Fatalf("wire URL %q is not on the public origin %q", wire, testCoverBaseURL)
}
if got := cf.calls.Load(); got != 1 {
t.Fatalf("fetcher called %d times, want 1 — the second read must come from the cache", got)
rr := getCover(t, srv, path, nil)
if rr.Code != http.StatusOK {
t.Fatalf("status = %d, want 200 without any credential", rr.Code)
}
if got := cf.lastID.Load(); got != testCoverID {
t.Fatalf("fetched image id = %v, want %s", got, testCoverID)
if got := rr.Body.String(); got != "\x00webp-bytes" {
t.Fatalf("body = %q, want the stored bytes", got)
}
if got := rr.Header().Get("Content-Type"); got != "image/webp" {
t.Fatalf("Content-Type = %q, want the stored one", got)
}
// Content-addressed bytes never change, so a client that has them must
// never need to ask again.
if got := rr.Header().Get("Cache-Control"); !strings.Contains(got, "immutable") {
t.Fatalf("Cache-Control = %q, want an immutable cache directive", got)
}
}
// The proxy reaches a headless browser, so it is not open to the internet.
func TestKaganeCoverRequiresSession(t *testing.T) {
cf := &fakeCovers{body: []byte("x"), contentType: "image/webp"}
cfg := testConfig()
cfg.Covers = cf
srv, _ := newWebTestServer(t, cfg)
rr := getCover(t, srv, "/img/kagane/"+testCoverID, nil)
if rr.Code != http.StatusUnauthorized {
t.Fatalf("status = %d, want 401", rr.Code)
func TestPublicCoverRejectsUnknownAddress(t *testing.T) {
srv, _ := newWebTestServer(t, testConfig())
cases := map[string]string{
"unknown": "/covers/" + store.CoverAddress("https://cdn.example/never-stored.jpg"),
"malformed": "/covers/not-an-address",
"traversal": "/covers/../../etc/passwd",
"empty": "/covers/",
}
if got := cf.calls.Load(); got != 0 {
t.Fatalf("fetcher called %d times for an unauthenticated request, want 0", got)
}
}
func TestKaganeCoverRejectsBadInput(t *testing.T) {
cases := []struct {
name string
id string
fetch *fakeCovers
}{
{
"an id that is not a uuid never reaches the browser",
"solo-leveling",
&fakeCovers{body: []byte("x"), contentType: "image/webp"},
},
{
"a uuid-shaped id with a trailing segment is rejected whole",
testCoverID + "x",
&fakeCovers{body: []byte("x"), contentType: "image/webp"},
},
{
"a challenged fetch is a missing cover",
testCoverID,
&fakeCovers{err: errors.New("challenge held")},
},
{
"a content type outside the image set is not echoed back",
testCoverID,
&fakeCovers{body: []byte("<script>"), contentType: "text/html"},
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
cfg := testConfig()
cfg.Covers = tc.fetch
srv, st := newWebTestServer(t, cfg)
rr := getCover(t, srv, "/img/kagane/"+tc.id, sessionCookie(t, st))
if rr.Code != http.StatusNotFound {
t.Fatalf("status = %d, want 404", rr.Code)
for name, path := range cases {
t.Run(name, func(t *testing.T) {
if rr := getCover(t, srv, path, nil); rr.Code == http.StatusOK {
t.Fatalf("%s: status = 200, want anything but a served body", path)
}
})
}
}
// ServeMux path-cleans a traversal into a redirect before the handler runs, so
// the guarantee to pin down is that no request shaped like one ever gets bytes.
func TestKaganeCoverTraversalServesNothing(t *testing.T) {
cf := &fakeCovers{body: []byte("secret"), contentType: "image/webp"}
cfg := testConfig()
cfg.Covers = cf
srv, st := newWebTestServer(t, cfg)
rr := getCover(t, srv, "/img/kagane/../../etc/passwd", sessionCookie(t, st))
// A content type outside the image set is never echoed back. The old kagane
// proxy could fetch text/html from a challenged fetch and had to refuse it;
// the general route's only input is the store, and the store refuses to
// record anything that is not an image — but the guarantee is pinned at the
// serving boundary, not the write gate, so a poisoned row (migrated data, a
// writer that skips the gate) is also never served.
func TestPublicCoverNeverEchoesNonImage(t *testing.T) {
const sourceURL = "https://cdn.example/cover"
st, dsn := newTestStoreURL(t)
// The write gate refuses non-image content types outright.
if err := st.PutCover(sourceURL, []byte("<script>"), "text/html"); err == nil {
t.Fatal("PutCover accepted a non-image content type")
}
// A legitimate row, then the content type flipped behind the store's back:
// the bytes exist at the address, so only the type is hostile.
address := store.CoverAddress(sourceURL)
if err := st.SetSeriesCover("asura", "solo", sourceURL, []byte("<script>"), "image/png"); err != nil {
t.Fatalf("seed row: %v", err)
}
db, err := sql.Open("pgx", dsn)
if err != nil {
t.Fatalf("open %s: %v", dsn, err)
}
defer db.Close()
if _, err := db.Exec(`UPDATE covers SET content_type = 'text/html' WHERE address = $1`, address); err != nil {
t.Fatalf("poison row: %v", err)
}
rr := getCover(t, newRouter(st, testConfig()), "/covers/"+address, nil)
if rr.Code == http.StatusOK {
t.Fatalf("status = 200, want anything but a served body")
}
if got := cf.calls.Load(); got != 0 {
t.Fatalf("fetcher called %d times for a traversal, want 0", got)
t.Fatalf("status = 200, want a refusal for a non-image row (body %q)", rr.Body.String())
}
}
// Without BROWSER_WS_URL there is no fetcher, and the endpoint must answer
// rather than reach for a nil one.
func TestKaganeCoverWithoutFetcher(t *testing.T) {
// The whole point of acquiring bytes is that the UI shows them: the card's
// <img> must carry the public address, not a third-party URL and not a
// placeholder.
func TestListRendersAcquiredCover(t *testing.T) {
const sourceURL = "https://static.comix.to/039d/i/1/34/6a6742bf15736@280.jpg"
srv, st := newWebTestServer(t, testConfig())
rr := getCover(t, srv, "/img/kagane/"+testCoverID, sessionCookie(t, st))
if rr.Code != http.StatusNotFound {
t.Fatalf("status = %d, want 404", rr.Code)
if _, err := st.Upsert(st.OwnerID(), store.Bookmark{
Key: "comix:n8we", Site: "comix", SeriesID: "n8we", Title: "Dungeons and Crayons",
SeriesURL: "https://comix.to/title/n8we", UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
if err := st.SetSeriesCover("comix", "n8we", sourceURL, []byte("\xff\xd8jpeg"), "image/jpeg"); err != nil {
t.Fatalf("SetSeriesCover: %v", err)
}
req := httptest.NewRequest(http.MethodGet, "/ui/list", nil)
req.AddCookie(sessionCookie(t, st))
rr := httptest.NewRecorder()
srv.ServeHTTP(rr, req)
if rr.Code != http.StatusOK {
t.Fatalf("status = %d, want 200", rr.Code)
}
want := `src="` + testCoverBaseURL + "/covers/" + store.CoverAddress(sourceURL) + `"`
if !strings.Contains(rr.Body.String(), want) {
t.Fatalf("rendered list does not contain %s", want)
}
}
+40
View File
@@ -50,6 +50,13 @@ func (h *Handler) Put(w http.ResponseWriter, r *http.Request) {
http.Error(w, "invalid JSON body", http.StatusBadRequest)
return
}
// A body may carry a cover, and it is discarded here rather than
// rejected: an older installed userscript may still send one, and
// ADR-0004's compatibility argument depends on those scripts continuing
// to work. The Cover is acquired server-side (ADR-0007), so the field is
// permanently inert - not pending removal, and not a value any later code
// should start reading.
b.Cover = ""
// Path key is authoritative; derive site/series_id from it when the body
// omits them so the stored row is always self-consistent.
@@ -124,3 +131,36 @@ func Healthz(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(http.StatusOK)
_, _ = w.Write([]byte("ok"))
}
// Cover serves stored cover bytes. GET /covers/{address}
//
// Public on purpose: the userscript renders these on Sites the deployment
// does not control, where no credential of ours may be sent, and the address
// is the SHA-256 of a URL the Site already publishes (ADR-0007). An unknown
// address is a 404 rather than an error - "no Cover yet" is a normal state,
// and the clients fall back to their placeholder.
func (h *Handler) Cover(w http.ResponseWriter, r *http.Request) {
body, contentType, ok, err := h.Store.CoverByAddress(r.PathValue("address"))
if err != nil {
log.Printf("cover: %v", err)
http.Error(w, "internal error", http.StatusInternalServerError)
return
}
if !ok {
http.NotFound(w, r)
return
}
// Refuse anything the write gate would not have recorded: a poisoned row
// (migrated data, a writer that skips the gate) must never be echoed back
// as bytes of a type no Cover may have.
if _, ok := store.CoverContentType(contentType); !ok {
log.Printf("cover %s: refusing non-image content type %q", r.PathValue("address"), contentType)
http.NotFound(w, r)
return
}
w.Header().Set("Content-Type", contentType)
// Content-addressed, so the bytes at this URL can never change. Public
// rather than private: no credential gates the route.
w.Header().Set("Cache-Control", "public, max-age=604800, immutable")
_, _ = w.Write(body)
}
+147
View File
@@ -0,0 +1,147 @@
package latest
import (
"context"
"log"
"sync"
"time"
"bookmarkmanager/backend/internal/store"
)
// acquireTimeout bounds one creation-time acquisition end to end: the series
// page plus the cover bytes. Nothing is waiting on it — the Reader's write has
// already returned — so this only stops a stalled Site from holding a
// goroutine and a connection open forever.
const acquireTimeout = 45 * time.Second
// Acquirer gives a Series its Latest Chapter and its Cover the moment the
// first Bookmark creates it, instead of leaving the Reader to wait out the
// poll queue — which is ordered by Reader count, so a Series with one Reader
// sits behind every popular one (ADR-0007).
//
// Both facts come from a single series-page fetch, which is also why no
// client-supplied cover hint is worth accepting: the page has to be fetched
// for the chapter signal regardless, so a hint would save no request while
// adding a client-controlled input to a server-side fetch.
//
// Every failure path is "log and move on". The Bookmark, its progress and its
// Latest Chapter are already committed; a Site that is down or a Cover that
// cannot be produced must not disturb any of them, and the Series is simply
// left blank until the poll's own cover pass (#61) fills it.
type Acquirer struct {
Store *store.Store
// Fetch retrieves the series page over plain TLS. Nil with a nil
// BrowserFetch disables acquisition entirely.
Fetch Fetcher
// BrowserFetch retrieves kagane and novelfull pages through the browser
// sidecar, the only thing that clears their Cloudflare challenge. The
// per-site fallback policy lives in fetcherFor. Nil leaves those Sites
// unacquired when no fallback applies.
BrowserFetch Fetcher
// Covers retrieves the cover bytes. Nil leaves the Cover blank and the
// chapter half working.
Covers CoverBytesFetcher
// BrowserCoverFetch retrieves browser-claimed cover bytes through the
// sidecar. Nil leaves those Covers blank; nothing falls back to a plain
// fetch, which would only ever retrieve a challenge page.
BrowserCoverFetch BrowserCoverFetcher
// Ctx cancels in-flight acquisitions at shutdown. A hook signature has
// nowhere to pass one, so it lives here; nil means context.Background.
Ctx context.Context
inflight sync.WaitGroup
}
// acquireSlots caps how many creation-time fetches run at once. A Reader whose
// userscript bulk-syncs creates many Series at once, and a burst of
// simultaneous requests from one server IP is the traffic shape most likely to
// move that IP's bot score — the same reason the poller staggers its batch.
var acquireSlots = make(chan struct{}, 2)
// Acquire starts one acquisition and returns immediately: a Reader's bookmark
// action may not block on a third-party Site's latency, nor fail with it. It
// is the store's OnSeriesCreated hook, so it only ever runs for a Series no
// Reader had bookmarked before.
func (a *Acquirer) Acquire(sr store.Series) {
a.inflight.Add(1)
go func() {
defer a.inflight.Done()
defer func() {
if r := recover(); r != nil {
log.Printf("acquire %q: recovered from panic: %v", sr.Key(), r)
}
}()
parent := a.Ctx
if parent == nil {
parent = context.Background()
}
select {
case acquireSlots <- struct{}{}:
defer func() { <-acquireSlots }()
case <-parent.Done():
return
}
ctx, cancel := context.WithTimeout(parent, acquireTimeout)
defer cancel()
a.acquire(ctx, sr)
}()
}
// Wait blocks until every started acquisition has finished. It exists for
// tests: an asynchronous side effect is otherwise unobservable without
// polling for it.
func (a *Acquirer) Wait() { a.inflight.Wait() }
func (a *Acquirer) acquire(ctx context.Context, sr store.Series) {
if a.Fetch == nil && a.BrowserFetch == nil {
return
}
// series_url arrives in a client-supplied PUT body, so the same gate the
// poller uses applies here — without it a token-holder chooses what the
// server fetches from its own network position.
if !fetchableSeriesURL(sr.Site, sr.SeriesURL) {
log.Printf("acquire %q: not fetchable: site=%q url=%q", sr.Key(), sr.Site, sr.SeriesURL)
return
}
f := fetcherFor(sr.Site, a.BrowserFetch, a.Fetch)
if f == nil {
log.Printf("acquire %q: no fetcher for site %q", sr.Key(), sr.Site)
return
}
body, status, err := f.Get(ctx, sr.SeriesURL)
if err != nil {
log.Printf("acquire %q: fetch %s: %v", sr.Key(), sr.SeriesURL, err)
return
}
if status != 200 {
log.Printf("acquire %q: fetch %s: status %d", sr.Key(), sr.SeriesURL, status)
return
}
// This page just served the same purpose a poll tick would have; without
// the stamp the row stays due and the poller refetches it immediately.
if err := a.Store.MarkLatestChecked(sr.Site, sr.SeriesID, time.Now().UnixMilli()); err != nil {
log.Printf("acquire %q: mark checked: %v", sr.Key(), err)
}
if latest, ok := latestChapterFrom(sr.Site, sr.SeriesURL, body); ok {
if err := a.Store.SetLatestChapter(sr.Site, sr.SeriesID, latest.Label, latest.Num); err != nil {
log.Printf("acquire %q: set latest chapter: %v", sr.Key(), err)
}
}
cover, ok := coverFrom(sr.Site, sr.SeriesURL, body)
if !ok {
return
}
bytes, contentType, err := fetchCoverBytes(ctx, cover, a.BrowserCoverFetch, a.Covers)
if err != nil {
log.Printf("acquire %q: fetch cover %s: %v", sr.Key(), cover, err)
return
}
if err := a.Store.SetSeriesCover(sr.Site, sr.SeriesID, cover, bytes, contentType); err != nil {
log.Printf("acquire %q: persist cover: %v", sr.Key(), err)
}
}
+430
View File
@@ -0,0 +1,430 @@
package latest
import (
"context"
"errors"
"testing"
"time"
"bookmarkmanager/backend/internal/store"
)
// The series page carries both facts, which is the whole argument for taking
// them from one fetch.
const asuraSeriesAndCoverFixture = asuraSeriesFixture + asuraCoverFixture
const (
acquireKey = "asura:chronicles-of-the-demon-faction-f886a8af"
acquireSeriesID = "chronicles-of-the-demon-faction-f886a8af"
acquireSeriesURL = "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af"
acquireCoverURL = "https://cdn.asurascans.com/asura-images/covers/chronicles-of-the-demon-faction.d4dcb8.webp"
)
const (
kaganeKey = "kagane:019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
kaganeSeriesID = "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
kaganeSeriesURL = "https://kagane.to/series/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
kaganeImageID = "019fe11a-84c3-7fc3-a84b-88787374b617"
kaganeCoverSrc = "https://kagane.to/api/v2/image/" + kaganeImageID + "/compressed"
)
// kagane's browser-fetched body is one JSON object carrying both the chapter
// list (series_books) and the cover image ids (series_covers), so the single
// acquisition fetch yields both facts.
const kaganeSeriesAndCoverFixture = `{"series_id":"019f84bc-9ba0-7ed9-86f5-8b905ec7c28b",` +
`"series_books":[{"book_id":"b","title":"Episode 41","chapter_no":"41","sort_no":41}],` +
`"series_covers":[{"cover_id":"019fe11a-84d1-714b-9cf4-2827f277f3c0","language":"en",` +
`"image_id":"019fe11a-84c3-7fc3-a84b-88787374b617"}]}`
const (
novelfullKey = "novelfull:reverend-insanity"
novelfullSeriesID = "reverend-insanity"
novelfullSeriesURI = "https://novelfull.com/reverend-insanity.html"
novelfullCoverURL = "https://novelfull.com/uploads/webp/novel/reverend-insanity-82661d911a.webp"
)
// newAcquirer wires an acquirer onto the store's creation hook, which is how
// main wires it: the write path is what starts an acquisition.
func newAcquirer(s *store.Store, page *fakeFetcher, covers *fakeBytesCoverFetcher) *Acquirer {
a := &Acquirer{Store: s, Fetch: page, Covers: covers}
s.OnSeriesCreated = a.Acquire
return a
}
func bookmarkNewSeries(t *testing.T, s *store.Store, seriesURL string) store.Bookmark {
t.Helper()
stored, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: acquireKey, Site: "asura", SeriesID: acquireSeriesID,
Title: "Chronicles of the Demon Faction", SeriesURL: seriesURL,
Cover: "https://evil.example/client-supplied.jpg", UpdatedAt: 1000,
})
if err != nil {
t.Fatalf("Upsert: %v", err)
}
return stored
}
func readBookmark(t *testing.T, s *store.Store, key string) store.Bookmark {
t.Helper()
b, ok, err := s.Get(s.OwnerID(), key)
if err != nil || !ok {
t.Fatalf("Get %q = %v, %v", key, ok, err)
}
return b
}
func bookmarkNewKaganeSeries(t *testing.T, s *store.Store) store.Bookmark {
t.Helper()
stored, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: kaganeKey, Site: "kagane", SeriesID: kaganeSeriesID,
Title: "Infinite Decryption", SeriesURL: kaganeSeriesURL, UpdatedAt: 1000,
})
if err != nil {
t.Fatalf("Upsert: %v", err)
}
return stored
}
func bookmarkNewNovelfullSeries(t *testing.T, s *store.Store) store.Bookmark {
t.Helper()
stored, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: novelfullKey, Site: "novelfull", SeriesID: novelfullSeriesID,
Title: "Reverend Insanity", SeriesURL: novelfullSeriesURI, UpdatedAt: 1000,
})
if err != nil {
t.Fatalf("Upsert: %v", err)
}
return stored
}
// The reported bug: a Reader bookmarks a Series nobody holds and expects the
// Cover, not a broken image. Both facts come from the one series-page fetch.
func TestAcquireFillsChapterAndCoverFromOneFetch(t *testing.T) {
s, _ := newTestStore(t)
page := &fakeFetcher{body: asuraSeriesAndCoverFixture, status: 200}
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"}
acq := newAcquirer(s, page, covers)
// The write itself must not carry the acquisition: it returns before the
// Cover exists, and the field is empty until the bytes land.
stored := bookmarkNewSeries(t, s, acquireSeriesURL)
if stored.Cover != "" {
t.Fatalf("Cover on the creating write = %q, want empty", stored.Cover)
}
acq.Wait()
if got := page.callCount(); got != 1 {
t.Fatalf("series page fetches = %d, want exactly 1", got)
}
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetches = %d, want 1", got)
}
got := readBookmark(t, s, acquireKey)
if got.LatestChapterNum == nil || *got.LatestChapterNum != 181 {
t.Fatalf("LatestChapterNum = %v, want 181", got.LatestChapterNum)
}
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(acquireCoverURL); got.Cover != want {
t.Fatalf("Cover = %q, want the absolute address %q", got.Cover, want)
}
body, contentType, ok, err := s.CoverByAddress(store.CoverAddress(acquireCoverURL))
if err != nil || !ok {
t.Fatalf("CoverByAddress = %v, %v", ok, err)
}
if string(body) != "cover-bytes" || contentType != "image/jpeg" {
t.Fatalf("stored cover = (%q, %q), want the fetched bytes", body, contentType)
}
}
// A Series that already exists is not re-acquired: no fetch, and the Cover it
// already has is left alone.
func TestAcquireSkipsAnExistingSeries(t *testing.T) {
s, _ := newTestStore(t)
page := &fakeFetcher{body: asuraSeriesAndCoverFixture, status: 200}
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"}
acq := newAcquirer(s, page, covers)
bookmarkNewSeries(t, s, acquireSeriesURL)
acq.Wait()
bookmarkNewSeries(t, s, acquireSeriesURL)
acq.Wait()
if got := page.callCount(); got != 1 {
t.Fatalf("series page fetches = %d, want 1 — an existing series is not re-acquired", got)
}
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetches = %d, want 1", got)
}
got := readBookmark(t, s, acquireKey)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(acquireCoverURL); got.Cover != want {
t.Fatalf("Cover = %q, want the acquired one %q", got.Cover, want)
}
}
// A Site that is down costs the Cover and nothing else.
func TestAcquireFailureLeavesTheBookmarkIntact(t *testing.T) {
cases := []struct {
name string
page *fakeFetcher
covers *fakeBytesCoverFetcher
// wantLatest is the chapter that still lands; 0 means none did.
wantLatest float64
}{
{
"the series page is unreachable",
&fakeFetcher{err: errors.New("connection reset")},
&fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"},
0,
},
{
"the series page answers with a challenge",
&fakeFetcher{body: challengeFixture, status: 200},
&fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"},
0,
},
{
"only the cover bytes fail",
&fakeFetcher{body: asuraSeriesAndCoverFixture, status: 200},
&fakeBytesCoverFetcher{err: errors.New("403")},
181,
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
s, _ := newTestStore(t)
acq := newAcquirer(s, tc.page, tc.covers)
stored := bookmarkNewSeries(t, s, acquireSeriesURL)
acq.Wait()
got := readBookmark(t, s, acquireKey)
if got.Cover != "" {
t.Fatalf("Cover = %q, want empty rather than an address that 404s", got.Cover)
}
if got.Title != stored.Title || got.UpdatedAt != stored.UpdatedAt {
t.Fatalf("bookmark = %+v, want it untouched by the failed acquisition", got)
}
if tc.wantLatest == 0 {
if got.LatestChapterNum != nil {
t.Fatalf("LatestChapterNum = %v, want none captured", *got.LatestChapterNum)
}
return
}
if got.LatestChapterNum == nil || *got.LatestChapterNum != tc.wantLatest {
t.Fatalf("LatestChapterNum = %v, want %v", got.LatestChapterNum, tc.wantLatest)
}
})
}
}
// series_url arrives in a client-supplied body, so the acquisition reuses the
// poller's gate rather than deriving a second one: a non-https scheme, a
// site the parsers do not know, or a host pinned to another site is refused
// before the server spends a request from its own network position.
func TestAcquireRefusesAnUnfetchableSeriesURL(t *testing.T) {
for _, seriesURL := range []string{
"http://asurascans.com/comics/x",
"file:///etc/passwd",
"",
} {
t.Run(seriesURL, func(t *testing.T) {
s, _ := newTestStore(t)
page := &fakeFetcher{body: asuraSeriesAndCoverFixture, status: 200}
acq := newAcquirer(s, page, &fakeBytesCoverFetcher{})
bookmarkNewSeries(t, s, seriesURL)
acq.Wait()
if got := page.callCount(); got != 0 {
t.Fatalf("fetches for %q = %d, want 0", seriesURL, got)
}
})
}
}
// blockingFetcher stands in for a Site that never answers, so a synchronous
// acquisition would be visible as a stalled write rather than a slow one.
type blockingFetcher struct {
release <-chan struct{}
body string
}
func (f *blockingFetcher) Get(ctx context.Context, _ string) (string, int, error) {
select {
case <-f.release:
return f.body, 200, nil
case <-ctx.Done():
return "", 0, ctx.Err()
}
}
// The Reader's write may not wait on a third-party Site: with the acquisition
// wedged on an unanswering page, the PUT still returns.
func TestAcquireDoesNotBlockTheWrite(t *testing.T) {
s, _ := newTestStore(t)
release := make(chan struct{})
acq := &Acquirer{Store: s, Fetch: &blockingFetcher{release: release, body: asuraSeriesAndCoverFixture}}
s.OnSeriesCreated = acq.Acquire
upserted := make(chan error, 1)
go func() {
_, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: acquireKey, Site: "asura", SeriesID: acquireSeriesID,
Title: "Chronicles of the Demon Faction", SeriesURL: acquireSeriesURL, UpdatedAt: 1000,
})
upserted <- err
}()
select {
case err := <-upserted:
if err != nil {
t.Fatalf("Upsert: %v", err)
}
case <-time.After(10 * time.Second):
t.Fatal("the creating write blocked on the acquisition")
}
close(release)
acq.Wait()
}
// The second symptom of #47: a kagane Series bookmarked from a chapter page
// gets its Cover at creation, with the bytes fetched through the browser
// sidecar — the only path that clears the challenge — into the
// content-addressed store.
func TestAcquireKaganeCoverThroughBrowser(t *testing.T) {
s, _ := newTestStore(t)
tlsPage := &fakeFetcher{body: "", status: 403}
browserPage := &fakeFetcher{body: kaganeSeriesAndCoverFixture, status: 200}
covers := &fakeCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
acq := &Acquirer{
Store: s, Fetch: tlsPage, BrowserFetch: browserPage,
BrowserCoverFetch: covers, Covers: &fakeBytesCoverFetcher{},
}
s.OnSeriesCreated = acq.Acquire
bookmarkNewKaganeSeries(t, s)
acq.Wait()
if got := tlsPage.callCount(); got != 0 {
t.Fatalf("plain-TLS page fetches = %d, want 0 — kagane pages are browser-only", got)
}
if got := browserPage.callCount(); got != 1 {
t.Fatalf("browser page fetches = %d, want 1", got)
}
if got := covers.callCount(); got != 1 {
t.Fatalf("browser cover fetches = %d, want 1", got)
}
if got := covers.calls[0]; got != kaganeCoverSrc {
t.Fatalf("browser cover fetched URL %q, want %q", got, kaganeCoverSrc)
}
got := readBookmark(t, s, kaganeKey)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(kaganeCoverSrc); got.Cover != want {
t.Fatalf("Cover = %q, want the content-addressed URL %q", got.Cover, want)
}
body, contentType, ok, err := s.CoverByAddress(store.CoverAddress(kaganeCoverSrc))
if err != nil || !ok {
t.Fatalf("CoverByAddress = %v, %v", ok, err)
}
if string(body) != "cover-bytes" || contentType != "image/webp" {
t.Fatalf("stored cover = (%q, %q), want the browser-fetched bytes", body, contentType)
}
}
// novelfull needs the browser only for its HTML: the cover URL comes out of
// the browser-fetched page, but the bytes go over plain TLS through the
// ordinary gated fetcher, never through the browser (issue #62).
func TestAcquireNovelfullCoverOverPlainTLS(t *testing.T) {
s, _ := newTestStore(t)
browserPage := &fakeFetcher{body: novelfullSeriesFixture + novelfullCoverFixture, status: 200}
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
acq := &Acquirer{
Store: s, Fetch: &fakeFetcher{body: "", status: 403},
BrowserFetch: browserPage, Covers: covers,
}
s.OnSeriesCreated = acq.Acquire
bookmarkNewNovelfullSeries(t, s)
acq.Wait()
if got := browserPage.callCount(); got != 1 {
t.Fatalf("browser page fetches = %d, want 1", got)
}
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetches = %d, want 1 — novelfull bytes never touch the browser", got)
}
if got := covers.calls[0]; got != novelfullCoverURL {
t.Fatalf("cover fetched from %q, want %q", got, novelfullCoverURL)
}
got := readBookmark(t, s, novelfullKey)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(novelfullCoverURL); got.Cover != want {
t.Fatalf("Cover = %q, want %q", got.Cover, want)
}
}
// With no browser sidecar configured, kagane is simply not acquired: no
// request is spent on a page that could only ever answer with a challenge,
// and nothing falls back to a plain fetch.
func TestAcquireKaganeSkippedWithoutBrowser(t *testing.T) {
s, _ := newTestStore(t)
tlsPage := &fakeFetcher{body: kaganeSeriesAndCoverFixture, status: 200}
acq := &Acquirer{
Store: s, Fetch: tlsPage,
Covers: &fakeBytesCoverFetcher{body: []byte("x"), contentType: "image/webp"},
}
s.OnSeriesCreated = acq.Acquire
bookmarkNewKaganeSeries(t, s)
acq.Wait()
if got := tlsPage.callCount(); got != 0 {
t.Fatalf("plain-TLS fetches for kagane = %d, want 0", got)
}
if got := readBookmark(t, s, kaganeKey); got.Cover != "" {
t.Fatalf("Cover = %q, want empty without a browser", got.Cover)
}
}
// The byte half of "nothing falls back to a plain fetch": with a browser for
// the page but none for the bytes, a kagane Cover stays absent and the TLS
// cover fetcher is never consulted.
func TestAcquireKaganeBytesNeverFallBackToPlainTLS(t *testing.T) {
s, _ := newTestStore(t)
browserPage := &fakeFetcher{body: kaganeSeriesAndCoverFixture, status: 200}
tlsCovers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
acq := &Acquirer{
Store: s, Fetch: &fakeFetcher{body: "", status: 403},
BrowserFetch: browserPage, Covers: tlsCovers,
}
s.OnSeriesCreated = acq.Acquire
bookmarkNewKaganeSeries(t, s)
acq.Wait()
if got := tlsCovers.callCount(); got != 0 {
t.Fatalf("plain-TLS cover fetches = %d, want 0 — kagane bytes are browser-only", got)
}
if got := readBookmark(t, s, kaganeKey); got.Cover != "" {
t.Fatalf("Cover = %q, want empty without a browser cover fetcher", got.Cover)
}
}
// novelfull's no-browser degradation differs from kagane's: only its HTML
// needs the sidecar, so when the page body is available — the challenge is a
// live time-varying fact that sometimes answers a plain request — the Cover
// still lands, bytes over plain TLS.
func TestAcquireNovelfullCoverWithoutBrowser(t *testing.T) {
s, _ := newTestStore(t)
tlsPage := &fakeFetcher{body: novelfullSeriesFixture + novelfullCoverFixture, status: 200}
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
acq := &Acquirer{Store: s, Fetch: tlsPage, Covers: covers}
s.OnSeriesCreated = acq.Acquire
bookmarkNewNovelfullSeries(t, s)
acq.Wait()
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetches = %d, want 1", got)
}
got := readBookmark(t, s, novelfullKey)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(novelfullCoverURL); got.Cover != want {
t.Fatalf("Cover = %q, want %q", got.Cover, want)
}
}
+66 -26
View File
@@ -24,12 +24,6 @@ const challengeTimeout = 45 * time.Second
var kaganeSeriesRe = regexp.MustCompile(`^/series/([0-9a-f-]{36})/?$`)
// kaganeImageIDRe pins the only path segment Image interpolates into an
// outbound URL. The id arrives from a stored cover URL, which a client
// supplied, so it is matched rather than trusted: a headless browser is a
// strong SSRF primitive.
var kaganeImageIDRe = regexp.MustCompile(`^[0-9a-f-]{36}$`)
// BrowserFetcher retrieves pages through a remote headless Chrome over the
// DevTools Protocol.
//
@@ -53,19 +47,22 @@ var kaganeImageIDRe = regexp.MustCompile(`^[0-9a-f-]{36}$`)
type BrowserFetcher struct {
allocCtx context.Context
cancel context.CancelFunc
// One page at a time: caps the sidecar's memory and keeps series from
// sharing page state.
// One page at a time: caps the browser's memory — it runs under a hard
// cgroup cap on a shared machine — and keeps series from sharing page state.
mu sync.Mutex
}
var _ Fetcher = (*BrowserFetcher)(nil)
// NewBrowserFetcher connects to a headless-shell over CDP. wsURL must name the
// sidecar by IP, e.g. ws://172.28.0.10:9222 — not by Docker DNS name. Chrome's
// DevTools HTTP handler 500s any /json/version request whose Host header
// isn't an IP or "localhost" (confirmed 2026-08-03 against
// chromedp/headless-shell:stable), so the compose network pins the sidecar's
// address for this to resolve at all.
// NewBrowserFetcher connects to a Chrome over CDP. The browser is not a
// sidecar: it runs on a separate machine and is reached over the tailnet
// (ADR-0006), so wsURL is that machine's tailnet address, e.g.
// ws://100.64.0.5:9222.
//
// It must be an IP, never a hostname — not MagicDNS, not a Docker service
// name. Chrome's DevTools HTTP handler 500s any /json/version request whose
// Host header isn't an IP or "localhost" (confirmed 2026-08-03), so a name
// fails at discovery and surfaces as a dead site rather than a bad URL.
//
// Do not add chromedp.NoModifyURL here: that option skips the /json/version
// discovery request entirely and dials wsURL as if it were already the full
@@ -73,8 +70,10 @@ var _ Fetcher = (*BrowserFetcher)(nil)
// /devtools/browser/<uuid>, a path chosen fresh at every Chrome start — dialing
// the bare host:port 404s. The default (discovery) path works precisely
// because Chrome's /json/version response echoes back the Host header of the
// discovery request in webSocketDebuggerUrl, so as long as wsURL is a
// container-reachable IP, the URL chromedp gets back already points at it.
// discovery request in webSocketDebuggerUrl, so as long as wsURL is an IP this
// process can reach, the URL chromedp gets back already points at it. That is
// also why a Chrome restarted behind a stable endpoint needs no reconnect
// here: the fresh UUID arrives with the next discovery.
func NewBrowserFetcher(wsURL string) (*BrowserFetcher, error) {
if wsURL == "" {
return nil, fmt.Errorf("empty browser websocket url")
@@ -129,12 +128,14 @@ func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int
return body, 200, nil
}
// Image retrieves one kagane cover as raw bytes and its content type.
// Image retrieves one cover's bytes through the browser sidecar, and its
// content type.
//
// It exists because kagane serves covers behind the same challenge as its
// pages *and* with `cross-origin-resource-policy: same-origin`, so an <img> on
// the web UI's origin cannot load one even from a browser that already holds
// the clearance cookie (verified 2026-08-08). Proxying is the only route.
// pages *and* with `cross-origin-resource-policy: same-origin`, so the bytes
// are only reachable from inside a browser that already holds the clearance
// cookie (verified 2026-08-08). Acquisition through the sidecar is the only
// route.
//
// The image URL is navigated to rather than fetched from some other kagane
// page: the challenge only runs on a top-level navigation, and once it clears
@@ -144,12 +145,14 @@ func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int
// The challenge is not solved by the first read: WaitReady("body") is satisfied
// by the interstitial too. run holds the tab open until the in-page fetch
// succeeds, which is what gives the challenge script the seconds it needs.
func (f *BrowserFetcher) Image(ctx context.Context, imageID string) ([]byte, string, error) {
if !kaganeImageIDRe.MatchString(imageID) {
return nil, "", fmt.Errorf("not a kagane image id: %q", imageID)
func (f *BrowserFetcher) Image(ctx context.Context, imageURL string) ([]byte, string, error) {
m := kaganeImageURLRe.FindStringSubmatch(imageURL)
if m == nil {
return nil, "", fmt.Errorf("not a browser-fetchable cover url: %q", imageURL)
}
imageID := m[1]
var dataURL string
err := f.run(ctx, "https://kagane.to/api/v2/image/"+imageID+"/compressed",
err := f.run(ctx, imageURL,
chromedp.Evaluate(`fetch(location.href).then(r => r.ok
? r.blob().then(b => new Promise(res => {
const fr = new FileReader();
@@ -178,6 +181,36 @@ func (f *BrowserFetcher) Image(ctx context.Context, imageID string) ([]byte, str
// the poller answers with a 403 and its ordinary cooldown.
var errChallengeHeld = errors.New("challenge held")
// errBrowserInterrupted distinguishes a remote Chrome restart from the
// caller's own deadline. chromedp reports both as context.Canceled.
var errBrowserInterrupted = errors.New("browser interrupted")
func classifyBrowserError(ctx context.Context, browserLost bool, err error) error {
if err == nil || ctx.Err() != nil {
return err
}
if !browserLost {
return err
}
if !errors.Is(err, context.Canceled) {
return err
}
return fmt.Errorf("%w: %w", errBrowserInterrupted, err)
}
func browserConnectionLost(ctx context.Context) bool {
c := chromedp.FromContext(ctx)
if c == nil || c.Browser == nil {
return true
}
select {
case <-c.Browser.LostConnection:
return true
default:
return false
}
}
// challengePollInterval paces re-reads while a challenge solves itself.
const challengePollInterval = 2 * time.Second
@@ -203,6 +236,7 @@ func (f *BrowserFetcher) run(ctx context.Context, target string, read chromedp.A
f.mu.Lock()
defer f.mu.Unlock()
callerCtx := ctx
ctx, cancel := context.WithTimeout(ctx, challengeTimeout)
defer cancel()
tabCtx, cancelTab := chromedp.NewContext(f.allocCtx)
@@ -219,20 +253,26 @@ func (f *BrowserFetcher) run(ctx context.Context, target string, read chromedp.A
chromedp.Navigate(target),
chromedp.WaitReady("body", chromedp.ByQuery),
); err != nil {
return err
return classifyBrowserError(callerCtx, browserConnectionLost(tabCtx), err)
}
var lastErr error
for {
// The challenge reloads the page when it passes, which tears down the
// execution context mid-read. That is a retry, not a failure.
if err := chromedp.Run(tabCtx, read); err != nil {
err = classifyBrowserError(callerCtx, browserConnectionLost(tabCtx), err)
if errors.Is(err, errBrowserInterrupted) {
return err
}
lastErr = err
} else if done() {
return nil
}
select {
case <-ctx.Done():
if err := callerCtx.Err(); err != nil {
return err
}
if lastErr != nil {
return fmt.Errorf("%w (last read: %v)", errChallengeHeld, lastErr)
}
+19 -1
View File
@@ -1,6 +1,10 @@
package latest
import "testing"
import (
"context"
"errors"
"testing"
)
func TestKaganeAPIURL(t *testing.T) {
const uuid = "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
@@ -57,3 +61,17 @@ func TestNovelfullSeriesURL(t *testing.T) {
})
}
}
func TestClassifyBrowserInterruption(t *testing.T) {
if err := classifyBrowserError(context.Background(), true, context.Canceled); !errors.Is(err, errBrowserInterrupted) {
t.Fatalf("classifyBrowserError(context.Canceled) = %v, want browser interruption", err)
}
if err := classifyBrowserError(context.Background(), false, context.Canceled); errors.Is(err, errBrowserInterrupted) {
t.Fatalf("ordinary cancellation misclassified as browser interruption: %v", err)
}
caller, cancel := context.WithCancel(context.Background())
cancel()
if err := classifyBrowserError(caller, true, context.Canceled); errors.Is(err, errBrowserInterrupted) {
t.Fatalf("caller cancellation misclassified as browser interruption: %v", err)
}
}
+204
View File
@@ -0,0 +1,204 @@
package latest
import (
"context"
"errors"
"fmt"
"io"
"mime"
"net"
"net/http"
"net/netip"
"net/url"
"strings"
"time"
"bookmarkmanager/backend/internal/store"
)
// CoverBytesFetcher retrieves one cover from its source URL. The caller owns
// persistence; this seam keeps network policy independent from the store.
type CoverBytesFetcher interface {
Fetch(ctx context.Context, sourceURL string) (body []byte, contentType string, err error)
}
// fetchCoverBytes routes a cover's byte retrieval by URL shape, not by Site
// name: the browser fetcher's module claims the addresses only it can fetch
// (kagane's image route answers a plain fetch with a challenge and
// `cross-origin-resource-policy: same-origin`), and everything else goes over
// plain TLS. Missing fetchers degrade to an error the caller logs, never a
// fallback onto a path that cannot succeed. One routing rule for the poll and
// the acquirer, so the two cannot drift apart.
func fetchCoverBytes(ctx context.Context, cover string, browser BrowserCoverFetcher, tls CoverBytesFetcher) ([]byte, string, error) {
if browserOnlyCoverURL(cover) {
if browser == nil {
return nil, "", errors.New("no cover fetcher")
}
return browser.Image(ctx, cover)
}
if tls == nil {
return nil, "", errors.New("no cover fetcher")
}
return tls.Fetch(ctx, cover)
}
// CoverResolver resolves a host before any connection is attempted. Tests
// inject it to exercise hostile DNS results without touching the live network.
type CoverResolver func(context.Context, string) ([]netip.Addr, error)
// TLSCoverFetcher retrieves image bytes with the standard HTTPS client. Unlike
// TLSFetcher, it does not need a browser fingerprint: cover hosts are public
// CDNs and the response is accepted only after the destination gate passes.
type TLSCoverFetcher struct {
client *http.Client
resolve CoverResolver
}
var _ CoverBytesFetcher = (*TLSCoverFetcher)(nil)
const coverRequestTimeout = 30 * time.Second
var carrierGradeNAT = netip.MustParsePrefix("100.64.0.0/10")
// NewCoverFetcher builds the production cover client with the real resolver.
func NewCoverFetcher() *TLSCoverFetcher {
return NewCoverFetcherWithResolver(nil)
}
// NewCoverFetcherWithResolver builds a cover client using resolve, or the real
// system resolver when resolve is nil.
func NewCoverFetcherWithResolver(resolve CoverResolver) *TLSCoverFetcher {
if resolve == nil {
resolve = defaultCoverResolver
}
return newCoverFetcher(newCoverHTTPClient(resolve), resolve)
}
func newCoverFetcher(client *http.Client, resolve CoverResolver) *TLSCoverFetcher {
f := &TLSCoverFetcher{client: client, resolve: resolve}
client.CheckRedirect = func(req *http.Request, _ []*http.Request) error {
if err := f.validateURL(req.Context(), req.URL); err != nil {
return fmt.Errorf("redirect destination: %w", err)
}
return nil
}
return f
}
func defaultCoverResolver(ctx context.Context, host string) ([]netip.Addr, error) {
return net.DefaultResolver.LookupNetIP(ctx, "ip", host)
}
func newCoverHTTPClient(resolve CoverResolver) *http.Client {
base, ok := http.DefaultTransport.(*http.Transport)
if !ok {
base = &http.Transport{}
}
transport := base.Clone()
// A proxy would make the dial target the proxy rather than the cover host,
// defeating destination classification. Cover fetching is direct by design.
transport.Proxy = nil
dialer := &net.Dialer{}
transport.DialContext = func(ctx context.Context, network, address string) (net.Conn, error) {
host, port, err := net.SplitHostPort(address)
if err != nil {
return nil, fmt.Errorf("split cover address %q: %w", address, err)
}
addrs, err := resolveCoverHost(ctx, host, resolve)
if err != nil {
return nil, err
}
for _, addr := range addrs {
if !publicCoverAddress(addr) {
return nil, fmt.Errorf("cover host resolves to refused address %s", addr)
}
conn, err := dialer.DialContext(ctx, network, net.JoinHostPort(addr.String(), port))
if err == nil {
return conn, nil
}
}
return nil, fmt.Errorf("cover host %q has no reachable address", host)
}
return &http.Client{Transport: transport, Timeout: coverRequestTimeout}
}
func (f *TLSCoverFetcher) Fetch(ctx context.Context, sourceURL string) ([]byte, string, error) {
u, err := url.Parse(sourceURL)
if err != nil {
return nil, "", fmt.Errorf("parse cover URL: %w", err)
}
if err := f.validateURL(ctx, u); err != nil {
return nil, "", err
}
req, err := http.NewRequestWithContext(ctx, http.MethodGet, u.String(), nil)
if err != nil {
return nil, "", fmt.Errorf("build cover request: %w", err)
}
resp, err := f.client.Do(req)
if err != nil {
return nil, "", fmt.Errorf("fetch cover: %w", err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return nil, "", fmt.Errorf("fetch cover: status %d", resp.StatusCode)
}
raw, _, err := mime.ParseMediaType(resp.Header.Get("Content-Type"))
if err != nil {
return nil, "", fmt.Errorf("fetch cover: unsupported content type %q", resp.Header.Get("Content-Type"))
}
contentType, ok := store.CoverContentType(raw)
if !ok {
return nil, "", fmt.Errorf("fetch cover: unsupported content type %q", raw)
}
if resp.ContentLength > maxBodyBytes {
return nil, "", fmt.Errorf("fetch cover: response exceeds %d bytes", maxBodyBytes)
}
body, err := io.ReadAll(io.LimitReader(resp.Body, maxBodyBytes+1))
if err != nil {
return nil, "", fmt.Errorf("read cover: %w", err)
}
if len(body) > maxBodyBytes {
return nil, "", fmt.Errorf("fetch cover: response exceeds %d bytes", maxBodyBytes)
}
return body, contentType, nil
}
// This gate deliberately differs from fetchableSeriesURL: cover hosts are
// site-independent CDNs, so a Site host allowlist would reject valid covers.
func (f *TLSCoverFetcher) validateURL(ctx context.Context, u *url.URL) error {
if u == nil || u.Scheme != "https" || u.Host == "" || u.User != nil {
return errors.New("cover URL must use HTTPS without credentials")
}
host := u.Hostname()
if host == "" {
return errors.New("cover URL has no host")
}
addrs, err := resolveCoverHost(ctx, host, f.resolve)
if err != nil {
return fmt.Errorf("resolve cover host %q: %w", host, err)
}
if len(addrs) == 0 {
return fmt.Errorf("resolve cover host %q: no addresses", host)
}
for _, addr := range addrs {
if !publicCoverAddress(addr) {
return fmt.Errorf("cover host %q resolves to refused address %s", host, addr)
}
}
return nil
}
func resolveCoverHost(ctx context.Context, host string, resolve CoverResolver) ([]netip.Addr, error) {
if literal, err := netip.ParseAddr(host); err == nil {
return []netip.Addr{literal.Unmap()}, nil
}
return resolve(ctx, strings.TrimSuffix(host, "."))
}
func publicCoverAddress(addr netip.Addr) bool {
addr = addr.Unmap()
return addr.IsValid() && addr.IsGlobalUnicast() &&
!addr.IsLoopback() && !addr.IsPrivate() && !addr.IsLinkLocalUnicast() &&
!carrierGradeNAT.Contains(addr)
}
+226
View File
@@ -0,0 +1,226 @@
package latest
import (
"bytes"
"context"
"crypto/tls"
"io"
"net"
"net/http"
"net/http/httptest"
"net/netip"
"testing"
)
func TestCoverFetcherFetchesPublicHTTPSImage(t *testing.T) {
server := httptest.NewTLSServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.TLS == nil {
t.Fatal("cover request was not made over TLS")
}
w.Header().Set("Content-Type", "image/jpeg")
io.WriteString(w, "cover-bytes")
}))
defer server.Close()
transport := server.Client().Transport.(*http.Transport).Clone()
transport.TLSClientConfig = &tls.Config{InsecureSkipVerify: true} // test server certificate
transport.DialContext = func(ctx context.Context, network, _ string) (net.Conn, error) {
return (&net.Dialer{}).DialContext(ctx, network, server.Listener.Addr().String())
}
client := &http.Client{Transport: transport}
fetcher := newCoverFetcher(client, func(context.Context, string) ([]netip.Addr, error) {
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
})
body, contentType, err := fetcher.Fetch(context.Background(), "https://cdn.example/cover.jpg")
if err != nil {
t.Fatalf("Fetch: %v", err)
}
if string(body) != "cover-bytes" || contentType != "image/jpeg" {
t.Fatalf("Fetch = (%q, %q), want (cover-bytes, image/jpeg)", body, contentType)
}
}
func TestNewCoverFetcherRechecksResolverBeforeConnection(t *testing.T) {
var requests int
server := httptest.NewTLSServer(http.HandlerFunc(func(http.ResponseWriter, *http.Request) {
requests++
}))
defer server.Close()
_, port, err := net.SplitHostPort(server.Listener.Addr().String())
if err != nil {
t.Fatalf("server address: %v", err)
}
resolves := 0
fetcher := NewCoverFetcherWithResolver(func(context.Context, string) ([]netip.Addr, error) {
resolves++
if resolves == 1 {
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
}
return []netip.Addr{netip.MustParseAddr("127.0.0.1")}, nil
})
_, _, err = fetcher.Fetch(context.Background(), "https://cdn.example:"+port+"/cover.jpg")
if err == nil {
t.Fatal("Fetch accepted a destination that became private")
}
if resolves != 2 {
t.Fatalf("resolver calls = %d, want preflight and dial checks", resolves)
}
if requests != 0 {
t.Fatalf("requests = %d, want 0", requests)
}
}
type roundTripFunc func(*http.Request) (*http.Response, error)
func (f roundTripFunc) RoundTrip(r *http.Request) (*http.Response, error) { return f(r) }
func coverResponse(status int, contentType, location string, body []byte) *http.Response {
header := make(http.Header)
if contentType != "" {
header.Set("Content-Type", contentType)
}
if location != "" {
header.Set("Location", location)
}
return &http.Response{
StatusCode: status,
Status: http.StatusText(status),
Header: header,
Body: io.NopCloser(bytes.NewReader(body)),
ContentLength: int64(len(body)),
}
}
func TestCoverFetcherRefusesUnsafeDestinationsBeforeRequest(t *testing.T) {
var calls int
client := &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) {
calls++
return coverResponse(http.StatusOK, "image/jpeg", "", []byte("must not reach network")), nil
})}
resolve := func(_ context.Context, host string) ([]netip.Addr, error) {
switch host {
case "loopback.example":
return []netip.Addr{netip.MustParseAddr("127.0.0.1")}, nil
case "private.example":
return []netip.Addr{netip.MustParseAddr("10.0.0.1")}, nil
case "linklocal.example":
return []netip.Addr{netip.MustParseAddr("169.254.1.1")}, nil
case "unique-local.example":
return []netip.Addr{netip.MustParseAddr("fc00::1")}, nil
case "cgnat.example":
return []netip.Addr{netip.MustParseAddr("100.64.0.1")}, nil
default:
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
}
}
fetcher := newCoverFetcher(client, resolve)
tests := []string{
"http://public.example/cover.jpg",
"https://127.0.0.1/cover.jpg",
"https://10.0.0.1/cover.jpg",
"https://169.254.1.1/cover.jpg",
"https://[fc00::1]/cover.jpg",
"https://100.64.0.1/cover.jpg",
"https://loopback.example/cover.jpg",
"https://private.example/cover.jpg",
"https://linklocal.example/cover.jpg",
"https://unique-local.example/cover.jpg",
"https://cgnat.example/cover.jpg",
}
for _, sourceURL := range tests {
t.Run(sourceURL, func(t *testing.T) {
calls = 0
if _, _, err := fetcher.Fetch(context.Background(), sourceURL); err == nil {
t.Fatal("Fetch accepted refused destination")
}
if calls != 0 {
t.Fatalf("network calls = %d, want 0", calls)
}
})
}
}
func TestCoverFetcherStopsRedirectIntoPrivateAddress(t *testing.T) {
var calls int
client := &http.Client{Transport: roundTripFunc(func(req *http.Request) (*http.Response, error) {
calls++
if req.URL.Hostname() != "cdn.example" {
t.Fatalf("redirect reached %s", req.URL)
}
return coverResponse(http.StatusFound, "", "https://internal.example/cover.jpg", nil), nil
})}
fetcher := newCoverFetcher(client, func(_ context.Context, host string) ([]netip.Addr, error) {
if host == "internal.example" {
return []netip.Addr{netip.MustParseAddr("192.168.1.1")}, nil
}
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
})
if _, _, err := fetcher.Fetch(context.Background(), "https://cdn.example/cover.jpg"); err == nil {
t.Fatal("Fetch followed redirect into private address")
}
if calls != 1 {
t.Fatalf("network calls = %d, want only public first hop", calls)
}
}
func TestCoverFetcherRejectsOversizedBody(t *testing.T) {
var calls int
client := &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) {
calls++
response := coverResponse(http.StatusOK, "image/webp", "", bytes.Repeat([]byte("x"), maxBodyBytes+1))
response.ContentLength = -1
return response, nil
})}
fetcher := newCoverFetcher(client, func(context.Context, string) ([]netip.Addr, error) {
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
})
if _, _, err := fetcher.Fetch(context.Background(), "https://cdn.example/large.webp"); err == nil {
t.Fatal("Fetch accepted oversized body")
}
if calls != 1 {
t.Fatalf("network calls = %d, want 1", calls)
}
}
func TestCoverFetcherRejectsNonImage(t *testing.T) {
var calls int
client := &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) {
calls++
return coverResponse(http.StatusOK, "text/html", "", []byte("challenge")), nil
})}
fetcher := newCoverFetcher(client, func(context.Context, string) ([]netip.Addr, error) {
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
})
if _, _, err := fetcher.Fetch(context.Background(), "https://cdn.example/challenge"); err == nil {
t.Fatal("Fetch accepted non-image response")
}
if calls != 1 {
t.Fatalf("network calls = %d, want 1", calls)
}
}
// comix labels its covers "image/jpg", which is not a registered type but is
// what the Site actually answers with; the bytes are stored under the real
// name so one image cannot land under two spellings.
func TestCoverFetcherCanonicalisesJpgAlias(t *testing.T) {
client := &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) {
return coverResponse(http.StatusOK, "image/jpg", "", []byte("cover-bytes")), nil
})}
fetcher := newCoverFetcher(client, func(context.Context, string) ([]netip.Addr, error) {
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
})
body, contentType, err := fetcher.Fetch(context.Background(), "https://static.comix.to/cover.jpg")
if err != nil {
t.Fatalf("Fetch: %v", err)
}
if string(body) != "cover-bytes" || contentType != "image/jpeg" {
t.Fatalf("Fetch = (%q, %q), want (cover-bytes, image/jpeg)", body, contentType)
}
}
+118 -25
View File
@@ -4,6 +4,7 @@ import (
"context"
"log"
"net/url"
"slices"
"time"
"bookmarkmanager/backend/internal/store"
@@ -15,6 +16,13 @@ type Fetcher interface {
Get(ctx context.Context, url string) (body string, status int, err error)
}
// BrowserCoverFetcher retrieves one cover's bytes through the browser-backed
// path — the only route that clears the challenge kagane's image URLs answer
// a plain fetch with. Satisfied by BrowserFetcher.
type BrowserCoverFetcher interface {
Image(ctx context.Context, imageURL string) (body []byte, contentType string, err error)
}
// Poller re-checks each bookmarked series' newest published chapter on a
// schedule, independent of the userscript's own in-browser checks. The two run
// in parallel and report the same observable fact, so whichever writes last wins
@@ -23,12 +31,12 @@ type Fetcher interface {
// Two clocks, deliberately independent:
//
// - Interval is how often this goroutine wakes up and looks.
// - Cooldown is how long one series rests since its own last check.
// - Cooldowns are how long a series rests since its own last check. Browser-
// backed sites use the longer BrowserCooldown.
//
// Only the cooldown is per series, and it is enforced by the WHERE clause in
// DueForLatestCheck rather than by any timer. Shortening Interval therefore
// cannot shorten anyone's cooldown; it only makes the poller wake up and find
// nothing due more often.
// Cooldowns are enforced by the WHERE clause in DueForLatestCheck rather than
// by any timer. Shortening Interval therefore cannot shorten anyone's cooldown;
// it only makes the poller wake up and find nothing due more often.
type Poller struct {
Store *store.Store
Fetch Fetcher
@@ -36,24 +44,103 @@ type Poller struct {
// cannot clear. Nil disables those sites entirely rather than falling back
// to Fetch, which would only ever retrieve a challenge page.
BrowserFetch Fetcher
Now func() time.Time // injected so tests can freeze it
Cooldown time.Duration
Interval time.Duration
Stagger time.Duration
Batch int
// CoverFetch is optional; failures are logged and never affect the chapter poll.
CoverFetch BrowserCoverFetcher
// CoverBytesFetch is optional; it handles plain-TLS sources through the
// same failure-isolated prefetch path.
CoverBytesFetch CoverBytesFetcher
Now func() time.Time // injected so tests can freeze it
Cooldown time.Duration
BrowserCooldown time.Duration
Interval time.Duration
Stagger time.Duration
Batch int
}
// fetcherFor returns the fetcher a site needs, or nil when the site cannot be
// fetched at all right now. kagane and novelfull both sit behind a Cloudflare
// JavaScript challenge that no TLS fingerprint clears — kagane verified
// 2026-08-03, novelfull verified 2026-08-05, both against the same Chrome_133
// profile TLSFetcher uses — so they are browser-only or nothing.
func (p *Poller) fetcherFor(site string) Fetcher {
switch site {
case "kagane", "novelfull":
return p.BrowserFetch
var browserBackedSites = []string{"kagane", "novelfull"}
// fillBlankCover gives a Series its Cover when it has none. The blank state is
// what "no Cover yet" means on the wire (ADR-0007): permanently-blank rows
// created before acquisition existed, and rows whose creation-time fetch
// failed, both heal here. A non-blank CoverAddress is left alone — refetching
// would add a request per Series per cycle and change artwork under the Reader
// for no visible reason. A row that already carries a source URL is owned by
// prefetchCover instead; this path only extracts from the series page.
//
// Failures are logged against the Series and never returned: the chapter poll
// must not notice. A failed fill is retried the next time this Series is due;
// there is no separate retry queue.
func (p *Poller) fillBlankCover(ctx context.Context, sr store.Series, body string) {
if sr.CoverAddress != "" || sr.Cover != "" {
return
}
cover, ok := coverFrom(sr.Site, sr.SeriesURL, body)
if !ok {
return
}
p.storeCover(ctx, sr, cover)
}
// prefetchCover heals Series that already carry a third-party source URL but
// no stored address — the state left by client-supplied covers before
// acquisition moved server-side. Every Site takes the same path; fetchCoverBytes
// routes by URL shape, so browser-claimed URLs still need the sidecar. New
// blanks have no source URL and go through fillBlankCover from the series page
// instead.
func (p *Poller) prefetchCover(ctx context.Context, sr store.Series) {
if sr.Cover == "" || sr.CoverAddress != "" {
return
}
body, contentType, found, err := p.Store.GetCover(sr.Cover)
if err != nil {
log.Printf("latest poll %q: read cover: %v", sr.Key(), err)
return
}
if found {
if err := p.Store.SetSeriesCover(sr.Site, sr.SeriesID, sr.Cover, body, contentType); err != nil {
log.Printf("latest poll %q: persist cover: %v", sr.Key(), err)
}
return
}
p.storeCover(ctx, sr, sr.Cover)
}
// storeCover fetches bytes for sourceURL and points the Series at them. Every
// failure is logged against the Series and swallowed so the chapter poll
// cannot see it.
func (p *Poller) storeCover(ctx context.Context, sr store.Series, sourceURL string) {
bytes, contentType, err := fetchCoverBytes(ctx, sourceURL, p.CoverFetch, p.CoverBytesFetch)
if err != nil {
log.Printf("latest poll %q: fetch cover %s: %v", sr.Key(), sourceURL, err)
return
}
if err := p.Store.SetSeriesCover(sr.Site, sr.SeriesID, sourceURL, bytes, contentType); err != nil {
log.Printf("latest poll %q: persist cover: %v", sr.Key(), err)
}
}
// fetcherFor returns the fetcher a site's page needs, or nil when the site
// cannot be fetched at all right now. kagane and novelfull pages sit behind a
// Cloudflare JavaScript challenge that no TLS fingerprint clears (kagane
// verified 2026-08-03, novelfull verified 2026-08-05, both against the same
// Chrome_133 profile TLSFetcher uses), so both prefer the browser; novelfull
// alone falls back to the plain-TLS fetcher when no browser is configured,
// because its challenge is a live time-varying fact (AGENTS.md) and its cover
// bytes never need the browser. kagane never falls back: a plain fetch of a
// kagane page or cover would only ever retrieve a challenge page. One routing
// rule for the poll and the acquirer, so the two cannot drift apart.
func fetcherFor(site string, browser, tls Fetcher) Fetcher {
switch {
case site == "kagane":
return browser
case slices.Contains(browserBackedSites, site): // novelfull
if browser != nil {
return browser
}
return tls
default:
return tls
}
return p.Fetch
}
// Run polls until ctx is cancelled.
@@ -63,8 +150,8 @@ func (p *Poller) fetcherFor(site string) Fetcher {
// failure mode for a misconfigured batch x stagger: a slower cadence, never
// concurrent fetch storms.
func (p *Poller) Run(ctx context.Context) {
log.Printf("latest-chapter poller: interval=%s cooldown=%s batch=%d stagger=%s",
p.Interval, p.Cooldown, p.Batch, p.Stagger)
log.Printf("latest-chapter poller: interval=%s cooldown=%s browser-cooldown=%s batch=%d stagger=%s",
p.Interval, p.Cooldown, p.BrowserCooldown, p.Batch, p.Stagger)
t := time.NewTicker(p.Interval)
defer t.Stop()
for {
@@ -80,8 +167,10 @@ func (p *Poller) Run(ctx context.Context) {
// runOnce processes one batch of due series.
func (p *Poller) runOnce(ctx context.Context) {
cutoff := p.Now().Add(-p.Cooldown).UnixMilli()
due, err := p.Store.DueForLatestCheck(cutoff, p.Batch)
now := p.Now()
cutoff := now.Add(-p.Cooldown).UnixMilli()
browserCutoff := now.Add(-p.BrowserCooldown).UnixMilli()
due, err := p.Store.DueForLatestCheck(cutoff, browserCutoff, browserBackedSites, p.Batch)
if err != nil {
log.Printf("latest poll: due query: %v", err)
return
@@ -147,11 +236,12 @@ func (p *Poller) checkOne(ctx context.Context, sr store.Series) {
return
}
f := p.fetcherFor(sr.Site)
f := fetcherFor(sr.Site, p.BrowserFetch, p.Fetch)
if f == nil {
log.Printf("latest poll %q: no fetcher for site %q", sr.Key(), sr.Site)
return
}
p.prefetchCover(ctx, sr)
body, status, err := f.Get(ctx, sr.SeriesURL)
if err != nil {
@@ -164,6 +254,9 @@ func (p *Poller) checkOne(ctx context.Context, sr store.Series) {
}
latest, ok := latestChapterFrom(sr.Site, sr.SeriesURL, body)
// Cover fill is independent of the chapter signal: a page that lost its
// chapter list may keep its og:image, and a blank Series heals either way.
p.fillBlankCover(ctx, sr, body)
if !ok {
// Most likely a challenge page or a layout change. Either way the row is
// already stamped, so this waits out a cooldown instead of hot-looping.
+659 -21
View File
@@ -3,7 +3,9 @@ package latest
import (
"context"
"crypto/sha256"
"database/sql"
"errors"
"log"
"os"
"strings"
"sync"
@@ -20,13 +22,16 @@ func TestMain(m *testing.M) { os.Exit(pgtest.Main(m)) }
// test needs one, is created by opening the same database as a second owner.
var testOwner = store.Owner{DiscordID: "test-owner", TokenHash: sha256.Sum256([]byte("owner-token-hash"))}
// testCoverBaseURL is the public origin every stored cover URL is built from.
const testCoverBaseURL = "https://bookmarks.test"
// newTestStore opens a store on a Postgres database of this test's own and
// returns the URL, for helpers that need a second connection to the same
// database (see TestRunOnceFetchesSharedSeriesOnce).
func newTestStore(t *testing.T) (*store.Store, string) {
t.Helper()
url := pgtest.URL(t)
s, err := store.Open(url, testOwner)
s, err := store.Open(url, testOwner, t.TempDir(), testCoverBaseURL)
if err != nil {
t.Fatalf("Open: %v", err)
}
@@ -103,18 +108,145 @@ func (f *fakeFetcher) callCount() int {
return len(f.calls)
}
type fakeCoverFetcher struct {
mu sync.Mutex
calls []string
body []byte
contentType string
err error
}
func (f *fakeCoverFetcher) Image(_ context.Context, imageURL string) ([]byte, string, error) {
f.mu.Lock()
f.calls = append(f.calls, imageURL)
f.mu.Unlock()
if f.err != nil {
return nil, "", f.err
}
return f.body, f.contentType, nil
}
func (f *fakeCoverFetcher) callCount() int {
f.mu.Lock()
defer f.mu.Unlock()
return len(f.calls)
}
type fakeBytesCoverFetcher struct {
mu sync.Mutex
calls []string
body []byte
contentType string
err error
}
func (f *fakeBytesCoverFetcher) Fetch(_ context.Context, sourceURL string) ([]byte, string, error) {
f.mu.Lock()
f.calls = append(f.calls, sourceURL)
f.mu.Unlock()
if f.err != nil {
return nil, "", f.err
}
return f.body, f.contentType, nil
}
func (f *fakeBytesCoverFetcher) callCount() int {
f.mu.Lock()
defer f.mu.Unlock()
return len(f.calls)
}
// newTestPoller wires a poller with a frozen clock and no stagger, so tests run
// instantly and deterministically.
func newTestPoller(t *testing.T, s *store.Store, f Fetcher, at time.Time) *Poller {
t.Helper()
return &Poller{
Store: s,
Fetch: f,
Now: func() time.Time { return at },
Cooldown: time.Hour,
Interval: 10 * time.Minute,
Stagger: 0,
Batch: 14,
Store: s,
Fetch: f,
Now: func() time.Time { return at },
Cooldown: time.Hour,
BrowserCooldown: 6 * time.Hour,
Interval: 10 * time.Minute,
Stagger: 0,
Batch: 14,
}
}
// seedCoverSource writes a Series' cover source address without any stored
// bytes. Nothing in production produces that state any more — a client cover
// is discarded and an acquired one arrives with its bytes — but rows created
// before covers moved server-side still carry one, and the prefetch is what
// heals them.
func seedCoverSource(t *testing.T, dbURL, site, seriesID, coverURL string) {
t.Helper()
db, err := sql.Open("pgx", dbURL)
if err != nil {
t.Fatalf("open %s: %v", dbURL, err)
}
defer db.Close()
if _, err := db.Exec(
`UPDATE series SET cover = $3 WHERE site = $1 AND series_id = $2`,
site, seriesID, coverURL); err != nil {
t.Fatalf("seed cover source %s:%s: %v", site, seriesID, err)
}
}
func TestRunOncePrefetchesPublicCover(t *testing.T) {
s, dbURL := newTestStore(t)
const (
key = "asura:chronicles-of-the-demon-faction-f886a8af"
seriesURL = "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af"
coverURL = "https://cdn.example/covers/chronicles.jpg"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "asura", SeriesID: "chronicles-of-the-demon-faction-f886a8af",
SeriesURL: seriesURL, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
seedCoverSource(t, dbURL, "asura", "chronicles-of-the-demon-faction-f886a8af", coverURL)
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"}
p := &Poller{
Store: s, Fetch: &fakeFetcher{body: asuraSeriesFixture, status: 200}, CoverBytesFetch: covers,
Now: func() time.Time { return time.UnixMilli(5_000_000) }, Cooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetch calls = %d, want 1", got)
}
body, contentType, found, err := s.GetCover(coverURL)
if err != nil || !found {
t.Fatalf("GetCover: %v found=%v", err, found)
}
if string(body) != "cover-bytes" || contentType != "image/jpeg" {
t.Fatalf("stored cover = (%q, %q), want (cover-bytes, image/jpeg)", body, contentType)
}
}
func TestRunOnceDoesNotStoreNonImagePublicCover(t *testing.T) {
s, dbURL := newTestStore(t)
const (
key = "asura:non-image-cover"
seriesURL = "https://asurascans.com/comics/non-image-cover"
coverURL = "https://cdn.example/covers/challenge"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "asura", SeriesID: "non-image-cover", SeriesURL: seriesURL,
UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
seedCoverSource(t, dbURL, "asura", "non-image-cover", coverURL)
p := &Poller{
Store: s, Fetch: &fakeFetcher{body: asuraSeriesFixture, status: 200},
CoverBytesFetch: &fakeBytesCoverFetcher{body: []byte("challenge"), contentType: "text/html"},
Now: func() time.Time { return time.UnixMilli(5_000_000) }, Cooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if _, _, found, err := s.GetCover(coverURL); err != nil || found {
t.Fatalf("non-image cover = found %v, err %v; want missing", found, err)
}
}
@@ -248,7 +380,7 @@ func TestRunOnceFetchesSharedSeriesOnce(t *testing.T) {
// A second reader tracks the same series. The seed is the only
// reader-creation path, so a second Open as a different owner is how a
// test gets a second reader on the same database.
other, err := store.Open(url, store.Owner{DiscordID: "second-reader", TokenHash: sha256.Sum256([]byte("second-token-hash"))})
other, err := store.Open(url, store.Owner{DiscordID: "second-reader", TokenHash: sha256.Sum256([]byte("second-token-hash"))}, t.TempDir(), testCoverBaseURL)
if err != nil {
t.Fatalf("Open second reader: %v", err)
}
@@ -304,6 +436,26 @@ func TestRunOnceOneBadSeriesDoesNotStallBatch(t *testing.T) {
}
}
func TestRunLogsCooldowns(t *testing.T) {
var logs strings.Builder
previous := log.Writer()
log.SetOutput(&logs)
t.Cleanup(func() { log.SetOutput(previous) })
ctx, cancel := context.WithCancel(context.Background())
cancel()
(&Poller{
Cooldown: time.Hour,
BrowserCooldown: 6 * time.Hour,
Interval: time.Hour,
}).Run(ctx)
if got := logs.String(); !strings.Contains(got, "cooldown=1h") ||
!strings.Contains(got, "browser-cooldown=6h") {
t.Fatalf("startup log = %q, want both cooldowns", got)
}
}
// The cooldown is enforced by the due query, so a second immediate pass must do
// nothing at all — this is what makes the tick interval independent of it.
func TestRunOnceHonoursCooldownAcrossPasses(t *testing.T) {
@@ -334,6 +486,44 @@ func TestRunOnceHonoursCooldownAcrossPasses(t *testing.T) {
}
}
func TestRunOnceUsesBrowserCooldown(t *testing.T) {
s, _ := newTestStore(t)
const browserKey = "kagane:019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
seedForCheck(t, s, "asura:plain", "https://asurascans.com/comics/plain", 0)
seedForCheck(t, s, browserKey, "https://kagane.to/series/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b", 0)
hour := time.Hour
now := time.Unix(2*int64(hour/time.Second), 0)
tls := &fakeFetcher{status: 200}
browser := &fakeFetcher{status: 200}
p := &Poller{
Store: s,
Fetch: tls,
BrowserFetch: browser,
Now: func() time.Time { return now },
Cooldown: hour,
BrowserCooldown: 6 * hour,
Batch: 10,
}
p.runOnce(context.Background())
if got := tls.callCount(); got != 1 {
t.Fatalf("plain-TLS fetches after 2h = %d, want 1", got)
}
if got := browser.callCount(); got != 0 {
t.Fatalf("browser fetches after 2h = %d, want 0", got)
}
now = time.Unix(7*int64(hour/time.Second), 0)
p.runOnce(context.Background())
if got := tls.callCount(); got != 2 {
t.Fatalf("plain-TLS fetches after 7h = %d, want 2", got)
}
if got := browser.callCount(); got != 1 {
t.Fatalf("browser fetches after 7h = %d, want 1", got)
}
}
// A site that retracts a chapter should correct the stored number downward,
// mirroring the userscript's equality check (L427) rather than a >.
func TestRunOnceCorrectsDownward(t *testing.T) {
@@ -465,12 +655,14 @@ func TestKaganeSkippedWhenNoBrowserFetcher(t *testing.T) {
}); err != nil {
t.Fatalf("seed: %v", err)
}
f := &fakeFetcher{body: kaganeAPIFixture, status: 200}
p := &Poller{
Store: s, Fetch: f,
Store: s,
Fetch: f,
Now: func() time.Time { return time.UnixMilli(5_000_000) },
Cooldown: time.Hour, Interval: time.Hour, Batch: 10,
Cooldown: time.Hour, BrowserCooldown: time.Hour,
Interval: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
@@ -479,6 +671,51 @@ func TestKaganeSkippedWhenNoBrowserFetcher(t *testing.T) {
}
}
// novelfull without a browser is not skipped outright: its challenge is a
// live time-varying fact, so the plain-TLS page fetch is attempted and — when
// the body answers — fills both the chapter and the Cover, exactly the
// client-scraped rows #62 wants healed.
func TestNovelfullUsesTLSWhenNoBrowserFetcher(t *testing.T) {
s, _ := newTestStore(t)
key := "novelfull:reverend-insanity"
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key,
Site: "novelfull",
SeriesID: "reverend-insanity",
SeriesURL: "https://novelfull.com/reverend-insanity.html",
UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
tlsF := &fakeFetcher{body: novelfullSeriesFixture + novelfullCoverFixture, status: 200}
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
p := &Poller{
Store: s, Fetch: tlsF, CoverBytesFetch: covers,
Now: func() time.Time { return time.UnixMilli(5_000_000) },
Cooldown: time.Hour, BrowserCooldown: time.Hour,
Interval: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if len(tlsF.calls) != 1 {
t.Fatalf("TLS fetcher calls = %d, want 1", len(tlsF.calls))
}
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetches = %d, want 1", got)
}
got, found, err := s.Get(s.OwnerID(), key)
if err != nil || !found {
t.Fatalf("Get: %v found=%v", err, found)
}
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(novelfullCoverURL); got.Cover != want {
t.Fatalf("Cover = %q, want %q", got.Cover, want)
}
if got.LatestChapterNum == nil || *got.LatestChapterNum != 2334 {
t.Fatalf("LatestChapterNum = %v, want 2334", got.LatestChapterNum)
}
}
// With a browser fetcher wired up, kagane goes to it and not to the TLS one.
func TestKaganeUsesBrowserFetcher(t *testing.T) {
s, _ := newTestStore(t)
@@ -498,7 +735,8 @@ func TestKaganeUsesBrowserFetcher(t *testing.T) {
p := &Poller{
Store: s, Fetch: tlsF, BrowserFetch: browserF,
Now: func() time.Time { return time.UnixMilli(5_000_000) },
Cooldown: time.Hour, Interval: time.Hour, Batch: 10,
Cooldown: time.Hour, BrowserCooldown: time.Hour,
Interval: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
@@ -517,23 +755,207 @@ func TestKaganeUsesBrowserFetcher(t *testing.T) {
}
}
func TestRunOncePrefetchesKaganeCover(t *testing.T) {
s, dbURL := newTestStore(t)
const (
key = "kagane:019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
coverURL = "https://kagane.to/api/v2/image/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b/compressed"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "kagane", SeriesID: "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b",
SeriesURL: "https://kagane.to/series/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b",
UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
seedCoverSource(t, dbURL, "kagane", "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b", coverURL)
covers := &fakeCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
p := &Poller{
Store: s, Fetch: &fakeFetcher{body: kaganeAPIFixture, status: 200},
BrowserFetch: &fakeFetcher{body: kaganeAPIFixture, status: 200}, CoverFetch: covers,
Now: func() time.Time { return time.UnixMilli(5_000_000) },
Cooldown: time.Hour, BrowserCooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
body, contentType, ok, err := s.CoverByAddress(store.CoverAddress(coverURL))
if err != nil || !ok {
t.Fatalf("CoverByAddress: %v found=%v", err, ok)
}
if string(body) != "cover-bytes" || contentType != "image/webp" {
t.Fatalf("stored cover = (%q, %q), want (cover-bytes, image/webp)", body, contentType)
}
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetch calls = %d, want 1", got)
}
if got := readBookmark(t, s, key); got.Cover != testCoverBaseURL+"/covers/"+store.CoverAddress(coverURL) {
t.Fatalf("wire Cover = %q, want content-addressed URL", got.Cover)
}
}
func TestRunOnceDoesNotRefetchKaganeCover(t *testing.T) {
s, dbURL := newTestStore(t)
const (
key = "kagane:019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
seriesID = "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
coverURL = "https://kagane.to/api/v2/image/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b/compressed"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "kagane", SeriesID: seriesID,
SeriesURL: "https://kagane.to/series/" + seriesID, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
seedCoverSource(t, dbURL, "kagane", seriesID, coverURL)
at := time.UnixMilli(5_000_000)
covers := &fakeCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
p := &Poller{
Store: s, BrowserFetch: &fakeFetcher{body: kaganeAPIFixture, status: 200}, CoverFetch: covers,
Now: func() time.Time { return at }, Cooldown: time.Hour, BrowserCooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
at = at.Add(2 * time.Hour)
p.runOnce(context.Background())
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetch calls = %d, want 1 after two due cycles", got)
}
}
func TestRunOnceCoverFailureDoesNotBlockChapter(t *testing.T) {
s, dbURL := newTestStore(t)
const (
key = "kagane:019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
seriesID = "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
coverURL = "https://kagane.to/api/v2/image/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b/compressed"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "kagane", SeriesID: seriesID,
SeriesURL: "https://kagane.to/series/" + seriesID, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
seedCoverSource(t, dbURL, "kagane", seriesID, coverURL)
now := time.UnixMilli(5_000_000)
p := &Poller{
Store: s, BrowserFetch: &fakeFetcher{body: kaganeAPIFixture, status: 200},
CoverFetch: &fakeCoverFetcher{err: errors.New("browser unavailable")},
Now: func() time.Time { return now }, Cooldown: time.Hour, BrowserCooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
got, found, err := s.Get(s.OwnerID(), key)
if err != nil || !found {
t.Fatalf("Get: %v found=%v", err, found)
}
if got.LatestChapterNum == nil || *got.LatestChapterNum != 41 {
t.Fatalf("LatestChapterNum = %v, want 41", got.LatestChapterNum)
}
if checked := readLatestCheckedAt(t, s, key); checked != now.UnixMilli() {
t.Fatalf("latest_checked_at = %d, want %d", checked, now.UnixMilli())
}
}
func TestRunOnceRejectsInvalidKaganeCover(t *testing.T) {
s, dbURL := newTestStore(t)
const (
key = "kagane:019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
seriesID = "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
coverURL = "https://kagane.to/api/v2/image/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b/compressed"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "kagane", SeriesID: seriesID,
SeriesURL: "https://kagane.to/series/" + seriesID, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
seedCoverSource(t, dbURL, "kagane", seriesID, coverURL)
p := &Poller{
Store: s, BrowserFetch: &fakeFetcher{body: kaganeAPIFixture, status: 200},
CoverFetch: &fakeCoverFetcher{body: []byte("not an image"), contentType: "text/html"},
Now: func() time.Time { return time.UnixMilli(5_000_000) }, Cooldown: time.Hour, BrowserCooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if _, _, found, err := s.CoverByAddress(store.CoverAddress(coverURL)); err != nil || found {
t.Fatalf("invalid cover persisted = %v, err %v; want missing", found, err)
}
}
func TestRunOnceWithoutCoverFetcherStillPollsKagane(t *testing.T) {
s, dbURL := newTestStore(t)
const (
key = "kagane:019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
seriesID = "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
coverURL = "https://kagane.to/api/v2/image/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b/compressed"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "kagane", SeriesID: seriesID,
SeriesURL: "https://kagane.to/series/" + seriesID, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
seedCoverSource(t, dbURL, "kagane", seriesID, coverURL)
p := &Poller{
Store: s, BrowserFetch: &fakeFetcher{body: kaganeAPIFixture, status: 200},
Now: func() time.Time { return time.UnixMilli(5_000_000) }, Cooldown: time.Hour, BrowserCooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if _, _, found, err := s.CoverByAddress(store.CoverAddress(coverURL)); err != nil || found {
t.Fatalf("cover after nil CoverFetch = found %v, err %v; want missing", found, err)
}
}
func TestRunOnceRoutesNonKaganeCoverToPublicFetcher(t *testing.T) {
s, dbURL := newTestStore(t)
const key = "asura:solo"
const coverURL = "https://asurascans.com/covers/solo.jpg"
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "asura", SeriesID: "solo", SeriesURL: "https://asurascans.com/comics/solo",
UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
seedCoverSource(t, dbURL, "asura", "solo", coverURL)
browserCovers := &fakeCoverFetcher{body: []byte("must not be fetched"), contentType: "image/webp"}
publicCovers := &fakeBytesCoverFetcher{body: []byte("public cover"), contentType: "image/webp"}
p := &Poller{
Store: s, Fetch: &fakeFetcher{body: asuraSeriesFixture, status: 200}, CoverFetch: browserCovers,
CoverBytesFetch: publicCovers,
Now: func() time.Time { return time.UnixMilli(5_000_000) }, Cooldown: time.Hour, BrowserCooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if got := publicCovers.callCount(); got != 1 {
t.Fatalf("public cover fetch calls = %d, want 1", got)
}
if got := browserCovers.callCount(); got != 0 {
t.Fatalf("browser cover fetch calls for asura = %d, want 0", got)
}
}
func TestFetcherForRoutesNovelSites(t *testing.T) {
tls := &fakeFetcher{}
browser := &fakeFetcher{}
p := &Poller{Fetch: tls, BrowserFetch: browser}
cases := []struct {
site string
want Fetcher
site string
browser Fetcher
tls Fetcher
want Fetcher
}{
{"asura", tls},
{"lightnovelworld", tls},
{"kagane", browser},
{"novelfull", browser},
{"asura", browser, tls, tls},
{"lightnovelworld", browser, tls, tls},
{"kagane", browser, tls, browser},
{"novelfull", browser, tls, browser},
// browser-less deployment: kagane is nothing, novelfull degrades to TLS
{"kagane", nil, tls, nil},
{"novelfull", nil, tls, tls},
}
for _, tc := range cases {
t.Run(tc.site, func(t *testing.T) {
if got := p.fetcherFor(tc.site); got != tc.want {
if got := fetcherFor(tc.site, tc.browser, tc.tls); got != tc.want {
t.Fatalf("fetcherFor(%q) = %v, want %v", tc.site, got, tc.want)
}
})
@@ -562,3 +984,219 @@ func TestFetchableSeriesURLPinsNovelHosts(t *testing.T) {
})
}
}
// A Series that has been blank since creation has no source URL to refetch.
// The poll extracts the Cover from the same series page it already fetched
// for the chapter signal and stores the bytes — every Site, both Libraries.
func TestRunOnceFillsBlankCoverFromSeriesPage(t *testing.T) {
cases := []struct {
name string
key string
site string
seriesID string
seriesURL string
kind string
body string
wantCover string
browser bool
}{
{
name: "asura manga",
key: "asura:chronicles-of-the-demon-faction-f886a8af", site: "asura",
seriesID: "chronicles-of-the-demon-faction-f886a8af",
seriesURL: "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af",
kind: store.KindManga, body: asuraSeriesFixture + asuraCoverFixture,
wantCover: "https://cdn.asurascans.com/asura-images/covers/chronicles-of-the-demon-faction.d4dcb8.webp",
},
{
name: "lightnovelworld novel",
key: "lightnovelworld:a-will-eternal", site: "lightnovelworld",
seriesID: "a-will-eternal", seriesURL: "https://lightnovelworld.net/novel/a-will-eternal/",
kind: store.KindNovel, body: lnwSeriesFixture + lnwCoverFixture,
wantCover: "https://lightnovelworld.net/wp-content/uploads/2026/03/a-will-eternal-1.webp",
},
{
name: "kagane manga",
key: "kagane:019fe11a-8670-7cf3-8343-0b02057d3787", site: "kagane",
seriesID: "019fe11a-8670-7cf3-8343-0b02057d3787",
seriesURL: "https://kagane.to/series/019fe11a-8670-7cf3-8343-0b02057d3787",
kind: store.KindManga, body: kaganeAPIFixtureWithCover, browser: true,
wantCover: "https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
s, _ := newTestStore(t)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: tc.key, Site: tc.site, SeriesID: tc.seriesID,
SeriesURL: tc.seriesURL, Kind: tc.kind, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
page := &fakeFetcher{body: tc.body, status: 200}
public := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"}
browser := &fakeCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
p := &Poller{
Store: s, Fetch: page, BrowserFetch: page,
CoverBytesFetch: public, CoverFetch: browser,
Now: func() time.Time { return time.UnixMilli(5_000_000) },
Cooldown: time.Hour, BrowserCooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
got := readBookmark(t, s, tc.key)
wantWire := testCoverBaseURL + "/covers/" + store.CoverAddress(tc.wantCover)
if got.Cover != wantWire {
t.Fatalf("Cover = %q, want %q", got.Cover, wantWire)
}
if tc.browser {
if got := browser.callCount(); got != 1 {
t.Fatalf("browser cover fetches = %d, want 1", got)
}
if got := public.callCount(); got != 0 {
t.Fatalf("public cover fetches = %d, want 0", got)
}
} else {
if got := public.callCount(); got != 1 {
t.Fatalf("public cover fetches = %d, want 1", got)
}
if got := browser.callCount(); got != 0 {
t.Fatalf("browser cover fetches = %d, want 0", got)
}
}
})
}
}
// Once a Cover exists the poll must leave it alone: refetching every cycle is
// noise for the Reader and a request per Series against Sites that already
// bot-score the deployment's single IP.
func TestRunOnceDoesNotReplaceExistingCover(t *testing.T) {
s, _ := newTestStore(t)
const (
key = "asura:chronicles-of-the-demon-faction-f886a8af"
seriesID = "chronicles-of-the-demon-faction-f886a8af"
seriesURL = "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af"
first = "https://cdn.example/covers/first.jpg"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "asura", SeriesID: seriesID, SeriesURL: seriesURL, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
if err := s.SetSeriesCover("asura", seriesID, first, []byte("first"), "image/jpeg"); err != nil {
t.Fatalf("seed cover: %v", err)
}
public := &fakeBytesCoverFetcher{body: []byte("second"), contentType: "image/jpeg"}
at := time.UnixMilli(5_000_000)
p := &Poller{
Store: s, Fetch: &fakeFetcher{body: asuraSeriesFixture + asuraCoverFixture, status: 200},
CoverBytesFetch: public,
Now: func() time.Time { return at }, Cooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
at = at.Add(2 * time.Hour)
p.runOnce(context.Background())
if got := public.callCount(); got != 0 {
t.Fatalf("cover fetch calls = %d, want 0", got)
}
got := readBookmark(t, s, key)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(first); got.Cover != want {
t.Fatalf("Cover = %q, want the first one %q", got.Cover, want)
}
}
// A blank Cover whose byte fetch fails is retried the next time the Series is
// polled. There is no separate retry queue — the due cycle is the queue.
func TestRunOnceRetriesFailedBlankCoverOnNextPoll(t *testing.T) {
s, _ := newTestStore(t)
const (
key = "asura:chronicles-of-the-demon-faction-f886a8af"
seriesID = "chronicles-of-the-demon-faction-f886a8af"
seriesURL = "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af"
coverURL = "https://cdn.asurascans.com/asura-images/covers/chronicles-of-the-demon-faction.d4dcb8.webp"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "asura", SeriesID: seriesID, SeriesURL: seriesURL, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
public := &fakeBytesCoverFetcher{err: errors.New("cdn down")}
at := time.UnixMilli(5_000_000)
p := &Poller{
Store: s, Fetch: &fakeFetcher{body: asuraSeriesFixture + asuraCoverFixture, status: 200},
CoverBytesFetch: public,
Now: func() time.Time { return at }, Cooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if got := readBookmark(t, s, key); got.Cover != "" {
t.Fatalf("Cover after failed fetch = %q, want blank", got.Cover)
}
if got := public.callCount(); got != 1 {
t.Fatalf("cover fetch calls after fail = %d, want 1", got)
}
public.err = nil
public.body = []byte("cover-bytes")
public.contentType = "image/jpeg"
at = at.Add(2 * time.Hour)
p.runOnce(context.Background())
if got := public.callCount(); got != 2 {
t.Fatalf("cover fetch calls after retry = %d, want 2", got)
}
got := readBookmark(t, s, key)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(coverURL); got.Cover != want {
t.Fatalf("Cover after retry = %q, want %q", got.Cover, want)
}
}
// Cover work is cosmetic: a failed blank fill must leave the chapter poll's
// result intact for every Site, not only kagane.
func TestRunOnceBlankCoverFailureDoesNotBlockChapter(t *testing.T) {
s, _ := newTestStore(t)
const (
key = "asura:chronicles-of-the-demon-faction-f886a8af"
seriesID = "chronicles-of-the-demon-faction-f886a8af"
seriesURL = "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "asura", SeriesID: seriesID, SeriesURL: seriesURL, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
now := time.UnixMilli(5_000_000)
var logs strings.Builder
prev := log.Writer()
log.SetOutput(&logs)
t.Cleanup(func() { log.SetOutput(prev) })
p := &Poller{
Store: s, Fetch: &fakeFetcher{body: asuraSeriesFixture + asuraCoverFixture, status: 200},
CoverBytesFetch: &fakeBytesCoverFetcher{err: errors.New("cdn down")},
Now: func() time.Time { return now }, Cooldown: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
got := readBookmark(t, s, key)
if got.LatestChapterNum == nil || *got.LatestChapterNum != 181 {
t.Fatalf("LatestChapterNum = %v, want 181", got.LatestChapterNum)
}
if got.Cover != "" {
t.Fatalf("Cover = %q, want blank after failed fetch", got.Cover)
}
if !strings.Contains(logs.String(), key) {
t.Fatalf("cover failure log missing series key %q; got %q", key, logs.String())
}
}
// kaganeAPIFixture carries chapter data only. The blank-fill path needs a
// cover image id in the same body the chapter poll already retrieved.
const kaganeAPIFixtureWithCover = `
{"series_id":"019fe11a-8670-7cf3-8343-0b02057d3787","title":"Infinite Decryption",
"series_covers":[{"cover_id":"019fe11a-84d1-714b-9cf4-2827f277f3c0","language":"en","volume_number":"1","chapter_number":null,"note":null,"image_id":"019fe11a-84c3-7fc3-a84b-88787374b617"}],
"series_books":[{"book_id":"a","title":"Episode 1","chapter_no":"1","sort_no":1},
{"book_id":"b","title":"Episode 41","chapter_no":"41","sort_no":41},
{"book_id":"c","title":"Episode 40.5","chapter_no":"40.5","sort_no":40}]}
`
+136 -6
View File
@@ -1,6 +1,8 @@
package latest
import (
"encoding/json"
"html"
"net/url"
"regexp"
"strconv"
@@ -34,6 +36,18 @@ var demonicChapterRe = regexp.MustCompile(`chaptered\.php\?manga=\d+&(?:amp;)?ch
// Only the id prefix is stable; the slug tail follows the title.
var comixSlugRe = regexp.MustCompile(`/title/([^/?#]+)`)
func comixSeriesID(seriesURL string) (string, bool) {
m := comixSlugRe.FindStringSubmatch(seriesURL)
if m == nil {
return "", false
}
id := m[1]
if i := strings.Index(id, "-"); i != -1 {
id = id[:i]
}
return id, true
}
// kaganeChapterRe matches the chapter numbers in a kagane API response. This
// branch is fed by the browser fetcher, so the body is JSON rather than HTML —
// there are no anchors to scan.
@@ -87,18 +101,14 @@ func latestChapterFrom(site, seriesURL, body string) (latestChapter, bool) {
case "demonic":
re = demonicChapterRe
case "comix":
m := comixSlugRe.FindStringSubmatch(seriesURL)
if m == nil {
id, ok := comixSeriesID(seriesURL)
if !ok {
return latestChapter{}, false
}
// comix ships an SPA: the served HTML carries a JSON state blob instead
// of chapter anchors, and latestChapterUrl is the only place the newest
// chapter appears. Scoping to this series' id prefix keeps a
// "recommended" strip's entries from winning the maximum.
id := m[1]
if i := strings.Index(id, "-"); i != -1 {
id = id[:i]
}
re = regexp.MustCompile(`"latestChapterUrl":"/title/` + regexp.QuoteMeta(id) + `-[^"]*-chapter-([0-9.]+)"`)
case "kagane":
re = kaganeChapterRe
@@ -146,3 +156,123 @@ func latestChapterFrom(site, seriesURL, body string) (latestChapter, bool) {
}
return best, found
}
var metaTagRe = regexp.MustCompile(`(?is)<meta\b[^>]*>`)
var doubleQuotedMetaAttrRe = regexp.MustCompile(`(?is)([a-z][a-z0-9:_-]*)\s*=\s*"([^"]*)"`)
var singleQuotedMetaAttrRe = regexp.MustCompile(`(?is)([a-z][a-z0-9:_-]*)\s*=\s*'([^']*)'`)
// comix's server-rendered page embeds query data in this JSON script; parsing
// the target detail entry avoids matching posters from recommended results.
var comixInitialDataRe = regexp.MustCompile(`(?is)<script\b[^>]*\bid\s*=\s*["']initial-data["'][^>]*>(.*?)</script>`)
// kaganeImageURLRe matches the canonical compressed image route kagane's API
// publishes — the only cover URL form the extractor emits and the browser
// fetcher accepts. The URL is matched in full (scheme, host, id shape) rather
// than trusted: the value a fetcher is pointed at may have been client-
// supplied, and a headless browser is a strong SSRF primitive.
var kaganeImageURLRe = regexp.MustCompile(`^https://kagane\.to/api/v2/image/([0-9a-f-]{36})/compressed$`)
// browserOnlyCoverURL reports whether the browser sidecar is the only fetcher
// for cover bytes at imageURL. kagane's image route answers a plain fetch with
// a challenge and `cross-origin-resource-policy: same-origin`, so a TLS fetch
// would only ever retrieve a challenge page and must not be attempted
// (ADR-0007). This is the byte-fetch router's per-Site knowledge; it lives in
// the extraction module, which owns kagane's URL shapes.
func browserOnlyCoverURL(imageURL string) bool {
return kaganeImageURLRe.MatchString(imageURL)
}
// kagane's browser-fetched series response publishes cover image IDs under
// series_covers. The API's canonical compressed image route is the only URL
// form accepted by the store and browser fetcher; no rendition is guessed.
func kaganeCoverURL(body string) string {
var response struct {
SeriesCovers []struct {
ImageID string `json:"image_id"`
} `json:"series_covers"`
}
if err := json.Unmarshal([]byte(body), &response); err != nil {
return ""
}
for _, cover := range response.SeriesCovers {
// Validate the assembled URL against the same regex the browser
// fetcher enforces, so the extractor can never emit an address the
// fetch would refuse.
imageURL := "https://kagane.to/api/v2/image/" + cover.ImageID + "/compressed"
if kaganeImageURLRe.MatchString(imageURL) {
return imageURL
}
}
return ""
}
func comixCoverURL(seriesURL, body string) string {
id, ok := comixSeriesID(seriesURL)
if !ok {
return ""
}
data := comixInitialDataRe.FindStringSubmatch(body)
if data == nil {
return ""
}
var state struct {
Queries map[string]json.RawMessage `json:"queries"`
}
if err := json.Unmarshal([]byte(data[1]), &state); err != nil {
return ""
}
raw := state.Queries[`["manga","detail","`+id+`"]`]
if len(raw) == 0 {
return ""
}
var detail struct {
Poster struct {
Medium string `json:"medium"`
} `json:"poster"`
}
if err := json.Unmarshal(raw, &detail); err != nil {
return ""
}
return publishedCoverURL(detail.Poster.Medium)
}
// coverFrom reports false for unknown sites, challenge bodies, and pages with
// no usable cover. Metadata extraction keeps scanning after an empty match so
// a later published cover is not hidden by an empty tag.
func coverFrom(site, seriesURL, body string) (string, bool) {
var cover string
switch site {
case "asura", "demonic", "lightnovelworld":
cover = metaContent(body, "property", "og:image")
case "novelfull":
cover = metaContent(body, "name", "image")
case "comix":
cover = comixCoverURL(seriesURL, body)
case "kagane":
cover = kaganeCoverURL(body)
}
return cover, cover != ""
}
func metaContent(body, attrName, attrValue string) string {
for _, tag := range metaTagRe.FindAllString(body, -1) {
attrs := make(map[string]string)
for _, m := range doubleQuotedMetaAttrRe.FindAllStringSubmatch(tag, -1) {
attrs[strings.ToLower(m[1])] = m[2]
}
for _, m := range singleQuotedMetaAttrRe.FindAllStringSubmatch(tag, -1) {
attrs[strings.ToLower(m[1])] = m[2]
}
if strings.EqualFold(attrs[strings.ToLower(attrName)], attrValue) {
if cover := publishedCoverURL(attrs["content"]); cover != "" {
return cover
}
}
}
return ""
}
func publishedCoverURL(value string) string {
value = strings.TrimSpace(html.UnescapeString(value))
return strings.ReplaceAll(value, " ", "%20")
}
+113
View File
@@ -79,6 +79,119 @@ const lnwSeriesFixture = `
<a href="https://lightnovelworld.net/overgeared-chapter-9999/">Chapter 9999</a>
`
// Trimmed from https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af
// (redirected to ...-00dcbf97) on 2026-08-10.
const asuraCoverFixture = `<meta property="og:image" content="https://cdn.asurascans.com/asura-images/covers/chronicles-of-the-demon-faction.d4dcb8.webp">`
// Trimmed from https://demonicscans.org/manga/Catastrophic-Necromancer on 2026-08-10.
// The source publishes the raw space in this URL.
const demonicCoverFixture = `<meta property="og:image" content="https://readermc.org/images/thumbnails/Catastrophic Necromancer.webp">`
// Trimmed from https://comix.to/title/n8we-dungeons-and-crayons on 2026-08-10.
// The state includes a recommended poster before the target detail object and
// nested IDs inside that object; no og:image is present.
const comixCoverFixture = `<script type="application/json" id="initial-data">{"queries":{"[\"manga\",\"recommended\",\"n8we\",1]":{"poster":{"medium":"https://static.comix.to/recommended@280.jpg","large":"https://static.comix.to/recommended.jpg"}},"[\"manga\",\"detail\",\"n8we\"]":{"chapters":[{"hid":"nested"}],"poster":{"medium":"https://static.comix.to/039d/i/1/34/6a6742bf15736@280.jpg","large":"https://static.comix.to/039d/i/1/34/6a6742bf15736.jpg"}}}}</script>`
// Trimmed from GET https://kagane.to/api/v2/series/019fe11a-8670-7cf3-8343-0b02057d3787 on 2026-08-10.
const kaganeCoverFixture = `{"series_covers":[{"cover_id":"019fe11a-84d1-714b-9cf4-2827f277f3c0","language":"en","volume_number":"1","chapter_number":null,"note":null,"image_id":"019fe11a-84c3-7fc3-a84b-88787374b617"}]}`
// Trimmed from https://novelfull.com/reverend-insanity.html on 2026-08-10.
const novelfullCoverFixture = `<meta name="image" content="https://novelfull.com/uploads/webp/novel/reverend-insanity-82661d911a.webp">`
// Trimmed from https://lightnovelworld.net/novel/a-will-eternal/ on 2026-08-10.
const lnwCoverFixture = `<meta content='https://lightnovelworld.net/wp-content/uploads/2026/03/a-will-eternal-1.webp' property='og:image'>`
func TestCoverFrom(t *testing.T) {
const comixURL = "https://comix.to/title/n8we-dungeons-and-crayons"
tests := []struct {
name string
site string
seriesURL string
body string
wantOK bool
wantCover string
}{
{
name: "asura uses published metadata URL",
site: "asura", body: asuraCoverFixture, wantOK: true,
wantCover: "https://cdn.asurascans.com/asura-images/covers/chronicles-of-the-demon-faction.d4dcb8.webp",
},
{
name: "demonic escapes raw spaces",
site: "demonic", body: demonicCoverFixture, wantOK: true,
wantCover: "https://readermc.org/images/thumbnails/Catastrophic%20Necromancer.webp",
},
{
name: "comix takes target medium poster",
site: "comix", seriesURL: comixURL, body: comixCoverFixture, wantOK: true,
wantCover: "https://static.comix.to/039d/i/1/34/6a6742bf15736@280.jpg",
},
{
name: "kagane reads API cover image ID",
site: "kagane", body: kaganeCoverFixture, wantOK: true,
wantCover: "https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
},
{
name: "novelfull reads image metadata",
site: "novelfull", body: novelfullCoverFixture, wantOK: true,
wantCover: "https://novelfull.com/uploads/webp/novel/reverend-insanity-82661d911a.webp",
},
{
name: "lightnovelworld reads og image",
site: "lightnovelworld", body: lnwCoverFixture, wantOK: true,
wantCover: "https://lightnovelworld.net/wp-content/uploads/2026/03/a-will-eternal-1.webp",
},
{
name: "later metadata cover survives empty match",
site: "asura",
body: `<meta property="og:image" content="">` + asuraCoverFixture,
wantOK: true,
wantCover: "https://cdn.asurascans.com/asura-images/covers/chronicles-of-the-demon-faction.d4dcb8.webp",
},
{
name: "page without cover is empty",
site: "asura", body: `<meta property="og:title" content="No Cover">`,
},
{
name: "unknown site is empty",
site: "unknown", body: asuraCoverFixture,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
got, ok := coverFrom(tt.site, tt.seriesURL, tt.body)
if ok != tt.wantOK {
t.Fatalf("ok = %v, want %v (got %q)", ok, tt.wantOK, got)
}
if got != tt.wantCover {
t.Errorf("cover = %q, want %q", got, tt.wantCover)
}
})
}
}
func TestCoverFromChallenge(t *testing.T) {
tests := []struct {
site string
seriesURL string
}{
{"asura", "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af"},
{"demonic", "https://demonicscans.org/manga/Catastrophic-Necromancer"},
{"comix", "https://comix.to/title/n8we-dungeons-and-crayons"},
{"kagane", "https://kagane.to/series/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"},
{"novelfull", "https://novelfull.com/reverend-insanity.html"},
{"lightnovelworld", "https://lightnovelworld.net/novel/a-will-eternal/"},
}
for _, tt := range tests {
t.Run(tt.site, func(t *testing.T) {
if got, ok := coverFrom(tt.site, tt.seriesURL, challengeFixture); ok || got != "" {
t.Fatalf("cover = %q, ok = %v, want empty", got, ok)
}
})
}
}
func TestLatestChapterFrom(t *testing.T) {
const asuraURL = "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af"
const demonicURL = "https://demonicscans.org/manga/Catastrophic-Necromancer"
+78 -12
View File
@@ -6,25 +6,30 @@ import (
"os"
"testing"
"time"
"bookmarkmanager/backend/internal/store"
)
// TestSmokeKaganeImage is the live proof that the cover proxy's fetch actually
// clears Cloudflare and returns image bytes. It needs a real headless Chrome
// with outbound network, so it runs only when SMOKE_BROWSER_WS_URL is set:
// TestSmokeKaganeImage is the live proof that the acquisition path's browser
// fetch actually clears Cloudflare and returns image bytes. It needs the real
// browser unit with outbound network, so it runs only when SMOKE_BROWSER_WS_URL
// is set:
//
// docker run --rm --shm-size=1gb -p 19222:9222 chromedp/headless-shell:stable
// SMOKE_BROWSER_WS_URL=ws://127.0.0.1:19222 go test -run TestSmokeKaganeImage ./internal/latest
// cd chrome && BROWSER_BIND_ADDR=127.0.0.1 docker compose up -d --build
// SMOKE_BROWSER_WS_URL=ws://127.0.0.1:9222 go test -run TestSmokeKaganeImage ./internal/latest
//
// Not chromedp/headless-shell: its challenge never clears (see chrome/Dockerfile),
// so a red run there proves nothing about kagane.
func TestSmokeKaganeImage(t *testing.T) {
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
if ws == "" {
t.Skip("SMOKE_BROWSER_WS_URL unset")
}
const imageID = "019fe11a-84c3-7fc3-a84b-88787374b617" // SP Baby's cover
const imageURL = "https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed" // SP Baby's cover
// The same URL through a plain client is what the web UI's <img> gets.
// The same URL through a plain client is what any other fetcher would get.
// Asserting on it keeps the test honest about why the browser is needed.
req, err := http.NewRequest(http.MethodGet,
"https://kagane.to/api/v2/image/"+imageID+"/compressed", nil)
req, err := http.NewRequest(http.MethodGet, imageURL, nil)
if err != nil {
t.Fatal(err)
}
@@ -43,7 +48,7 @@ func TestSmokeKaganeImage(t *testing.T) {
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
defer cancel()
body, contentType, err := f.Image(ctx, imageID)
body, contentType, err := f.Image(ctx, imageURL)
if err != nil {
t.Fatalf("Image: %v", err)
}
@@ -59,8 +64,10 @@ func TestSmokeKaganeImage(t *testing.T) {
}
t.Logf("fetched %d bytes of %s", len(body), contentType)
if _, _, err := f.Image(ctx, "not-a-uuid"); err == nil {
t.Fatal("Image accepted a non-uuid id")
// The browser module claims only the cover URL shape it can clear a
// challenge for; anything else must be refused before any navigation.
if _, _, err := f.Image(ctx, "https://cdn.example/cover.jpg"); err == nil {
t.Fatal("Image accepted a cover URL the browser module does not claim")
}
}
@@ -89,3 +96,62 @@ func TestSmokeKaganeGet(t *testing.T) {
t.Fatalf("status = %d, want 200 — the sidecar is not clearing the challenge", status)
}
}
// TestSmokeAcquireKaganeCover proves the #62 acquisition path end to end
// against the real browser: a kagane Series bookmarked at creation gets its
// Cover, bytes fetched through the sidecar into the content-addressed store.
// Same SMOKE_BROWSER_WS_URL gate as the tests above; a red run means the
// challenge is not clearing from this IP (a live fact to re-check), not
// necessarily a defect in the pipeline.
func TestSmokeAcquireKaganeCover(t *testing.T) {
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
if ws == "" {
t.Skip("SMOKE_BROWSER_WS_URL unset")
}
const (
seriesID = "019fe11a-8670-7cf3-8343-0b02057d3787"
coverURL = "https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed"
)
s, _ := newTestStore(t)
bf, err := NewBrowserFetcher(ws)
if err != nil {
t.Fatalf("NewBrowserFetcher: %v", err)
}
defer bf.Close()
tlsF, err := NewTLSFetcher()
if err != nil {
t.Fatalf("NewTLSFetcher: %v", err)
}
acq := &Acquirer{
Store: s, Fetch: tlsF, BrowserFetch: bf,
BrowserCoverFetch: bf, Covers: NewCoverFetcher(),
}
s.OnSeriesCreated = acq.Acquire
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: "kagane:" + seriesID, Site: "kagane", SeriesID: seriesID,
Title: "smoke", SeriesURL: "https://kagane.to/series/" + seriesID, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("Upsert: %v", err)
}
acq.Wait()
got, found, err := s.Get(s.OwnerID(), "kagane:"+seriesID)
if err != nil || !found {
t.Fatalf("Get: %v found=%v", err, found)
}
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(coverURL); got.Cover != want {
t.Fatalf("Cover = %q, want %q — the acquire path did not store the browser-fetched bytes", got.Cover, want)
}
body, contentType, ok, err := s.CoverByAddress(store.CoverAddress(coverURL))
if err != nil || !ok {
t.Fatalf("CoverByAddress: %v found=%v", err, ok)
}
if len(body) < 1000 {
t.Fatalf("stored cover is %d bytes, want a real image", len(body))
}
if contentType != "image/webp" {
t.Fatalf("content type = %q, want image/webp", contentType)
}
t.Logf("stored %d bytes of %s", len(body), contentType)
}
@@ -0,0 +1,8 @@
-- Kagane cover bytes belong in their own table so image blobs never enter the
-- series queries that drive the latest-chapter poller.
CREATE TABLE covers (
image_id text PRIMARY KEY,
body bytea NOT NULL,
content_type text NOT NULL,
fetched_at timestamptz NOT NULL DEFAULT now()
);
@@ -0,0 +1,9 @@
-- Cover bytes move out of Postgres. Existing rows are intentionally dropped:
-- the old kagane path already refetches missing Covers on demand.
DROP TABLE covers;
CREATE TABLE covers (
address text PRIMARY KEY,
path text NOT NULL,
content_type text NOT NULL
);
@@ -0,0 +1,8 @@
-- The Cover splits into two facts. `cover` keeps the third-party address the
-- bytes come from, which is what the acquisition path refetches and dedupes
-- on; `cover_address` is the content address of the bytes once they are
-- actually stored, and is what the wire's absolute URL is built from.
--
-- Empty `cover_address` therefore means "no Cover yet" rather than "a Cover
-- that 404s", which is the distinction the API and the UI both depend on.
ALTER TABLE series ADD COLUMN cover_address text NOT NULL DEFAULT '';
+264 -51
View File
@@ -1,17 +1,22 @@
package store
import (
"crypto/sha256"
"database/sql"
"embed"
"encoding/hex"
"errors"
"fmt"
"io/fs"
"os"
"path"
"path/filepath"
"regexp"
"slices"
"strconv"
"strings"
"github.com/jackc/pgx/v5/pgtype"
_ "github.com/jackc/pgx/v5/stdlib"
)
@@ -25,12 +30,16 @@ import (
// between readers: progress, favourite, lifecycle bucket, updated_at. The wire
// format stays flat regardless — see ADR-0004.
type Bookmark struct {
Key string `json:"key"`
Site string `json:"site"`
SeriesID string `json:"series_id"`
Title string `json:"title"`
SeriesURL string `json:"series_url"`
Cover string `json:"cover"`
Key string `json:"key"`
Site string `json:"site"`
SeriesID string `json:"series_id"`
Title string `json:"title"`
SeriesURL string `json:"series_url"`
// Cover is the wire value: an absolute URL on this deployment's own
// origin once the bytes exist, and "" until they do — never a third-party
// address and never an address that 404s (ADR-0007). A client may still
// send this field and it is discarded on the way in; see Upsert.
Cover string `json:"cover"`
LastChapter string `json:"last_chapter"`
LastChapterNum float64 `json:"last_chapter_num"`
LastChapterURL string `json:"last_chapter_url"`
@@ -57,11 +66,16 @@ type Bookmark struct {
// bookmark's own fields. Never serialized: the wire format is the flat
// Bookmark (ADR-0004).
type Series struct {
Site string
SeriesID string
Title string
SeriesURL string
Site string
SeriesID string
Title string
SeriesURL string
// Cover is the third-party source address the bytes come from, and
// CoverAddress the content address they are stored under. A blank
// CoverAddress is what "no Cover yet" means: the poll fills it and never
// replaces a filled one (ADR-0007).
Cover string
CoverAddress string
Kind string
LatestChapter string
LatestChapterNum *float64 // nil until first captured
@@ -137,20 +151,19 @@ func (b Bookmark) Initial() string {
return "?"
}
// kaganeCoverRe matches the cover URL kagane's og:image carries, which is what
// the userscript stores for that site.
var kaganeCoverRe = regexp.MustCompile(`^https://kagane\.to/api/v2/image/([0-9a-f-]{36})/compressed$`)
// CoverURL is the src the web UI puts in an <img>. For every site but kagane
// that is Cover as stored. kagane serves its images behind a Cloudflare
// challenge *and* with `cross-origin-resource-policy: same-origin`, so no page
// on another origin can load one however it asks (verified 2026-08-08); those
// go through the backend's own proxy instead.
func (b Bookmark) CoverURL() string {
if m := kaganeCoverRe.FindStringSubmatch(b.Cover); m != nil {
return "/img/kagane/" + m[1]
// CoverContentType canonicalises a fetched response's media type and reports
// whether the bytes are safe to store and serve. comix answers "image/jpg",
// which no standard lists but browsers accept; it is stored as the real name
// rather than passed through, so one image never lands under two spellings.
func CoverContentType(contentType string) (string, bool) {
switch contentType {
case "image/jpg":
return "image/jpeg", true
case "image/webp", "image/jpeg", "image/png", "image/avif", "image/gif":
return contentType, true
default:
return "", false
}
return b.Cover
}
// Library buckets. A bookmark is in exactly one. This cannot be derived from
@@ -175,14 +188,14 @@ var migrations embed.FS
// compile-time constant; every request value is bound as a parameter. The
// series-owned fields are joined in from the series table, in scanBookmark
// order, so the flat Bookmark reads back whole despite the split (ADR-0004).
const bookmarkColumns = `b.site, b.series_id, s.title, s.series_url, s.cover,
const bookmarkColumns = `b.site, b.series_id, s.title, s.series_url, s.cover_address,
b.last_chapter, b.last_chapter_num, b.last_chapter_url,
b.favorite, s.latest_chapter, s.latest_chapter_num, b.updated_at, b.status, s.kind`
// seriesColumns is the series row in scanSeries order, used by the poller's
// due query. latest_checked_at lives only on series — see MarkLatestChecked
// for why it stays off every client-visible write.
const seriesColumns = `s.site, s.series_id, s.title, s.series_url, s.cover,
const seriesColumns = `s.site, s.series_id, s.title, s.series_url, s.cover, s.cover_address,
s.kind, s.latest_chapter, s.latest_chapter_num, s.latest_checked_at`
// Owner is the person running the service: the first Reader, seeded at startup
@@ -203,7 +216,18 @@ type Store struct {
// ownerID is the seeded owner Reader (issue #22) — the only Reader with
// administrative reach (revoking another Reader's sessions). Every store
// method takes a reader id explicitly, so ownership is never implicit.
ownerID int64
ownerID int64
coverDir string
// coverBaseURL is this deployment's public origin. Cover addresses are
// absolute because the userscript renders them on third-party origins,
// where a relative path would resolve against the Site (ADR-0007).
coverBaseURL string
// OnSeriesCreated fires once, after commit, for a Series no Reader had
// bookmarked before. It is how creation-time Cover and Latest Chapter
// acquisition is triggered without the write waiting on a third-party
// Site; nil disables it, which is what every test that does not care
// about acquisition leaves it as.
OnSeriesCreated func(Series)
}
// OwnerID returns the seeded owner Reader's id: the administrator, and the
@@ -338,8 +362,31 @@ const allMigrations = 0
// Open connects to Postgres at url — a libpq connection URL such as
// "postgres://user:pass@host:5432/bookmarks?sslmode=disable" — brings its
// schema up to date, and seeds the owner Reader.
func Open(url string, owner Owner) (*Store, error) {
// schema up to date, seeds the owner Reader, and prepares cover storage.
func Open(url string, owner Owner, coverDir, coverBaseURL string) (*Store, error) {
if strings.TrimSpace(coverDir) == "" {
return nil, errors.New("cover directory is required")
}
// Every wire Cover is this string with a path glued on, rendered by a
// userscript on a Site's own origin: anything but an absolute origin
// produces addresses no client can load, silently (ADR-0007).
base := strings.TrimRight(coverBaseURL, "/")
if host, ok := strings.CutPrefix(base, "https://"); !ok || host == "" {
if host, ok := strings.CutPrefix(base, "http://"); !ok || host == "" {
return nil, fmt.Errorf("cover base URL %q is not an absolute http(s) origin", coverBaseURL)
}
}
if err := os.MkdirAll(coverDir, 0o755); err != nil {
return nil, fmt.Errorf("create cover directory: %w", err)
}
info, err := os.Stat(coverDir)
if err != nil {
return nil, fmt.Errorf("stat cover directory: %w", err)
}
if !info.IsDir() {
return nil, fmt.Errorf("cover directory %q is not a directory", coverDir)
}
db, err := sql.Open("pgx", url)
if err != nil {
return nil, fmt.Errorf("open postgres: %w", err)
@@ -374,7 +421,9 @@ func Open(url string, owner Owner) (*Store, error) {
db.Close()
return nil, fmt.Errorf("resolve owner: %w", err)
}
return &Store{db: db, ownerID: ownerID}, nil
return &Store{
db: db, ownerID: ownerID, coverDir: coverDir, coverBaseURL: base,
}, nil
}
// seedOwner makes sure the configured owner exists as exactly one readers row.
@@ -477,18 +526,20 @@ func applyMigration(db *sql.DB, version int64, body string) error {
// scanBookmark reads one row in bookmarkColumns order. Every column is NOT
// NULL except latest_chapter_num, where NULL means "never captured" — a
// distinct state from chapter zero, and the reason for the pointer.
func scanBookmark(scan func(...any) error) (Bookmark, error) {
func (s *Store) scanBookmark(scan func(...any) error) (Bookmark, error) {
var (
b Bookmark
coverAddress string
latestChapterNum sql.NullFloat64
)
if err := scan(
&b.Site, &b.SeriesID, &b.Title, &b.SeriesURL, &b.Cover,
&b.Site, &b.SeriesID, &b.Title, &b.SeriesURL, &coverAddress,
&b.LastChapter, &b.LastChapterNum, &b.LastChapterURL,
&b.Favorite, &b.LatestChapter, &latestChapterNum, &b.UpdatedAt, &b.Status, &b.Kind,
); err != nil {
return Bookmark{}, err
}
b.Cover = s.CoverWireURL(coverAddress)
if latestChapterNum.Valid {
b.LatestChapterNum = &latestChapterNum.Float64
}
@@ -513,7 +564,7 @@ func scanSeries(scan func(...any) error) (Series, error) {
latestChapterNum sql.NullFloat64
)
if err := scan(
&sr.Site, &sr.SeriesID, &sr.Title, &sr.SeriesURL, &sr.Cover,
&sr.Site, &sr.SeriesID, &sr.Title, &sr.SeriesURL, &sr.Cover, &sr.CoverAddress,
&sr.Kind, &sr.LatestChapter, &latestChapterNum, &sr.LatestCheckedAt,
&sr.readerCount,
); err != nil {
@@ -528,6 +579,146 @@ func scanSeries(scan func(...any) error) (Series, error) {
// Close releases the underlying database handle.
func (s *Store) Close() error { return s.db.Close() }
func coverSourceAddress(sourceURL string) string {
sum := sha256.Sum256([]byte(sourceURL))
return hex.EncodeToString(sum[:])
}
func coverRelativePath(address string) string {
return address[:2] + "/" + address[2:4] + "/" + address
}
func (s *Store) getCover(sourceURL string) ([]byte, string, bool, error) {
return s.getCoverByAddress(coverSourceAddress(sourceURL))
}
func (s *Store) getCoverByAddress(address string) ([]byte, string, bool, error) {
var relativePath, contentType string
err := s.db.QueryRow(
`SELECT path, content_type FROM covers WHERE address = $1`, address,
).Scan(&relativePath, &contentType)
if errors.Is(err, sql.ErrNoRows) {
return nil, "", false, nil
}
if err != nil {
return nil, "", false, fmt.Errorf("get cover %q: %w", address, err)
}
expectedPath := coverRelativePath(address)
if relativePath != expectedPath {
return nil, "", false, fmt.Errorf("cover %q has unexpected path %q", address, relativePath)
}
body, err := os.ReadFile(filepath.Join(s.coverDir, filepath.FromSlash(relativePath)))
if errors.Is(err, fs.ErrNotExist) {
return nil, "", false, nil
}
if err != nil {
return nil, "", false, fmt.Errorf("read cover %q: %w", address, err)
}
return body, contentType, true, nil
}
func (s *Store) putCover(sourceURL string, body []byte, contentType string) error {
stored, ok := CoverContentType(contentType)
if !ok {
return fmt.Errorf("put cover %q: unsupported content type %q", sourceURL, contentType)
}
contentType = stored
address := coverSourceAddress(sourceURL)
relativePath := coverRelativePath(address)
coverPath := filepath.Join(s.coverDir, filepath.FromSlash(relativePath))
if err := os.MkdirAll(filepath.Dir(coverPath), 0o755); err != nil {
return fmt.Errorf("create cover shard: %w", err)
}
tmp, err := os.CreateTemp(filepath.Dir(coverPath), ".cover-*")
if err != nil {
return fmt.Errorf("create cover temp file: %w", err)
}
tmpName := tmp.Name()
defer os.Remove(tmpName)
if _, err := tmp.Write(body); err != nil {
tmp.Close()
return fmt.Errorf("write cover temp file: %w", err)
}
if err := tmp.Sync(); err != nil {
tmp.Close()
return fmt.Errorf("sync cover temp file: %w", err)
}
if err := tmp.Close(); err != nil {
return fmt.Errorf("close cover temp file: %w", err)
}
if err := os.Link(tmpName, coverPath); err != nil && !errors.Is(err, fs.ErrExist) {
return fmt.Errorf("install cover file: %w", err)
}
if _, err := s.db.Exec(`
INSERT INTO covers (address, path, content_type)
VALUES ($1, $2, $3)
ON CONFLICT (address) DO NOTHING`, address, relativePath, contentType); err != nil {
return fmt.Errorf("record cover %q: %w", address, err)
}
return nil
}
// GetCover returns the immutable object addressed by its source URL. Missing
// files are reported with ok=false so callers can retry acquisition later.
func (s *Store) GetCover(sourceURL string) ([]byte, string, bool, error) {
return s.getCover(sourceURL)
}
// PutCover persists bytes under the source URL's content address. A later
// write for the same URL cannot replace the immutable object.
func (s *Store) PutCover(sourceURL string, body []byte, contentType string) error {
return s.putCover(sourceURL, body, contentType)
}
// CoverAddress is the content address bytes fetched from sourceURL are stored
// under. It is a pure function of the URL, so the acquisition path can name a
// Cover before it has the bytes.
func CoverAddress(sourceURL string) string { return coverSourceAddress(sourceURL) }
// coverAddressRe is the shape of a stored address: the hex SHA-256 of a source
// URL. Request paths reach CoverByAddress, so the shape is checked before the
// value is ever turned into a filesystem path.
var coverAddressRe = regexp.MustCompile(`^[0-9a-f]{64}$`)
// CoverByAddress returns the immutable object at one content address. An
// address that is not a stored one - malformed, unknown, or recorded but with
// its file gone - is reported with ok=false rather than as an error.
func (s *Store) CoverByAddress(address string) ([]byte, string, bool, error) {
if !coverAddressRe.MatchString(address) {
return nil, "", false, nil
}
return s.getCoverByAddress(address)
}
// CoverWireURL is the absolute URL a client renders for a stored Cover, and ""
// for a Series that has none yet. A blank is a real state, not a placeholder
// address: it is what tells both clients to draw their own fallback instead of
// requesting bytes that do not exist (ADR-0007).
func (s *Store) CoverWireURL(address string) string {
if address == "" {
return ""
}
return s.coverBaseURL + "/covers/" + address
}
// SetSeriesCover stores the bytes and points the Series at them, but only
// while the Series has no Cover: acquisition at creation and the poll both
// call this, and whichever arrives second must not overwrite the first. The
// bytes themselves are content-addressed and immutable, so storing them twice
// is free.
func (s *Store) SetSeriesCover(site, seriesID, sourceURL string, body []byte, contentType string) error {
if err := s.putCover(sourceURL, body, contentType); err != nil {
return err
}
if _, err := s.db.Exec(`
UPDATE series SET cover = $3, cover_address = $4
WHERE site = $1 AND series_id = $2 AND cover_address = ''`,
site, seriesID, sourceURL, coverSourceAddress(sourceURL)); err != nil {
return fmt.Errorf("set cover for %q: %w", site+":"+seriesID, err)
}
return nil
}
// List returns every bookmark of one reader, newest activity first.
// Series-owned fields are joined in, so each Bookmark reads back whole and
// flat (ADR-0004).
@@ -544,7 +735,7 @@ func (s *Store) List(readerID int64) ([]Bookmark, error) {
out := []Bookmark{}
for rows.Next() {
b, err := scanBookmark(rows.Scan)
b, err := s.scanBookmark(rows.Scan)
if err != nil {
return nil, fmt.Errorf("scan bookmark: %w", err)
}
@@ -561,7 +752,7 @@ func (s *Store) Get(readerID int64, key string) (Bookmark, bool, error) {
if !ok {
return Bookmark{}, false, nil
}
b, err := scanBookmark(s.db.QueryRow(
b, err := s.scanBookmark(s.db.QueryRow(
`SELECT `+bookmarkColumns+` FROM bookmarks b
JOIN series s ON s.site = b.site AND s.series_id = b.series_id
WHERE b.reader_id = $1 AND b.site = $2 AND b.series_id = $3`,
@@ -618,18 +809,28 @@ func (s *Store) Upsert(readerID int64, b Bookmark) (Bookmark, error) {
// The ::text casts are load-bearing: inside COALESCE/NULLIF there is no
// target column to infer the parameter type from, and Postgres rejects the
// statement rather than guessing.
if _, err := tx.Exec(`
INSERT INTO series (site, series_id, title, series_url, cover, kind,
//
// The cover columns are absent on purpose: the Cover is acquired
// server-side (ADR-0007), so a client-supplied one is not written even
// when the row is brand new.
//
// xmax is zero only on a row this statement inserted, which is how a
// Series nobody had bookmarked before is told apart from one that already
// existed — DO UPDATE returns a row either way.
var created bool
if err := tx.QueryRow(`
INSERT INTO series (site, series_id, title, series_url, kind,
latest_chapter, latest_chapter_num)
VALUES ($1, $2, $3, $4, $5,
COALESCE(NULLIF($6::text, ''), (SELECT kind FROM series WHERE site = $1 AND series_id = $2), 'manga'),
$7, $8)
VALUES ($1, $2, $3, $4,
COALESCE(NULLIF($5::text, ''), (SELECT kind FROM series WHERE site = $1 AND series_id = $2), 'manga'),
$6, $7)
ON CONFLICT (site, series_id) DO UPDATE SET
kind=excluded.kind,
latest_chapter=excluded.latest_chapter,
latest_chapter_num=excluded.latest_chapter_num`,
b.Site, b.SeriesID, b.Title, b.SeriesURL, b.Cover, b.Kind,
b.LatestChapter, latestNum); err != nil {
latest_chapter_num=excluded.latest_chapter_num
RETURNING xmax = 0`,
b.Site, b.SeriesID, b.Title, b.SeriesURL, b.Kind,
b.LatestChapter, latestNum).Scan(&created); err != nil {
return Bookmark{}, fmt.Errorf("upsert series for %q: %w", b.Key, err)
}
@@ -659,7 +860,7 @@ func (s *Store) Upsert(readerID int64, b Bookmark) (Bookmark, error) {
return Bookmark{}, fmt.Errorf("upsert %q: %w", b.Key, err)
}
stored, err := scanBookmark(tx.QueryRow(
stored, err := s.scanBookmark(tx.QueryRow(
`SELECT `+bookmarkColumns+` FROM bookmarks b
JOIN series s ON s.site = b.site AND s.series_id = b.series_id
WHERE b.reader_id = $1 AND b.site = $2 AND b.series_id = $3`,
@@ -670,6 +871,14 @@ func (s *Store) Upsert(readerID int64, b Bookmark) (Bookmark, error) {
if err := tx.Commit(); err != nil {
return Bookmark{}, fmt.Errorf("commit %q: %w", b.Key, err)
}
// After commit, never inside the transaction: the hook reaches a
// third-party Site, and the Reader's write must not wait on it.
if created && s.OnSeriesCreated != nil {
s.OnSeriesCreated(Series{
Site: b.Site, SeriesID: b.SeriesID, Title: stored.Title,
SeriesURL: stored.SeriesURL, Kind: stored.Kind,
})
}
return stored, nil
}
@@ -689,16 +898,17 @@ func (s *Store) Delete(readerID int64, key string) error {
}
// DueForLatestCheck returns series whose server-side latest-chapter check has
// aged past cutoffMs, ordered by how many bookmarks reference them (descending)
// then least-recently-checked first, at most limit of them.
// aged past the appropriate cutoff, ordered by how many bookmarks reference
// them (descending) then least-recently-checked first, at most limit of them.
// Browser-backed sites use browserCutoffMs; every other site uses cutoffMs.
//
// The reader_count ordering is the point of the split (ADR-0003): a series
// shared by several readers is fetched once per due cycle, and the popular
// ones stay freshest while the long tail absorbs any shortfall. Within one
// reader count, oldest-first keeps the poll fair when the backlog outgrows its
// reader count, oldest-first keeps the poll fair when the backlog outgrows
// throughput: the most neglected series is always next, so a large collection
// refreshes uniformly slower rather than leaving a tail that never refreshes at
// all. The userscript sorts its own queue the same way (L453).
// refreshes uniformly slower rather than leaving a tail that never refreshes
// at all. The userscript sorts its own queue the same way (L453).
//
// Series with no series_url are skipped — there is nothing to fetch, which is
// the same filter the userscript applies at L452. Series whose only bookmarks
@@ -706,17 +916,20 @@ func (s *Store) Delete(readerID int64, key string) error {
// burns requests. Archived bookmarks still count — knowing what a shelved
// series is up to is the whole reason for archiving instead of deleting.
// A series with no bookmarks at all never appears: the join excludes it.
func (s *Store) DueForLatestCheck(cutoffMs int64, limit int) ([]Series, error) {
func (s *Store) DueForLatestCheck(cutoffMs, browserCutoffMs int64, browserSites []string, limit int) ([]Series, error) {
rows, err := s.db.Query(`SELECT `+seriesColumns+`, COUNT(*) AS reader_count
FROM series s
JOIN bookmarks b ON b.site = s.site AND b.series_id = s.series_id
WHERE s.series_url <> ''
AND s.latest_checked_at <= $1
AND s.latest_checked_at <= CASE
WHEN s.site = ANY($3::text[]) THEN $2::bigint
ELSE $1::bigint
END
GROUP BY s.site, s.series_id, s.title, s.series_url, s.cover,
s.kind, s.latest_chapter, s.latest_chapter_num, s.latest_checked_at
HAVING COUNT(*) FILTER (WHERE b.status <> 'finished') > 0
ORDER BY reader_count DESC, s.latest_checked_at ASC
LIMIT $2`, cutoffMs, limit)
LIMIT $4`, cutoffMs, browserCutoffMs, pgtype.FlatArray[string](browserSites), limit)
if err != nil {
return nil, fmt.Errorf("query due series: %w", err)
}
+298 -64
View File
@@ -4,7 +4,9 @@ import (
"bytes"
"crypto/sha256"
"database/sql"
"encoding/hex"
"os"
"path/filepath"
"strconv"
"strings"
"testing"
@@ -19,9 +21,12 @@ func TestMain(m *testing.M) { os.Exit(pgtest.Main(m)) }
// reader register one (see secondReader).
var testOwner = Owner{DiscordID: "test-owner", TokenHash: sha256.Sum256([]byte("owner-token-hash"))}
// testCoverBaseURL is the public origin every stored cover URL is built from.
const testCoverBaseURL = "https://bookmarks.test"
func newTestStore(t *testing.T) *Store {
t.Helper()
store, err := Open(pgtest.URL(t), testOwner)
store, err := Open(pgtest.URL(t), testOwner, t.TempDir(), testCoverBaseURL)
if err != nil {
t.Fatalf("Open: %v", err)
}
@@ -46,7 +51,8 @@ func secondReader(t *testing.T, s *Store) int64 {
// error, and must leave the rows alone.
func TestOpenIsIdempotent(t *testing.T) {
url := pgtest.URL(t)
first, err := Open(url, testOwner)
coverDir := t.TempDir()
first, err := Open(url, testOwner, coverDir, testCoverBaseURL)
if err != nil {
t.Fatalf("Open: %v", err)
}
@@ -57,7 +63,7 @@ func TestOpenIsIdempotent(t *testing.T) {
}
first.Close()
second, err := Open(url, testOwner)
second, err := Open(url, testOwner, coverDir, testCoverBaseURL)
if err != nil {
t.Fatalf("reopen: %v", err)
}
@@ -109,7 +115,8 @@ func TestReaderTokenInfo(t *testing.T) {
// epoch-0 hash only while the row has never been rotated.
func TestRotateTokenInvalidatesOldAndSurvivesRestart(t *testing.T) {
url := pgtest.URL(t)
store, err := Open(url, testOwner)
coverDir := t.TempDir()
store, err := Open(url, testOwner, coverDir, testCoverBaseURL)
if err != nil {
t.Fatalf("Open: %v", err)
}
@@ -141,7 +148,7 @@ func TestRotateTokenInvalidatesOldAndSurvivesRestart(t *testing.T) {
}
store.Close()
reopened, err := Open(url, testOwner)
reopened, err := Open(url, testOwner, coverDir, testCoverBaseURL)
if err != nil {
t.Fatalf("reopen: %v", err)
}
@@ -291,7 +298,7 @@ func TestDueForLatestCheck(t *testing.T) {
s := newTestStore(t)
seedForCheck(t, s, "asura:x", tt.seriesURL, tt.checkedAt)
due, err := s.DueForLatestCheck(now-hour, 10)
due, err := s.DueForLatestCheck(now-hour, now-hour, nil, 10)
if err != nil {
t.Fatalf("DueForLatestCheck: %v", err)
}
@@ -309,7 +316,7 @@ func TestDueForLatestCheckOldestFirstAndLimited(t *testing.T) {
seedForCheck(t, s, "asura:b", "https://asurascans.com/comics/b", 200)
seedForCheck(t, s, "asura:a", "https://asurascans.com/comics/a", 100)
due, err := s.DueForLatestCheck(1000, 2)
due, err := s.DueForLatestCheck(1000, 1000, nil, 2)
if err != nil {
t.Fatalf("DueForLatestCheck: %v", err)
}
@@ -489,7 +496,7 @@ func TestDueForLatestCheckSkipsFinishedKeepsArchived(t *testing.T) {
}
}
due, err := store.DueForLatestCheck(time.Now().UnixMilli(), 10)
due, err := store.DueForLatestCheck(time.Now().UnixMilli(), time.Now().UnixMilli(), nil, 10)
if err != nil {
t.Fatalf("DueForLatestCheck: %v", err)
}
@@ -543,38 +550,6 @@ func TestDisplayChapter(t *testing.T) {
}
}
func TestCoverURL(t *testing.T) {
cases := []struct {
name string
cover string
want string
}{
{
"kagane routes through the proxy",
"https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
"/img/kagane/019fe11a-84c3-7fc3-a84b-88787374b617",
},
{
"another site is served as stored",
"https://gg.asuracomic.net/storage/media/1/conversions/cover.webp",
"https://gg.asuracomic.net/storage/media/1/conversions/cover.webp",
},
{
"a lookalike host is not rewritten",
"https://evil.example/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
"https://evil.example/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
},
{"no cover stays empty", "", ""},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
if got := (Bookmark{Cover: tc.cover}).CoverURL(); got != tc.want {
t.Errorf("CoverURL() = %q, want %q", got, tc.want)
}
})
}
}
func TestUpsertKindDefaultsToManga(t *testing.T) {
store := newTestStore(t)
got, err := store.Upsert(store.OwnerID(), Bookmark{
@@ -668,7 +643,7 @@ func TestMigration0002BackfillsExistingBookmarks(t *testing.T) {
// Bring it current through the production path: Open runs the schema to
// 0003, seeds the owner, then applies 0004 which attaches this row. 0002
// must have backfilled the series row, not lost data.
st, err := Open(url, testOwner)
st, err := Open(url, testOwner, t.TempDir(), testCoverBaseURL)
if err != nil {
t.Fatalf("Open after migrate: %v", err)
}
@@ -697,6 +672,48 @@ func TestMigration0002BackfillsExistingBookmarks(t *testing.T) {
}
}
func TestMigration0008DropsLegacyCoverRows(t *testing.T) {
db, err := sql.Open("pgx", pgtest.URL(t))
if err != nil {
t.Fatalf("open: %v", err)
}
defer db.Close()
if err := migrate(db, readersMigration); err != nil {
t.Fatalf("migrate readers: %v", err)
}
if err := seedOwner(db, testOwner); err != nil {
t.Fatalf("seed owner: %v", err)
}
if err := migrate(db, 7); err != nil {
t.Fatalf("migrate legacy covers: %v", err)
}
if _, err := db.Exec(`
INSERT INTO covers (image_id, body, content_type)
VALUES ('legacy-image', 'legacy-bytes', 'image/jpeg')`); err != nil {
t.Fatalf("seed legacy cover: %v", err)
}
if err := migrate(db, 0); err != nil {
t.Fatalf("migrate filesystem covers: %v", err)
}
var count int
if err := db.QueryRow(`SELECT count(*) FROM covers`).Scan(&count); err != nil {
t.Fatalf("count covers: %v", err)
}
if count != 0 {
t.Fatalf("legacy covers = %d, want 0", count)
}
var bodyColumn int
if err := db.QueryRow(`SELECT count(*) FROM information_schema.columns
WHERE table_name = 'covers' AND column_name = 'body'`).Scan(&bodyColumn); err != nil {
t.Fatalf("cover columns: %v", err)
}
if bodyColumn != 0 {
t.Fatal("legacy covers table still has body column")
}
}
// readSeries reads the series row directly, for asserting on what Upsert
// actually stored rather than what the joined Bookmark reports.
func readSeries(t *testing.T, s *Store, site, seriesID string) Series {
@@ -710,39 +727,143 @@ func readSeries(t *testing.T, s *Store, site, seriesID string) Series {
return sr
}
// The first PUT for a series creates its row from the client's title, cover
// and URL — there is no other source for them (ADR-0003).
// The first PUT for a series creates its row from the client's title and URL —
// there is no other source for them (ADR-0003). The Cover is not among them:
// it is acquired server-side, so a client-supplied one is dropped even on a
// brand-new row (ADR-0007).
func TestUpsertCreatesSeriesFromClient(t *testing.T) {
store := newTestStore(t)
if _, err := store.Upsert(store.OwnerID(), Bookmark{
stored, err := store.Upsert(store.OwnerID(), Bookmark{
Key: "asura:solo", Site: "asura", SeriesID: "solo",
Title: "Solo Leveling", SeriesURL: "https://asurascans.com/comics/solo",
Cover: "https://asurascans.com/covers/solo.jpg", Kind: KindManga,
UpdatedAt: 1000,
}); err != nil {
})
if err != nil {
t.Fatalf("Upsert: %v", err)
}
if stored.Cover != "" {
t.Fatalf("Cover = %q, want empty — a client cover is never stored", stored.Cover)
}
sr := readSeries(t, store, "asura", "solo")
if sr.Title != "Solo Leveling" || sr.SeriesURL != "https://asurascans.com/comics/solo" ||
sr.Cover != "https://asurascans.com/covers/solo.jpg" {
t.Fatalf("series = %+v, want client title/url/cover stored", sr)
if sr.Title != "Solo Leveling" || sr.SeriesURL != "https://asurascans.com/comics/solo" {
t.Fatalf("series = %+v, want client title/url stored", sr)
}
if sr.Cover != "" {
t.Fatalf("series cover = %q, want empty", sr.Cover)
}
}
// A PUT naming an existing series must not overwrite its title, cover or URL:
// the row is shared, and those values are scraped page content (ADR-0003).
// The hook is what starts creation-time acquisition, so it must fire exactly
// once per Series — on the PUT that created it, and on no later one, whichever
// Reader sends it.
func TestOnSeriesCreatedFiresOnceForANewSeries(t *testing.T) {
store := newTestStore(t)
var created []Series
store.OnSeriesCreated = func(sr Series) { created = append(created, sr) }
b := Bookmark{
Key: "comix:solo", Site: "comix", SeriesID: "solo", Title: "Solo Leveling",
SeriesURL: "https://comix.to/series/solo", Kind: KindManga, UpdatedAt: 1000,
}
if _, err := store.Upsert(store.OwnerID(), b); err != nil {
t.Fatalf("Upsert: %v", err)
}
b.LastChapterNum = 12
b.UpdatedAt = 2000
if _, err := store.Upsert(store.OwnerID(), b); err != nil {
t.Fatalf("second Upsert: %v", err)
}
if _, err := store.Upsert(secondReader(t, store), b); err != nil {
t.Fatalf("second reader Upsert: %v", err)
}
if len(created) != 1 {
t.Fatalf("hook fired %d times, want 1: %+v", len(created), created)
}
if created[0].Site != "comix" || created[0].SeriesID != "solo" ||
created[0].SeriesURL != "https://comix.to/series/solo" {
t.Fatalf("hook got %+v, want the created series' identity and URL", created[0])
}
}
// Acquisition at creation and the poll both write covers, and whichever
// arrives second must leave the first one alone: a Cover is replaced by
// nothing short of the series row being rebuilt.
func TestSetSeriesCoverDoesNotOverwrite(t *testing.T) {
store := newTestStore(t)
if _, err := store.Upsert(store.OwnerID(), Bookmark{
Key: "asura:solo", Site: "asura", SeriesID: "solo", UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
first := "https://asurascans.com/covers/first.jpg"
if err := store.SetSeriesCover("asura", "solo", first, []byte("first"), "image/jpeg"); err != nil {
t.Fatalf("SetSeriesCover: %v", err)
}
if err := store.SetSeriesCover("asura", "solo", "https://asurascans.com/covers/second.jpg",
[]byte("second"), "image/jpeg"); err != nil {
t.Fatalf("second SetSeriesCover: %v", err)
}
got, ok, err := store.Get(store.OwnerID(), "asura:solo")
if err != nil || !ok {
t.Fatalf("Get = %v, %v", ok, err)
}
if want := "https://bookmarks.test/covers/" + CoverAddress(first); got.Cover != want {
t.Fatalf("Cover = %q, want the first one %q", got.Cover, want)
}
}
// The address comes straight off a public request path, so anything that is
// not a stored address must be a miss rather than a filesystem lookup.
func TestCoverByAddress(t *testing.T) {
store := newTestStore(t)
if _, err := store.Upsert(store.OwnerID(), Bookmark{
Key: "asura:solo", Site: "asura", SeriesID: "solo", UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
source := "https://asurascans.com/covers/solo.jpg"
if err := store.SetSeriesCover("asura", "solo", source, []byte("bytes"), "image/jpeg"); err != nil {
t.Fatalf("SetSeriesCover: %v", err)
}
body, contentType, ok, err := store.CoverByAddress(CoverAddress(source))
if err != nil || !ok {
t.Fatalf("CoverByAddress = %v, %v", ok, err)
}
if string(body) != "bytes" || contentType != "image/jpeg" {
t.Fatalf("CoverByAddress = %q, %q, want the stored bytes", body, contentType)
}
for _, address := range []string{"", "../../etc/passwd", "ZZ" + CoverAddress(source)[2:],
CoverAddress("never stored")} {
_, _, ok, err := store.CoverByAddress(address)
if err != nil || ok {
t.Fatalf("CoverByAddress(%q) = %v, %v, want a clean miss", address, ok, err)
}
}
}
// A PUT naming an existing series must not overwrite its title or URL: the row
// is shared, and those values are scraped page content (ADR-0003). An acquired
// Cover is likewise untouched by any client.
func TestUpsertExistingSeriesIgnoresClientTitleCoverURL(t *testing.T) {
store := newTestStore(t)
base := Bookmark{
Key: "asura:solo", Site: "asura", SeriesID: "solo",
Title: "Solo Leveling", SeriesURL: "https://asurascans.com/comics/solo",
Cover: "https://asurascans.com/covers/solo.jpg", LastChapterNum: 10,
UpdatedAt: 1000,
LastChapterNum: 10, UpdatedAt: 1000,
}
if _, err := store.Upsert(store.OwnerID(), base); err != nil {
t.Fatalf("seed: %v", err)
}
acquired := "https://asurascans.com/covers/solo.jpg"
if err := store.SetSeriesCover("asura", "solo", acquired, []byte("bytes"), "image/jpeg"); err != nil {
t.Fatalf("SetSeriesCover: %v", err)
}
// Same series, hostile/compromised values, real progress advance.
base.Title = "Scraped Rename"
@@ -753,8 +874,9 @@ func TestUpsertExistingSeriesIgnoresClientTitleCoverURL(t *testing.T) {
if err != nil {
t.Fatalf("Upsert: %v", err)
}
wantCover := "https://bookmarks.test/covers/" + CoverAddress(acquired)
if got.Title != "Solo Leveling" || got.SeriesURL != "https://asurascans.com/comics/solo" ||
got.Cover != "https://asurascans.com/covers/solo.jpg" {
got.Cover != wantCover {
t.Fatalf("stored = %+v, want original title/url/cover kept", got)
}
if got.LastChapterNum != 11 {
@@ -790,16 +912,19 @@ func TestUpsertExistingSeriesAcceptsKindAndLatest(t *testing.T) {
}
// Deleting the last bookmark must leave the series row behind, so a later
// re-bookmark shows title and cover immediately instead of waiting for a poll.
// re-bookmark shows title and cover immediately instead of re-acquiring them.
func TestDeleteKeepsSeriesRow(t *testing.T) {
store := newTestStore(t)
if _, err := store.Upsert(store.OwnerID(), Bookmark{
Key: "asura:solo", Site: "asura", SeriesID: "solo",
Title: "Solo Leveling", Cover: "https://asurascans.com/covers/solo.jpg",
UpdatedAt: 1000,
Title: "Solo Leveling", UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
acquired := "https://asurascans.com/covers/solo.jpg"
if err := store.SetSeriesCover("asura", "solo", acquired, []byte("bytes"), "image/jpeg"); err != nil {
t.Fatalf("SetSeriesCover: %v", err)
}
if err := store.Delete(store.OwnerID(), "asura:solo"); err != nil {
t.Fatalf("Delete: %v", err)
}
@@ -817,7 +942,8 @@ func TestDeleteKeepsSeriesRow(t *testing.T) {
if err != nil {
t.Fatalf("re-upsert: %v", err)
}
if stored.Title != "Solo Leveling" || stored.Cover != "https://asurascans.com/covers/solo.jpg" {
wantCover := "https://bookmarks.test/covers/" + CoverAddress(acquired)
if stored.Title != "Solo Leveling" || stored.Cover != wantCover {
t.Fatalf("re-bookmark = %+v, want title/cover from the surviving series row", stored)
}
}
@@ -844,7 +970,7 @@ func TestDueForLatestCheckOrdersByReaderCountThenAge(t *testing.T) {
seedSecondReader(t, s, "asura:pop:2", "asura", "pop", 1001)
seedForCheck(t, s, "asura:solo", "https://asurascans.com/comics/solo", 100)
due, err := s.DueForLatestCheck(1000, 10)
due, err := s.DueForLatestCheck(1000, 1000, nil, 10)
if err != nil {
t.Fatalf("DueForLatestCheck: %v", err)
}
@@ -870,7 +996,7 @@ func TestDueForLatestCheckExcludesOrphanSeries(t *testing.T) {
t.Fatalf("seed orphan series: %v", err)
}
due, err := s.DueForLatestCheck(1000, 10)
due, err := s.DueForLatestCheck(1000, 1000, nil, 10)
if err != nil {
t.Fatalf("DueForLatestCheck: %v", err)
}
@@ -893,14 +1019,15 @@ func TestDueForLatestCheckExcludesOrphanSeries(t *testing.T) {
// rotations.
func TestSeedOwnerIdempotentAndRefreshesTokenHash(t *testing.T) {
url := pgtest.URL(t)
first, err := Open(url, Owner{DiscordID: "owner", TokenHash: sha256.Sum256([]byte("hash-v1"))})
coverDir := t.TempDir()
first, err := Open(url, Owner{DiscordID: "owner", TokenHash: sha256.Sum256([]byte("hash-v1"))}, coverDir, testCoverBaseURL)
if err != nil {
t.Fatalf("Open: %v", err)
}
ownerID := first.OwnerID()
first.Close()
second, err := Open(url, Owner{DiscordID: "owner", TokenHash: sha256.Sum256([]byte("hash-v2"))})
second, err := Open(url, Owner{DiscordID: "owner", TokenHash: sha256.Sum256([]byte("hash-v2"))}, coverDir, testCoverBaseURL)
if err != nil {
t.Fatalf("reopen: %v", err)
}
@@ -957,7 +1084,7 @@ func TestMigration0004AttachesBookmarksToOwner(t *testing.T) {
t.Fatalf("migrate to 0002: %v", err)
}
st, err := Open(url, testOwner)
st, err := Open(url, testOwner, t.TempDir(), testCoverBaseURL)
if err != nil {
t.Fatalf("Open: %v", err)
}
@@ -1207,7 +1334,7 @@ func TestTwoReadersShareOneSeriesWithIndependentProgress(t *testing.T) {
t.Fatalf("series rows = %d, want 1 shared row for two bookmarks", series)
}
due, err := s.DueForLatestCheck(time.Now().UnixMilli(), 10)
due, err := s.DueForLatestCheck(time.Now().UnixMilli(), time.Now().UnixMilli(), nil, 10)
if err != nil {
t.Fatalf("DueForLatestCheck: %v", err)
}
@@ -1223,7 +1350,7 @@ func TestTwoReadersShareOneSeriesWithIndependentProgress(t *testing.T) {
if b, ok, err := s.Get(s.OwnerID(), "asura:solo"); err != nil || !ok || b.LastChapterNum != 200 {
t.Fatalf("owner's bookmark after the other's delete = %+v ok=%v err=%v, want it intact", b, ok, err)
}
due, err = s.DueForLatestCheck(time.Now().UnixMilli(), 10)
due, err = s.DueForLatestCheck(time.Now().UnixMilli(), time.Now().UnixMilli(), nil, 10)
if err != nil {
t.Fatalf("DueForLatestCheck after delete: %v", err)
}
@@ -1231,3 +1358,110 @@ func TestTwoReadersShareOneSeriesWithIndependentProgress(t *testing.T) {
t.Fatalf("due after one Reader left = %+v, want the series still polled", due)
}
}
func TestCoverPersistsAcrossReopen(t *testing.T) {
url := pgtest.URL(t)
coverDir := t.TempDir()
first, err := Open(url, testOwner, coverDir, testCoverBaseURL)
if err != nil {
t.Fatalf("Open: %v", err)
}
body := []byte("stored-cover")
const sourceURL = "https://cdn.example/covers/series.jpg"
if err := first.PutCover(sourceURL, body, "image/webp"); err != nil {
t.Fatalf("PutCover: %v", err)
}
if err := first.Close(); err != nil {
t.Fatalf("close first store: %v", err)
}
second, err := Open(url, testOwner, coverDir, testCoverBaseURL)
if err != nil {
t.Fatalf("reopen: %v", err)
}
defer second.Close()
got, contentType, ok, err := second.GetCover(sourceURL)
if err != nil {
t.Fatalf("GetCover: %v", err)
}
if !ok || !bytes.Equal(got, body) || contentType != "image/webp" {
t.Fatalf("stored cover = (%q, %q, %v), want (%q, image/webp, true)", got, contentType, ok, body)
}
}
func TestOpenRequiresCoverDirectory(t *testing.T) {
if _, err := Open(pgtest.URL(t), testOwner, "", testCoverBaseURL); err == nil || !strings.Contains(err.Error(), "cover directory is required") {
t.Fatalf("Open without cover directory = %v, want required-directory error", err)
}
}
// A base URL without a scheme reads like a hostname and starts cleanly, but
// every Cover it puts on the wire is an address no browser can resolve.
func TestOpenRequiresAbsoluteCoverBaseURL(t *testing.T) {
for _, base := range []string{"", "bookmarks.test", "https://", "ftp://bookmarks.test"} {
if _, err := Open(pgtest.URL(t), testOwner, t.TempDir(), base); err == nil ||
!strings.Contains(err.Error(), "absolute http(s) origin") {
t.Fatalf("Open with base %q = %v, want absolute-origin error", base, err)
}
}
}
func TestCoverIsContentAddressedOnFilesystem(t *testing.T) {
url := pgtest.URL(t)
coverDir := t.TempDir()
first, err := Open(url, testOwner, coverDir, testCoverBaseURL)
if err != nil {
t.Fatalf("Open: %v", err)
}
body := []byte("stored-cover")
const sourceURL = "https://cdn.example/covers/series.jpg"
if err := first.PutCover(sourceURL, body, "image/webp"); err != nil {
first.Close()
t.Fatalf("PutCover: %v", err)
}
defer first.Close()
addressBytes := sha256.Sum256([]byte(sourceURL))
address := hex.EncodeToString(addressBytes[:])
wantPath := filepath.Join(address[:2], address[2:4], address)
var gotPath, contentType string
if err := first.db.QueryRow(`SELECT path, content_type FROM covers WHERE address = $1`, address).Scan(&gotPath, &contentType); err != nil {
t.Fatalf("cover row: %v", err)
}
if gotPath != wantPath || contentType != "image/webp" {
t.Fatalf("cover row = (%q, %q), want (%q, image/webp)", gotPath, contentType, wantPath)
}
if got, err := os.ReadFile(filepath.Join(coverDir, gotPath)); err != nil || !bytes.Equal(got, body) {
t.Fatalf("cover file = (%q, %v), want (%q, nil)", got, err, body)
}
var bodyColumn int
if err := first.db.QueryRow(`SELECT count(*) FROM information_schema.columns WHERE table_name = 'covers' AND column_name = 'body'`).Scan(&bodyColumn); err != nil {
t.Fatalf("cover columns: %v", err)
}
if bodyColumn != 0 {
t.Fatalf("covers still has body column")
}
}
func TestCoverStoreAcceptsAnySourceURL(t *testing.T) {
s := newTestStore(t)
const sourceURL = "https://cdn.example/covers/series.jpg"
want := []byte("cover-bytes")
if err := s.PutCover(sourceURL, want, "image/jpeg"); err != nil {
t.Fatalf("PutCover: %v", err)
}
got, contentType, ok, err := s.GetCover(sourceURL)
if err != nil {
t.Fatalf("GetCover: %v", err)
}
if !ok || !bytes.Equal(got, want) || contentType != "image/jpeg" {
t.Fatalf("GetCover = (%q, %q, %v), want (%q, image/jpeg, true)", got, contentType, ok, want)
}
if err := s.PutCover("https://cdn.example/not-image", []byte("html"), "text/html"); err == nil {
t.Fatal("PutCover accepted a non-image")
}
if _, _, ok, err := s.GetCover("https://cdn.example/not-image"); err != nil || ok {
t.Fatalf("rejected cover = found %v, err %v; want missing", ok, err)
}
}
-129
View File
@@ -1,129 +0,0 @@
package web
import (
"context"
"log"
"net/http"
"regexp"
"sync"
"time"
)
// CoverFetcher retrieves one kagane cover by image id. Satisfied by
// latest.BrowserFetcher, and nil when BROWSER_WS_URL is unset — which leaves
// kagane covers exactly as unavailable as they were before this endpoint
// existed, rather than hanging a request on a fetcher that cannot run.
type CoverFetcher interface {
Image(ctx context.Context, imageID string) (body []byte, contentType string, err error)
}
// coverIDRe matches the request path segment that becomes part of an outbound
// URL. The proxy is session-gated, but the id still reaches a headless browser,
// so it is validated at the boundary rather than passed through.
var coverIDRe = regexp.MustCompile(`^[0-9a-f-]{36}$`)
// coverTypes is the set of content types the proxy will echo back. A response
// header sourced from a third party is not repeated verbatim: anything outside
// this set is treated as "not a cover".
var coverTypes = map[string]bool{
"image/webp": true,
"image/jpeg": true,
"image/png": true,
"image/avif": true,
"image/gif": true,
}
// coverTimeout bounds one proxied cover. Shorter than the fetcher's own
// challenge budget on purpose: a browser page is waiting on this, and a cover
// that has not arrived by now is better left as a broken slot than as a request
// holding a connection open.
const coverTimeout = 20 * time.Second
// coverCacheMax caps the in-memory cover cache. Covers are immutable per image
// id and a library holds tens of series, so this is a ceiling that is never
// reached in practice; reaching it clears the map rather than evicting by age.
//
// ponytail: flush-on-full, not LRU. Swap it for an LRU if a library ever grows
// past this and the flush starts costing refetches.
const coverCacheMax = 500
type cachedCover struct {
body []byte
contentType string
}
type coverCache struct {
mu sync.Mutex
m map[string]cachedCover
}
func (c *coverCache) get(id string) (cachedCover, bool) {
c.mu.Lock()
defer c.mu.Unlock()
v, ok := c.m[id]
return v, ok
}
func (c *coverCache) put(id string, v cachedCover) {
c.mu.Lock()
defer c.mu.Unlock()
if c.m == nil || len(c.m) >= coverCacheMax {
c.m = make(map[string]cachedCover, coverCacheMax)
}
c.m[id] = v
}
// kaganeCover serves a kagane cover from the backend's own origin.
//
// kagane answers image requests with a Cloudflare challenge and
// `cross-origin-resource-policy: same-origin`, so the web UI cannot render one
// directly under any combination of referrer policy or crossorigin attribute
// (verified 2026-08-08). Fetching it through the headless browser that already
// clears the challenge, and re-serving it here, is what puts the bytes on an
// origin the page may load from.
//
// ponytail: covers are fetched on first view, one browser navigation at a time
// behind the fetcher's mutex, so a first load of a large kagane library
// trickles in over a few seconds. The cache makes it a one-off. Prefetching
// during the poll cycle is the upgrade if that ever grates.
func (h *Handler) kaganeCover(w http.ResponseWriter, r *http.Request) {
id := r.PathValue("id")
if !coverIDRe.MatchString(id) {
http.NotFound(w, r)
return
}
if h.covers == nil {
http.NotFound(w, r)
return
}
if v, ok := h.coverCache.get(id); ok {
writeCover(w, v)
return
}
ctx, cancel := context.WithTimeout(r.Context(), coverTimeout)
defer cancel()
body, contentType, err := h.covers.Image(ctx, id)
if err != nil {
log.Printf("kagane cover %s: %v", id, err)
http.NotFound(w, r)
return
}
if !coverTypes[contentType] {
log.Printf("kagane cover %s: unexpected content type %q", id, contentType)
http.NotFound(w, r)
return
}
v := cachedCover{body: body, contentType: contentType}
h.coverCache.put(id, v)
writeCover(w, v)
}
// writeCover sends the bytes with a long cache life: an image id names one
// immutable rendering, so a client that has it never needs to ask again.
func writeCover(w http.ResponseWriter, v cachedCover) {
w.Header().Set("Content-Type", v.contentType)
w.Header().Set("Cache-Control", "private, max-age=604800, immutable")
w.Write(v.body)
}
+1 -1
View File
@@ -6,7 +6,7 @@
<div class="row">
<a class="cover" href="{{.ContinueURL}}" target="_blank" rel="noopener noreferrer"
tabindex="-1" aria-hidden="true">
{{if .CoverURL}}<img src="{{.CoverURL}}" alt="" loading="lazy">
{{if .Cover}}<img src="{{.Cover}}" alt="" loading="lazy">
{{/* aria-hidden on the cover link is not enough — Chromium still exposes
the letter because the link is programmatically focusable — so the
monogram carries its own, same as the recent strip's. */}}
+1 -1
View File
@@ -14,7 +14,7 @@
<a class="recent-card {{if .HasNewChapter}}is-new{{end}}" href="{{.ContinueURL}}"
target="_blank" rel="noopener noreferrer">
<span class="recent-cover">
{{if .CoverURL}}<img src="{{.CoverURL}}" alt="" loading="lazy">
{{if .Cover}}<img src="{{.Cover}}" alt="" loading="lazy">
{{else}}<span class="monogram" aria-hidden="true">{{.Initial}}</span>{{end}}
{{if .HasNewChapter}}<span class="foot-rule"></span>
{{else if .Favorite}}<span class="foot-rule brass"></span>{{end}}
+1 -10
View File
@@ -50,10 +50,6 @@ type Handler struct {
// httpClient is the plain stdlib client that talks to Discord. It is not
// an injected interface: tests point APIBase at a stub server instead.
httpClient *http.Client
// covers proxies kagane cover images, which no browser can load directly.
// Nil disables the endpoint — see CoverFetcher.
covers CoverFetcher
coverCache coverCache
}
// listView is what every list-rendering template receives.
@@ -115,7 +111,7 @@ type loginView struct {
// New parses every template up front so a broken one kills the process at
// startup rather than the first request that touches it.
func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, novelPath string, covers CoverFetcher) (*Handler, error) {
func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, novelPath string) (*Handler, error) {
tmpl, err := template.ParseFS(templateFS, "templates/*.html")
if err != nil {
return nil, err
@@ -130,7 +126,6 @@ func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, nove
states: newOAuthStates(),
limiter: session.NewLoginLimiter(),
httpClient: &http.Client{Timeout: discordTimeout},
covers: covers,
}, nil
}
@@ -147,10 +142,6 @@ func (h *Handler) Register(mux *http.ServeMux) {
mux.HandleFunc("POST /ui/bookmarks/{key}/chapter", h.requireSession(h.uiChapter))
mux.HandleFunc("DELETE /ui/bookmarks/{key}", h.requireSession(h.uiDelete))
// Session-gated like every other UI route: the deployment proxies kagane's
// images for its own Readers, not for the internet.
mux.HandleFunc("GET /img/kagane/{id}", h.requireSession(h.kaganeCover))
// Install endpoints render the script directly under the session: the
// credential travels inside the served bytes, never in the address bar or
// the page markup. Updates after install use the credential-bearing /u/
+105 -40
View File
@@ -30,7 +30,17 @@ type Config struct {
// DatabaseURL is the Postgres connection URL; required, no default,
// because a wrong guess would silently start on an empty database.
DatabaseURL string
Port string
// CoverDir is the filesystem volume for immutable cover bytes. Required:
// serving a stored address without durable bytes would be worse than a
// startup failure.
CoverDir string
// PublicBaseURL is the origin this deployment answers on, e.g.
// "https://bookmarks.example.com". Required: cover URLs go out absolute
// because the userscript renders them on third-party origins, where a
// relative path would resolve against the Site (ADR-0007), and there is
// no way to guess it from a request the poller never sees.
PublicBaseURL string
Port string
// OwnerDiscordID identifies the seeded owner Reader (issue #22). Required:
// bookmarks are scoped to a Reader, and a fresh deployment needs one
// before anybody logs in. The owner is also the only Reader who can revoke
@@ -47,10 +57,6 @@ type Config struct {
NovelUserscriptPath string
// LatestPoll configures the background latest-chapter fetcher.
LatestPoll LatestPoll
// Covers proxies kagane cover images for the web UI. Not from the
// environment: it is the shared headless browser, wired in main once it
// connects, and nil in every test router.
Covers web.CoverFetcher
}
// LatestPoll configures the background latest-chapter poller.
@@ -60,18 +66,20 @@ type Config struct {
// deployment. Past that nothing breaks; the effective cadence stretches to
// N x interval / batch and the oldest-checked-first ordering keeps it uniform.
type LatestPoll struct {
Enabled bool
Cooldown time.Duration
Interval time.Duration
Stagger time.Duration
Batch int
Enabled bool
Cooldown time.Duration
BrowserCooldown time.Duration
Interval time.Duration
Stagger time.Duration
Batch int
}
const (
defaultPollCooldown = time.Hour
defaultPollInterval = 10 * time.Minute
defaultPollStagger = 20 * time.Second
defaultPollBatch = 14
defaultPollCooldown = time.Hour
defaultBrowserPollCooldown = 6 * time.Hour
defaultPollInterval = 10 * time.Minute
defaultPollStagger = 20 * time.Second
defaultPollBatch = 14
// minPollCooldown keeps a typo from turning a polite background check into
// a hammer against sites that are already bot-scoring us.
minPollCooldown = 15 * time.Minute
@@ -129,20 +137,27 @@ func envInt(key string, def int) int {
return n
}
func clampPollCooldown(name string, d time.Duration) time.Duration {
if d < minPollCooldown {
log.Printf("config: %s %s is below the %s floor, clamping", name, d, minPollCooldown)
return minPollCooldown
}
return d
}
// loadLatestPoll reads the poller's settings, clamping anything that would make
// it antisocial.
func loadLatestPoll() LatestPoll {
p := LatestPoll{
Enabled: envBool("LATEST_CHAPTER_POLL_ENABLED", true),
Cooldown: envDuration("LATEST_CHAPTER_POLL_COOLDOWN", defaultPollCooldown),
Interval: envDuration("LATEST_CHAPTER_POLL_INTERVAL", defaultPollInterval),
Stagger: envDuration("LATEST_CHAPTER_POLL_STAGGER", defaultPollStagger),
Batch: envInt("LATEST_CHAPTER_POLL_BATCH", defaultPollBatch),
}
if p.Cooldown < minPollCooldown {
log.Printf("config: cooldown %s is below the %s floor, clamping", p.Cooldown, minPollCooldown)
p.Cooldown = minPollCooldown
Enabled: envBool("LATEST_CHAPTER_POLL_ENABLED", true),
Cooldown: envDuration("LATEST_CHAPTER_POLL_COOLDOWN", defaultPollCooldown),
BrowserCooldown: envDuration("LATEST_CHAPTER_POLL_BROWSER_COOLDOWN", defaultBrowserPollCooldown),
Interval: envDuration("LATEST_CHAPTER_POLL_INTERVAL", defaultPollInterval),
Stagger: envDuration("LATEST_CHAPTER_POLL_STAGGER", defaultPollStagger),
Batch: envInt("LATEST_CHAPTER_POLL_BATCH", defaultPollBatch),
}
p.Cooldown = clampPollCooldown("cooldown", p.Cooldown)
p.BrowserCooldown = clampPollCooldown("browser cooldown", p.BrowserCooldown)
// batch x stagger has to fit inside one tick or a batch is still running
// when the next one is due. Run() serialises them, so this degrades to a
// slower cadence rather than to overlapping fetches — worth a warning, not
@@ -158,6 +173,8 @@ func loadConfig() Config {
c := Config{
TokenKey: os.Getenv("TOKEN_KEY"),
DatabaseURL: os.Getenv("DATABASE_URL"),
CoverDir: os.Getenv("COVER_DIR"),
PublicBaseURL: os.Getenv("PUBLIC_BASE_URL"),
Port: envOr("PORT", "8080"),
OwnerDiscordID: os.Getenv("OWNER_DISCORD_ID"),
UserscriptPath: envOr("USERSCRIPT_PATH", "/userscript/manga-bookmark.user.js"),
@@ -185,7 +202,12 @@ func loadConfig() Config {
// /healthz is public.
func newRouter(s *store.Store, cfg Config) http.Handler {
mux := http.NewServeMux()
h := &api.Handler{Store: s}
mux.HandleFunc("GET /healthz", api.Healthz)
// Public: cover bytes are rendered by the userscript on origins that may
// not send our credentials, and the address is the hash of a URL the Site
// already publishes (ADR-0007).
mux.HandleFunc("GET /covers/{address}", h.Cover)
// Outside httpmw.Auth (the updater sends no Authorization header) and
// outside the web UI's Discord auth (the script must be installable
@@ -197,7 +219,6 @@ func newRouter(s *store.Store, cfg Config) http.Handler {
mux.HandleFunc("GET /u/{token}/novel-bookmark.user.js",
userscript.Handler(s, cfg.NovelUserscriptPath))
h := &api.Handler{Store: s}
protected := http.NewServeMux()
protected.HandleFunc("GET /bookmarks", h.List)
protected.HandleFunc("PUT /bookmarks/{key}", h.Put)
@@ -210,7 +231,7 @@ func newRouter(s *store.Store, cfg Config) http.Handler {
// The browser UI is always registered; signing in is Discord OAuth, so
// there is no password to forget and no gate to leave unset.
wh, err := web.New(s, cfg.Discord, []byte(cfg.TokenKey),
cfg.UserscriptPath, cfg.NovelUserscriptPath, cfg.Covers)
cfg.UserscriptPath, cfg.NovelUserscriptPath)
if err != nil {
log.Fatalf("web handler: %v", err)
}
@@ -245,6 +266,12 @@ func main() {
if cfg.DatabaseURL == "" {
log.Fatal("DATABASE_URL is required")
}
if cfg.CoverDir == "" {
log.Fatal("COVER_DIR is required")
}
if cfg.PublicBaseURL == "" {
log.Fatal("PUBLIC_BASE_URL is required")
}
// The web UI signs in through Discord, so a deployment without the OAuth
// application is misconfigured rather than passwordless.
for key, v := range map[string]string{
@@ -265,7 +292,7 @@ func main() {
TokenHash: token.Hash(token.Token([]byte(cfg.TokenKey), cfg.OwnerDiscordID, 0)),
}
s, err := store.Open(cfg.DatabaseURL, owner)
s, err := store.Open(cfg.DatabaseURL, owner, cfg.CoverDir, cfg.PublicBaseURL)
if err != nil {
log.Fatalf("open store: %v", err)
}
@@ -276,9 +303,9 @@ func main() {
// chapters on its own.
//
// One headless browser serves both consumers that need a Cloudflare
// challenge cleared: the poller's kagane/novelfull fetches and the web
// UI's kagane cover proxy. Optional — unset leaves both degraded to what
// they were before the sidecar existed.
// challenge cleared: the poller's kagane/novelfull page fetches and
// kagane's cover bytes. Optional — unset leaves kagane unpolled and its
// Covers blank until the bytes exist.
var browser latest.Fetcher
pollCtx, stopPoll := context.WithCancel(context.Background())
defer stopPoll()
@@ -288,11 +315,37 @@ func main() {
log.Printf("browser fetcher disabled: %v", err)
} else {
browser = bf
cfg.Covers = bf
context.AfterFunc(pollCtx, bf.Close)
log.Printf("browser fetcher at %s", ws)
}
}
// A Series nobody had bookmarked before gets its Latest Chapter and its
// Cover from one fetch, at creation, instead of waiting out a poll queue
// ordered by Reader count. Off the write path: the hook returns as soon
// as the goroutine is started.
var tlsFetch latest.Fetcher
if f, err := latest.NewTLSFetcher(); err != nil {
log.Printf("creation-time acquisition: plain-TLS Sites disabled, cannot build client: %v", err)
} else {
tlsFetch = f
}
var browserCover latest.BrowserCoverFetcher
if b, ok := browser.(latest.BrowserCoverFetcher); ok {
browserCover = b
}
// The Acquirer must survive a TLS client failure: kagane needs only the
// sidecar, and novelfull degrades to whatever is left.
if tlsFetch != nil || browser != nil {
acq := &latest.Acquirer{
Store: s,
Fetch: tlsFetch,
BrowserFetch: browser,
BrowserCoverFetch: browserCover,
Covers: latest.NewCoverFetcher(),
Ctx: pollCtx,
}
s.OnSeriesCreated = acq.Acquire
}
startLatestPoller(pollCtx, s, cfg.LatestPoll, browser)
srv := &http.Server{
@@ -324,6 +377,27 @@ func main() {
}
}
// newLatestPoller wires the configured cooldowns and fetchers into the poller.
func newLatestPoller(s *store.Store, cfg LatestPoll, fetch, browser latest.Fetcher) *latest.Poller {
var covers latest.BrowserCoverFetcher
if f, ok := browser.(latest.BrowserCoverFetcher); ok {
covers = f
}
return &latest.Poller{
Store: s,
Fetch: fetch,
BrowserFetch: browser,
CoverFetch: covers,
CoverBytesFetch: latest.NewCoverFetcher(),
Now: time.Now,
Cooldown: cfg.Cooldown,
BrowserCooldown: cfg.BrowserCooldown,
Interval: cfg.Interval,
Stagger: cfg.Stagger,
Batch: cfg.Batch,
}
}
// startLatestPoller launches the background poller unless it is disabled or its
// HTTP client cannot be built. Any problem here is logged and skipped: this
// feature going missing degrades the service to userscript-only latest-chapter
@@ -341,16 +415,7 @@ func startLatestPoller(ctx context.Context, s *store.Store, cfg LatestPoll, brow
// Nil browser: sites behind a JavaScript challenge are simply not polled,
// and their latest_chapter comes from the userscript alone — which is how
// the service behaved before the sidecar existed.
p := &latest.Poller{
Store: s,
Fetch: f,
BrowserFetch: browser,
Now: time.Now,
Cooldown: cfg.Cooldown,
Interval: cfg.Interval,
Stagger: cfg.Stagger,
Batch: cfg.Batch,
}
p := newLatestPoller(s, cfg, f, browser)
go p.Run(ctx)
}
+49 -7
View File
@@ -17,25 +17,33 @@ import (
func TestLoadLatestPollDefaults(t *testing.T) {
for _, k := range []string{
"LATEST_CHAPTER_POLL_ENABLED", "LATEST_CHAPTER_POLL_COOLDOWN",
"LATEST_CHAPTER_POLL_INTERVAL", "LATEST_CHAPTER_POLL_STAGGER",
"LATEST_CHAPTER_POLL_BATCH",
"LATEST_CHAPTER_POLL_BROWSER_COOLDOWN", "LATEST_CHAPTER_POLL_INTERVAL",
"LATEST_CHAPTER_POLL_STAGGER", "LATEST_CHAPTER_POLL_BATCH",
} {
t.Setenv(k, "")
}
got := loadLatestPoll()
want := LatestPoll{
Enabled: true,
Cooldown: time.Hour,
Interval: 10 * time.Minute,
Stagger: 20 * time.Second,
Batch: 14,
Enabled: true,
Cooldown: time.Hour,
BrowserCooldown: 6 * time.Hour,
Interval: 10 * time.Minute,
Stagger: 20 * time.Second,
Batch: 14,
}
if got != want {
t.Fatalf("loadLatestPoll() = %+v, want %+v", got, want)
}
}
func TestLoadConfigReadsCoverDirectory(t *testing.T) {
t.Setenv("COVER_DIR", "/covers")
if got := loadConfig().CoverDir; got != "/covers" {
t.Fatalf("CoverDir = %q, want /covers", got)
}
}
func TestLoadLatestPollEnabledParsing(t *testing.T) {
tests := []struct {
raw string
@@ -74,6 +82,30 @@ func TestLoadLatestPollClampsAndFallsBack(t *testing.T) {
wantFrom: func(p LatestPoll) any { return p.Cooldown },
want: 15 * time.Minute,
},
{
name: "browser cooldown below the floor is clamped up",
env: map[string]string{"LATEST_CHAPTER_POLL_BROWSER_COOLDOWN": "1m"},
wantFrom: func(p LatestPoll) any { return p.BrowserCooldown },
want: 15 * time.Minute,
},
{
name: "browser cooldown at the floor is kept",
env: map[string]string{"LATEST_CHAPTER_POLL_BROWSER_COOLDOWN": "15m"},
wantFrom: func(p LatestPoll) any { return p.BrowserCooldown },
want: 15 * time.Minute,
},
{
name: "browser cooldown override is honoured",
env: map[string]string{"LATEST_CHAPTER_POLL_BROWSER_COOLDOWN": "8h"},
wantFrom: func(p LatestPoll) any { return p.BrowserCooldown },
want: 8 * time.Hour,
},
{
name: "browser cooldown unparseable value falls back",
env: map[string]string{"LATEST_CHAPTER_POLL_BROWSER_COOLDOWN": "six hours"},
wantFrom: func(p LatestPoll) any { return p.BrowserCooldown },
want: 6 * time.Hour,
},
{
name: "a valid override is honoured",
env: map[string]string{"LATEST_CHAPTER_POLL_INTERVAL": "5m"},
@@ -123,6 +155,16 @@ func TestLoadLatestPollClampsAndFallsBack(t *testing.T) {
}
}
func TestNewLatestPollerWiresCooldowns(t *testing.T) {
p := newLatestPoller(nil, LatestPoll{
Cooldown: time.Hour,
BrowserCooldown: 6 * time.Hour,
}, nil, nil)
if p.Cooldown != time.Hour || p.BrowserCooldown != 6*time.Hour {
t.Fatalf("poller cooldowns = %s/%s, want 1h/6h", p.Cooldown, p.BrowserCooldown)
}
}
func TestPutStatusValidation(t *testing.T) {
cases := []struct {
name string
+28
View File
@@ -0,0 +1,28 @@
# Copy to chrome/.env on the home machine. Never commit the real .env.
#
# This file configures the browser unit only. It is separate from the API
# stack's ../.env on purpose: the two run on different machines.
# The address the CDP port is published on — required, no default.
#
# Use this machine's **tailnet IP**, e.g. 100.x.y.z (`tailscale ip -4`). Not
# 0.0.0.0, not the LAN address: CDP has no authentication of its own, so
# anything that can reach this port has full control of the browser and a
# foothold on this host. Tailscale device identity plus an ACL is the access
# control; the bind address is what enforces it.
#
# For a throwaway local test, 127.0.0.1 is fine — but then only this machine
# can reach it, so the API must run here too.
# Left commented so `cp .env.example .env && docker compose up` fails with the
# variable's own message telling you what to set, rather than Docker rejecting
# "100.x.y.z" as an invalid IP.
# BROWSER_BIND_ADDR=100.x.y.z
# Clock zone the browser reports. Any real zone works and it need not match
# the egress IP's country — but it must not be UTC, which is itself the bot
# signal that stops the challenge clearing. The measurement is in entrypoint.sh.
#
# Unset falls back to the host's /etc/timezone, which is a real zone whenever
# the host clock is set to local time. Set this when the host runs UTC — a UTC
# server is exactly the case that fails.
# BROWSER_TZ=Asia/Jakarta
+8 -2
View File
@@ -18,7 +18,7 @@ FROM debian:trixie-slim
# stale, and a stale browser is exactly what Cloudflare turns away — the 124 in
# alpine-chrome is the worked example. Rebuild is the upgrade path.
RUN apt-get update \
&& apt-get install -y --no-install-recommends ca-certificates wget gnupg \
&& apt-get install -y --no-install-recommends ca-certificates wget gnupg util-linux \
&& wget -qO- https://dl.google.com/linux/linux_signing_key.pub \
| gpg --dearmor -o /usr/share/keyrings/google-chrome.gpg \
&& echo "deb [arch=amd64 signed-by=/usr/share/keyrings/google-chrome.gpg] https://dl.google.com/linux/chrome/deb/ stable main" \
@@ -27,9 +27,15 @@ RUN apt-get update \
&& apt-get install -y --no-install-recommends google-chrome-stable socat \
&& rm -rf /var/lib/apt/lists/*
# Keep the profile path present so Docker initializes the named volume with
# the unprivileged user's ownership.
RUN useradd --create-home --shell /usr/sbin/nologin chrome \
&& mkdir -p /home/chrome/profile /home/chrome/state \
&& chown -R chrome:chrome /home/chrome
# Unprivileged: Chrome refuses to run as root, and the CDP endpoint is a shell
# on whatever user owns it.
RUN useradd --create-home --shell /usr/sbin/nologin chrome
USER chrome
WORKDIR /home/chrome
+62
View File
@@ -0,0 +1,62 @@
# The browser, as its own deployable unit.
#
# This does NOT run beside the API. It runs on the home machine, reached from
# the VPS over the tailnet, and is updated without touching the API stack:
#
# cd chrome && docker compose up -d --build
#
# Set BROWSER_BIND_ADDR in chrome/.env to this machine's tailnet IP. See
# ../DEPLOY.md §7 for the full first-time procedure and ../docs/adr/
# 0006-browser-on-the-home-machine.md for why the browser lives here at all.
name: bookmark-browser
services:
browser:
build: .
image: bookmarkmanager-chrome:latest
container_name: bookmark-browser
restart: unless-stopped
environment:
# Any real zone works, but a UTC clock is itself the bot signal and the
# challenge then never clears — measurement in entrypoint.sh. Unset falls
# back to the host's /etc/timezone below, which is a real zone whenever
# the host clock is local; set BROWSER_TZ when the host runs UTC.
TZ: ${BROWSER_TZ:-}
volumes:
# The zone *name*, which is what Chrome's ICU needs — see entrypoint.sh.
# Absent on a non-Debian host, which the entrypoint handles by falling back to UTC.
- /etc/timezone:/etc/timezone:ro
# Cloudflare clearance must survive Chrome reaping and image recreation.
- chrome-profile:/home/chrome/profile
# Bound to the tailnet address only, never 0.0.0.0. CDP authenticates
# nothing: whatever reaches this port drives the browser and, through it,
# this host. On the VPS the safety was Docker network membership; here the
# machine has a real LAN, so the bind address *is* the access control,
# backed by Tailscale device identity. No default — an unset variable must
# fail the deploy rather than silently publish CDP to the LAN.
ports:
- "${BROWSER_BIND_ADDR:?set BROWSER_BIND_ADDR to this machine's tailnet IP}:9222:9222"
# Reaps zombie renderer processes, which otherwise accumulate for the
# container's lifetime.
init: true
# Chrome allocates shared memory per tab and dies on Docker's 64MB default.
# 128MB against a measured 19MB peak: the old 1GB reservation was sized by
# superstition, and this box has 1.8GB total.
shm_size: '128mb'
# The browser is the newcomer on a machine where a Gitea runner already
# holds ~1.2GiB of 1.8GiB. Load-bearing, not decorative: untuned Chrome
# peaked at 645MiB cgroup, which is more than is free here.
#
# memswap_limit is memory+swap combined, so this allows 512MiB of swap —
# Chrome reclaims its own cold pages onto this box's 5.9GiB of SATA swap
# instead of taking resident memory from the runner.
mem_limit: 512m
memswap_limit: 1g
# If the box does run out, the kernel takes the browser and never CI.
oom_score_adj: 800
# A challenge solve yields to a running build. Cold start degrades to ~3s
# at half a CPU, immaterial against a 45-second challenge budget.
cpu_shares: 512
volumes:
chrome-profile:
+203 -39
View File
@@ -16,46 +16,210 @@ set -eu
[ -n "${TZ:-}" ] || TZ=$(cat /etc/timezone 2>/dev/null || echo UTC)
export TZ
# Chrome's own UA advertises "HeadlessChrome" under --headless=new, and that one
# token is the difference between kagane.to's challenge clearing in ~4s and
# never clearing at all (measured 2026-08-08, same host, same Chrome, only the
# UA changed). Overriding it does not touch the Sec-CH-UA client hints, which
# report the real version, so the version is read back out of the binary rather
# than hardcoded: a hardcoded one would drift out of step with the hints on the
# next Chrome update and become a fresh tell.
state=/home/chrome/state
profile=/home/chrome/profile
lock_file=$state/lock
pid_file=$state/chrome.pid
connections_dir=$state/connections
last_use_file=$state/last-use
idle_seconds=300
mkdir -p "$state" "$profile" "$connections_dir"
exec 9>>"$lock_file"
# Chrome's own UA advertises "HeadlessChrome" under --headless=new, and that
# one token is the difference between kagane.to's challenge clearing in ~4s and
# never clearing at all. Read the installed major version so client hints and
# the UA stay aligned after an image rebuild.
major=$(google-chrome-stable --version | sed -E 's/[^0-9]*([0-9]+)\..*/\1/')
ua="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/${major}.0.0.0 Safari/537.36"
# Chrome binds its DevTools port to loopback and silently ignores
# --remote-debugging-address (verified 2026-08-08: Chrome 151 with
# --remote-debugging-address=0.0.0.0 still listened on 127.0.0.1 only), so the
# caller — another container — cannot reach it directly. socat fronting the
# loopback port is how chromedp/headless-shell solved the same problem and is
# why this image is a drop-in for it.
#
# Nothing publishes 9222; reachability is the `browser` network in
# docker-compose.yml, and an exposed CDP endpoint is remote code execution.
socat TCP-LISTEN:9222,fork,reuseaddr TCP:127.0.0.1:9223 &
lock() {
flock 9
}
# Chrome stays in the foreground so that its death takes the container down and
# compose's restart policy applies; a backgrounded browser behind a live socat
# would leave the sidecar looking healthy while answering nothing.
#
# No --enable-automation: it sets navigator.webdriver, the first thing a bot
# check reads.
#
# --no-sandbox because Chrome's zygote wants user namespaces, which Docker's
# default profile does not hand out; the alternative is --cap-add=SYS_ADMIN,
# which gives the container strictly more than it takes away. Containment here
# is the unprivileged user, the isolated network, and the fact that this
# browser only ever navigates to kagane.to and novelfull.com.
exec google-chrome-stable \
--headless=new \
--no-sandbox \
--remote-debugging-port=9223 \
--user-agent="$ua" \
--user-data-dir=/home/chrome/profile \
--no-first-run \
--no-default-browser-check \
--disable-gpu \
about:blank
unlock() {
flock -u 9
}
browser_alive() {
[ -s "$pid_file" ] || return 1
pid=$(cat "$pid_file")
[ -n "$pid" ] && kill -0 "$pid" 2>/dev/null
}
has_connections() {
for marker in "$connections_dir"/*; do
[ -e "$marker" ] || continue
pid=${marker##*/}
if kill -0 "$pid" 2>/dev/null; then
return 0
fi
# A SIGKILLed helper cannot run its cleanup trap. Reconcile its marker
# here so one dead client cannot pin Chrome forever.
rm -f "$marker"
done
return 1
}
start_browser() {
# No --enable-automation: it sets navigator.webdriver, the first thing a
# bot check reads. setsid gives Chrome a process group so the reaper can
# terminate its renderer children with the browser.
# --no-sandbox avoids granting SYS_ADMIN solely for Docker's unavailable
# user namespaces; containment is the unprivileged user and private network.
setsid google-chrome-stable \
--headless=new \
--no-sandbox \
--remote-debugging-port=9223 \
--user-agent="$ua" \
--user-data-dir="$profile" \
--no-first-run \
--no-default-browser-check \
--disable-gpu \
about:blank >/dev/null &
printf '%s\n' "$!" >"$pid_file"
}
stop_browser() {
pid=$(cat "$pid_file")
kill -TERM -- "-$pid" 2>/dev/null || kill -TERM "$pid" 2>/dev/null || true
i=0
while kill -0 "$pid" 2>/dev/null && [ "$i" -lt 100 ]; do
i=$((i + 1))
sleep 0.1
done
if kill -0 "$pid" 2>/dev/null; then
kill -KILL -- "-$pid" 2>/dev/null || kill -KILL "$pid" 2>/dev/null || true
fi
rm -f "$pid_file"
}
wait_for_browser() {
i=0
while [ "$i" -lt 300 ]; do
if wget -qO /dev/null http://127.0.0.1:9223/json/version; then
return 0
fi
browser_alive || return 1
i=$((i + 1))
sleep 0.1
done
return 1
}
finish_connection() {
lock
rm -f "$connection_marker"
date +%s >"$last_use_file"
unlock
}
connection_signal() {
trap - INT TERM HUP
finish_connection
exit 143
}
connection() {
connection_marker=$connections_dir/$$
lock
: >"$connection_marker"
if ! browser_alive; then
rm -f "$pid_file"
start_browser
fi
date +%s >"$last_use_file"
unlock
trap connection_signal INT TERM HUP
if wait_for_browser; then
if socat STDIO TCP:127.0.0.1:9223; then
result=0
else
result=$?
fi
else
# The client only ever sees a bare connection reset here, so this is
# the sole record that the browser, not the network, was the problem.
echo "browser did not come up; dropping connection" >&2
result=1
fi
finish_connection
return "$result"
}
reaper() {
while :; do
sleep 10
lock
if ! has_connections && browser_alive; then
now=$(date +%s)
last=$(cat "$last_use_file" 2>/dev/null || printf '%s' "$now")
if [ $((now - last)) -ge "$idle_seconds" ]; then
stop_browser
fi
fi
unlock
done
}
if [ "${1:-}" = connection ]; then
connection
exit $?
fi
# The files are process state, not the Chrome profile. The profile is a named
# volume in Compose, so clearance survives both a reap and a container rebuild.
for marker in "$connections_dir"/*; do
[ -e "$marker" ] || continue
rm -f "$marker"
done
rm -f "$pid_file" "$last_use_file"
# Chrome's singleton lock names the hostname and pid that took it, and a
# container rebuild changes both — so a Chrome killed uncleanly (OOM, docker
# kill) leaves a lock the next container reads as "another computer holds this
# profile" and refuses to start behind, permanently, with the only symptom a
# bare connection reset at 9222. Clearing it here is safe precisely because
# container_name pins this volume to one container: nothing can be holding the
# profile at the moment this line runs. The lock is process state; the
# clearance cookies it sits beside are not, and are left alone.
rm -f "$profile"/Singleton*
# Chrome binds DevTools to loopback and silently ignores
# --remote-debugging-address. socat remains the network front-end, but each
# accepted connection now starts a browser on demand and is tracked by a
# per-helper marker. A connection held by Go's transport delays reap by its
# idle timeout; the 300-second threshold starts once the last connection closes.
reaper &
reaper_pid=$!
socat TCP-LISTEN:9222,reuseaddr,fork EXEC:'/entrypoint.sh connection',nofork &
front_pid=$!
stop_browser_gracefully() {
lock
if browser_alive; then
# Chrome is a separate session, so stop its process group explicitly;
# this gives its cookie batch time to flush before the container exits.
stop_browser
fi
unlock
}
shutdown() {
trap - INT TERM HUP
stop_browser_gracefully
kill "$front_pid" "$reaper_pid" 2>/dev/null || true
exit 143
}
trap shutdown INT TERM HUP
if wait "$front_pid"; then
status=0
else
status=$?
fi
stop_browser_gracefully
kill "$reaper_pid" 2>/dev/null || true
exit "$status"
+10 -20
View File
@@ -17,24 +17,12 @@ services:
bookmark-api:
# Traffic arrives over the Traefik network, not a published port.
ports: !reset []
environment:
# Must be an IP, not the DNS name — see the base file's comment on this
# same key: Chrome's DevTools HTTP handler 500s any Host header that
# isn't an IP or "localhost".
BROWSER_WS_URL: ${BROWSER_WS_URL:-ws://172.28.0.10:9222}
depends_on:
headless-shell:
condition: service_started
postgres:
condition: service_healthy
# `networks:` here replaces the base file's list entirely, so all three must
# be named: `proxy` for Traefik routing, and `browser` / `db` (defined in
# the base file) to keep reaching headless-shell and Postgres without
# putting either on `proxy`.
# Compose *merges* this list with the base file's, so the service ends up on
# `default`, `db` and `proxy` — only the addition is named here. Do not
# "tidy" the base file down to `db` on the strength of `proxy` being present:
# `db` is `internal: true`, and egress comes from `default`.
networks:
- proxy
- browser
- db
labels:
- "traefik.enable=true"
- "traefik.docker.network=${PROXY_NETWORK:-proxy}"
@@ -52,10 +40,12 @@ services:
- "traefik.http.routers.bmweb.tls.certresolver=${TRAEFIK_CERTRESOLVER:-le}"
- "traefik.http.routers.bmweb.service=bmapi"
# headless-shell is untouched here: it keeps its `browser` network membership
# from the base file and must never join `proxy` — that network is shared
# with whatever else sits behind Traefik on this host, and an exposed
# CDP endpoint on it would be remote code execution for any of them.
# No browser service here. It runs on the home machine as its own unit
# (chrome/docker-compose.yml) and is reached over the tailnet — see
# docs/adr/0006-browser-on-the-home-machine.md. It must never be given a
# service on this host: `proxy` is shared with whatever else sits behind
# Traefik, and an unauthenticated CDP endpoint on it is remote code
# execution for any of them.
networks:
proxy:
+36 -56
View File
@@ -5,10 +5,16 @@
# If your proxy runs in Docker on its own network, use the prod override which
# attaches to that network instead of publishing a port:
# docker compose -f docker-compose.yml -f docker-compose.prod.yml up -d
#
# The browser is not here. It is its own unit on the home machine —
# chrome/docker-compose.yml — reached over the tailnet via BROWSER_WS_URL.
services:
bookmark-api:
build: ./backend
build:
context: ./backend
args:
COVER_DIR: ${COVER_DIR:?set COVER_DIR in .env}
image: bookmarkmanager-backend:latest
container_name: bookmark-api
restart: unless-stopped
@@ -23,11 +29,18 @@ services:
# The bookmarks database. Host is the compose service name; the password
# comes from .env so it is never committed.
DATABASE_URL: ${DATABASE_URL:-postgres://bookmarks:${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD in .env}@postgres:5432/bookmarks?sslmode=disable}
# Required path inside the API. The build seeds ownership at this path
# and the named volume below mounts there.
COVER_DIR: ${COVER_DIR:?set COVER_DIR in .env}
# Public origin of this deployment, no trailing slash. Required: the
# Cover URLs on the wire are absolute, since the userscript renders them
# on a Site's origin rather than ours (ADR-0007).
PUBLIC_BASE_URL: ${PUBLIC_BASE_URL:?set PUBLIC_BASE_URL in .env}
PORT: "8080"
# Log timestamps only. Go's `log` stamps lines in local time, and this
# service has no other use for a zone: bookmark timestamps are unix ms
# and the two real time columns are timestamptz, both absolute instants.
# Purely so these lines read on the same clock as the sidecar's. Named
# Purely so these lines read on the same clock as the browser's. Named
# API_TZ rather than TZ so an operator's exported shell TZ cannot leak
# in; distroless already carries tzdata, so the name just resolves.
TZ: ${API_TZ:-Asia/Jakarta}
@@ -49,21 +62,24 @@ services:
# switch; it only takes effect because these are listed here.
LATEST_CHAPTER_POLL_ENABLED: ${LATEST_CHAPTER_POLL_ENABLED:-1}
LATEST_CHAPTER_POLL_COOLDOWN: ${LATEST_CHAPTER_POLL_COOLDOWN:-1h}
LATEST_CHAPTER_POLL_BROWSER_COOLDOWN: ${LATEST_CHAPTER_POLL_BROWSER_COOLDOWN:-6h}
LATEST_CHAPTER_POLL_INTERVAL: ${LATEST_CHAPTER_POLL_INTERVAL:-10m}
LATEST_CHAPTER_POLL_BATCH: ${LATEST_CHAPTER_POLL_BATCH:-14}
LATEST_CHAPTER_POLL_STAGGER: ${LATEST_CHAPTER_POLL_STAGGER:-20s}
# CDP endpoint for sites behind a JavaScript challenge (kagane). Unset
# disables browser polling for those sites; the userscript still covers them.
# Must be an IP, not the "headless-shell" DNS name: Chrome's DevTools HTTP
# handler rejects the discovery request (GET /json/version) with a 500
# unless the Host header is an IP address or "localhost" — confirmed
# 2026-08-03 against chromedp/headless-shell:stable, independent of
# chromedp's own dial logic. The sidecar's static address below exists so
# this URL survives container recreation.
BROWSER_WS_URL: ${BROWSER_WS_URL:-ws://172.28.0.10:9222}
# CDP endpoint for sites behind a JavaScript challenge (kagane,
# novelfull). The browser is not part of this stack — it runs on the home
# machine as its own unit (chrome/docker-compose.yml) and is reached over
# the tailnet. Unset disables browser polling for those sites and serves
# 404 from the cover proxy for covers not already stored; the userscript
# still covers them. Set it in .env to ws://<home machine tailnet IP>:9222.
#
# Must be an IP, not a MagicDNS hostname: Chrome's DevTools HTTP handler
# rejects the discovery request (GET /json/version) with a 500 unless the
# Host header is an IP address or "localhost" — confirmed 2026-08-03,
# independent of chromedp's own dial logic. The same trap that used to
# force a pinned Docker IP now forbids the tailnet name.
BROWSER_WS_URL: ${BROWSER_WS_URL:-}
depends_on:
headless-shell:
condition: service_started
# The migration runner is the first thing the binary does, so a Postgres
# that is still initialising means a crash-loop until it is not.
postgres:
@@ -74,12 +90,17 @@ services:
# no rebuild, no restart. `git pull` restores the committed version, which
# is why a redeploy always ships the repo's script.
- ./userscript:/userscript:ro
# Content-addressed cover bytes survive API restarts and redeploys.
- cover-data:${COVER_DIR:?set COVER_DIR in .env}
# Bound to loopback only: the proxy (or curl during smoke test) reaches it,
# the public internet does not.
ports:
- "127.0.0.1:8080:8080"
# `default` is not decoration: `db` is `internal: true`, and a container on
# nothing but an internal network gets neither a published port nor egress
# — which would silently kill every poller fetch.
networks:
- browser
- default
- db
postgres:
@@ -101,56 +122,15 @@ services:
networks:
- db
headless-shell:
# Real Google Chrome, not chromedp/headless-shell — see chrome/Dockerfile.
# The service name is kept so existing overrides and BROWSER_WS_URL stay put.
build: ./chrome
image: bookmarkmanager-chrome:latest
restart: unless-stopped
environment:
# A UTC clock is itself the bot signal: Cloudflare treats it as the
# datacenter default, and kagane's challenge then never clears. Measured
# 2026-08-08, identical container, one Indonesian egress IP: UTC never
# cleared in 60s (twice); Asia/Jakarta and America/New_York both cleared
# in 4s. So any real zone works and it need not match the IP's country —
# only UTC fails. Unset falls back to the host's /etc/timezone below,
# which is a real zone whenever the host clock is set to local time; set
# BROWSER_TZ when the host runs UTC.
TZ: ${BROWSER_TZ:-}
volumes:
# The zone *name*, which is what Chrome's ICU needs — see chrome/entrypoint.sh.
# Absent on a non-Debian host, which the entrypoint handles by falling back to UTC.
- /etc/timezone:/etc/timezone:ro
# Chrome allocates shared memory per tab and dies on Docker's 64MB default.
shm_size: '1gb'
# Reaps zombie renderer processes, which otherwise accumulate for the
# container's lifetime.
init: true
# Deliberately no `ports:` — an exposed CDP endpoint is remote code
# execution. Only bookmark-api, via the `browser` network below, may reach it.
# No `command:` either: every flag this browser needs is in its entrypoint,
# and the UA override there is load-bearing for the challenge.
networks:
browser:
# Pinned so BROWSER_WS_URL can name an IP (required, see above) that
# survives `docker compose up` recreating this container.
ipv4_address: 172.28.0.10
volumes:
postgres-data:
cover-data:
# The pre-Postgres SQLite volume (bookmarks-data) is deliberately no longer
# declared here: undeclared means `docker compose down -v` cannot take it
# with the rest, so the old database survives the cutover until someone
# removes it by hand.
networks:
# Not `internal: true`: headless Chrome still needs outbound access to reach
# kagane.to. Isolation here comes from membership (only bookmark-api and
# headless-shell join it), not from cutting egress.
browser:
ipam:
config:
- subnet: 172.28.0.0/24
# Postgres needs no egress and nothing outside bookmark-api needs to reach
# it, so this one really can be cut off from the outside world.
db:
+51
View File
@@ -0,0 +1,51 @@
# ADR-0005: On-demand browser sidecar
Date: 2026-08-09
Status: accepted
Superseded in part by ADR-0006: the lifecycle below is unchanged, but the
service no longer lives in the API stack and the name `headless-shell` is gone.
## Decision
Keep the `headless-shell` service and its CDP port alive, but start Google Chrome
only when the first CDP connection arrives. The entrypoint supervises a `socat`
front-end, serializes browser start/reap state with `flock`, and tracks each
connection with a marker named for its helper PID. A reaper stops Chrome after
300 seconds with no live markers. Marker reconciliation covers a helper killed
before its cleanup trap runs.
Chrome runs in its own process group so reap sends the termination signal to
Chrome and its renderer children. The explicit `/home/chrome/profile` user-data
directory remains: Chrome remaps remote debugging to loopback on modern builds,
and Chrome ignores remote-debugging flags on a default profile. `socat` therefore
continues to front Chrome's loopback CDP port.
The profile is a named Compose volume. Clearance cookies survive both a reap and
`docker compose up --build`; the browser still starts with a fresh debugger UUID,
so chromedp must keep endpoint discovery enabled and must not use
`chromedp.NoModifyURL`.
The socat front-end and explicit profile are retained because Chromium remaps a
non-loopback debugging address to loopback since M113, while Chrome ignores the
remote-debugging flags on a default profile since Chrome 136. Flag tuning is
deliberately not adopted: its roughly 30% idle-footprint saving is irrelevant
to a browser that exists for seconds per wake and risks an untested fingerprint.
## Constraints
The 300-second floor is deliberate. Chromium batches cookie persistence on a
roughly 31-second timer, and Go's default HTTP transport can keep the discovery
connection parked for about 90 seconds after use. Reaping only with zero live
connections holds Chrome through both windows and through the poller's staggered
batch plus cover prefetch.
The anti-bot properties remain unchanged: a plausible non-UTC timezone, a
Chrome-version-derived User-Agent without `HeadlessChrome`, and no automation
flag. A remote browser restart can surface as `context.Canceled`, the same error
as a caller deadline, so the backend wraps cancellation observed with a closed
CDP connection as `browser interrupted`; the focused test asserts that
classification without killing a real browser.
The same process-group stop runs during supervisor shutdown, not only during
idle reap, so Chrome can flush its cookie batch before a container rebuild or
graceful stop.
@@ -0,0 +1,74 @@
# ADR-0006: The browser runs on the home machine, over the tailnet
Date: 2026-08-09
Status: accepted
## Decision
The headless browser is no longer part of the API stack. It is its own compose
unit (`chrome/docker-compose.yml`), deployed on the home machine, and the API on
the VPS reaches it over the existing tailnet through `BROWSER_WS_URL`. No
fallback sidecar remains on the VPS.
The backend needs no code change for this. The CDP endpoint was already a
configuration seam and the fetcher only ever holds the endpoint URL, so
relocation — and reversal — is one environment variable.
## Why
The sidecar held 471 MiB working set (645 MiB peak) on a 1974 MiB VPS with no
swap, which also hosts Traefik, Gitea and its Postgres. That is 24% of the host
and 86% of this project's memory, for a service that at the time answered zero
requests: the poller's due query joins bookmarks, production held four kagane
series and no bookmarks on any of them, and with no kagane bookmark the web UI
never rendered a kagane cover either.
The home machine has 5.9 GiB of swap and a residential egress, which Cloudflare
scores better than a datacenter IP. Both machines were already on the tailnet.
This move is only safe because covers are persisted (ADR-0005's sibling work,
issue #43/#45) and the browser is on-demand (ADR-0005). Without stored covers a
sleeping home machine would blank the library; without on-demand start the CI
runner that already holds ~1.2 GiB of that box's 1.8 GiB would be squeezed
around the clock.
## Constraints
**`BROWSER_WS_URL` must be the tailnet IP, never a MagicDNS hostname.** Chrome's
DevTools HTTP handler answers `/json/version` with a 500 for any `Host` header
that is not an IP or `localhost`. This is the same trap that previously forced a
pinned Docker IP; the pinned subnet is gone, the constraint is not.
**The CDP port binds to the tailnet address only, never `0.0.0.0`.** CDP
authenticates nothing: whatever reaches the port drives the browser and, through
it, the host. On the VPS the safety came from Docker network membership; the
home machine has a real LAN, so a `0.0.0.0` bind is a hole punched into it. The
bind address is the enforcement and Tailscale device identity plus a per-device
ACL is the policy. `BROWSER_BIND_ADDR` deliberately has no default, so an unset
value fails the deploy instead of publishing CDP to the LAN.
No bearer-token proxy is added in front of CDP. It would only defend against a
device already inside the tailnet, and it would be one more thing between the
poller and a browser that is already hard enough to keep clearing challenges.
**Resource limits are load-bearing, not decorative.** The browser is the
newcomer on that box, not the incumbent. A hard 512 MiB cap with 1 GiB
memory+swap makes Chrome reclaim its own cold pages onto the machine's SATA swap
instead of taking resident memory from the runner; untuned Chrome peaked at
645 MiB cgroup, which is more than is free there. `oom_score_adj` biases the
kernel to kill the browser first and never CI. Reduced CPU weight makes a
challenge solve yield to a running build — cold start degrades to about 3 s at
half a CPU, immaterial against a 45-second challenge budget. The shared-memory
reservation drops from 1 GiB to 128 MiB against a measured 19 MiB peak.
## Consequences
An unreachable browser degrades exactly as an unset `BROWSER_WS_URL` already
does: plain-TLS libraries are unaffected, kagane and novelfull log and skip, the
series waits out its cooldown, and stored covers keep serving. A power outage at
home costs chapter freshness on two sites, never the appearance of the library.
The two units are deployed and updated independently. `REDEPLOY.md` §8 covers
the browser; everything before it covers the API stack. A local `docker compose
up` now brings up two services, not three, and polls kagane only if
`BROWSER_WS_URL` is pointed somewhere.
+105
View File
@@ -0,0 +1,105 @@
# ADR-0007: The backend hosts every Site's Cover bytes
Date: 2026-08-09
Status: accepted
## Decision
A Cover is the image a Reader's browser can display for a Series. Where the Site
keeps the picture is the backend's problem, not the client's: the backend fetches
the bytes, stores them, and serves them from its own origin. No client ever
renders a third-party URL, and no client ever supplies one.
Concretely:
- **Acquisition is server-side.** The Cover is extracted from the same series-page
fetch that already yields Latest Chapter. It runs once at Series creation rather
than waiting for the poll queue, so a newly bookmarked Series has both facts in
seconds instead of up to a queue's depth. The poll fills a blank Cover and never
overwrites a non-blank one.
- **Bytes live on a filesystem volume**, content-addressed by the SHA-256 of the
source URL, sharded `${COVER_DIR}/ab/cd/<sha256>`. The database holds the path
and content type, not the bytes.
- **One public route** serves them. No session, no credential.
- **The wire carries an absolute URL** built from a configured public base, and
carries `""` until the bytes exist.
## Why a future reader will find this surprising
Four of the six Sites let anyone hot-link their covers — `static.comix.to` even
answers `access-control-allow-origin: *`. Hosting copies looks like work we were
not obliged to do.
We were obliged. kagane serves covers with `cross-origin-resource-policy:
same-origin` behind a JavaScript challenge (measured 2026-08-08), so no `<img>`
outside kagane.to can load one under any combination of referrer policy and
`crossorigin` attribute. The first fix for that was a kagane-only proxy applied in
the web templates — and it produced issue #47, because the JSON API kept emitting
the raw kagane URL and the userscript rendered it into a broken-image glyph. A
per-Site exception that only one of two clients knows about is not a fix; it is a
bug with a delay on it. Uniformity is the property being bought: every client
renders every Cover the same way, and a Site changing its CORP header or its CDN
cannot break a client again.
## Considered options
**Per-Site exceptions, proxying only what must be proxied.** Cheapest, and what we
had. Rejected: it is what produced #47, and it requires every current and future
client to know which Sites are special.
**A host allowlist for the outbound fetch**, mirroring `fetchableSeriesURL`.
Rejected in favour of destination-class control — see below.
**Cover bytes in Postgres `bytea`**, extending the existing `covers` table.
Rejected: covers are immutable blobs served straight to browsers, which is what a
filesystem is for. The cost is real and accepted — durability is now two things to
back up instead of one, against ADR-0001's grain.
**Per-Reader Cover overrides.** Rejected, consistent with ADR-0003's rejection of
per-Reader title overrides. A Cover is a fact about the Series.
## Two deliberate relaxations
**Destination control is deny-class, not an allowlist.** The outbound fetch
requires `https`, resolves DNS first and refuses loopback, private, link-local and
CGNAT addresses, re-checks on every redirect hop, and caps body size and content
type. It does *not* pin a host set, which is what `fetchableSeriesURL` does for
`series_url`. Cover hosts are CDNs that move: `demonicscans.org` serves its covers
from `readermc.org`, a host with no visible relationship to the Site. An allowlist
would silently stop producing Covers the day a Site switched CDN, and the failure
would look like this bug. The resolved-IP check is the load-bearing part; without
it, an attacker-controlled page need only publish a DNS name pointing at
`127.0.0.1`.
**The cover route is public, where the kagane proxy was session-gated.** An `<img>`
in the userscript panel cannot send a bearer token, and it cannot be given one: the
panel's shadow root is `mode: "open"`, so the host page's own JavaScript can read
any `src` we set. A credential in an image URL is a credential handed to a
third-party site. The route serves public artwork from public Sites and its path
reveals nothing about which Reader holds what. The residual cost is that we can be
hot-linked by others.
## Consequences
- The poll's cover prefetch, today guarded by `sr.Site != "kagane"`, applies to
every Site in both Libraries. Nothing about Covers is conditioned on `kind`.
- Only kagane still needs the CDP browser for its bytes. The other five Sites fetch
over plain TLS — including novelfull, whose HTML answers `cf-mitigated: challenge`
while its image paths answer 200 with `access-control-allow-origin: *`
(measured 2026-08-09).
- Client-side cover scraping is deleted from both userscripts. It could not help: a
scraped URL has no render path left, and it is absent exactly when a Series is
created — neither comix nor lightnovelworld exposes a cover on a chapter page,
which is where a Reader bookmarks mid-read.
- `PUT /bookmarks/{key}` still accepts a `cover` field and ignores it. This extends
ADR-0003's "ignored after creation" to "ignored always", and keeps the flat wire
contract ADR-0004 requires so installed scripts keep working. The field is
therefore permanently inert rather than pending removal, and says so at the
decode site.
- A Cover that fails to load falls back to the placeholder in both clients. The
broken-image glyph reported in #47 is not a state we render.
- The existing kagane `covers` rows are dropped rather than migrated; that path
re-fetches on demand already.
- Two Sites deserve a note for whoever writes the extractor: asura's `.webp` cover
URL answers `Content-Type: image/jpeg`, so trust the header; demonic's `og:image`
carries a raw unencoded space and must be percent-encoded before fetching.
+11 -10
View File
@@ -2,7 +2,7 @@ Guidance for OpenCode (and Claude Code) working under `userscript/`. See root `A
### Userscript structure (single IIFE, `manga-bookmark.user.js`)
1. **Site adapters** — one per host, `detect(location, document)` return page `type` + IDs. Identify type/IDs from **URL regex** (most stable); pull `title`/`cover` from **`og:title`/`og:image` meta tags**, not CSS classes.
1. **Site adapters** — one per host, `detect(location, document)` return page `type` + IDs. Identify type/IDs from **URL regex** (most stable); pull `title` from **`og:title`** (or the page heading where a site ships no og: tags), not CSS classes. **No adapter reads a cover**: the backend acquires, stores and serves every Cover from its own origin (ADR-0007), the wire's `cover` is already an address on our origin, and `apiPut` strips any `cover` off an outgoing body.
2. **API client** — `apiGet/apiPut/apiDelete` with bearer header; `localStorage` key `bmgr:manga:cache` for instant render + offline fallback.
3. **Progress logic** — auto-upsert `last_chapter` only when `chapterNum >= stored last_chapter_num` (re-reading old chapters must not regress progress; unparseable -> set current). Manual panel override forces any value.
4. **Retry queue** — every write go through `pushBookmark`/`pushDelete`, so
@@ -52,26 +52,27 @@ Guidance for OpenCode (and Claude Code) working under `userscript/`. See root `A
"Comix — Read Comics online for free" and after an in-page hop it is the
*previous* series' name. `document.title` is the one thing client routing does
update, so titles come from there, with the chapter page's `" · Ch.<n>"` tail
stripped. Covers likewise: `og:image` is absent, so the cover is the `img`
whose `alt` matches the cleaned title — verified live 2026-08-08.
stripped. It publishes no `og:image` either, which is one of the reasons cover
acquisition moved to the backend.
- **kagane.to**: series `/series/<uuid>`, reader
`/series/<uuid>/reader/<bookUuid>`. Reader URLs carry no chapter number, so
the number comes out of `og:title`. Two shapes exist: `"<Series> - Chapter
<n>[ - Episode <n>]"` and, for volume-numbered series, `"<Series> - Volume <v>
Chapter <n>"` with no episode name — both must yield a bare series title, or
the volume tail lands in the bookmark's title. Its covers are challenge- and
CORP-protected, so the web UI proxies them; the userscript still stores the
raw `og:image`. Behind a Cloudflare JS challenge, so the backend polls it
the volume tail lands in the bookmark's title.
Its covers are challenge- and CORP-protected, so nothing outside kagane.to can
load one directly; the panel renders the backend's own cover address like every
other Site. Behind a Cloudflare JS challenge, so the backend polls it
through the headless browser.
- **novelfull.com** (novel script): series `/<slug>.html`, chapter
`/<slug>/chapter-<n>[-<title-slug>].html`. No `og:*` tags at all — title from
`h3.title` (series) or `a.truyen-title` (chapter), cover from
`meta[name="image"]`. Behind a Cloudflare JS challenge no TLS fingerprint
`h3.title` (series) or `a.truyen-title` (chapter); the script reads no cover.
Behind a Cloudflare JS challenge no TLS fingerprint
clears, so the backend polls it through the headless browser.
- **lightnovelworld.net** (novel script): series `/novel/<slug>/`, chapter
`/<slug>-chapter-<n>/` — flat, at the site root. `h1.entry-title` is the clean
title on a series page and `<Title> Chapter <n>` on a chapter page. Chapter
pages carry no `og:image`. Its series page lists every chapter with an
title on a series page and `<Title> Chapter <n>` on a chapter page. Its series
page lists every chapter with an
absolute href, so the backend polls it with the plain TLS client.
### Second script: `novel-bookmark.user.js`
+17 -25
View File
@@ -56,9 +56,10 @@
// ============================================================
// Site adapters
//
// Page type + IDs come from URL regex (most stable); title/cover come from
// og: meta tags. Verified live 2026-07-24 against asurascans.com and
// demonicscans.org — see README "Adapter reference".
// Page type + IDs come from URL regex (most stable); the title comes from
// og: meta tags. Covers are never read here: the backend acquires and serves
// them itself (ADR-0007). Verified live 2026-07-24 against asurascans.com
// and demonicscans.org — see README "Adapter reference".
// ============================================================
function meta(prop) {
@@ -128,7 +129,6 @@
site: this.site,
seriesId: stripBuildHash(m[1]),
title: cleanTitle(meta("og:title")),
cover: meta("og:image") || "",
seriesUrl: loc.origin + "/comics/" + m[1],
chapterLabel: "Chapter " + m[2],
chapterNum: isNaN(num) ? null : num,
@@ -143,7 +143,6 @@
site: this.site,
seriesId: stripBuildHash(m[1]),
title: cleanTitle(meta("og:title")),
cover: meta("og:image") || "",
seriesUrl: loc.origin + "/comics/" + m[1],
chapterLabel: null,
chapterNum: null,
@@ -192,7 +191,6 @@
site: this.site,
seriesId: decodeURIComponent(m[1]),
title: cleanTitle(meta("og:title")),
cover: meta("og:image") || "",
seriesUrl: loc.origin + "/manga/" + m[1],
chapterLabel: "Chapter " + m[2],
chapterNum: isNaN(num) ? null : num,
@@ -207,7 +205,6 @@
site: this.site,
seriesId: decodeURIComponent(m[1]),
title: cleanTitle(meta("og:title")),
cover: meta("og:image") || "",
seriesUrl: loc.origin + "/manga/" + m[1],
chapterLabel: null,
chapterNum: null,
@@ -259,7 +256,6 @@
site: this.site,
seriesId: comixSeriesId(m[1]),
title: pageTitle,
cover: coverFromPage(pageTitle),
seriesUrl: loc.origin + "/title/" + m[1],
chapterLabel: "Chapter " + m[2],
chapterNum: isNaN(num) ? null : num,
@@ -274,7 +270,6 @@
site: this.site,
seriesId: comixSeriesId(m[1]),
title: pageTitle,
cover: coverFromPage(pageTitle),
seriesUrl: loc.origin + "/title/" + m[1],
chapterLabel: null,
chapterNum: null,
@@ -289,17 +284,6 @@
return t.replace(/\s*·\s*Ch\.[\d.]+\s*$/i, "").trim();
}
// comix serves no og:image, so this is the one adapter that has to read
// the DOM for a cover. Matching on alt rather than a class keeps it off
// the site's styling: the cover is the image whose alt is the title.
// Do not "simplify" this into meta("og:image") — that returns null.
function coverFromPage(title) {
if (!title || !document.querySelectorAll) return "";
for (const img of document.querySelectorAll("img[alt]")) {
if (img.getAttribute("alt") === title) return img.getAttribute("src") || "";
}
return "";
}
},
// Scoped to this series' own id prefix so a recommendation strip's links
// cannot win the maximum. seriesId is passed in because the anchors alone
@@ -342,7 +326,6 @@
site: this.site,
seriesId: m[1],
title: cleanTitle(meta("og:title")),
cover: meta("og:image") || "",
seriesUrl: loc.origin + "/series/" + m[1],
chapterLabel: num === null ? null : "Chapter " + num,
chapterNum: num,
@@ -357,7 +340,6 @@
site: this.site,
seriesId: m[1],
title: cleanTitle(meta("og:title")),
cover: meta("og:image") || "",
seriesUrl: loc.origin + "/series/" + m[1],
chapterLabel: null,
chapterNum: null,
@@ -498,6 +480,10 @@
async function apiPut(key, obj, { sendStatus = false } = {}) {
const body = Object.assign({}, obj);
if (!sendStatus) delete body.status;
// Covers belong to the backend, which acquires and serves them itself
// (ADR-0007) and ignores an incoming one; a third-party address must never
// go back on the wire.
delete body.cover;
const res = await fetch(API_BASE + "/bookmarks/" + encodeURIComponent(key), {
method: "PUT",
headers: authHeaders({ "Content-Type": "application/json" }),
@@ -855,7 +841,6 @@
series_id: p.seriesId,
title: p.title || (existing && existing.title) || p.seriesId,
series_url: p.seriesUrl || (existing && existing.series_url) || "",
cover: p.cover || (existing && existing.cover) || "",
last_chapter: p.chapterLabel || (existing && existing.last_chapter) || "",
last_chapter_num:
p.chapterNum != null ? p.chapterNum : existing ? existing.last_chapter_num : null,
@@ -878,7 +863,6 @@
series_id: p.seriesId,
title: existing.title || p.title || p.seriesId,
series_url: existing.series_url || p.seriesUrl || "",
cover: existing.cover || p.cover || "",
last_chapter: p.chapterLabel || existing.last_chapter || "",
last_chapter_num: p.chapterNum != null ? p.chapterNum : existing.last_chapter_num,
last_chapter_url: p.chapterUrl || "",
@@ -1394,7 +1378,15 @@
return el("div", { class: "item" + heat }, [
el("a", { class: "go", href: cont }, [
b.cover
? el("img", { class: "cover", src: b.cover, loading: "lazy", alt: "" })
? el("img", {
class: "cover",
src: b.cover,
loading: "lazy",
alt: "",
// A Cover that will not load shows the designed placeholder
// rather than the browser's broken-image glyph (#47).
onerror: (e) => e.target.replaceWith(el("div", { class: "cover ph" })),
})
: el("div", { class: "cover ph" }),
]),
el("div", { class: "meta" }, [
+23 -23
View File
@@ -25,6 +25,13 @@
// This script owns the novel library; the manga script is a separate install
// with its own prefix, so the two never share a cache, a queue or a panel.
const STORE_PREFIX = "bmgr:novel:";
const CACHE_KEY = STORE_PREFIX + "cache";
// Per-device record of when each series was last checked for new chapters.
// Deliberately not synced: each device does its own checking.
const LASTCHECKED_KEY = STORE_PREFIX + "lastchecked";
const LATEST_CHECK_THROTTLE_MS = 4 * 60 * 60 * 1000;
const LATEST_CHECK_BATCH = 1; // series fetched per navigation // series fetched per navigation
// Which library this script's rows belong to. The manga script is a separate
// install that declares "manga"; the backend keeps whichever it is told.
@@ -33,16 +40,11 @@
// ============================================================
// Site adapters
//
// Page type + IDs come from URL regex (most stable); title/cover come from
// og: meta tags (with the novelfull name= meta as the exception).
// Page type + IDs come from URL regex (most stable); the title comes from
// og: meta tags or the page's own heading. Covers are never read here: the
// backend acquires and serves them itself (ADR-0007).
// ============================================================
function meta(prop) {
const el = document.querySelector('meta[property="' + prop + '"]');
return el ? el.getAttribute("content") : null;
}
// Chapter lists are read from two places: the page we are standing on, and
// series pages fetched in the background. Both are reduced to {href, text}
// pairs so each adapter needs only one rule for picking the latest chapter.
@@ -66,12 +68,6 @@
return out;
}
// novelfull ships no og: tags at all — its cover lives on a name= meta.
function metaName(name) {
const el = document.querySelector('meta[name="' + name + '"]');
return el ? el.getAttribute("content") : null;
}
function escapeRe(s) {
return s.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
}
@@ -108,7 +104,6 @@
// h3.title on a chapter page is the *chapter's* title; the breadcrumb
// link back to the series page carries the series name.
title: back ? (back.textContent || "").trim() : "",
cover: metaName("image") || "",
seriesUrl: loc.origin + "/" + m[1] + ".html",
chapterLabel: "Chapter " + m[2],
chapterNum: isNaN(num) ? null : num,
@@ -124,7 +119,6 @@
site: this.site,
seriesId: m[1],
title: h3 ? (h3.textContent || "").trim() : "",
cover: metaName("image") || "",
seriesUrl: loc.origin + "/" + m[1] + ".html",
chapterLabel: null,
chapterNum: null,
@@ -157,9 +151,6 @@
seriesId: m[1],
// The heading is "<Series> Chapter <n>"; drop the suffix.
title: heading.replace(/\s*Chapter\s+[0-9.]+\s*$/i, "").trim(),
// Chapter pages carry no og:image. Empty is safe: every write merges
// against the cached row, which keeps the cover the series page gave.
cover: "",
seriesUrl: "https://lightnovelworld.net/novel/" + m[1] + "/",
chapterLabel: "Chapter " + m[2],
chapterNum: isNaN(num) ? null : num,
@@ -175,7 +166,6 @@
site: this.site,
seriesId: m[1],
title: h1 ? (h1.textContent || "").trim() : "",
cover: meta("og:image") || "",
seriesUrl: "https://lightnovelworld.net/novel/" + m[1] + "/",
chapterLabel: null,
chapterNum: null,
@@ -286,6 +276,10 @@
async function apiPut(key, obj, { sendStatus = false } = {}) {
const body = Object.assign({}, obj);
if (!sendStatus) delete body.status;
// Covers belong to the backend, which acquires and serves them itself
// (ADR-0007) and ignores an incoming one; a third-party address must never
// go back on the wire.
delete body.cover;
const res = await fetch(API_BASE + "/bookmarks/" + encodeURIComponent(key), {
method: "PUT",
headers: authHeaders({ "Content-Type": "application/json" }),
@@ -613,7 +607,6 @@
series_id: p.seriesId,
title: p.title || (existing && existing.title) || p.seriesId,
series_url: p.seriesUrl || (existing && existing.series_url) || "",
cover: p.cover || (existing && existing.cover) || "",
last_chapter: p.chapterLabel || (existing && existing.last_chapter) || "",
last_chapter_num:
p.chapterNum != null ? p.chapterNum : existing ? existing.last_chapter_num : null,
@@ -636,7 +629,6 @@
series_id: p.seriesId,
title: existing.title || p.title || p.seriesId,
series_url: existing.series_url || p.seriesUrl || "",
cover: existing.cover || p.cover || "",
last_chapter: p.chapterLabel || existing.last_chapter || "",
last_chapter_num: p.chapterNum != null ? p.chapterNum : existing.last_chapter_num,
last_chapter_url: p.chapterUrl || "",
@@ -1147,7 +1139,15 @@
return el("div", { class: "item" + heat }, [
el("a", { class: "go", href: cont }, [
b.cover
? el("img", { class: "cover", src: b.cover, loading: "lazy", alt: "" })
? el("img", {
class: "cover",
src: b.cover,
loading: "lazy",
alt: "",
// A Cover that will not load shows the designed placeholder
// rather than the browser's broken-image glyph (#47).
onerror: (e) => e.target.replaceWith(el("div", { class: "cover ph" })),
})
: el("div", { class: "cover ph" }),
]),
el("div", { class: "meta" }, [
+6 -44
View File
@@ -34,9 +34,6 @@ let metaTags = {};
// document.title. comix's SPA rewrites this on client routing but never
// og:title, so the comix adapter reads it instead. Reassigned per test.
let docTitle = "";
// img[alt] elements comix's coverFromPage() scans. Reassigned per test; each
// entry is {alt, src}.
let pageImages = [];
globalThis.document = {
querySelector(sel) {
const m = sel.match(/^meta\[property="([^"]+)"\]$/);
@@ -44,12 +41,6 @@ globalThis.document = {
const v = metaTags[m[1]];
return v == null ? null : { getAttribute: () => v };
},
querySelectorAll(sel) {
if (sel !== "img[alt]") return [];
return pageImages.map((img) => ({
getAttribute: (attr) => img[attr] ?? null,
}));
},
addEventListener() {},
get title() {
return docTitle;
@@ -100,7 +91,7 @@ test("stripBuildHash ignores suffixes that are not exactly 8 hex chars", () => {
// ============================================================
test("asura.detect reads a series page, stripping the hash from the id only", () => {
metaTags = { "og:title": "Solo Leveling | Asura Scans", "og:image": "https://cdn.example/x.jpg" };
metaTags = { "og:title": "Solo Leveling | Asura Scans" };
const p = asura.detect(loc("https://asurascans.com/comics/solo-leveling-059befe1"));
assert.equal(p.type, "series");
assert.equal(p.site, "asura");
@@ -108,12 +99,11 @@ test("asura.detect reads a series page, stripping the hash from the id only", ()
// seriesUrl keeps the hash: navigation needs the current one (stale ones 302).
assert.equal(p.seriesUrl, "https://asurascans.com/comics/solo-leveling-059befe1");
assert.equal(p.title, "Solo Leveling");
assert.equal(p.cover, "https://cdn.example/x.jpg");
assert.equal(p.chapterNum, null);
});
test("asura.detect reads a chapter page including a decimal number", () => {
metaTags = { "og:title": "Solo Leveling Chapter 12.5 - Read Online | Asura Scans", "og:image": "" };
metaTags = { "og:title": "Solo Leveling Chapter 12.5 - Read Online | Asura Scans" };
const url = "https://asurascans.com/comics/solo-leveling-059befe1/chapter/12.5";
const p = asura.detect(loc(url));
assert.equal(p.type, "chapter");
@@ -150,7 +140,7 @@ test("asura.latestChapterFromAnchors returns null when nothing matches", () => {
// ============================================================
test("demonic.detect reads a series page", () => {
metaTags = { "og:title": "The World After The Fall", "og:image": "https://cdn.example/y.jpg" };
metaTags = { "og:title": "The World After The Fall" };
const p = demonic.detect(loc("https://demonicscans.org/manga/the-world-after-the-fall"));
assert.equal(p.type, "series");
assert.equal(p.site, "demonic");
@@ -159,7 +149,7 @@ test("demonic.detect reads a series page", () => {
});
test("demonic.detect reads a chapter page and strips the suffix from the title", () => {
metaTags = { "og:title": "The World After The Fall Chapter 3", "og:image": "" };
metaTags = { "og:title": "The World After The Fall Chapter 3" };
const p = demonic.detect(loc("https://demonicscans.org/title/the-world-after-the-fall/chapter/3/1"));
assert.equal(p.type, "chapter");
assert.equal(p.chapterNum, 3);
@@ -247,26 +237,6 @@ test("comix parses decimal chapter numbers", () => {
assert.equal(p.chapterNum, 80.5);
});
test("comix.detect reads the cover from an img whose alt matches the cleaned title", () => {
metaTags = { "og:title": COMIX_STALE_HOME };
docTitle = "Dungeons and Crayons";
pageImages = [
{ alt: "Some Other Series", src: "https://cdn.example/other.jpg" },
{ alt: "Dungeons and Crayons", src: "https://cdn.example/cover.jpg" },
];
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
assert.equal(p.cover, "https://cdn.example/cover.jpg");
});
test("comix.detect leaves cover empty when no img alt matches the title", () => {
metaTags = {};
docTitle = "Dungeons and Crayons";
pageImages = [{ alt: "Some Other Series", src: "https://cdn.example/other.jpg" }];
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
assert.equal(p.cover, "");
pageImages = [];
});
test("comix ignores unrelated paths", () => {
assert.equal(comix.detect(loc("https://comix.to/browse")).type, "other");
});
@@ -308,23 +278,18 @@ const KAGANE_SERIES = "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b";
const KAGANE_BOOK = "019fa2e0-6dbd-73ca-b40b-fe06ab75eb0e";
test("kagane detects a series page", () => {
metaTags = {
"og:title": "Infinite Decryption: The Strongest Level 0",
"og:image": "https://kagane.to/api/v2/image/abc/compressed",
};
metaTags = { "og:title": "Infinite Decryption: The Strongest Level 0" };
const p = kagane.detect(loc("https://kagane.to/series/" + KAGANE_SERIES));
assert.equal(p.type, "series");
assert.equal(p.site, "kagane");
assert.equal(p.seriesId, KAGANE_SERIES);
assert.equal(p.title, "Infinite Decryption: The Strongest Level 0");
assert.equal(p.cover, "https://kagane.to/api/v2/image/abc/compressed");
assert.equal(p.seriesUrl, "https://kagane.to/series/" + KAGANE_SERIES);
});
test("kagane reads the chapter number out of og:title", () => {
metaTags = {
"og:title": "Infinite Decryption: The Strongest Level 0 - Chapter 41 - Episode 41",
"og:image": "https://kagane.to/api/v2/image/abc/compressed",
};
const p = kagane.detect(
loc("https://kagane.to/series/" + KAGANE_SERIES + "/reader/" + KAGANE_BOOK)
@@ -341,10 +306,7 @@ test("kagane reads the chapter number out of og:title", () => {
// no episode name, because the book carries volume_no and an empty title.
// Captured live 2026-08-08 from SP Baby.
test("kagane reads through a Volume-numbered chapter suffix", () => {
metaTags = {
"og:title": "SP Baby - Volume 1 Chapter 1",
"og:image": "https://kagane.to/api/v2/image/abc/compressed",
};
metaTags = { "og:title": "SP Baby - Volume 1 Chapter 1" };
const p = kagane.detect(
loc("https://kagane.to/series/" + KAGANE_SERIES + "/reader/" + KAGANE_BOOK)
);
+27 -17
View File
@@ -4,8 +4,7 @@ const test = require("node:test");
const assert = require("node:assert");
// ============================================================
// Minimal browser stub. Same shape as logic.test.js, plus a meta[name=...]
// branch: novelfull ships no og: tags, so its cover comes from name="image".
// Minimal browser stub. Same shape as logic.test.js.
// document.body stays UNDEFINED so the boot block waits for a DOMContentLoaded
// that never fires and no network call is ever made.
// ============================================================
@@ -20,20 +19,14 @@ globalThis.localStorage = {
globalThis.location = { href: "about:blank", hostname: "", pathname: "/", origin: "" };
let metaTags = {};
let namedMetas = {};
let elements = {};
globalThis.document = {
querySelector(sel) {
let m = sel.match(/^meta\[property="([^"]+)"\]$/);
const m = sel.match(/^meta\[property="([^"]+)"\]$/);
if (m) {
const v = metaTags[m[1]];
return v == null ? null : { getAttribute: () => v };
}
m = sel.match(/^meta\[name="([^"]+)"\]$/);
if (m) {
const v = namedMetas[m[1]];
return v == null ? null : { getAttribute: () => v };
}
const text = elements[sel];
return text == null ? null : { textContent: text };
},
@@ -58,7 +51,6 @@ function loc(href) {
function reset() {
metaTags = {};
namedMetas = {};
elements = {};
}
@@ -68,21 +60,18 @@ function reset() {
test("novelfull.detect reads a series page", () => {
reset();
namedMetas = { image: "https://novelfull.com/uploads/thumbs/ri.jpg" };
elements = { "h3.title": "Reverend Insanity" };
const p = novelfull.detect(loc("https://novelfull.com/reverend-insanity.html"));
assert.equal(p.type, "series");
assert.equal(p.site, "novelfull");
assert.equal(p.seriesId, "reverend-insanity");
assert.equal(p.title, "Reverend Insanity");
assert.equal(p.cover, "https://novelfull.com/uploads/thumbs/ri.jpg");
assert.equal(p.seriesUrl, "https://novelfull.com/reverend-insanity.html");
assert.equal(p.chapterNum, null);
});
test("novelfull.detect reads a chapter page and points seriesUrl at the series", () => {
reset();
namedMetas = { image: "https://novelfull.com/uploads/thumbs/ri.jpg" };
elements = { "a.truyen-title": "Reverend Insanity" };
const url = "https://novelfull.com/reverend-insanity/chapter-2334-fang-yuan.html";
const p = novelfull.detect(loc(url));
@@ -120,14 +109,12 @@ test("novelfull.latestChapterFromAnchors takes the max and ignores other series"
test("lightnovelworld.detect reads a series page", () => {
reset();
metaTags = { "og:image": "https://lightnovelworld.net/wp-content/uploads/awe.webp" };
elements = { "h1.entry-title": "A Will Eternal" };
const p = lightnovelworld.detect(loc("https://lightnovelworld.net/novel/a-will-eternal/"));
assert.equal(p.type, "series");
assert.equal(p.site, "lightnovelworld");
assert.equal(p.seriesId, "a-will-eternal");
assert.equal(p.title, "A Will Eternal");
assert.equal(p.cover, "https://lightnovelworld.net/wp-content/uploads/awe.webp");
});
test("lightnovelworld.detect strips the chapter suffix off the heading", () => {
@@ -141,8 +128,6 @@ test("lightnovelworld.detect strips the chapter suffix off the heading", () => {
assert.equal(p.chapterLabel, "Chapter 1298");
assert.equal(p.title, "A Will Eternal");
assert.equal(p.seriesUrl, "https://lightnovelworld.net/novel/a-will-eternal/");
// Chapter pages have no cover; the merge in bookmarkCurrent keeps the stored one.
assert.equal(p.cover, "");
});
test("lightnovelworld.detect returns other for non-series paths", () => {
@@ -180,3 +165,28 @@ test("kindOf defaults a missing kind to manga", () => {
test("kindOf passes through novel", () => {
assert.equal(kindOf({ kind: "novel" }), "novel");
});
// ============================================================
// Source guard
//
// Issue #74: the novel script was split off the manga one and lost four
// module-scope constants. The reads sit inside try/catch or a fire-and-forget
// promise, so the ReferenceError never surfaced — nothing but a static check
// catches this class.
// ============================================================
test("every SCREAMING_CASE constant the script uses is declared in it", () => {
const fs = require("node:fs");
for (const f of ["novel-bookmark.user.js", "manga-bookmark.user.js"]) {
const src = fs.readFileSync(require.resolve("../" + f), "utf8")
// comments and strings carry prose and SVG path data in the same shape
.replace(/\/\/[^\n]*|\/\*[\s\S]*?\*\/|"[^"\n]*"|'[^'\n]*'|`[\s\S]*?`/g, " ");
const declared = new Set(
[...src.matchAll(/\b(?:const|let|var|function)\s+([A-Z][A-Z0-9_]{2,})\b/g)].map((m) => m[1]),
);
for (const name of new Set(src.match(/\b[A-Z][A-Z0-9_]{2,}\b/g) || [])) {
if (name.startsWith("GM_") || name in globalThis) continue;
assert.ok(declared.has(name), `${f} uses ${name} but never declares it`);
}
}
});