Compare commits

..

47 Commits

Author SHA1 Message Date
sulthan 2be9256e5e chore: refresh the graphify map after the #102 review fixes 2026-08-16 14:31:11 +07:00
sulthan 58014eb8dd review: verdigris patina, honest Lane figures, tests that can fail (#102)
Two-axis review found the shipped colour and three weak tests.

- --patina was a warm gold at the same hue family as --brass; the spec
  asked for a cool blue-green so an unhealthy Lane is unmistakably not
  ember. Now verdigris in both branches, with docs/design-system.md
  stating why the far side of the wheel is the point.
- A Lane pass that returns before computing its figures carries the
  previous pass's due count and gap forward instead of recording zeroes,
  and Checked rides beside Due so a stopped Lane is distinguishable from
  a quiet one.
- TestAdminPageWithoutAPollerSaysSo now separates the two causes it
  conflated, TestOwnerClearsReaderMarks seeds real counters so the
  clearing assertion can fail, and TestLaneStatus asserts Checked and
  the carry-forward.
- backend/AGENTS.md records the one deliberate owner comparison outside
  requireOwner and the carry-forward rule.
2026-08-16 14:31:00 +07:00
sulthan d878580df8 chore: refresh the graphify map for the admin page (#102) 2026-08-16 14:18:08 +07:00
sulthan 0b946ee2c4 test, docs: admin page gate tests and the page's written rules (#102)
- web_test.go walks web.AdminPatterns() rather than naming routes by hand, so a
  new administrative route that forgets requireOwner fails the gate test
  instead of shipping open.
- The harnesses take a LaneReporter; a fake one keeps the page's tests free of
  a poller and a Site.
- backend/AGENTS.md records the adminRoutes/requireOwner rule and the nil-poller
  trap; design-system.md records --patina and the admin page's shape.
2026-08-16 14:17:36 +07:00
sulthan e69da2a99b web: owner-only /admin with the Reader roster and Poll Lane status (#102)
The only operational surface was /healthz and a fold-out roster inside the
owner's own reading page. This gives the owner a page: two sections of facts on
the same measured sheet, no cards.

- admin.go holds every route that reaches past the acting Reader, listed once
  in adminRoutes() and wrapped in requireOwner at registration - a missing gate
  is visible in the route list rather than hidden inside a handler. A non-owner
  gets 404, the same answer revoke already gave.
- Lane figures arrive through the LaneReporter seam, so the page reads the
  running poller rather than a table. newRouter converts a nil *Poller to a nil
  interface: a typed nil would make the page claim a poller exists.
- The roster moves out of the reading page and gains the Sighting counters, the
  blocked verdict and Clear marks. Clearing restores a privilege, so it is a
  plain ghost button; --danger stays with revocation.
- --patina is the page's one accent, held at the weight of the other action
  accents. Neither --ember (new chapter) nor --danger (destruction) is borrowed
  for system health.
2026-08-16 14:17:30 +07:00
sulthan d555f54529 latest: report each Poll Lane's last pass to the owner's page (#102)
A Lane's pace, its refusal backoff and whether its Site needs the browser were
only ever visible in the log. The admin page (issue #102) has to state them, so
the Poller keeps one snapshot per Lane.

- LaneState/Status (status.go) is the page's view of a Lane: due, gap, clamped,
  refusing, browser.
- runOnce records a snapshot on every return path, including the pass that
  refuses, so a cooling Lane does not read as one that never ran.
- Refusing is derived at snapshot read time from the backoff, not stored: a
  Lane that cooled down between passes must report false without a new pass.
- LaneStatus reports only Sites that have completed a pass, so a fresh restart
  renders "no data yet" instead of confident zeroes.
2026-08-16 14:17:22 +07:00
sulthan cdd4eb7c25 store: Sighting counters and the owner's clear control (#102)
The owner's page (issue #102) shows each Reader's Sighting record and offers
one action to wipe it. The counters ship before the Sighting feature that
moves them (issue #103) on purpose: the remedy for a false mark must exist
before marks can be made, or the first broken adapter is fixed with SQL
against production.

- 0010 adds sighting_agreements / sighting_disagreements, both zero-defaulted,
  so every existing Reader reads as trusted.
- ReaderSummary carries both, plus Blocked() against SightingDisagreementLimit,
  so the page states the verdict rather than making the owner derive it.
- ClearReaderMarks zeroes one Reader's pair.
2026-08-16 14:17:10 +07:00
sulthan 3303a55b20 feat: one Poll Lane per Site, replacing the shared pace (#100) (#106)
Closes #100.

Each Site runs its own Poll Lane: an independent goroutine with its own rest
and pace from the registry (`backend/internal/latest/sites.go`), replacing the
shared cooldown/interval/stagger/batch configuration. Rest (1h, all six Sites
including the browser trio) is enforced by the due query's WHERE clause; the
Lane sleeps its effective gap between fetches — the registry 10s, or
rest/eligible when a Site holds enough Series, floored at 1s with a
Site-naming warning when the floor engages.

Lane-local failure handling:
- Two challenge-held results stop that Site's Lane for 15m; the probes keep
  their stamp, untried Series stay due.
- A lost browser sets a shared Poller flag: the other browser Lanes skip
  their passes for the same 15m (no stamp-per-pass-per-Lane on a dead tab),
  then decay and probe again.
- Browser wake gate preserved (5 due, or one waiting 15m, ADR-0005); one tab
  shared by the three browser Sites; "browser lane behind by X" logged every
  pass.
- Cover work (healing a stored source URL and filling a blank from the series
  page) runs in the background so a slow CDN cannot consume a Lane's gap.

Removed: `LATEST_CHAPTER_POLL_{COOLDOWN,BROWSER_COOLDOWN,INTERVAL,BATCH,STAGGER}`
and the 6h browser rest. Only `LATEST_CHAPTER_POLL_ENABLED` remains; DEPLOY.md
documents the exact `.env` edit. ADR-0010 records the decisions.

Reviewed-on: #106
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-16 13:16:19 +07:00
sulthan ddbd57070d Poll comix.to through the browser sidecar (#98) (#105)
Closes #98.

comix.to began answering plain-TLS fetches with a Cloudflare JavaScript
challenge on 2026-08-12, so every poll got a 403 interstitial. Its cover host
`static.comix.to` is gated the same way. comix therefore joins kagane and
novelfull as a browser-backed Site.

## What changed

- **Registry** (`internal/latest/sites.go`): comix gains a `Browser` entry —
  `comixRead`, `Done: body != "" && !isInterstitial(body)`, `Fallback: false`.
  Skip-when-no-browser falls out of the existing routing; no site-string compare
  was added anywhere.
- **Read shape** (`internal/latest/browser.go`): an in-tab `fetch()` of the
  Series URL, not a DOM render. comix is an SPA — rendering it costs ~65
  requests for the same server-rendered HTML one fetch returns (24.5 KB,
  ~480 ms measured). `comixSeriesPageURL` pins scheme + host + `/title/<slug>`
  and rebuilds the address, so a client-supplied `series_url` cannot aim the
  browser anywhere else.
- **Cover bytes**: `comixImageURLRe` pins `https://static.comix.to/<path>.<ext>`;
  `BrowserFetcher.Image` now gates on `browserOnlyCoverURL` rather than a
  kagane-only regex, so both Sites' image URLs route through the one path.
  Bytes come from direct navigation, not a page-context fetch — comix's Series
  page sets `cross-origin-embedder-policy: require-corp`, which fails one.
- **Parsers and stored Series identity: untouched.** The in-tab body is the same
  server-rendered HTML the existing fixtures were cut from.

## Verification

- `go test ./...` green (needs Docker).
- New seam tests: comix routes to the browser when one is configured, and is
  not fetched at all when none is (`TestComixUsesBrowserFetcher`,
  `TestComixSkippedWhenNoBrowserFetcher`); URL-pin and cover-gate table tests.
- Live proof against the real browser unit, `TestSmokeComix` (env-gated):
  page 24793 bytes in one in-tab fetch, chapter 53, cover accepted by the pin,
  26862 bytes of `image/jpg` retrieved.
- Two-axis review run; findings were stale comments on `BrowserFetcher`, `Get`
  and the `Fallback` field, fixed in f000cc7.

Docs updated: root `AGENTS.md` (constraint + smoke command, including the note
that this dev machine's ISP DNS-hijacks `comix.to`), `backend/AGENTS.md`
(poller, cover pipeline, `BROWSER_WS_URL`), `REDEPLOY.md` §8 degrade note.

Reviewed-on: #105
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-16 12:13:59 +07:00
sulthan 3f53c79cf4 docs: add Poll Lane and Sighting to the shared vocabulary (#104)
Two new glossary terms in `CONTEXT.md`, settled in a design session, plus the graphify refresh.

- **Poll Lane** — one Site's own stream of Polls, carrying the pace at which that Site is willing to be asked. No Lane can slow, block or borrow from another's; a Reader never has one.
- **Sighting** — what a Reader's browser happened to see of a Series's Latest Chapter. Reports the same fact as a Poll, carries none of its authority.
- **Latest Chapter** amended: it no longer claims to be discovered without the reader present, since a Sighting establishes it between Polls.

No code. The work these terms describe is specified in #98 (comix), #101 (Poll Lanes), #102 (admin page) and #103 (Sightings).

Reviewed-on: #104
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-16 11:42:16 +07:00
sulthan 17ee0bd3f8 docs: correct the bot-score claims behind the browser poller (#99)
Docs only. No code changes - `git diff origin/main --stat` touches five Markdown files and adds one research note.

## What was wrong

Several docs explained Cloudflare challenges as a "bot score" that our request rate could worsen. That mechanism does not exist on these sites.

Researched live on 2026-08-12 against Cloudflare's own documentation and blog plus RFC 9309 - 22 primary pages, every claim carrying a source URL and read date, seven areas explicitly marked `Not publicly documented`. The note is `docs/research/cloudflare-bot-scoring-and-poll-cadence.md`.

- The 1-99 bot score is **Enterprise Bot Management only**. A free-plan zone has no score at all; it gets Bot Fight Mode, which matches *signatures* (headless browsers, cloud-hosting IPs).
- **No per-IP request rate is documented as an input to challenge issuance.** Volume is policed by Rate Limiting Rules, a separate opt-in product: one rule, IP-only counting, 10-second windows on Free. Published DDoS thresholds are ~1,000 errors/sec.
- **`cf_clearance` defaults to 30 minutes**, so every cadence at or above 1 hour re-solves the challenge anyway. Cadence changes how many ~4s solves happen per day and nothing else.
- The documented risk is **fingerprint quality**, which this repo already solved (real Chrome, stock UA, non-UTC clock).

## What changed

| File | Correction |
|---|---|
| `AGENTS.md` | The block is per-zone configuration plus request fingerprint, not IP reputation. comix.to turning its gate on 2026-08-12 is the worked example. Residential egress avoids the cloud-hosting-IP *signature* rather than earning a better score. The UTC measurement stands; its mechanism is now marked undocumented. |
| `backend/AGENTS.md` | Says why `_BROWSER_COOLDOWN` is longer: cost, not safety. |
| `docs/adr/0003` | Dated correction - the sites do not "bot-score" the VPS IP. Decision stands on its sweep-depth argument. |
| `docs/adr/0006` | Dated correction - no score to be better at. Decision stands on VPS memory. |
| `DEPLOY.md` | A red kagane smoke run means the Site's settings or this Chrome's fingerprint moved, not "Cloudflare's scoring". |

ADRs got dated `Corrected 2026-08-12:` paragraphs rather than silent rewrites - the record of what was decided stays intact, only the wrong mechanism is retracted.

## Deliberately not in this PR

- **The 6h browser cooldown is unchanged.** I had lowered it to 1h and reverted that; cadence is a behaviour change and belongs with the comix work in #98, not in a docs correction.
- **Two code comments still carry the myth**: `backend/main.go:83-84` ("a hammer against sites that are already bot-scoring us"). Left alone to keep this diff docs-only.

Related: #98.
Reviewed-on: #99
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-12 09:32:07 +07:00
sulthan c62c3bb07b chore: drop asuracomic.net from the userscript, CORS allowlist and docs (#97)
Closes #96.

## What

Removes every reference that still invites a Reader onto `asuracomic.net`.
The domain's deep links 301 to the `asurascans.com` **root**, discarding the
path (re-checked 2026-07-25), so a page on it never yields a series document
client-side and a stored address on it never yields a series page server-side.
#95 already pinned each Site to one hostname, so the backend rejects such an
address cleanly; this is the cleanup around that.

| File | Change |
|---|---|
| `userscript/manga-bookmark.user.js` | drops the `@match`, narrows the asura adapter to `/(^\|\.)asurascans\.com$/` |
| `userscript/test/logic.test.js` | new test pinning the narrowed host match |
| `.env.example`, `docker-compose.yml` | origin dropped from the `ALLOWED_ORIGINS` default |
| `DEPLOY.md` | same, and the sample list gains the two novel origins it was missing |
| `backend/api_test.go` | CORS fixtures and round-trip seed move to `asurascans.com` |
| `README.md`, `AGENTS.md` | notes say the host is dropped, not "stays matched" |

## Behaviour

- A Reader landing on `asuracomic.net` gets no userscript UI. Previously the
  script loaded and could do nothing useful — the redirect had already
  discarded the path.
- A request whose `Origin` is `https://asuracomic.net` is no longer reflected
  by a deployment using the shipped defaults.
- No backend logic changed: the CORS rule, the address gate and the poller are
  untouched. `AllowedOrigins` is data, not code.

## Security invariant preserved

CORS still reflects `Origin` only when it appears in `ALLOWED_ORIGINS`, with
`GET,PUT,DELETE,OPTIONS` and a `204` preflight — `TestCORSPreflight` and
`TestCORSDisallowedOrigin` still pin both halves, now against a live origin.
This change only removes a value from the allowlist, which is a narrowing.

## Verification

- `go test ./...` — full backend suite green (real Postgres per package).
- `node --test test/*.test.js` — 66/66 green, up one from the new match test.

## Deploy note (does not happen on merge)

The live allowlist comes from the VPS `.env`, not from these defaults, so the
origin must be dropped there in the same deploy. The one-off row repair for any
stored `asuracomic.net` address is in #96.

Reviewed-on: #97
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-12 05:53:17 +07:00
sulthan 21615be2bd feat: one registry entry per Site, one shared Series-page read (#95)
Closes #94.

## What

Two phases per the spec, in three feature commits plus two review-fix commits:

**Phase one — one registry entry per Site** (`2d134fb`)
The six per-site comparison points that used to live across three files collapse
into one `sites` map in `backend/internal/latest/sites.go`: Latest Chapter parse,
Cover parse, browser-backed list, fetcher route, host pins, and the browser
payload read all become lookups into it. `browserBackedSites()` is derived from
the registry (sorted, deterministic); `fetcherFor` and `fetchableSeriesURL` keep
their signatures and become lookups; `BrowserFetcher.Get` dispatches through the
entries' `Read`/`Done` while the tab lifecycle stays in `BrowserFetcher.run`.

**Phase two — one shared Series-page read** (`f215130`)
`readSeriesPage` (new `read.go`) performs the read the Poll and the Acquisition
have in common: gate, route, fetch, parse Latest Chapter, parse Cover address.
It returns facts only — polling and persistence policies (stamp order, cooldowns,
cover policy) stay with the callers; `acquire.go` gained the comment naming the
deliberate post-fetch stamp order. The poll's legacy cover heal and the
no-chapter byte-count diagnostic were restored after review (`d998f87`) so the
claims "the Poll keeps its own Cover policy" and "pinning is the only
behavioural change" both hold.

## Behaviour

- All six Sites now pin their host exactly; asura/demonic/comix previously
  accepted any https host. For asura this is a strict improvement: its dead old
  domain redirects deep links to the site root and would parse the wrong
  document.
- Everything else is unchanged: existing parse tables, the challenge-body table
  and the gate table pass unmodified except the one deliberate exception — the
  gate table gains the three new pin cases.

## Security invariants preserved

- The address gate is recognisably the same rule, now a single registry lookup:
  `https` + exact hostname match, all callers route through it. No fetch path
  was widened; asura/demonic/comix were narrowed.
- The second host pin inside each browser entry's Read is retained deliberately
  (browser = strong SSRF primitive, `series_url` is client-supplied) and is not
  deduplicated against the shared gate.
- Review hardening: `fetcherFor` now fails closed for unknown site strings
  (previously fell through to the TLS fetcher on an unreachable path), and the
  browser dispatch iterates a sorted list so outcomes cannot depend on map order.
- The security review's log-injection finding was checked against Go's
  `url.Parse` and does not hold: control characters are rejected anywhere in a
  URL, so a client-supplied value in a log line cannot carry a newline.

## Review

Reviewed on three axes (spec, standards, security) by read-only subagents over
`672c16f..f1b26f4`. No blocking findings; all minor/nit findings addressed in
`d998f87` and `700de20`. Verified end to end with `go test ./...` (Docker
Postgres per test package) on every commit.

## Out of scope (tracked separately)

- Dropping asuracomic.net (CORS allowlist, userscript match, API fixtures,
  live env) — separate issue, per spec.

Reviewed-on: #95
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-12 05:43:37 +07:00
sulthan 672c16ffbf Remove the client latest-chapter scan for lightnovelworld (#91) (#93)
Follows the spec published on #91: remove the client-side latest-chapter scan for lightnovelworld rather than porting #87's truncation into a second codebase.

- lightnovelworld.latestChapterFromAnchors deleted, not stubbed: absence is what the background-fetch guard keys off.
- computeLatestChapter tolerates an adapter with no scanner (yields null) and is exported as the test seam.
- backgroundRefreshLatest skips a scanner-less Site before the due filter: no Series page fetched, no freshness timestamp recorded, no batch slot consumed. The on-page path (maybeCaptureLatestOnSeriesPage) routes through the same null-tolerant computation.
- novelfull's scanner, the shared max-chapter helper and all four manga Sites untouched.
- userscript/AGENTS.md records the Poll-only contract for this Site and why.

All seven acceptance criteria from the spec met. node --check clean; novel suite 30/30 (the regression pin fails if a lnw scan is reintroduced, scoped or not); manga suite 35/35, manga userscript byte-for-byte unchanged.

Two-axis code review: no hard standard violations, spec-clean; one follow-up commit matching the sibling adapter guard from the manga script.

Reviewed-on: #93
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-11 20:58:18 +07:00
sulthan e7e22a12a5 lightnovelworld Series identity is read from the chapter page (#80) (#92)
Implements spec #80 / ADR-0008 — Gitea issues #86, #87, #88, #89, #90, all closed.

A Reader bookmarks a novel on lightnovelworld and it never shows a New Chapter, because the Series identity was derived from the chapter address instead of read from the page. One Series can publish under several Chapter Slugs, so the derived key points at a slug that 404s.

- **#89** — the userscript's lnw adapter stops deriving `seriesUrl`/`seriesId` from the path. It reads the page's own pointer (`a[aria-label='All Chapter']`), falling back to the microdata breadcrumb's second crumb, and carries `chapterSlug` on the page object, stored nowhere.
- **#87** — the Poll's lnw chapter scan is unscoped (no stored-slug pattern can cover a Series' whole list) and truncated at the `wpd-threads` comment thread, the one region a visitor can write to. Marker absent means skip and log with the body length, never scan whole. Corrects the `maxBodyBytes` headroom comment to the measured 3.5x.
- **#86** — the scan fixture is now text trimmed from a real, wholly-fetched Series page instead of a hand-written cross-series anchor that no live page carries.
- **#90** — stale stored rows repair themselves on the next chapter visit: a pure transform over cache, queue and last-checked map, silent to the Reader, with progress, favourite and lifecycle bucket preserved when two rows merge.
- **#88** — an env-gated live canary (`SMOKE_LNW_SERIES_URL`) proving the marker still occurs exactly once and still follows the last chapter anchor, asserted against the production symbols themselves.

Verified on the merged branch: `go test ./...` green, `gofmt -l internal/latest/` silent, both userscripts `node --check` clean, 35/35 + 29/29 logic tests. Live canary green (marker once at byte 612,182 of 651,795). #90 verified on device with Playwright.

Open follow-up: **#91** — the userscript's client-side latest-chapter scan is still scoped to the derived slug.

Reviewed-on: #92
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-11 18:21:50 +07:00
sulthan 90d8ab72ad chore: refresh graphify map and tidy AGENTS.md (#84)
graphify update regeneration: semantic hashes now populated in manifest.json, graph rebuilt (1634 nodes, 3179 edges). Track the map (5 curated files + .graphify_root) so a fresh checkout starts with it; cost.json, cache/, dated snapshots and .rebuild.lock stay ignored.

AGENTS.md: drop stale Relevant skills and Notes sections; graphify rule now says to always query the graph first.

Reviewed-on: #84
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-11 11:19:47 +07:00
sulthan c400c91a80 Implement a batch of tickets through per-ticket subagents (#82)
## What this adds

Two files that turn the one-ticket-at-a-time `/implement` loop into an orchestrated batch.

**`.claude/skills/implement-tickets/SKILL.md`** — user-invoked (`disable-model-invocation: true`, so it costs no context until typed). The agent that runs it is an orchestrator, not an implementer:

1. Collect the tickets over `tea`, reading each `Blocked by` line.
2. Plan waves from the blocking edges, three tickets wide, and fix every cross-ticket contract (shared signature, JSON shape, column, token) before anything is dispatched.
3. Present the plan and stop for approval.
4. Per ticket: `git worktree add ../ticket-<n>`, copy the gitignored `.env`, claim the issue, write a brief to `.scratch/`, then dispatch the whole wave as one `task` batch.
5. Land each result — merge `--no-ff`, comment the report, close, remove the worktree. Textual conflicts are the orchestrator's; a semantic clash goes back to whichever ticket owns the contract.
6. Full suite once on the merged base.

**`.omp/agents/ticket-implementer.md`** — the worker. Brief-driven, worktree-bound, and gated on review before it reports: it runs the `code-review` skill over its own diff with `cr-spec` and `cr-standards` on the two axes, fixes Critical and Important findings in at most two rounds, and returns a short status contract (`DONE` / `DONE_WITH_CONCERNS` / `BLOCKED` / `NEEDS_CONTEXT` / `REVIEW_BLOCKED`).

The brief template makes the subagent read `tea issue <n> --comments` for its ticket and for the issue that ticket refers to — the comments carry decisions the body never got updated with — and names the `tdd` skill at each seam where a test comes first. Briefs are written in the ubiquitous language of `CONTEXT.md`; a brief that says "scrape" where the domain says Poll hands the subagent the wrong model of the system.

## Verification

Dispatched a real `ticket-implementer` as a probe. The agent resolved from `.omp/agents`, and it spawned `cr-spec`, which replied. That was the one thing that could have silently killed the design: `task.maxRecursionDepth` defaults to 2, and the chain is session to orchestrator to implementer to reviewer. It clears. If that ever changes, the implementer returns `REVIEW_BLOCKED` and the orchestrator runs the review itself.

Confirmed against the omp binary that `autoloadSkills: code-review, tdd` is split by `parseArrayOrCSV`, not swallowed as one unknown name.

## Notes

- Agents are discovered from `.omp/agents`, never `.claude/agents` — the latter is deliberately skipped by omp because its frontmatter is a different contract.
- No product code changes. `.gitignore` gains `.scratch/`, where briefs and reports live.
- Not included: retry after a failed dispatch, a state file for resuming a crashed wave, a cheap model tier for mechanical tickets. Add them when a real batch needs them.

Reviewed-on: #82
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-11 10:08:09 +07:00
sulthan f1eb7d514c Record the lightnovelworld series-identity decision (#77) (#81)
Docs only. No code, no tests, nothing to run. Implementation is specified in #80.

Outcome of a grilling session on 2026-08-11 against #77, backed by live measurement of lightnovelworld over 2026-08-10/11.

## What changed

**`docs/adr/0008-series-identity-is-discovered-not-derived.md`** (new)

A Series identity is discovered from the Site's own links, never derived from an address.
On lightnovelworld the userscript reads the chapter page's `All Chapter` anchor instead of
building a `/novel/<slug>/` address by string manipulation. A Chapter Slug is not an
identity and is not stored. The backend's chapter scan drops its per-Series scoping and
runs against the body truncated before the visitor comment thread.

Evidence in the ADR: 3 of 41 sampled novels serve chapters under a slug that differs from
their series slug, divergence runs in both directions, one novel serves chapters under two
slugs, and neither slug is computable from the other. The pointer was checked on 8 chapter
pages and agreed every time. Three narrower selectors are recorded as rejected, each with
the measurement that killed it.

Three rejected options are recorded with reasons: correcting the stored address only, which
keeps an identity the Site does not guarantee; scoping the scan to a container, which the
probe refuted; and a SQL migration, which is impossible because the database holds no
source for the correct slug.

**`CONTEXT.md`**

- **Series** - identity is the canonical slug the Site publishes, never the title and never a Chapter Slug.
- **Chapter Slug** - new term. A slug a Site builds its chapter addresses from. Not an identity: one Series may have several, and none is computable from another.
- **Latest Chapter** - now the highest-numbered chapter, explicitly not a date and not the Site's own newest-chapter banner. Settles #79.

**`docs/research/lightnovelworld-chapter-vs-series-slug.md`** (new, committed with its corrections)

The 41-novel survey behind the ADR. Two claims are struck through and corrected in place,
with the date and sample size of the probe that refuted each: the `ul.clstyle` container it
named is the hidden, empty "Latest Reading" template rather than the chapter list, and its
caveat about the comment region understated the risk, because that region is writable by
any visitor while the scan takes an unbounded maximum into a Series row shared by every
Reader (ADR-0003).

## Review notes

Nothing here constrains code that exists today - the ADR describes work not yet written.
The part worth disagreeing with, if any of it is wrong, is the fail-closed rule: a missing
truncation marker means skip the Series and log, never scan the whole page.

Related: #77 (the defect), #80 (the spec), #79 (the numbering anomaly, closed by decision),
#71 (the same size cap seen from the cover side).

Reviewed-on: #81
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-11 09:34:26 +07:00
sulthan 1ee5eb67ea Clear the stale Chrome singleton lock at browser boot (#76)
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 19:11:24 +07:00
sulthan b22ae82897 Restore the novel script's lost module-scope constants (#74) (#75)
> *This was generated by AI during triage.*

Fixes #74.

`novel-bookmark.user.js` was split out of `manga-bookmark.user.js` and lost four module-scope constants. Every use of them is behind a `try/catch` or a fire-and-forget promise, so the `ReferenceError`s were swallowed rather than reported.

| constant | used at | effect while missing |
| --- | --- | --- |
| `LATEST_CHECK_THROTTLE_MS` | `:722` | `backgroundRefreshLatest()` throws before computing `due` — no background latest-check ever runs for novels (the symptom in #74) |
| `LATEST_CHECK_BATCH` | `:724` | same throw |
| `CACHE_KEY` | `:215`, `:224` | `loadCache()` always returns `[]`, `saveCache()` silently no-ops — the local cache never persists |
| `LASTCHECKED_KEY` | `:232`, `:241` | last-checked map never persists, so the throttle would not hold even once the first two are defined |

#74 named only the two throttle constants. The two cache keys are the same lost lines with the same root cause, so they are restored here too — fixing only the pair the issue named would leave `backgroundRefreshLatest()` re-fetching every series on every navigation, because `saveLastChecked()` would still be a no-op.

Values and comments copied verbatim from `manga-bookmark.user.js:37-43`; throttle 4h, batch 1.

## Verification

- `node --check userscript/novel-bookmark.user.js` — clean.
- `node --test userscript/test/logic.test.js userscript/test/novel-logic.test.js` — 47/47 pass.
- New test `every SCREAMING_CASE constant the script uses is declared in it` scans both scripts (comments and string literals stripped first, so prose and SVG path data do not trip it). Confirmed it fails — 1 failing test — when `LATEST_CHECK_BATCH` is deleted again, and passes when restored.

A behavioural test cannot reach this: the storage helpers and the background refresh are exactly the layers the harness does not cover (see the `testing-the-userscript` skill), and the errors are swallowed anyway. A static guard is the only instrument that sees this bug class.

On-device confirmation that the ember now lights for novels is still outstanding — that needs Violentmonkey against a live novelfull/lightnovelworld page.

Reviewed-on: #75
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 18:12:03 +07:00
sulthan 7c7d597019 Delete the kagane-specific cover path (#63) (#73)
Closes #63

Deletes the second way to reach a Cover. Since #62, every Site's cover bytes land in the content-addressed store at creation or on the poll, and the one public route serves them all — nothing needs the kagane proxy anymore.

## What went

- **Template-level rewrite:** `Bookmark.CoverURL()` and both templates' use of it. Cards and chrome now render `.Cover` — the wire value — and nothing else. `Bookmark.CoverSource` was dead once `CoverURL` went, so it and its `bookmarkColumns` entry are gone too.
- **Kagane-only cover route and its identifier validation:** `GET /img/kagane/{id}`, `web.CoverFetcher`, `coverIDRe`, and the whole `internal/web/cover.go`.
- **The proxy's persistence:** `store.KaganeImageID`, `GetKaganeCover`, `PutKaganeCover`, `kaganeCoverSourceURL`, `kaganeCoverRe`.
- **The kagane-shaped branch in the byte-fetch routing:** `fetchCoverBytes` no longer takes a `site` argument and no longer names a Site. The URL shape kagane's API publishes is claimed by the browser module itself — `kaganeImageURLRe` + `browserCoverURL` live in `latest/browser.go` with the rest of the per-Site knowledge — and `BrowserFetcher.Image` is now URL-driven (it validates the URL it will navigate to, same SSRF discipline as before). The no-plain-TLS-fallback rule for a claimed URL is preserved: a claimed address with no browser is an error, never a challenge-page fetch.

## What stayed (deliberately)

- `BrowserFetcher.Image` and the browser-backed acquisition path: kagane genuinely serves cover bytes behind the challenge + `cross-origin-resource-policy: same-origin`, so the sidecar remains the only fetcher for them — it just routes by URL claim now instead of by Site name.
- `fetcherFor`'s per-Site page routing (kagane/novelfull page fetches) — that is the page path, not a cover path.

## Acceptance criteria

- [x] Template-level kagane cover rewrite gone
- [x] Kagane-only cover route and its identifier validation gone
- [x] Tests removed/rewritten against the general route, guarantees kept: unstored + traversal-shaped addresses serve nothing (`TestPublicCoverRejectsUnknownAddress`), non-image content types never echoed (`TestPublicCoverNeverEchoesNonImage` — new; the store-side gate was already pinned by `TestCoverStoreAcceptsAnySourceURL`). Store reopen-persistence and filesystem content-addressing tests rewritten against `PutCover`/`GetCover`, no guarantee lost.
- [x] No Site name in a cover code path outside the acquisition module (`grep kagane backend`: store/web/templates/api are clean; remaining hits are `latest/browser.go` + `latest/sites.go`, tests, docs)
- [x] Web UI and panel render Covers for all six Sites (templates render the wire address; panel renders `b.cover` — untouched, it never had a kagane path)
- [x] `go test ./...` green

## Verification

- `go vet ./...` clean
- `go test ./...` — all packages pass (root 16.9s, latest 12.7s, store 12.7s, web 0.004s)
- `CGO_ENABLED=0 go build` produces the static binary
- Cover-path tests run verbosely: `TestPublicCoverServesStoredBytesUnauthenticated`, `TestPublicCoverRejectsUnknownAddress` (unknown/malformed/traversal/empty), `TestPublicCoverNeverEchoesNonImage`, `TestListRendersAcquiredCover`, `TestAcquireKaganeCoverThroughBrowser`, `TestRunOncePrefetchesKaganeCover`, `TestRunOnceRoutesNonKaganeCoverToPublicFetcher` all pass; the three `SMOKE_*` tests skip without the browser sidecar, as designed

Live browser verification of the "web UI and panel render Covers for all six Sites" criterion is being run separately with Playwright against real Site pages and a locally mocked backend.

Reviewed-on: #73
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 18:02:47 +07:00
sulthan 78234f3c19 Browser-backed Sites join the Cover pipeline (#62) (#72)
Fixes #62

Browser-backed Sites join the Cover pipeline: kagane and novelfull Series now get their Covers at creation, through the same acquisition path as every other Site, instead of waiting for a poll pass.

## What changed

`latest.Acquirer` (creation-time acquisition, fired by the first Bookmark of a Series) previously skipped kagane and novelfull entirely — their pages only yield a Cloudflare challenge to the TLS client, so the request was spent for nothing. It now routes them like the poller does, with the two Sites split exactly as the issue demands:

- **kagane** — page fetched through the browser sidecar, cover URL extracted from the API JSON, bytes fetched through the browser sidecar (the only path that clears the challenge) into the content-addressed store. With no `BROWSER_WS_URL` configured, acquisition is skipped entirely and nothing falls back to a plain fetch.
- **novelfull** — page fetched through the browser sidecar, cover URL extracted from the HTML, bytes fetched over plain TLS through the ordinary gated fetcher (its image paths answer 200 with `access-control-allow-origin: *`, measured 2026-08-09). With no browser configured, the page fetch falls back to the TLS client — novelfull's challenge is a live time-varying fact (AGENTS.md), so when the page body answers, the Cover still lands; when it is challenged, nothing happens.

The byte-routing rule (kagane → browser, every other Site → TLS) is now one shared function (`latest.fetchCoverBytes`) used by both the Poller and the Acquirer, so the two cannot drift apart.

## Acceptance criteria

- [x] kagane cover bytes are fetched through the browser sidecar and stored in the content-addressed store — `TestAcquireKaganeCoverThroughBrowser`
- [x] novelfull cover URLs are extracted from the browser-fetched HTML, and its bytes are fetched over plain TLS — `TestAcquireNovelfullCoverOverPlainTLS`
- [x] With no browser sidecar configured, kagane Covers are absent and nothing falls back to a plain fetch — `TestAcquireKaganeSkippedWithoutBrowser`
- [x] With no browser sidecar configured, novelfull Covers still work if its page body is available — `TestAcquireNovelfullCoverWithoutBrowser`
- [x] Manually verified on-device: a kagane Series shows its Cover in the panel, not a broken-image glyph — being run by a separate manual-verification agent against a mocked scenario (no prod data); not part of this PR
- [x] `go test ./...` is green, with live-network checks gated behind `SMOKE_BROWSER_WS_URL` like the existing kagane image smoke test — new `TestSmokeAcquireKaganeCover` proves the end-to-end acquire path against the real browser when the env var is set

## Verification

- `go test ./...` green across all packages
- New unit tests exercise every routing decision with fakes — no network in the default suite
- Smoke test gated behind `SMOKE_BROWSER_WS_URL`, skipped by default

## Post-review changes (a66491a)

- **One routing rule for pages too** — `fetcherFor` is now a shared function used by both the Poller and the Acquirer; novelfull falls back to the plain-TLS fetcher in *both* when no browser is configured, so pre-existing (client-scraped) novelfull rows get healed by the poll as well, not just Series created after this change (`TestNovelfullUsesTLSWhenNoBrowserFetcher`).
- **Byte-level no-fallback proof** — `TestAcquireKaganeBytesNeverFallBackToPlainTLS` pins that kagane cover bytes never route to the TLS fetcher even when the page came through a browser.
- **Acquirer wired independent of the TLS client** — if `NewTLSFetcher` fails, kagane/novelfull acquisition still works via the sidecar (`main.go`).
- AGENTS.md (root + backend) updated for the novelfull plain-TLS fallback.

Reviewed-on: #72
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 11:06:36 +07:00
sulthan b9220b3dfc The Poll fills blank Covers for every Site and both Libraries (#70)
Closes #61.

## Summary

Permanently-blank Series (the half of #47 that creation-time acquisition cannot reach) heal on the next due poll cycle. The cover path is no longer kagane-only: every Site and both Libraries fill a blank Cover from the series page the chapter poll already fetched, and never replace a Cover that already exists.

## What changed

### `backend/internal/latest/poller.go`

- **`fillBlankCover`** — when `Cover` and `CoverAddress` are both blank, extract a source URL via `coverFrom` from the series-page body and store bytes through `SetSeriesCover`. Skips any Series that already has a source URL (owned by prefetch) or a stored address (never overwrite).
- **`prefetchCover`** — source-URL healing path, now site-uniform. Kagane no longer special-cases into `PutKaganeCover` alone; every Site lands on `SetSeriesCover`, so the wire Cover becomes a content-addressed public URL. Reuses already-stored bytes when present.
- **`storeCover` / `fetchCoverBytes`** — shared fetch+persist. Only kagane routes image bytes through the browser fetcher; every other Site uses plain TLS `CoverBytesFetch`. Failures log with the Series key and never return to the chapter path.
- **`checkOne`** — after a successful series-page fetch, calls `fillBlankCover` once regardless of whether chapter extraction succeeded (cover fill is independent of the chapter signal).

### `backend/internal/latest/poller_test.go`

Extended the existing poller harness (real store, fake fetchers) rather than a new one:

- `TestRunOnceFillsBlankCoverFromSeriesPage` — asura manga, lightnovelworld novel, kagane manga; asserts wire Cover + correct fetcher routing.
- `TestRunOnceDoesNotReplaceExistingCover` — second poll does not refetch.
- `TestRunOnceRetriesFailedBlankCoverOnNextPoll` — failed fill stays blank, next due cycle retries (no separate queue).
- `TestRunOnceBlankCoverFailureDoesNotBlockChapter` — chapter still lands; failure log carries the Series key.
- Kagane prefetch test now also asserts the content-addressed wire Cover.

## Acceptance criteria (#61)

| Criterion | Status |
|---|---|
| Cover prefetch runs for every Site | done |
| Cover prefetch runs for both Libraries | done |
| Poll fills a blank Cover | done |
| Poll never replaces an existing Cover | done |
| Failed cover fetch does not fail/block chapter poll | done |
| Failed cover fetch retried next poll, no separate queue | done |
| Failures logged with the Series | done |
| Existing poller tests extended | done |
| `go test ./...` green | done |
| Manually verified: blank Series gets Cover after a poll cycle | **left for you** |

## Out of scope / not closed

- Does **not** close #47 or #55 (per ticket).
- No migration/backfill script — the Poll walks every Series already.
- No admin refetch (#54).

## Review notes addressed

- Removed the kagane-only `PutKaganeCover` branch from prefetch so source-URL healing also sets `CoverAddress` (wire Cover).
- Guard so `fillBlankCover` does not double-fetch after `prefetchCover` healed the same snapshot.
- Single `fillBlankCover` call site after the series-page fetch.

## Test plan

- [x] `go test ./...` (backend; needs Docker/Postgres via `pgtest`)
- [ ] After deploy: pick a Series that was blank, wait one poll cycle, confirm Cover in web UI and userscript panel

Reviewed-on: #70
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 10:14:27 +07:00
sulthan e2c054e7ce Covers render in the userscript panel, from a public route (#60) (#69)
Closes #60.

Spec: #55. Originating bug: #47. Architecture: `docs/adr/0007-backend-hosts-cover-bytes.md`. Neither #47 nor #55 is closed from here.

## What this branch does

The panel now renders Covers from the deployment's own origin, and both userscripts stop having an opinion about where a Cover lives.

**The public route was already in place.** `GET /covers/{address}` landed with #59 (`92eba07`) and is registered on the bare mux, outside `httpmw.Auth` and outside the web UI's Discord session — `backend/main.go:210-214`, handler `backend/internal/api/handlers.go:142-158`. It reads no cookie and no header, answers `404` for an address that was never stored (and for a row whose file has gone missing — recorded-but-gone is not-found, never a fabricated body), refuses anything that is not `^[0-9a-f]{64}$` *before* the value becomes a path, and sets `Cache-Control: public, max-age=604800, immutable`. Those four properties are asserted by `backend/cover_test.go:231-278`. This branch re-verified them rather than re-implementing them; the only backend line it touches is a comment.

**Both userscripts lose cover scraping entirely.** Every adapter's `cover:` field is gone, along with the two helpers that fed them: the manga script's `coverFromPage()` (the `img[alt]` DOM scan comix needed, because comix publishes no `og:image`) and the novel script's `metaName()` plus the now-callerless module-level `meta()`. Nothing under `userscript/` reads `og:image`, `meta[name=image]`, or `img[alt]` any more.

**Nothing sends a cover either.** `delete body.cover` sits in `apiPut` — `manga-bookmark.user.js:486`, `novel-bookmark.user.js:275` — which is the single chokepoint every write passes through (`pushBookmark`, the retry-queue flush, `toggleFavorite`, `toggleArchive`). It operates on the `Object.assign` copy, so the in-memory row keeps the cover it renders with. This matters beyond tidiness: a Reader upgrading from an older copy has `localStorage` rows carrying third-party scraped URLs, and without the strip those would ride back up on the next write. The handler discards the field regardless (`handlers.go:53-59`) — it is permanently inert, not pending removal.

**Failed loads get the designed empty state, not the broken-image glyph.** `onerror: (e) => e.target.replaceWith(el("div", { class: "cover ph" }))` on the cover `<img>` in both card renderers (`manga:1380-1390`, `novel:1134-1144`). The replacement is byte-identical to the existing no-cover branch on the very next line, so it picks up the `.cover.ph` styling already in the panel CSS — no new tokens, no new rule. `el()` routes any `on*` prop through `addEventListener`, so this is a listener, not an inline attribute string, and the swap is a `createElement` + DOM call with no markup parsing anywhere near it. This is the half of #47 that was visible on kagane.

**The deleted scraping's tests went with it**: the two comix cover cases, the `pageImages` and `namedMetas` fixtures, the `img[alt]` and `meta[name=...]` stub branches, the now-dead `querySelectorAll` stub member, and every stale `og:image` fixture and `p.cover` assertion across both suites. The export lists needed no change and that was checked, not assumed — `coverFromPage` and `metaName` were module-private on `origin/main` and no cover symbol ever appeared in `module.exports`.

Docs that described the deleted behaviour were corrected in the same breath, because leaving them would instruct the next agent to put the scraping back: `userscript/AGENTS.md` (adapter contract + the per-site notes for comix, kagane and novelfull), the README's adapter reference, and the userscript testing skill's stub table.

## Verification

- `go test -count=1 ./...` — green across all nine packages (`backend` 29.8s, `latest`, `store`, `session`, `token`, `userscript`, `web`).
- `node --check` clean on both userscripts; `node --test` on both logic suites — 46 tests, 46 pass.
- `gofmt -l` clean; `go build ./...` clean.
- The `onerror` swap is DOM behaviour and deliberately has no coverage in the Node harness — that harness stubs a browser precisely so it never needs a DOM, and #60 says not to invent coverage for it. It was instead exercised for real: the `el()` helper and the exact render expression were loaded into a headless Chromium with a deliberately unloadable `src`, and the resulting DOM was `<div class="cover ph"></div>`. Ad hoc, not committed.
- **Not done, needs you:** the on-device criterion — a comix Series bookmarked mid-chapter showing its Cover in the panel. That needs a real install against the deployment and is the one box left unticked on #60.

## Reviewed

Both `/code-review` axes ran against `cc0fa92`. Spec found no missed requirement and no scope creep; standards found the diff clean on the four areas it scrutinised (the `delete body.cover` placement, the `onerror` handler's DOM safety, comment quality, dead-code removal). Their combined findings — the dead `querySelectorAll` stub, the stale README and skill text, and the handler comment whose premise this change invalidates — are fixed in `8b58019`.

## Out of scope, deliberately

The kagane-specific cover proxy still exists and still carries its session gate (#63 deletes it). The poll's blank-Cover fill (#61) and browser-backed Sites joining the pipeline (#62) are untouched.

Reviewed-on: #69
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 08:52:57 +07:00
sulthan 92eba07da7 A newly bookmarked Series acquires its Cover at creation (#59) (#68)
Closes #59.

Part of spec #55, and the ticket that fixes the reported bug #47. Architecture: `docs/adr/0007-backend-hosts-cover-bytes.md`. Does not close #47 or #55.

## What changed

A Reader bookmarks a Series nobody holds yet — the exact case in #47 — and within seconds the list shows its artwork instead of a broken image. The first Bookmark to create a Series fires `Store.OnSeriesCreated` after commit, and the new `latest.Acquirer` turns that into **one** series-page fetch that yields both the Latest Chapter and the cover URL. The bytes go through the gated cover fetcher from #57 and are stored content-addressed through #56, so the wire carries an absolute URL on this deployment's own origin — never a third-party address, and never one that 404s.

### Store

- Migration `0009_series_cover_address.sql` adds `series.cover_address`. The two facts are now split: `series.cover` is the third-party source address the bytes came from (the acquisition path's dedupe key), `series.cover_address` is the SHA-256 they are stored under. An empty `cover_address` is precisely what "no Cover yet" means, which is the distinction both the API and the UI depend on.
- `SetSeriesCover` writes the address only after the bytes are on disk, so the wire can never name an object that is not there.
- `CoverWireURL` builds `PUBLIC_BASE_URL + /covers/<sha256>` for every scanned row, and returns `""` for a blank address.
- The cover columns are gone from `Upsert`'s `INSERT` and its `DO UPDATE`. A client-supplied cover cannot reach the shared Series row on any path, not just the creation path.
- `Open` now rejects a base URL that is not an absolute `http(s)` origin: `PUBLIC_BASE_URL=bookmarks.example.com` would otherwise start cleanly and emit addresses no browser can load.

### Acquisition

- `internal/latest/acquire.go`: one fetch, gated by the poller's own `fetchableSeriesURL` (a `series_url` arrives in a client-supplied PUT body, so without the gate a token-holder chooses what the server fetches from its own network position).
- Asynchronous and log-and-drop. The Bookmark, its progress and its Latest Chapter are already committed; a Site that is down or a cover that cannot be produced disturbs none of them.
- Bounded by a two-slot semaphore. A bulk sync creating N Series would otherwise fire N simultaneous requests from one IP — the traffic shape the poller's stagger exists to avoid.
- Cancelled at shutdown (shares the poller's context) and stamps `latest_checked_at`, so the poller does not refetch the same page a tick later.
- Browser-backed Sites (kagane, novelfull) are deliberately skipped: their pages only yield a Cloudflare challenge to the TLS client, so the request would be spent for nothing. They arrive in #62.

### Wire and route

- `GET /covers/{address}` serves the bytes publicly and uncredentialed with `Cache-Control: public, max-age=604800, immutable`. The address is gated by a `^[0-9a-f]{64}$` pattern and cross-checked against a pure function of itself before any filesystem read, so no request shaped like a traversal reaches disk.
- `PUT /bookmarks/{key}` still accepts a `cover` field and discards it, permanently. Rejecting it would break every installed userscript the moment this deploys, and ADR-0004's compatibility argument depends on those scripts continuing to work. The decode site says so in place of a TODO nobody intends to keep.
- `store.CoverContentType` canonicalises comix's non-standard `image/jpg` to `image/jpeg`, so one image cannot land under two spellings. This one was found by the live smoke test, not by reading.

### Config

`PUBLIC_BASE_URL` is new and required (cover URLs must go out absolute — the userscript renders them on third-party origins, where a relative path resolves against the Site). Documented in `.env.example`, `docker-compose.yml` (`:?` so compose fails too), `DEPLOY.md` and `backend/AGENTS.md`.

## Acceptance criteria

All twelve of #59's criteria are met; the checklist on the issue is ticked with the evidence.

## Verification

- `go test ./...` green (Docker-backed Postgres suite).
- Live smoke against a real backend + Postgres: bookmarking `comix:n8we-dungeons-and-crayons` produced `"cover": "http://127.0.0.1:8099/covers/8ce74d80…"` and `"latest_chapter": "Chapter 81"` within seconds of the PUT; `curl` on that address returned `200`, `Content-Type: image/jpeg`, `Cache-Control: public, max-age=604800, immutable`, and a 280x420 JPEG. That run is what surfaced the `image/jpg` content type.
- Mutation-checked the asynchrony test: removing the `go` from `Acquire` turns `TestAcquireDoesNotBlockTheWrite` red.

## Reviewed

Both axes of `/code-review` were run against this diff before commit. Their findings that were actionable here are folded in: the concurrency bound, the shutdown tie, the `PUBLIC_BASE_URL` validation, the missing `latest_checked_at` stamp, and a test that could not fail.

## Known sequencing

A kagane/novelfull Series created between this deploy and #62 has no cover source at all: the acquisition skips those Sites and `Upsert` no longer persists the userscript-scraped address. This is #59's stated boundary rather than a defect, but it is a user-visible gap on two Sites and should order #62 accordingly.

Reviewed-on: #68
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 04:07:53 +07:00
sulthan b6b88bde8a feat(latest): extract per-site covers (#58) (#67)
Closes #58

## Summary

- Add pure per-Site cover extraction beside latest-chapter parsing for all six Sites.
- Read Asura, Demonic, LightNovelWorld, and NovelFull metadata; read the Comix target detail state; read Kagane's browser-fetched `series_covers[].image_id` JSON.
- Preserve published cover URLs, percent-encode Demonic raw spaces, select Comix's smaller published `medium`, and avoid thumbnail rendition URL synthesis.
- Add live-source fixtures plus no-cover and Cloudflare challenge coverage for every Site.

## Correctness

- Scope Comix extraction to the requested series detail key, avoiding recommended posters.
- Parse Kagane's current live API shape and emit its canonical compressed image route from the published image ID; unrelated JSON fields are ignored.
- Validate Kagane image IDs against the existing UUID-shaped route constraint.
- Keep extraction pure; storage, polling, and wire integration remain outside issue #58.

## Acceptance criteria

- [x] Cover extraction exists for all six Sites in the existing latest parser module.
- [x] Each Site has a live-source fixture with source URL and date.
- [x] Comix reads the state blob, not metadata.
- [x] Demonic raw spaces are percent-encoded.
- [x] Comix returns the smaller published rendition.
- [x] No-cover pages return empty.
- [x] Cloudflare challenge pages return empty.
- [x] No thumbnail URL is synthesized by editing a published URL.
- [x] `go test ./...` passes.

## Verification

- `go test ./...`
- `go vet ./...`
- `git diff --check`

Parent issues #47 and #55 remain open as requested.

Reviewed-on: #67
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 01:43:57 +07:00
sulthan 9d6d3bde72 Add gated cover byte fetcher (#66)
## Summary

Adds a plain-TLS cover byte fetcher with a destination-class SSRF gate and wires public cover sources through the content-addressed filesystem store.

## Changes

- Resolve hostnames before connecting; refuse non-HTTPS, loopback, private, link-local, unique-local, CGNAT, credentials, and mixed public/private DNS answers.
- Re-check every redirect and resolve/classify again at dial time to close DNS rebinding.
- Reuse `maxBodyBytes`; reject oversized responses and non-image content types before persistence.
- Add generic `Store.GetCover`/`PutCover` source-URL storage while preserving the browser-backed kagane path.
- Keep cover prefetch failures isolated from chapter polling.
- Add observable tests for TLS, no-connection refusals, all refused address classes, redirect blocking, streaming body caps, non-image rejection, content-addressed persistence, DNS rebinding, and poller routing.

## Verification

- `go test -count=1 ./...`
- `go vet ./...`

Both pass. No test touches the live network.

Closes #57

Reviewed-on: #66
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 00:45:55 +07:00
sulthan e8d1cba6c5 Move cover bytes to content-addressed filesystem storage (#65)
Refs #56

## Summary

Moves Kagane cover bytes out of Postgres bytea storage into an immutable, content-addressed filesystem store. Reader-visible behavior remains unchanged: the existing session-gated route serves stored bytes, missing bytes use the existing browser fetch path, and no browser still returns a missing cover.

## Changes

- Added migration 0008, which drops the legacy `covers` table and recreates it with only `address`, `path`, and `content_type`. Existing byte rows are intentionally dropped.
- Added SHA-256 source-URL addressing with two-level sharding (`ab/cd/<sha256>`). Writes use a temp file plus atomic link; reads validate the stored relative path before opening it.
- Made `COVER_DIR` required in runtime config and Compose. Compose passes it as a Docker build argument and volume target, so custom durable paths keep image ownership, runtime config, and the named `cover-data` volume aligned.
- Updated every `store.Open` caller and documented configuration, deployment, backup, and troubleshooting behavior.
- Added filesystem, restart, migration-drop, no-browser, and content-addressing coverage.

## Verification

- `go test ./...`
- `CGO_ENABLED=0 go build ./...`
- `docker build --build-arg COVER_DIR=/data/covers -t manga-bookmark-cover-check-custom ./backend`
- `docker compose config --format json` confirms custom `COVER_DIR` is the volume target
- `git diff --check origin/main`
- LSP diagnostics clean for touched Go files

Parents #47 and #55 remain open as required by #56.

Reviewed-on: #65
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-10 00:11:20 +07:00
sulthan 30c57bd39c Define Cover and record hosting its bytes (#47) (#64)
Defines **Cover** in the glossary and records ADR-0007, the decision behind #47's fix.

## Why these two files, and why now

`CONTEXT.md` named Cover inside the **Series** entry — "facts true regardless of who is reading — title, cover, Latest Chapter" — but never said *what* one is. That gap is the bug. Nothing in the model distinguished "an address on a Site" from "an image a Reader's browser can display", so both clients were left to work it out independently, and one of them got it wrong. kagane serves covers with `cross-origin-resource-policy: same-origin`, the web UI rewrote them to a proxy in its templates, the JSON API did not, and the panel rendered a broken-image glyph. The new entry closes the ambiguity: *an address no client can load is not a Cover, it is a missing one.*

ADR-0007 records what follows from that — the backend fetches, stores and serves every Site's cover bytes — plus the alternatives that were rejected and, more importantly, the two places this deliberately departs from existing precedent:

- **Destination-class control instead of a host allowlist.** `fetchableSeriesURL` sets the allowlist precedent for `series_url`, and covers do not follow it. Cover hosts are CDNs that move independently of their Site — demonicscans serves its covers from `readermc.org` — so an allowlist would stop producing Covers the day a Site switched CDN, and that failure would look exactly like #47. The resolve-then-classify step is what actually stops the SSRF.
- **A public cover route where the kagane proxy is session-gated.** An `<img>` cannot send a bearer token, and it cannot be given one either: the panel's shadow root is `mode: "open"`, so the host page's JavaScript can read any `src` the script sets.

Both are security-adjacent departures, which is precisely why they are written down rather than left in a commit message.

## Scope

Documentation only — no code, no schema, no behaviour. The implementation is #56–#63.

## Why this should merge promptly rather than sit

All eight implementation tickets cite `docs/adr/0007-backend-hosts-cover-bytes.md` as the authority for decisions they must not relitigate, and they are written in the vocabulary this glossary entry defines. An agent picking up #56 reads both from `main`. Until this lands they get a 404 and either invent a rationale or stall — so this PR gates the tickets, not the other way round.

Related: #47 (bug), #55 (spec), #54 (deferred admin refetch).
Reviewed-on: #64
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 23:19:37 +07:00
sulthan 8081a0a5d8 Give the Tailscale ACL step a working policy file (#53)
Follow-up to #52, which merged before this landed. Docs only — no code, no compose changes.

`DEPLOY.md` §7 told the operator to "tag the two machines" and showed a bare `acls` fragment. Following it literally does not work and is actively harmful:

- the fragment references `tag:bookmark-api` / `tag:bookmark-browser` without a `tagOwners` section, so the policy is rejected on save;
- it never says how a tag gets onto a device (`tailscale up --advertise-tags=...`, which re-authenticates);
- replacing the tailnet's default allow-all with only that one rule **removes the operator's own SSH access to the browser machine**.

Replaced with a complete, saveable policy file: `tagOwners`, the CDP rule, a second rule preserving own-device access including `:22`, and a `tests` block so a later edit that widens 9222 is rejected rather than silently applied.

Also records two things that were assumed rather than stated:

- **Why tagging is load-bearing.** Tailscale has no `deny`, so restricting 9222 means removing the blanket accept and enumerating what remains. That is only expressible if the browser machine falls outside a selector that still covers your own devices — which is exactly what a tag does, since a tagged device has no user and stops matching `autogroup:member` / `autogroup:self`. Without that, the whole step reads as arbitrary ceremony.
- **Tagging replaces a device's user identity**, so it suits a dedicated box and disrupts a daily driver. Both paths are now written down.

Finally, separates two checks the old text conflated: the existing `curl` runs on the home machine and proves only the **bind**, because node-local traffic is not filtered. Proving the **ACL** needs a third device, so that check is now its own step.

Verified: `tailscale.com/docs/reference/syntax/policy-file` and `/docs/features/tags` (validated Apr 2026 / Dec 2025) for `tagOwners`, `autogroup:self` semantics vs tagged devices, `--advertise-tags` re-auth and key-expiry behaviour. Markdown fences balanced.
Reviewed-on: #53
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 16:12:03 +07:00
sulthan 2a3bb6922d Move the browser off the VPS to its own unit (#46) (#52)
Closes #46 once deployed.

The headless browser leaves the API stack and becomes its own compose unit
(`chrome/docker-compose.yml`) intended for the home machine, reached over the
tailnet. No fallback sidecar is left on the VPS.

The backend needs no code change — `BROWSER_WS_URL` was already the only
coupling. Its default is now empty rather than a pinned Docker IP, so an
unconfigured or unreachable browser degrades exactly as it always has: plain-TLS
libraries unaffected, kagane/novelfull logged and skipped, stored covers still
served.

### What shipped

- `chrome/docker-compose.yml` + `chrome/.env.example` — the browser unit, with
  the CDP port bound to `${BROWSER_BIND_ADDR}` (no default) and the resource
  limits from the epic: 512 MiB / 1 GiB memory+swap, `oom_score_adj 800`,
  halved CPU weight, shm 1 GiB -> 128 MiB.
- API stack drops the service, its `depends_on` and the `browser` network.
- `bookmark-api` gains the `default` network. Dropping `browser` had left it on
  `db` alone, which is `internal: true` — no published port and, worse, no
  egress for the poller at all. Caught by actually bringing the stack up.
- ADR-0006 for the topology; `DEPLOY.md` §7 for first-time setup of the browser
  machine; `REDEPLOY.md` §8 for its independent update cadence; architecture
  diagrams, config tables and troubleshooting rows across README/AGENTS/env.

### Verified locally

- Browser unit builds and runs: Chrome 151, UA carries no `HeadlessChrome`,
  all limits applied as declared.
- **Live smoke passes through the new unit**: `TestSmokeKaganeImage` fetched
  56710 bytes of `image/webp`, `TestSmokeKaganeGet` got a 200 with a real
  chapter list. The challenge cleared under the reduced 128 MiB shm.
- Bind isolation proven: refused on the host's non-loopback address, accepted
  on the configured one.
- 321 MiB peak of the 512 MiB cap after a full solve; 0 restarts, no OOM kill.
- API stack comes up clean, `/healthz` 200; egress confirmed present on
  `default` and absent on `db`.
- `go test ./...`, `go vet`, `gofmt` clean.

### Left to the operator

Provisioning the home machine, the Tailscale ACL, setting `BROWSER_WS_URL` in
production, and observing acceptance criteria 5-7 (covers with the machine off,
several days of zero OOM/restarts, VPS memory improvement). `DEPLOY.md` §7 now
carries the before/after `free -m` reading those need.

Reviewed-on: #52
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 15:28:21 +07:00
sulthan d1800d0707 Prefetch Kagane covers during latest polling (#51)
## Summary
- Add an optional browser-backed cover fetcher to the latest-chapter poller.
- Prefetch missing Kagane covers during the existing due-series cycle and persist them before a Reader opens the web UI.
- Keep chapter polling, cooldown stamping, and on-first-view fallback independent from cover failures.

## Behavior and safety
- Stored Kagane covers are detected before browser work, so later poll cycles do not refetch them.
- Nil cover fetchers and non-Kagane series retain the existing behavior.
- Shared Kagane image-id and content-type validation prevents challenge or non-image responses from poisoning persistent cover storage.
- The browser is wired into both the chapter and cover poller paths from the composition root.

## Verification
- `go test ./...`
- Focused latest, store, and web package tests
- Deterministic tests cover missing covers, cached covers, failed fetches, invalid content types, nil fetchers, and non-Kagane series.

Closes #45

Reviewed-on: #51
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 14:58:59 +07:00
sulthan 84cfd1b2c1 Make browser sidecar on-demand (#44) (#50)
Closes #44. Chrome now starts on first CDP connection, tracks concurrent helpers, reaps after 300 seconds idle, preserves the named profile, and classifies reap interruptions. Shutdown stops Chrome's process group so cookie batches flush. ADR-0005 records the measured constraints and decisions. Verification: docker build, live CDP wake, graceful stop cleanup, sh -n, and go test ./... (7 packages, 3 no tests).

Reviewed-on: #50
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 07:39:07 +07:00
sulthan bfae84c5c3 Persist kagane covers in Postgres (#49)
Closes #43

Persist kagane cover bytes in a dedicated Postgres covers table keyed by image ID. The web handler reads storage before the browser, writes validated fetches through, and no longer keeps an in-process cover cache. Added migration, store persistence tests including reopen, handler coverage for stored/miss/rejected paths, and corrected repository guidance.

Verification:
- go test ./...
- CGO_ENABLED=0 go build ./...

Reviewed-on: #49
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 06:52:02 +07:00
sulthan cd3a7e3d01 feat(latest): split browser poll cooldown (#48)
## Summary

Split latest-chapter polling cooldowns by fetch cost. Browser-backed kagane and novelfull series now rest longer without changing the cadence of plain-TLS sites.

## Behavior

- Plain-TLS series keep the 1h default cooldown.
- Browser-backed series use `LATEST_CHAPTER_POLL_BROWSER_COOLDOWN`, defaulting to 6h.
- Both cooldowns share the existing 15m minimum floor; invalid values retain the existing fallback behavior.
- The poller still selects both classes in one due query per cycle.
- Existing ordering and exclusions remain unchanged: reader-count precedence, least-recently-checked ordering, finished exclusion, archived polling, and orphan exclusion.

## Implementation

- Added the browser cooldown to backend configuration and passed it through production poller construction.
- Added the browser-site list as the single routing source used for both due-query cutoff selection and fetcher choice.
- Kept all query values parameterized; the site list is passed as a bound PostgreSQL array parameter.
- Updated startup logging to report interval, plain cooldown, browser cooldown, batch, and stagger.
- Documented the variable, default, and floor in `README.md`, `.env.example`, `backend/AGENTS.md`, and `docker-compose.yml`.

## Review findings addressed

The first review found that configuration parsing was correct but `startLatestPoller` did not pass `BrowserCooldown` into `latest.Poller`; every browser-backed row would therefore have been due immediately. Production construction now goes through `newLatestPoller`, with a regression test covering both cooldown fields.

The review also identified duplicated browser-site knowledge in fetch routing. `slices.Contains(browserBackedSites, site)` now reuses the same list already supplied to the store query.

## Verification

- Focused backend tests pass: `go test ./internal/latest ./internal/store .`.
- Full suite passes: `go test ./...`.
- `graphify update .` completed.
- Issue #42 was updated and closed.

Reviewed-on: #48
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 06:25:48 +07:00
sulthan 741b23322b Fix comix titles and covers, kagane volume chapters, and kagane cover rendering (#37)
Fixes five reported symptoms across comix.to and kagane.to. Diagnosing them turned up two latent bugs underneath, both of which had to be fixed for the kagane cover work to function at all.

## Reported symptoms and their causes

| # | Symptom | Cause |
|---|---------|-------|
| 1 | comix bookmark titled `Comix - Read Comics online for free` | comix is an SPA that rewrites `document.title` on client routing but never touches the server-rendered `og:title`. The adapter read `og:title`, so a cold load stored the homepage's title. |
| 2 | next comix bookmark gets the *previous* series' title | Same cause. After an in-page hop, `og:title` still holds whatever page loaded first. |
| 3 | comix cover shows the placeholder | comix serves no `og:image` at all, so `coverFromPage()` had nothing to read. |
| 4 | kagane chapter never appears in the bookmark list | Reader URLs carry no chapter number, so it is parsed out of `og:title`. Volume-numbered series render `"<Series> - Volume <v> Chapter <n>"`, which the suffix regex did not match, so `chapterNum` came back null and nothing was recorded. |
| 5 | kagane title includes the chapter, e.g. `SP Baby - Volume 1 Chapter 1` | Same unmatched regex — the tail was never stripped. One fix covers 4 and 5. |
| 6 | kagane cover blocked in the web UI | kagane serves covers behind its Cloudflare challenge **and** with `cross-origin-resource-policy: same-origin`. No `<img>` on the UI's origin can load one even from a browser holding the clearance cookie. Hot-linking cannot be made to work. |

## What changed

**Userscript.** comix titles now come from `document.title` with the chapter page's `" - Ch.<n>"` tail stripped, and the cover is the `img` whose `alt` matches the cleaned title. comix fills `document.title` a beat *after* the URL changes — later than the nav watcher's 300 ms snapshot — so the watcher also re-detects when the `detect()` signature changes, not only when the URL does. The kagane suffix regex takes an optional `Volume <v> ` segment. All three page shapes were captured live on 2026-08-08 and pinned as regression tests.

**Cover proxy.** `Bookmark.CoverURL()` rewrites a stored kagane `og:image` to `/img/kagane/{id}`; templates render `.CoverURL` instead of `.Cover`. The endpoint is session-gated like every other UI route and fetches through the shared headless browser, which is same-origin with kagane and so satisfies both the challenge and the CORP header. Results are memoised in-process, so a cover costs one navigation per deployment lifetime. With `BROWSER_WS_URL` unset the endpoint answers 404 rather than reaching for a nil fetcher — the same degrade-to-userscript behaviour the poller already has.

The image id is matched against a UUID regex before it reaches the browser. That gate is load-bearing rather than tidiness: the cover is a stored client-supplied string, so an unvalidated one turns this endpoint into an SSRF primitive aimed at the deployment's own network. `ServeMux` path-cleans a traversal into a redirect before the handler runs, but the handler does not depend on that, and a test pins it.

## Two latent bugs found underneath

**`BrowserFetcher.run` never let a challenge solve.** It navigated, waited for `body`, read once, and closed the tab — roughly half a second end to end. The Cloudflare interstitial has a `body` too, so `WaitReady` was satisfied by the challenge page itself. This made the challenge *unclearable* rather than merely slow: an interstitial needs several seconds of a live page to solve itself and write clearance into the browser's shared cookie jar, so tearing the tab down first means every subsequent call is challenged exactly like the one before it. `run` now holds one tab and re-reads until the caller's predicate reports an answer, bounded by `challengeTimeout` and the caller's own deadline. Exhausting the budget maps back to the 403 the poller already expects, keeping a challenged site distinct from a broken transport.

**`chromedp/headless-shell` cannot clear kagane's challenge at all.** It is a stripped Chrome build and the tells are structural rather than a header: `navigator.webdriver` is true, the plugin list is empty, and the client hints are Chromium- rather than Chrome-branded. Overriding `webdriver` through CDP was tried on its own and changed nothing.

All measured 2026-08-08 from one IP against the same cover, so the comparisons are like for like:

| Browser | Result |
|---------|--------|
| `chromedp/headless-shell:stable` | never cleared (90 s) |
| `zenika/alpine-chrome` | never cleared — ships Chrome 124, old enough that Cloudflare refuses it and old enough to break chromedp's CDP structs |
| `google-chrome`, default UA | never cleared (60 s) — `--headless=new` advertises `HeadlessChrome` |
| `google-chrome`, stock UA, `TZ=UTC` | never cleared (90 s) |
| `google-chrome`, stock UA, any non-UTC `TZ` | **cleared in ~4 s** |

Both remaining tells are load-bearing, and each was tested in isolation. `chrome/` is a Debian image with `google-chrome-stable`, a UA whose version is read back out of the binary at startup (a hardcoded one would drift out of step with the `Sec-CH-UA` hints on the next Chrome update and become a fresh tell), and no `--enable-automation`.

### The timezone tell: UTC, not a country mismatch

The first pass concluded the zone had to match the egress IP's country. Re-measuring against the actual deployment case shows that was wrong, and the correction is in `1552dd1`.

The original inference read the host's `/etc/timezone` (`Asia/Bangkok`) and assumed a Thai egress. It isn't — this host egresses from an Indonesian IP. `Asia/Bangkok` cleared not because it matched a country but because it simply isn't UTC, and the two share +07, which hid the distinction. Same container, same Indonesian IP:

| `TZ` | Result |
|------|--------|
| `UTC` | never cleared (60 s, **twice**) |
| `Asia/Jakarta` | cleared in 4 s |
| `America/New_York` | cleared in 4 s |

`America/New_York` matches neither the country nor the offset nor the hemisphere and clears just as fast. A UTC clock is itself the bot signal — Cloudflare scores it as the datacenter default — and any real zone satisfies the check. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one, and a deployment that changes region need not keep it in sync.

One sharp edge remains: the usual `-v /etc/localtime:/etc/localtime:ro` does **not** work. Chrome resolves the zone through ICU, which takes the name from that path's symlink target and ignores the file's contents, so glibc reports the host zone while Chrome still reports UTC. `/etc/timezone` carries the name and is mounted instead.

Chrome also binds its DevTools port to loopback and silently ignores `--remote-debugging-address`, which is why headless-shell fronted it with socat. This image does the same, so it stays a drop-in: the compose service keeps the `headless-shell` name and its pinned address, and `BROWSER_WS_URL` is unchanged.

## Verification

```
go test ./...        all packages ok
node --test          37 + 12 pass, 0 fail

SMOKE_BROWSER_WS_URL=... go test -run TestSmokeKagane ./internal/latest
  TestSmokeKaganeImage  PASS (5.29s)  fetched 56710 bytes of image/webp
  TestSmokeKaganeGet    PASS (1.17s)  status=200, real chapter-list JSON
```

The smoke test ran against the exact compose configuration — built image, empty `BROWSER_TZ`, `/etc/timezone` mounted, cold profile — hitting real kagane.to. It skips unless `SMOKE_BROWSER_WS_URL` names a sidecar, so `go test ./...` stays hermetic and Docker-only.

A red smoke run means the challenge is not clearing from that IP, which is a live, time-varying fact to re-check rather than necessarily a defect.

## Security invariants

- Auth unchanged. `/img/kagane/{id}` is session-gated by `requireSession`, the same guard as every other UI route.
- Outbound fetch gated: the id is UUID-validated before it reaches the browser, keeping the existing rule that a client-supplied string never selects a fetch target unchecked.
- No new secrets, no new logging of credentials, no change to CORS, sessions, or crypto.
- Templates still escape everything; `.CoverURL` returns a plain string and is not wrapped in `template.HTML`/`URL`.
- One new dependency-free image (`chrome/`) built from Debian plus Google's own apt repo; no new Go modules.

## Deploying

Needs `docker compose build headless-shell`.

**A UTC host must set `BROWSER_TZ`, or kagane silently stops working.** With it unset the sidecar falls back to the host's `/etc/timezone`; on a UTC server that yields UTC, which is the one value that never clears. Any real zone works — `BROWSER_TZ=Asia/Jakarta` for the current deployment. `.env.example` now documents this; it previously did not mention the knob at all.

Only the browser sidecar reads `BROWSER_TZ`. The backend keeps its UTC clock, and stored timestamps are unix ms, so nothing else shifts.

## Deliberately not done

Retry/backoff around the cover proxy, and a panel-side cover fix. The panel renders no covers, and covers cache in-process after the first fetch. Worth adding if kagane starts rate-limiting.

## Correction after review of the deployment case

`1552dd1` was added after the branch was first pushed: the deployment host runs UTC with an Indonesian egress IP, which prompted re-measuring the timezone claim and falsifying it. The earlier commits' reasoning is left intact rather than rebased away, so the diagnostic trail — including the wrong turn and what disproved it — stays readable.

Reviewed-on: #37
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-08 23:27:32 +07:00
sulthan 2ef769d421 Open registration to guild members (#27) (#36)
Closes #27.

Guild membership is now the whole gate. `discordCallback` checks membership
(and `DISCORD_REQUIRED_ROLE` when set), then `Store.EnsureReader` creates the
Reader on first sight and returns the same row on every later login. The
refusal returns before `EnsureReader`, so a turned-away sign-in leaves no row
behind. `OWNER_DISCORD_ID` still seeds the owner, but only as the
administrator — it no longer gates login.

The cutover grace path goes with it: `API_TOKEN`, `API_TOKEN_GRACE_UNTIL` and
the legacy branch in `httpmw.ResolveReader` are deleted, so a credential
authenticates exactly one Reader or nothing. `userscript.Handler` drops its
re-derivation too — the resolved path segment is already the credential.

New surfaces: an empty library offers both install links (behind the
tab-specific empty states, so "No favourites yet" still wins), and the owner
alone gets a Readers panel with `POST /readers/{id}/revoke`. The owner's own
row is not revocable — 404, not a self-logout.

Isolation is asserted from both directions for read, modify and delete, and
the shared-series invariant is pinned: two Readers on one series produce one
series row, two independent progresses, one poll per due cycle, and one
Reader's delete leaves the other's bookmark and the poll intact.

Verified: `go test ./...` green; live smoke against a throwaway Postgres —
empty-library state in both colour branches, roster rendering, a real revoke
through the panel (target 401s next request, owner untouched), owner
self-revoke refused 404, per-Reader `/u/<cred>` and bearer auth both 200 with
404 for an unknown credential.

Reviewed-on: #36
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-08 20:23:17 +07:00
sulthan c2b47eb05b Offer the userscripts as a download for mobile Violentmonkey (#26) (#35)
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-08 16:39:30 +07:00
sulthan 1b1820d85a Cut production over: runbook corrections, env contract, Discord OAuth endpoint fix (#26) (#34)
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-08 16:06:47 +07:00
sulthan 2cc1e69f5d Prove the import against a copy of the real library (#25) (#33)
Closes #25.

Retires the biggest risk in #18 — losing the owner's reading history — on a copy, before production is anywhere near it.

## What was run

A throwaway generator (python3 stdlib `sqlite3`, ~20 lines, **not committed**) read a copy of `bookmarks-20260807-213515.db` and emitted plain SQL: 29 distinct Series first, then 29 Bookmarks referencing them, each `INSERT ... SELECT id FROM owner` so the reader id is resolved rather than hardcoded. The target was a scratch Postgres whose schema and owner Reader were built by the real binary (`go run .` against a throwaway container), not by hand-written DDL. Production was not touched.

## Verified

| check | result |
|---|---|
| Bookmarks total | 29 |
| reading / archived / other | 18 / 11 / 0 |
| Series | 29, equal to the distinct `(site, series_id)` count in the source |
| Readers | 1; Bookmarks not owned by the owner: 0 |
| Field-by-field diff, all 29 rows x 15 columns | 0 differences |
| `GET /bookmarks` over the real read path | 29 rows, values match source |
| `TRUNCATE bookmarks, series;` then re-apply | clean, 29 again |

The spot-check the ticket asked for was widened to a full row-by-row comparison — 29 rows is small enough that sampling was the more expensive option.

## What is committed

`CUTOVER.md` only, plus two cross-links from `REDEPLOY.md`. The generator stays out of the repository: its output is the owner's reading history, and it reads SQLite, which the backend module dropped in ADR-0001. So the runbook specifies the transformation — column mapping, ordering, nullability, quoting, the temp-table ownership trick — rather than shipping a script. `backend/go.mod` gains nothing.

## Review

Two-axis review ran on the diff; six findings applied, all in the runbook:

- Six source columns (`title`, `series_url`, `cover`, `last_chapter`, `last_chapter_url`, `last_chapter_num`) are nullable in SQLite but `NOT NULL` in Postgres and must be coalesced — the opposite of `latest_chapter_num`, the one column where `NULL` is meaningful. The 2026-08-07 export had none; a fresh one is not promised the same.
- `CREATE TEMP TABLE ... ON COMMIT DROP` must sit *inside* the transaction, or psql's autocommit drops it instantly.
- The ownership check now resolves the Reader by Discord id; comparing against `ORDER BY id LIMIT 1` was true by construction and could never fail.
- The spot-check now samples archived and favourite rows explicitly instead of hoping they fall inside `ORDER BY updated_at DESC LIMIT 5`.
- `git pull --ff-only` before `up -d --build`, or a pre-cutover server rebuilds the SQLite image.
- `python3` and `jq` named as prerequisites; column count corrected to sixteen.

Every query in the runbook was executed against the scratch database as written.

`go vet`, `CGO_ENABLED=0 go build ./...` and `go test ./...` all pass — no Go code changed.

Reviewed-on: #33
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-08 15:15:33 +07:00
sulthan 27cf0955de Per-Reader userscript credential with UI install and rotation (#24) (#32)
Closes #24. Child of #18; based on current main (includes Postgres, Reader table, Discord OAuth).

## What

Each Reader's userscript credential is derived from `TOKEN_KEY`, their Discord id and a token epoch (HMAC-SHA256, hex); only its SHA-256 sits in `readers.token_sha256` (new `token_epoch` column, migration 0006). One credential authenticates the script download path and the API bearer header.

- `internal/token`: derivation + hashing; the seed refreshes the owner's epoch-0 hash only before first rotation, so a restart can never resurrect a rotated-away credential
- `httpmw.Auth`/`ResolveReader`: acting Reader resolved from the credential hash, stashed in request context; the retired global `API_TOKEN` resolves to the owner until `API_TOKEN_GRACE_UNTIL` (enforced in code, logged per use) on both the bearer and script-download paths
- Userscript handler renders the bindmounted file with the resolved Reader's credential substituted for `__API_TOKEN__`; a legacy-path request during grace serves the derived credential, so installed devices self-migrate on their next update poll
- Web UI: "Userscripts" panel — session-gated install endpoints render the script directly (credential never in markup, address bar, or a redirect), confirm-gated rotation with an atomic epoch bump + hash rewrite and a reinstall warning
- Both userscripts carry `__API_TOKEN__` placeholders; the committed global-token literal is removed

## Design note

Credentials are derived rather than stored-random because the server must rebuild install URLs after restarts while the DB holds only hashes. HMAC output is high-entropy and unbrute-forceable; the AC's intent (unguessable, DB-leak-proof) is met.

## Deploy (also in DEPLOY.md)

1. Add `TOKEN_KEY` (`openssl rand -hex 32`) — required; changing it later invalidates every credential.
2. Keep `API_TOKEN` + set `API_TOKEN_GRACE_UNTIL` for the 14-day window.
3. After deploy, sign in → Userscripts → reinstall both scripts on every device. This also retires the old global credential for real — its literal survives in git history (present since 0ef5286), so rotation is what kills it.

## Verification

- Full Go suite green against real Postgres per test; userscript JS suite 45/45
- New router-level tests: per-Reader isolation (read/write/delete), grace expiry on bearer + script path, self-migrating legacy path, install serving, rotation (old cred 401/404, new cred works, install renders new credential), app page leaks no credential
- Store tests: hash lookup, token info, atomic rotation with stale-epoch rejection, rotation survives restart
- Live smoke of the built binary: grace acceptance logged, derived auth, substitution, restart resilience, stored hash = SHA-256 of derived credential

Reviewed-on: #32
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-08 14:54:03 +07:00
sulthan bcc6b45515 feat(backend): Discord OAuth login with DB-backed sessions (#23) (#31)
Implements #23 per ADR-0002.

- Discord authorization code grant (identify + guilds.members.read), form-encoded token exchange
- Guild membership gate via the single-guild endpoint; optional DISCORD_REQUIRED_ROLE (empty default)
- Owner Discord ID is the only identity allowed to sign in
- Sessions are DB rows with opaque random ids; cookie carries only the id; expiry enforced; delete = revoke
- HMAC session signing, derived key, and WEB_PASSWORD removed; no replacement signing secret
- Login rate limiting preserved on the callback
- Full flow tested through the real router against a local Discord stub (DISCORD_API_BASE)
- Env: DISCORD_CLIENT_ID/_CLIENT_SECRET/_GUILD_ID/_REQUIRED_ROLE/_API_BASE/_REDIRECT_URI; docs updated

go test ./... passes.

Reviewed-on: #31
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-08 08:51:22 +07:00
sulthan 8cebb94b92 Give every Bookmark an owner (Reader table) (#30)
Closes #22

## What

A `readers` table appears; every Bookmark belongs to one. The owner is seeded as the first and only Reader, and all existing rows are attached to them.

- **Migration 0003**: `readers` (discord_id UNIQUE, token_sha256 UNIQUE, created_at).
- **Migration 0004** (run-once, version-table-gated): attaches existing bookmarks to the seeded owner, drops the surrogate `key` column, composite PK `(reader_id, site, series_id)`, FK to readers `ON DELETE CASCADE` — a duplicate Bookmark for one Reader and Series is impossible at the database level.
- **Seed**: `Store.Open` runs schema to 0003, seeds exactly one owner row from `OWNER_DISCORD_ID` (hash = SHA-256 of `API_TOKEN`, refreshed on every start so rotation stays current), then migrates the rest.
- **Scoping**: `List/Get/Upsert/Delete` take `readerID`; the wire `key` is derived as `site:series_id` on read. Handlers act as `Store.OwnerID()` while the global token remains the only credential.
- **Unchanged**: authentication and the flat wire format — nothing observable changes from outside.
- **New env** `OWNER_DISCORD_ID` (required): compose, .env.example, DEPLOY.md, README.md, backend/AGENTS.md updated.

Series-level methods (due queue, mark-checked, set-latest-chapter) stay unscoped deliberately: series are shared rows polled once per due cycle, and the reader_count ordering requires cross-reader visibility (ADR-0003).

## Verification

- `go test ./...` green, including new tests: seed idempotency + hash refresh, 0004 attach migration, DB-level duplicate impossibility, per-reader scoping, reader-delete cascade.
- Live smoke test on fresh Postgres: seed → PUT/GET (flat wire intact) → restart idempotent; stored hash matches SHA-256 of the token.

## Deploy note

`OWNER_DISCORD_ID` is required after this lands — the backend refuses to start without it. Set it to the owner's Discord snowflake (Settings → Advanced → Developer Mode → right-click name → Copy User ID).

Reviewed-on: #30
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-08 08:05:17 +07:00
sulthan 984965ed9f Split Series from Bookmark, keeping the wire format flat (#21) (#29)
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-08 07:19:54 +07:00
sulthan 08749df050 feat(backend)!: run on Postgres with a migration-owned schema (#28)
Swap modernc.org/sqlite for jackc/pgx/v5 with no observable change:
same endpoints, same wire format, same updated_at ordering rule.

The schema now comes from numbered SQL embedded in the binary and
applied on startup, one transaction each, recorded in
schema_migrations. That replaces two pieces of SQLite-era machinery,
both deleted rather than ported: the column probing (Postgres has ADD
COLUMN IF NOT EXISTS, and there is no legacy database left to probe)
and the Asura key rewrite, which has run clean on every start for
months now that the userscripts strip build hashes before writing. Its
regexp survives as latest.asuraBuildHash, where the poller still needs
it to scope chapter links to a series whose slug carries a rotating
hash.

Types get real: favorite is a boolean, chapter numbers double
precision, timestamps stay unix-ms bigint. SQLite's null-safe IS NOT
becomes IS DISTINCT FROM, which is what implements the rule that only
reading progress reorders a list. Inside COALESCE/NULLIF the status
and kind parameters need an explicit ::text -- there is no target
column to infer from and Postgres refuses to guess.

Tests lose their free t.TempDir() database, so Docker is now a hard
prerequisite for `go test ./...`: internal/pgtest starts one
postgres:17-alpine per test binary and hands each test a database of
its own.

Also lands CONTEXT.md and the four ADRs written while scoping #18.

BREAKING CHANGE: DB_PATH is retired for DATABASE_URL, which is
required and has no default. Compose gains a postgres service on an
internal network with its own volume; POSTGRES_PASSWORD joins .env.
The old bookmarks-data volume is deliberately left undeclared so
`docker compose down -v` cannot take the pre-migration database with
it. main is not deployable until #25 and #26 land.

Closes #20

Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-08 06:52:20 +07:00
sulthan b9f9aea82c docs: secure-coding rules for agents, and refresh stale AGENTS.md top-matter (#16)
Two doc commits: a new secure-coding rules section, plus a fix for top-matter the rebrand left stale.

## `9156525` — secure-coding rules

`AGENTS.md` carried two security invariants (bearer auth, CORS) but nothing about the code an agent actually writes here. That is the gap worth closing: measured rates for AI-generated web/backend code are ~40% vulnerable (Pearce et al.), 45% failing security tests (Veracode 2025), and users *with* assistants shipped SQLi at 36% vs 7% for the control group (Perry et al., Stanford). The failure classes cluster on broken access control, injection, session/error handling and invented dependencies — all live surfaces in this repo.

Rules were **extracted, not pasted**. Every one names a guard that already exists in-tree, so the instruction is *match this*, not *invent something*:

| Rule | Existing anchor |
| --- | --- |
| parameterized SQL only; constants may concatenate | `store.go` — all queries use `?` |
| `html/template` only; no `template.HTML` on stored data | `web.go:85` |
| client-supplied URLs pass the fetch gate | `poller.go:144 fetchableSeriesURL` |
| cap remote bodies | `fetch.go:16 maxBodyBytes` |
| `subtle.ConstantTimeCompare`, never `==` | `middleware.go:22`, `session.go:61` |
| generic error out, detail to log, never log the token | `web.go:151` |
| `X-Forwarded-Proto` for Secure; **rightmost** XFF for IP | `session.go:68,101` |
| cookie flags; expiry checked before signature | `session.go:72-94`, `Verify` |
| validate at handler boundary | `handlers.go:48` `MaxBytesReader` 64 KB, 400 on bad key/status/kind |
| site strings via `el({text})`, never `{html}` | `el()` in both userscripts |
| `fetch()`/`authHeaders()` → `API_BASE` only | existing `authHeaders` |
| `localStorage` = cache/queue, never credentials | shared with site JS |

Plus a dependency rule (stdlib first; verify a package exists before adding — ~20% of LLM-proposed packages don't resolve, which is the slopsquatting vector) and a review gate marking auth/CORS/session/crypto/fetch-gate as security-critical.

Deliberately **excluded**: container signing, k8s admission control, IaC scanning, PII/HIPAA/PCI, C/C++ memory safety. Per OpenSSF's guide for AI assistant instructions, irrelevant rules make a model generate code compensating for attacks that cannot happen. None of those apply to a single-user Go + SQLite + userscript stack.

Sources: OWASP AISVS 1.0 Appendix C, OWASP Top 10 / ASVS v5, OpenSSF *Security-Focused Guide for AI Code Assistant Instructions* (2025-08-01).

## `5d4d890` — stale top-matter

The rebrand rewrote root `AGENTS.md` as a compression pass and switched Bromite -> Violentmonkey, but left the project described as a manga-only tracker over two sites. Six sites, two libraries and two userscripts now exist.

Root `AGENTS.md`:

- *What this is* names both scripts with their site lists, the `kind` column, and the `<site>:<series_id>` key shape.
- Origins constraint generalised past Asura/Demonic.
- Records that **kagane and novelfull are reliably Cloudflare-challenged** and browser-polled over CDP. Without it that bullet list reads as contradicting the code, since the paragraph above asserts blocking is "not universal — and not reliably reproducible".
- Diagram says two userscripts.

`backend/AGENTS.md` — two instances of the same defect, found while verifying the above:

- Store key list gained `novelfull|lightnovelworld` and the `kind` column.
- **`NOVEL_USERSCRIPT_PATH` documented** — it shipped in `main.go:150` undocumented.

The dated Cloudflare paragraph is left verbatim: it is a timestamped observation ("Verified 2026-07-26"), so rewriting it would falsify a record rather than update it. `userscript/AGENTS.md` is untouched; it already documents both novel adapters and the `LIBRARY`/`STORE_PREFIX` split.

## Verification

Docs-only, no code touched. Every code reference above was read at `4229c17` before being cited — the fetch gate, body cap, constant-time compares, cookie flags, handler validation, `el()` helper and both route registrations. No invented line numbers.

## Not addressed here

The `API_TOKEN` literal is committed in plaintext in both userscripts and in their `@downloadURL`/`@updateURL` lines. The new rules say not to propagate it, but the actual remedy is rotation plus build-time substitution, since the value is already in git history. Separate change; flagging it so it does not get lost.

Reviewed-on: #16
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-06 20:07:34 +07:00
sulthan 4229c179b0 rebrand: MangaBM → BookmarkManager, add novel library support (#15)
Two intertwined changes — the rebrand and the novel library were developed on
the same branch because the novel UI plumbing is part of the new "Bookmark
Manager" wordmark in the web shell.

## What it does

- **Rebrand**: MangaBM → BookmarkManager across the Go module, compose stack,
  env vars, Traefik hostnames, container/image names, userscript storage
  prefixes (`mangabm:cache` → `bmgr:manga:cache`, `mangabm:queue` → `bmgr:manga:queue`),
  and docs.
- **Novel library**: same backend, two libraries. New `kind` column splits
  bookmarks into `manga` / `novel`; PUT validates it. Two userscripts:
  - `manga-bookmark.user.js` — unchanged behaviour, just stamps its own `kind`.
  - `novel-bookmark.user.js` — separate Violentmonkey install with adapters
    for **novelfull.com** (polled via headless browser — Cloudflare JS
    challenge) and **lightnovelworld.net** (polled via plain TLS).
- **Web UI**: library switch on the app shell. Login art, libswitch, and
  novel-site colours from the Cinder design snapshot.

## Plumbing

- `addedColumns` ALTER for `kind` runs on first start after upgrade; every
  pre-existing row is backfilled to `'manga'`. No manual SQL, no down-time.
- `ALLOWED_ORIGINS` gains the two novel sites.
- New `NOVEL_USERSCRIPT_PATH` env (default `/userscript/novel-bookmark.user.js`),
  bindmounted alongside the manga script.
- Traefik router names `mangabm*` → `bmapi*` / `bmweb*`.

## Test status

- `go test ./...` — green
- `node --test userscript/test/logic.test.js` — 34 pass
- `node --test userscript/test/novel-logic.test.js` — 11 pass
- `node --check` on both userscripts — clean

## Notes for the redeploy

.env keys were renamed (`MANGA_API_HOST` → `BOOKMARK_API_HOST`,
`MANGA_WEB_HOST` → `BOOKMARK_WEB_HOST`). Update DNS / Traefik labels on the
prod override before pulling, otherwise the public hostnames go dark.
See the redeploy instructions I'll post next to this PR.

Reviewed-on: #15
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-06 03:58:23 +07:00
109 changed files with 70818 additions and 2576 deletions
+134
View File
@@ -0,0 +1,134 @@
---
name: implement-tickets
description: "Orchestrate a batch of tickets: plan the briefs, then hand each ticket to its own implementer subagent in its own worktree."
disable-model-invocation: true
---
# Implement tickets
You are the **orchestrator**. You write briefs, dispatch, land results, and talk
to the tracker. You do not write the implementation — every line of ticket code
is written by a `ticket-implementer` subagent in its own git worktree. Reach for
the editor yourself only for a merge conflict resolution.
Ticket source and `tea` usage: `docs/agents/issue-tracker.md`. Codebase
questions: `graphify query "<question>"` before grepping.
## 1. Collect the tickets
The user's argument is the selector: issue numbers, a label, a parent issue, or
nothing. With nothing, take the open issues labelled `ready-for-agent`.
Fetch each with `tea issue <n> --comments`, and read the **whole** body —
acceptance criteria and the `Blocked by` line are what the rest of this skill
runs on. A ticket whose blockers are still open is out of this batch unless a
blocker is also in it.
## 2. Plan the batch
Explore enough of the codebase to write briefs a fresh context can act on: the
files each ticket lands in, the patterns it must follow, the `AGENTS.md`
invariants it touches.
Then decide three things:
- **Waves.** Blocking edges set the order; tickets with no open blocker inside
the batch share a wave. Cap each wave at **3** concurrent tickets unless the
user set another width.
- **Contracts.** Two tickets in one wave that meet at a function signature, a
JSON shape, a table column, or a token name: you decide the shape now and
write the identical wording into both briefs. A contract left for the
subagents to negotiate is a merge conflict you scheduled.
- **Splits.** A ticket too big for one fresh context window goes into the wave
as two briefs, or back to the user.
## 3. Get the plan approved
Present, and stop:
- the wave list, and for each ticket: number, title, one-line brief summary,
the files or areas it will touch, its verification commands
- every cross-ticket contract, verbatim as it will appear in the briefs
- anything you had to assume
Wait for approval. Apply the user's edits to the plan, do not relitigate them.
## 4. Run a wave
Per ticket, before dispatch:
```bash
git worktree add ../ticket-<n> -b ticket/<n>-<slug> <base> # base = the branch you are on
cp .env ../ticket-<n>/ 2>/dev/null # gitignored, worktrees do not get it
tea issue edit <n> --add-assignees <your gitea username> # tea login list has it
```
Write the brief to `.scratch/<batch-slug>/t<n>-brief.md` using the template
below, in the ubiquitous language of `CONTEXT.md` — a brief that says "scrape"
where the domain says Poll hands the subagent the wrong model of the system.
Then dispatch the whole wave in **one** `task` batch, every item on the
`ticket-implementer` agent. Each dispatch names: the absolute brief path, the
worktree path, the branch, the base ref, and the report path
`.scratch/<batch-slug>/t<n>-report.md`.
<brief-template>
# Ticket #<n> — <title>
**Read first.** `tea issue <n> --comments` for this ticket, then the issue it
refers to — the parent or spec — the same way. The comments carry decisions the
body never got updated with. This brief stays the requirements; those two reads
are the intent behind them.
**Goal.** The end-to-end behaviour this ticket makes work, from the user's side.
**Acceptance criteria.** Verbatim from the ticket.
**Contract.** The exact shared signatures / shapes / names this ticket must
implement or consume, and which sibling ticket is on the other end. Omit when
the ticket touches nothing shared.
**Where it lands.** The files and packages, and the existing pattern to follow
in each.
**Binding invariants.** The `AGENTS.md` rules this change can break — name them.
**TDD seams.** Where a test comes first — run the `tdd` skill at each one and
follow its red → green loop. Or "none — verify after".
**Verify.** The exact commands, e.g. `cd backend && go test ./...`,
`node --test userscript/test/logic.test.js`.
**Out of scope.** What not to touch, especially a sibling ticket's files.
</brief-template>
## 5. Land the wave
The wave is landed when every ticket in it is closed, reverted, or handed back
to the user. Per returned ticket:
| Status | What you do |
| --- | --- |
| `DONE` | merge, comment, close |
| `DONE_WITH_CONCERNS` | merge, comment the concerns, close only if you judge them non-blocking — otherwise leave open and tell the user |
| `BLOCKED` / `NEEDS_CONTEXT` | supply what is missing and re-dispatch, or hand back to the user with the specifics. Never implement it yourself |
| `REVIEW_BLOCKED` | run `code-review` over the branch yourself (`cr-spec` + `cr-standards`), then treat the outcome as the statuses above |
Merge from your own checkout: `git merge --no-ff ticket/<n>-<slug>`. A textual
conflict is yours to resolve (`resolving-merge-conflicts`). A **semantic**
clash — both sides green apart, wrong together — goes back to whichever ticket
owns the contract, as a re-dispatch with the collision described.
Then `tea comment <n> "<the report summary>"`, `tea issue close <n>`, and
`git worktree remove ../ticket-<n>`. Keep the report file.
Only once the whole wave is landed does the next wave start — its briefs may
need what this one changed.
## 6. Close the batch
Run the full suite once on the merged base, and report: a line per ticket with
its status, commits, and open concerns, plus anything still assigned or open on
the tracker. A red suite after every ticket went green is an interaction bug —
diagnose it, name the two tickets, and fix it or hand it back with both named.
@@ -14,7 +14,7 @@ parsers, helpers. UI, network, and storage behaviour are verified on-device.
```bash
node --check userscript/manga-bookmark.user.js # parse check, silent on success
node --test userscript/test/logic.test.js # 14 tests as of 2026-07-28
node --test userscript/test/logic.test.js # 35 tests as of 2026-08-10
```
Run both before every commit that touches the userscript.
@@ -31,7 +31,7 @@ The test file installs four globals **before** requiring the userscript:
|---|---|---|
| `localStorage` | `Map`-backed stub | `loadCache`, `loadQueue`, and the key-migration IIFE touch it at module scope |
| `location` | `{href, hostname, pathname, origin}` | read during boot |
| `document` | `querySelector` for `meta[property="…"]` only, plus a no-op `addEventListener` | adapters read `og:title`/`og:image` |
| `document` | `querySelector` for `meta[property="…"]` only, plus a no-op `addEventListener` | adapters read `og:title` (covers are the backend's, never scraped) |
| `document.body` | **left `undefined`** | this is the whole trick |
`document.body === undefined` sends the userscript's boot block down its `else`
+81 -36
View File
@@ -1,13 +1,38 @@
# Copy to .env and fill in. Never commit the real .env.
# Long random secret shared with the userscript's API_TOKEN. Generate one:
# Secret every Reader's userscript credential is derived from (issue #24):
# the backend rebuilds install URLs from it, and only SHA-256 hashes of the
# credentials ever touch the database. Generate one:
# openssl rand -hex 32
API_TOKEN=changeme-generate-a-long-random-token
TOKEN_KEY=changeme-generate-a-long-random-token
# The owner's Discord user ID — seeded at startup as the first Reader, the
# administrator (the only one who can revoke another Reader's sessions), and
# the owner of every bookmark that predates registration. Discord snowflake,
# e.g. 1046923170000000000.
OWNER_DISCORD_ID=changeme-your-discord-user-id
# Comma-separated origins allowed to call the API (CORS). Both Asura domains
# plus Demonic, Comix, Kagane, and the two novel sites. Add/remove as the
# sites' hostnames change.
ALLOWED_ORIGINS=https://asuracomic.net,https://asurascans.com,https://demonicscans.org,https://comix.to,https://kagane.to,https://novelfull.com,https://lightnovelworld.net
ALLOWED_ORIGINS=https://asurascans.com,https://demonicscans.org,https://comix.to,https://kagane.to,https://novelfull.com,https://lightnovelworld.net
# Password for the bundled Postgres container, and therefore half of the
# DATABASE_URL compose builds for the backend. Generate one:
# openssl rand -hex 24
POSTGRES_PASSWORD=changeme-generate-a-long-random-password
# Override only to point the backend at a Postgres compose does not run.
# DATABASE_URL=postgres://user:pass@host:5432/bookmarks?sslmode=require
# Directory inside bookmark-api for immutable, content-addressed Cover bytes.
# Compose builds the image and mounts its named volume at this path.
COVER_DIR=/covers
# Public origin this deployment answers on, no trailing slash. Required: Cover
# URLs go out absolute, because the userscript renders them on a Site's own
# origin where a relative path would resolve against the Site (ADR-0007).
PUBLIC_BASE_URL=https://bookmark-api.example.com
# --- Prod override (Traefik) only ---
# Subdomain Traefik routes to this service (required by the prod override).
@@ -18,46 +43,66 @@ ALLOWED_ORIGINS=https://asuracomic.net,https://asurascans.com,https://demonicsca
# TRAEFIK_ENTRYPOINT=websecure
# TRAEFIK_CERTRESOLVER=le
# --- Web UI ---
# Password for the browser UI at https://$BOOKMARK_WEB_HOST. Leave unset to
# disable the web UI entirely (the routes are not registered at all).
# Generate one: openssl rand -base64 18
WEB_PASSWORD=
# --- Web UI (Discord OAuth) ---
# Sign-in is a Discord authorization code grant (ADR-0002), and it is also
# registration: any member of the configured guild becomes a Reader on their
# first successful login, with their own empty library. Create the application
# at https://discord.com/developers/applications and register the exact
# callback URL ($BOOKMARK_WEB_HOST/auth/discord/callback) as an OAuth2
# redirect.
DISCORD_CLIENT_ID=
DISCORD_CLIENT_SECRET=
# The guild whose membership gates sign-in (Developer Mode -> right-click the
# server -> Copy Server ID).
DISCORD_GUILD_ID=
# Exact callback URL, e.g. https://bookmark.example.com/auth/discord/callback.
# Discord matches it verbatim, so it must equal the registered redirect.
DISCORD_REDIRECT_URI=
# Optional: a role snowflake members must hold on top of guild membership.
# Empty (the default) means membership alone suffices.
# DISCORD_REQUIRED_ROLE=
# Subdomain Traefik routes to the browser UI (required by the prod override,
# whether or not WEB_PASSWORD is set). Left commented on purpose: an example
# value here would be a silent wrong-hostname fallback, and Traefik would
# publish the UI router on a domain you do not own. The same container also
# answers on BOOKMARK_API_HOST for the userscript's API.
# Subdomain Traefik routes to the browser UI (required by the prod override).
# Left commented on purpose: an example value here would be a silent
# wrong-hostname fallback, and Traefik would publish the UI router on a domain
# you do not own. The same container also answers on BOOKMARK_API_HOST for the
# userscript's API.
# BOOKMARK_WEB_HOST=bookmark.example.com
# --- Latest-chapter poller ---
# The backend re-checks each bookmarked series' newest published chapter on its
# own schedule, so latest_chapter stays fresh even when you never open the manga
# sites. This runs in parallel with the userscript's own in-browser check.
# Set to 0 to turn it off entirely.
# Set to 0 to turn it off entirely. Pace is per Site (one Poll Lane per Site,
# issue #100) and lives in the backend registry, not here — there is nothing
# else to configure.
# LATEST_CHAPTER_POLL_ENABLED=1
#
# Two independent clocks. COOLDOWN is how long one series rests between checks;
# INTERVAL is how often the poller wakes up and looks for series past that
# cooldown. Shortening INTERVAL cannot shorten a COOLDOWN.
# LATEST_CHAPTER_POLL_COOLDOWN=1h # per series, floor 15m
# LATEST_CHAPTER_POLL_INTERVAL=10m # how often to wake
# LATEST_CHAPTER_POLL_BATCH=14 # series per wake
# LATEST_CHAPTER_POLL_STAGGER=20s # delay between fetches in a batch
#
# Uses a ticker, not an immediate first run: the first poll happens one
# INTERVAL after startup, not at startup. A container restarting more often
# than INTERVAL never polls.
#
# BATCH x (COOLDOWN / INTERVAL) series hold the cooldown cadence — 84 with these
# defaults. Beyond that the cadence stretches uniformly rather than breaking;
# raise BATCH or lower INTERVAL. Keep BATCH x STAGGER under INTERVAL.
# Every Site rests an hour between checks and gaps ten seconds between fetches;
# a Site with many Series tightens its own gap. See backend/internal/latest/sites.go.
# Headless-shell CDP endpoint for sites behind a JavaScript challenge (kagane).
# Unset disables browser polling; those sites then rely on the userscript alone.
# Leave commented — the compose files' own default (ws://172.28.0.10:9222) is
# correct. Do NOT set this to the "headless-shell" DNS name: Chrome's DevTools
# HTTP handler 500s any /json/version request whose Host header isn't an IP or
# "localhost", which silently breaks every kagane poll.
# BROWSER_WS_URL=ws://172.28.0.10:9222
# CDP endpoint of the browser, used for the two sites behind a Cloudflare
# JavaScript challenge (kagane, novelfull) and by the web UI's kagane cover
# proxy. Unset disables browser polling and serves 404 for covers not already
# stored; those sites then rely on the userscript alone. That is also exactly
# how an unreachable browser degrades, so a home machine that is off costs
# chapter freshness and nothing else.
#
# The browser does NOT run in this stack. It is its own compose unit on the
# home machine (chrome/docker-compose.yml, chrome/.env.example) and is reached
# over the tailnet, so set this to that machine's tailnet address:
#
# BROWSER_WS_URL=ws://100.x.y.z:9222
#
# It must be the tailnet **IP**, never a MagicDNS hostname and never the old
# Docker service name: Chrome's DevTools HTTP handler 500s any /json/version
# request whose Host header isn't an IP or "localhost", which silently breaks
# every kagane poll. Left unset here on purpose — a wrong default would poll a
# stranger's address, and "no browser" is a safe, self-announcing state.
# BROWSER_WS_URL=ws://100.x.y.z:9222
# Zone the backend stamps its log lines in. Cosmetic only. Nothing else in
# the service has a zone: bookmark timestamps are unix ms, and the two real
# time columns are timestamptz. Defaults to Asia/Jakarta; set to UTC for the
# conventional server default.
# API_TZ=Asia/Jakarta
+6 -1
View File
@@ -5,8 +5,13 @@
backend/server
backend/backend
.playwright-mcp/
graphify-out/
# graphify map is committed; only regenerable/local parts are ignored
graphify-out/cost.json
graphify-out/cache/
graphify-out/[0-9][0-9][0-9][0-9]-[0-9][0-9]-[0-9][0-9]/
graphify-out/.rebuild.lock
plans/
.scratch/
docs/superpowers/
.superpowers/
go.work
+90
View File
@@ -0,0 +1,90 @@
---
name: ticket-implementer
description: Implements one ticket end to end inside its own git worktree - reads a brief file, implements, tests, commits, runs the two-axis code review through cr-spec and cr-standards, fixes findings, writes a report file, returns a short status contract. Dispatched by the implement-tickets skill.
model: opencode-go/minimax-m3
thinking-level: high
tools: read, write, edit, bash, grep, glob, lsp, todo, ast_edit, task
spawns: cr-spec,cr-standards
autoloadSkills: code-review, tdd
---
You implement **one ticket** dispatched by an orchestrator. Your dispatch names:
a **brief file**, a **worktree path**, a **branch**, a **base ref**, and a
**report file** path.
## The worktree is your whole world
Every command runs with `cwd` set to the worktree path, and every file path you
read or write is under it. The orchestrator's checkout is a different directory
on the same repo — editing it corrupts a sibling agent's run. If a command must
run elsewhere, say so in the report instead of doing it.
Your branch is already checked out there. Never `git checkout`, `git switch`,
`git rebase`, or `git worktree` anything.
## Order of work
1. Read the brief file. It is the single source of requirements — use its exact
values verbatim.
2. Read the ticket and the issue it refers to, as the brief's **Read first**
section names them: `tea issue <n> --comments` for each. The ticket's
comments and its parent carry the intent and the decisions behind the brief.
Read no other ticket and no other brief.
3. Read `AGENTS.md` in the worktree, plus the nested `AGENTS.md` for the area
you touch. Its invariants bind you: security rules, design system, comment
policy.
4. Ask before writing code if requirements, acceptance criteria, approach, or
dependencies are unclear. Asking is free; guessing is not.
5. Implement exactly what the brief specifies. At each TDD seam the brief names,
run the `tdd` skill and follow its red → green loop.
Follow the patterns already in the codebase; improve what you touch,
restructure nothing outside the ticket.
6. Verify. Focused tests while iterating, the brief's full verification commands
once at the end. Test output must be pristine.
7. Commit to your branch. Reference the ticket number in the subject.
8. Review (below), fix, re-verify, commit the fixes.
9. Write the report file, then return the status contract.
## Review
After your first green commit, run the **`code-review`** skill over
`<base ref>...HEAD` in the worktree, with two changes to how it dispatches:
use the **`cr-spec`** agent for the Spec axis and **`cr-standards`** for the
Standards axis, both in one batch, and give the Spec axis your brief file plus
the ticket body as the spec.
Fix every Critical and Important finding, then re-run the tests that cover the
amended code. Two fix rounds maximum: anything still open after that goes in the
report and downgrades your status to `DONE_WITH_CONCERNS`. Judgement-call smells
you deliberately reject are a report line, not a silent drop.
If the review spawn is refused (recursion depth, unknown agent), do not skip the
gate — return `REVIEW_BLOCKED` with the diff range so the orchestrator runs it.
## Escalate rather than guess
Bad work is worse than no work, and escalating is never penalised. Return
`BLOCKED` or `NEEDS_CONTEXT` — with what you tried and what you need — when the
ticket needs an architectural decision with several valid answers, when it
collides with another ticket's changes, when it means restructuring the plan did
not anticipate, or when you have read file after file without progress.
## Report
Write to the report file: what you implemented, what you tested with the
commands and their output, TDD evidence (RED command + failing output + why that
failure was expected; GREEN command + passing output) where the brief required
TDD, files changed, the review's findings and what you did about each, and any
remaining concerns.
Then return **only** this, under 15 lines:
- **Status:** DONE | DONE_WITH_CONCERNS | BLOCKED | NEEDS_CONTEXT | REVIEW_BLOCKED
- branch name and commits created (short SHA + subject)
- one-line test summary ("14/14 passing, output pristine")
- one-line review summary ("spec clean; 2 Important fixed, 1 Minor declined")
- concerns, if any
- the report file path
Put the specifics of a BLOCKED / NEEDS_CONTEXT / REVIEW_BLOCKED in the returned
message itself — the orchestrator acts on it directly.
+82 -17
View File
@@ -4,34 +4,61 @@ Guidance for OpenCode (and Claude Code) working in this repo.
## What this is
Manga read-progress tracker, user read on **asurascans.com** (current domain; asuracomic.net 301s here) and **demonicscans.org** via **Violentmonkey**. Userscript inject on-page UI (floating button + slide-in panel), sync progress to self-hosted Go backend so bookmarks unify across both sites and devices.
Read-progress tracker for two libraries — manga and novels — behind one self-hosted Go backend. Two separate Violentmonkey userscripts inject on-page UI (floating button + slide-in panel) and sync progress, so bookmarks unify across sites and devices:
- `manga-bookmark.user.js` — **asurascans.com** (asuracomic.net is dropped: its deep links 301 to the asurascans.com root, discarding the path), **demonicscans.org**, **comix.to**, **kagane.to**.
- `novel-bookmark.user.js` — **novelfull.com**, **lightnovelworld.net**.
One backend, one `bookmarks` table: a `kind` column (`manga`|`novel`) splits the libraries and the web UI switches between them. Rows are keyed `<site>:<series_id>`.
## Hard constraints (drive design — don't violate)
Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free where plain web APIs suffice — keeps portability across engines:
- **Avoid `GM_*` unless needed.** Prefer page `localStorage` over `GM_setValue`/`GM_getValue`, on-page UI over `GM_registerMenuCommand`, plain `fetch()` over `GM_xmlhttpRequest` for cross-origin.
- Cross-origin `fetch()` work **only** against CORS-enabled backend. Manga sites `https://`, so backend **must be HTTPS** (else mixed-content block).
- Asura and Demonic are **separate origins with separate `localStorage`** — shared remote store only way to unify bookmarks. Cloud sync required, not optional.
- Every site is its **own origin with its own `localStorage`** — a shared remote store is the only way to unify bookmarks. Cloud sync required, not optional.
- Userscript run in **isolated world**, so embedded API token safe from site's JS.
- Cloudflare's block on manga sites **IP-reputation-based, not universal — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — Cloudflare's bot scoring can flip previously-clean IP without notice. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test.
- Cloudflare's block on manga sites is **per-zone configuration plus request fingerprint, not IP reputation — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — a Site can turn its protection on overnight, which is exactly what comix.to did on 2026-08-12. An earlier version of this line blamed "Cloudflare's bot scoring"; that was wrong. The 1-99 bot score is Enterprise Bot Management only and does not exist for a free-plan zone, and no per-IP request rate is documented as an input to challenge issuance — `docs/research/cloudflare-bot-scoring-and-poll-cadence.md`. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test.
- **kagane.to, comix.to and novelfull.com are the exception to the above** — all three sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`). When that's unset, kagane and comix are skipped entirely (a plain fetch would only retrieve a challenge page) while novelfull pages are still attempted over plain TLS — its challenge is a live time-varying fact and its cover bytes never need the browser. comix turned hostile on 2026-08-12 (#98): its cover host `static.comix.to` is gated too, so its cover bytes go through the browser as well, and its page is read as an in-tab `fetch()` of the series URL rather than a rendered DOM — comix is an SPA, and rendering costs ~65 requests for the same server-rendered HTML one fetch returns. The three other sites poll fine over plain TLS.
- **The CDP browser must look like a real browser, and stock headless images don't.** Measured 2026-08-08 against kagane.to, all from the same IP: `chromedp/headless-shell:stable` never cleared the challenge in 90s (`navigator.webdriver` true, empty plugin list, Chromium-branded client hints — suppressing `webdriver` alone changed nothing); `zenika/alpine-chrome` ships Chrome 124, refused outright; real Chrome with the default `--headless=new` UA never cleared, because the UA says `HeadlessChrome`; real Chrome with a stock UA **and** a non-UTC clock zone cleared in ~4s. Hence `chrome/` — a Debian image with `google-chrome-stable`, a version-derived UA, and `TZ`/`BROWSER_TZ`. Chrome reads the zone *name* through ICU from `/etc/localtime`'s symlink target, ignoring the file's contents, so mounting the host's `/etc/localtime` does **not** work; `/etc/timezone` is mounted instead.
- **The browser is not in the API stack and must not be put back.** It's its own compose unit (`chrome/docker-compose.yml`) on a second machine, reached over the tailnet — it held 471 MiB on a 1974 MiB swapless VPS, and a residential egress avoids the cloud-hosting-IP signature Bot Fight Mode documentedly challenges (ADR-0006; not a better "score" — free-plan zones have no score). Consequences that constrain code: `BROWSER_WS_URL` must be a tailnet **IP** (a MagicDNS name 500s at `/json/version`, same trap as the old Docker service name); the CDP port binds to the tailnet address only, since CDP authenticates nothing and that host has a real LAN; and the browser is on-demand (ADR-0005), so an unreachable or asleep one must degrade exactly as an unset `BROWSER_WS_URL` — plain-TLS libraries unaffected, kagane/comix logged and skipped, stored covers still served. Never add `chromedp.NoModifyURL`: discovery per fetch is what makes a restarted Chrome invisible.
- **UTC is the tell, not a country mismatch.** A UTC clock is the datacenter default, and the challenge refuses it; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while `Asia/Jakarta` **and** `America/New_York` both cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (`Asia/Bangkok`) rather than the measured egress. A second earlier claim, that Cloudflare "scores" a UTC clock, was also wrong: the measurement is real but the mechanism is not documented anywhere — Cloudflare publishes no timezone signal, and free-plan zones carry no score at all. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one.
- **A challenged page needs the tab kept open.** The interstitial takes seconds to solve and only then writes clearance into the browser's shared cookie jar. Navigate-read-close never clears anything; `BrowserFetcher.run` holds one tab and re-reads until the payload arrives.
## Architecture
```
Violentmonkey userscript (isolated world, per-site adapters, localStorage cache)
-- fetch() HTTPS --> reverse proxy (TLS + CORS) --> Go net/http --> SQLite (volume)
Two Violentmonkey userscripts (isolated world, per-site adapters, localStorage cache)
-- fetch() HTTPS --> reverse proxy (TLS + CORS) --> Go net/http --> Postgres (volume)
|
| CDP over tailnet
v
on-demand Chrome, separate machine (chrome/)
```
Backend-specific architecture (packages, endpoints, poller, config env vars) lives in `backend/AGENTS.md`. Userscript-specific structure (adapters, retry queue, UI, live URL shapes) lives in `userscript/AGENTS.md`.
Two deployable units on two machines: the API stack (`docker-compose.yml` + `docker-compose.prod.yml`, on the VPS) and the browser (`chrome/docker-compose.yml`, on the home machine). They share nothing but `BROWSER_WS_URL` and update independently. Backend-specific architecture (packages, endpoints, poller, config env vars) lives in `backend/AGENTS.md`. Userscript-specific structure (adapters, retry queue, UI, live URL shapes) lives in `userscript/AGENTS.md`. Deploy order `DEPLOY.md` (§7 for the browser), redeploy `REDEPLOY.md` (§8 for the browser).
## Commands
Backend (`cd backend`):
- Test all: `go test ./...`
- Test all: `go test ./...` — **needs Docker.** Each test package starts a throwaway `postgres:17-alpine` container (`internal/pgtest`).
- Single test: `go test -run TestName ./...`
- Build static binary: `CGO_ENABLED=0 go build`
Local stack: `docker compose up` (named volume mounted at `/data`, `restart: unless-stopped`).
Local stack: `docker compose up` (bookmark-api + postgres only; `postgres-data` named volume, `restart: unless-stopped`). No browser — without `BROWSER_WS_URL` the poller logs and skips kagane and comix. To run one: `cd chrome && BROWSER_BIND_ADDR=172.17.0.1 docker compose up -d --build`, then `BROWSER_WS_URL=ws://172.17.0.1:9222` in the root `.env` (bridge gateway, so the API container can name it by IP).
Live CDP proof (needs that browser and network, skipped otherwise):
`SMOKE_BROWSER_WS_URL=ws://<ip>:<port> go test -run 'TestSmokeKagane|TestSmokeComix' ./internal/latest`
— fetches a real kagane and comix cover and chapter list. A red run means the challenge is
not clearing from this IP, which is a live fact to re-check, not necessarily a defect.
A red `TestSmokeComix` reporting `ERR_CERT_COMMON_NAME_INVALID` is not the
challenge: it means the resolver the browser container uses hijacks `comix.to`.
Observed 2026-08-16 on one Indonesian ISP, which CNAMEs it to a block page
(`aduankonten.id`). Check with `docker exec <browser> getent hosts comix.to`,
and if it is hijacked, run the container with
`--add-host comix.to:<ip> --add-host static.comix.to:<ip>` from a DoH lookup
(`curl -H 'accept: application/dns-json' 'https://1.1.1.1/dns-query?name=comix.to&type=A'`).
Machine-local, so don't put those hosts in `chrome/docker-compose.yml`.
Smoke test: `curl` endpoints with `Authorization: Bearer <token>`; confirm `OPTIONS` preflight return CORS headers and `/healthz` return 200.
@@ -62,9 +89,41 @@ instantly.
## Security invariants
- Auth on `/bookmarks*`: require `Authorization: Bearer <API_TOKEN>`, **constant-time compare**, 401 otherwise.
Existing guarantees — don't regress:
- Auth on `/bookmarks*`: require `Authorization: Bearer <credential>` — the acting Reader's credential, matched by SHA-256 against `readers.token_sha256` — **constant-time compare** (via the hash, never the secret itself), 401 otherwise.
- CORS: reflect `Origin` only when in `ALLOWED_ORIGINS`; allow `GET,PUT,DELETE,OPTIONS` + headers `Authorization,Content-Type`; answer preflight `OPTIONS` with `204`.
## Secure coding rules (code you write here)
Anchored to OWASP Top 10 / ASVS. Every rule below already has a working example in-tree — match it, don't start a second convention. AI-written backends fail on exactly these: broken access control, injection, weak session/error handling, invented dependencies.
Go backend:
- SQL always parameterized (`$N`). Only compile-time constants (`bookmarkColumns`) may be concatenated into query text — never a request value, not even a validated one.
- `html/template` only for anything a browser parses, never `text/template`. Never wrap stored or fetched strings in `template.HTML`/`JS`/`URL`; that switches off the escaping every template depends on.
- Any outbound fetch of a client-supplied URL passes `fetchableSeriesURL` (site + `https` + host check) first. `series_url` arrives in a PUT body, so without the gate the poller will probe arbitrary hosts from the server's own network position. New fetch path reuses the gate rather than re-deriving one.
- Cap every remote body with `io.LimitReader` (`maxBodyBytes`). An unbounded read is an OOM handed to whatever is on the other end.
- Compare secrets with `hmac.Equal` / `subtle.ConstantTimeCompare`, never `==`. A credential is matched by the SHA-256 the `readers` table holds, which is already a fixed-width equality — a new secret comparison must not regress to `==`.
- Errors: generic text to the client (`http.Error(w, "internal error", 500)`), detail to `log.Printf`. Never log `TOKEN_KEY`, a Reader's credential, `DISCORD_CLIENT_SECRET`, a session id, or a whole `Authorization` header.
- Proxy headers are trusted only where they already are: `X-Forwarded-Proto` for the Secure cookie flag, **rightmost** `X-Forwarded-For` for client IP (leftmost is attacker-supplied). Don't read either anywhere else.
- Session cookies keep `HttpOnly`, `SameSite`, `Secure`-when-HTTPS; expiry is enforced by the `sessions` table lookup, not a signature.
- Stdlib crypto only. No hand-rolled hashing, no MD5/SHA-1 anywhere security-bearing.
- Validate at the handler boundary before storing: body capped by `http.MaxBytesReader` (64 KB), empty `key` and unknown `status`/`kind` rejected with `400`. A bad value that reaches the store becomes every later reader's problem.
Userscript:
- Site-derived and stored strings render via `el(..., {text})` / `textContent`. `{html}` and `innerHTML` are for author-written literal markup only (`TEMPLATE`, `CSS`) — never a title, chapter label, or API response field. The page DOM belongs to a third-party site; treat it as attacker-controlled.
- Isolated world protects the credential from the site's JS. It does not protect anything from an `innerHTML` sink you add yourself.
- The userscripts carry `__API_TOKEN__` placeholders, substituted at serve time with the requesting Reader's credential (`internal/userscript`). Never put a real credential in the repo, docs, commit messages, or issues. Rotation is a web-UI action (epoch bump, `internal/token`); `TOKEN_KEY` in backend env is what derives every credential — never log it.
- `fetch()` targets `API_BASE` only — no dynamic origin, no site-supplied URL. `authHeaders()` goes nowhere but the backend.
- `localStorage` is shared with the site's own JS: cache and queue live there, credentials never do.
- Wrap every `localStorage` read/write and `JSON.parse` in try/catch (quota, private mode, corrupt entry), as the existing helpers do.
Dependencies: stdlib first; a new module needs a stated reason. Confirm a package actually exists before adding it — a plausible name may be fiction (~20% of LLM-proposed packages don't resolve, which is how slopsquatting lands). Pin exact versions.
Review gate: auth, CORS, session, crypto, and the fetch gate are security-critical. Editing one is not a drive-by change — say which invariant you preserved and run `go test ./...` before calling it done.
## Comments
Comment only if code alone can't carry info. Cost per read — must earn spot.
@@ -89,22 +148,28 @@ Style: one dense comment over function beats one per line inside. Tight, no work
Test: "competent reader get this from code in few sec?" Yes → skip. Needs detour through another file/spec/git-blame → write it.
## Relevant skills
## Agent skills
`multi-stage-dockerfile` and `docker-compose-orchestration` for container work (referenced in plan).
`AGENTS.md` is the single source of truth for agent guidance; every `CLAUDE.md` in this repo is a symlink to the `AGENTS.md` beside it. Edit `AGENTS.md`.
`golang-code-style`, `golang-error-handling`, `golang-performance`, `golang-testing` for backend Go work.
### Issue tracker
Issues live as Gitea issues on `gitea.violetcrown.my.id` (`sulthan/mangaBookmark`), driven by the `tea` CLI — not `gh`. See `docs/agents/issue-tracker.md`.
### Triage labels
Default five-role vocabulary, label strings unchanged (`needs-triage`, `needs-info`, `ready-for-agent`, `ready-for-human`, `wontfix`). See `docs/agents/triage-labels.md`.
### Domain docs
Single-context: one root `CONTEXT.md` plus `docs/adr/`, both created lazily. See `docs/agents/domain.md`.
## graphify
Project has knowledge graph at graphify-out/ with god nodes, community structure, cross-file relationships.
Rules:
- For codebase questions, first run `graphify query "<question>"` when graphify-out/graph.json exists. Use `graphify path "<A>" "<B>"` for relationships and `graphify explain "<concept>"` for focused concepts. Return scoped subgraph, usually much smaller than GRAPH_REPORT.md or raw grep output.
- For codebase questions and exploration, always first run `graphify query "<question>"` when graphify-out/graph.json exists. Use `graphify path "<A>" "<B>"` for relationships and `graphify explain "<concept>"` for focused concepts. Return scoped subgraph, usually much smaller than GRAPH_REPORT.md or raw grep output.
- If graphify-out/wiki/index.md exists, use for broad navigation instead of raw source browsing.
- Read graphify-out/GRAPH_REPORT.md only for broad architecture review or when query/path/explain don't surface enough context.
- After modifying code, run `graphify update .` to keep graph current (AST-only, no API cost).
## Notes
- Keep comms terse — drop articles, fluff, pleasantries. Code/commits/security written normally.
-106
View File
@@ -1,106 +0,0 @@
# CLAUDE.md
Guidance for Claude Code (claude.ai/code) working in this repo.
## What this is
Manga read-progress tracker, user read on **asurascans.com** (current domain; asuracomic.net 301s here) and **demonicscans.org** via **Violentmonkey**. Userscript inject on-page UI (floating button + slide-in panel), sync progress to self-hosted Go backend so bookmarks unify across both sites and devices.
## Hard constraints (drive design — don't violate)
Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free where plain web APIs suffice — keeps portability across engines:
- **Avoid `GM_*` unless needed.** Prefer page `localStorage` over `GM_setValue`/`GM_getValue`, on-page UI over `GM_registerMenuCommand`, plain `fetch()` over `GM_xmlhttpRequest` for cross-origin.
- Cross-origin `fetch()` work **only** against CORS-enabled backend. Manga sites `https://`, so backend **must be HTTPS** (else mixed-content block).
- Asura and Demonic are **separate origins with separate `localStorage`** — shared remote store only way to unify bookmarks. Cloud sync required, not optional.
- Userscript run in **isolated world**, so embedded API token safe from site's JS.
- Cloudflare's block on manga sites **IP-reputation-based, not universal — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — Cloudflare's bot scoring can flip previously-clean IP without notice. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, or direct probe) before finalize, not assumed from single earlier test.
## Architecture
```
Violentmonkey userscript (isolated world, per-site adapters, localStorage cache)
-- fetch() HTTPS --> reverse proxy (TLS + CORS) --> Go net/http --> SQLite (volume)
```
Backend-specific architecture (packages, endpoints, poller, config env vars) lives in `backend/CLAUDE.md`. Userscript-specific structure (adapters, retry queue, UI, live URL shapes) lives in `userscript/CLAUDE.md`.
## Commands
Backend (`cd backend`):
- Test all: `go test ./...`
- Single test: `go test -run TestName ./...`
- Build static binary: `CGO_ENABLED=0 go build`
Local stack: `docker compose up` (named volume mounted at `/data`, `restart: unless-stopped`).
Smoke test: `curl` endpoints with `Authorization: Bearer <token>`; confirm `OPTIONS` preflight return CORS headers and `/healthz` return 200.
## Forge: Gitea, not GitHub
`origin` is self-hosted Gitea instance (`gitea.violetcrown.my.id`), so **`gh` don't work here — use `tea` (Gitea CLI) for anything past plain git.** Common ones:
- Open PR: `tea pr create --head <branch> --base main --title "..." --description "..."`
- List / view / check out: `tea pr list`, `tea pr <n>`, `tea pr checkout <n>`
- Issues: `tea issue create`, `tea issue list`
- Auth lives in `tea login`, not `GH_TOKEN` env var.
`tea` print output as rendered boxes rather than plain text; PR URL lands on last line.
## Design system
Web UI + userscript panel follow **Cinder**, rules in `docs/design-system.md`
— source of truth Claude Design project `BookmarkManager Web UI`
(`969ac210-fe02-4c01-ae1b-9a271dcc779a`). Read it before touching
`backend/internal/web/static/style.css`, `backend/internal/web/templates/*`, or userscript
`TEMPLATE`/`CSS`. Core law: **ember means new chapter only** — no other
state (busy, error, destruction) may use `--ember`; destruction gets
`--danger`. No cards/corners/shadows, one `--measure: 760px` column, tokens
only (never hardcode hex outside `:root`), both colour branches touched
together. Any move that pulls series out of list (archive/finish/remove)
must be confirm-gated via its own `.confirm-row`; only restore fires
instantly.
## Security invariants
- Auth on `/bookmarks*`: require `Authorization: Bearer <API_TOKEN>`, **constant-time compare**, 401 otherwise.
- CORS: reflect `Origin` only when in `ALLOWED_ORIGINS`; allow `GET,PUT,DELETE,OPTIONS` + headers `Authorization,Content-Type`; answer preflight `OPTIONS` with `204`.
## Comments
Comment only if code alone can't carry info. Cost per read — must earn spot.
Write for:
- Why not what. Tradeoffs, non-obvious decisions.
- Load-bearing detail looking incidental — say so if "simplify" breaks it.
- Non-local consequence, invisible from function alone.
- Wire format / encoding / interface contract — save callers re-deriving.
- Gotcha/workaround, with ref if exists.
- Domain/business rule not derivable from code.
Skip:
- Restating code (no `// increment i` above `i++`).
- Trivial getter/setter/pass-through.
- Banners, dividers, `// helpers`.
- Change narration (`// fix bug`, `// as requested`, `// new impl`) — git's job.
- Commented-out code — delete.
- TODO without concrete action.
Style: one dense comment over function beats one per line inside. Tight, no worked example unless bug subtle. Wrong comment worse than none — update/delete on change. Default fewer — sparse+high-signal beats comprehensive.
Test: "competent reader get this from code in few sec?" Yes → skip. Needs detour through another file/spec/git-blame → write it.
## Relevant skills
`multi-stage-dockerfile` and `docker-compose-orchestration` for container work (referenced in plan).
`golang-code-style`, `golang-error-handling`, `golang-performance`, `golang-testing` for backend Go work.
## graphify
Project has knowledge graph at graphify-out/ with god nodes, community structure, cross-file relationships.
Rules:
- For codebase questions, first run `graphify query "<question>"` when graphify-out/graph.json exists. Use `graphify path "<A>" "<B>"` for relationships and `graphify explain "<concept>"` for focused concepts. Return scoped subgraph, usually much smaller than GRAPH_REPORT.md or raw grep output.
- If graphify-out/wiki/index.md exists, use for broad navigation instead of raw source browsing.
- Read graphify-out/GRAPH_REPORT.md only for broad architecture review or when query/path/explain don't surface enough context.
- After modifying code, run `graphify update .` to keep graph current (AST-only, no API cost).
Symlink
+1
View File
@@ -0,0 +1 @@
AGENTS.md
+109
View File
@@ -0,0 +1,109 @@
# Bookmark Manager
Read-progress tracker for serialised fiction. A reader browses third-party manga and
novel sites; userscripts capture where they got to and sync it to a self-hosted backend,
so progress survives across sites and devices.
## Language
**Series**:
One ongoing work — a manga or a novel — as published by a Site. Identified by the canonical
slug the Site itself publishes for it, never by its title and never by a Chapter Slug. A
Series exists once and is shared by every Reader who bookmarks it; it owns the facts that
are true regardless of who is reading — title, cover, Latest Chapter. A Reader cannot
change them; they describe the Series, not anyone's relationship to it.
_Avoid_: manga, title, book, comic
**Site**:
One third-party source a Series is published on. A Series on two Sites is two Series.
_Avoid_: source, host, provider, domain
**Chapter Slug**:
A slug a Site builds its chapter addresses from. Not an identity: one Series may have
several, any of them may differ from the slug that identifies the Series, and none is
computable from another. Only the Site's own links say which ones a Series uses, so a
Chapter Slug is always discovered, never derived.
_Avoid_: series slug, url slug, permalink, chapter path
**Cover**:
The image that stands for a Series wherever it is listed. A fact about the Series like
its title — one Cover per Series, shared by every Reader, never per-Reader. Defined by
what a Reader's browser can display, not by where the Site keeps the picture: an address
no client can load is not a Cover, it is a missing one.
_Avoid_: thumbnail, poster, image URL, artwork
**Reader**:
A person with their own Progress. Exactly one per set of credentials, so there is no
separate "account" concept to model — the credential belongs to the Reader.
_Avoid_: user, account, member, subscriber
**Bookmark**:
One Reader's tracked relationship with one Series, holding only what differs between
Readers: Progress, Favourite, Lifecycle bucket. Facts about the Series itself belong
to the Series, not here.
_Avoid_: entry, item, record, subscription
**Library**:
One of the two halves of the collection — manga or novel — selected by a Bookmark's
`kind`. The web UI and the userscripts each address exactly one Library at a time.
Not a per-person concept: "everything one person has bookmarked" is a different idea
and must not be called a Library.
_Avoid_: section, tab, category
**Progress**:
The furthest chapter a reader has actually read in a Series. Only a change in Progress
is real activity, so only Progress reorders the list.
_Avoid_: position, bookmark (the noun is taken), last read
**Latest Chapter**:
The highest-numbered chapter a Site has published for a Series. The number is what
ranks it, never a date and never the Site's own "newest chapter" banner — where a Site
disagrees with itself, its list of chapters is the record and its summary of that list
is not. Established by a Poll and, between Polls, by a Sighting. Distinct from Progress
in every way that matters: it is a fact about the Site, not about the reader, and it
must never reorder the list.
_Avoid_: newest, current chapter, update
**Poll**:
The backend's own check of a Site for a Series's Latest Chapter, made without the
Reader present. Performed once per Series no matter how many Readers bookmarked it —
a Poll is work done on behalf of the Series, never on behalf of a Reader.
_Avoid_: scrape, refresh, check, sync
**Poll Lane**:
One Site's own stream of Polls, carrying the pace at which that Site is willing to be
asked. Every Site has exactly one and no Lane can slow, block or borrow from another's;
a Reader never has one and never influences one.
_Avoid_: worker, queue, scheduler, batch, wave
**Sighting**:
What a Reader's browser happened to see of a Series's Latest Chapter while that Reader
was on the page. It reports the same fact as a Poll but carries none of its authority:
a Poll always overrules it, and only a Sighting on a Series no Reader else holds may
defer one. A Sighting a later Poll contradicts downwards is a false Sighting, and
enough of those cost the Reader the right to defer at all.
_Avoid_: client report, user poll, observation, claim
**Acquisition**:
The single read of a Series page made the moment the Series first exists, giving it
both its Latest Chapter and its Cover without waiting for the Lane's pace. Distinct
from a Poll in the two ways that matter: a Reader is present — it is triggered by
their first Bookmark of that Series — and it is the only read that establishes a
Cover rather than refreshing facts. It happens once in a Series's life; every later
read of the same page is a Poll.
_Avoid_: initial poll, first fetch, prefetch, warm-up
**New Chapter**:
The state where Latest Chapter is ahead of Progress. The single condition the ember
accent is permitted to signal.
_Avoid_: unread, update available
**Lifecycle bucket**:
Which of three mutually exclusive states a Bookmark sits in — reading, archived, or
finished. A Bookmark is in exactly one. Orthogonal to being a favourite.
_Avoid_: state, status (as a domain word), list
**Favourite**:
A reader's manual pin on a Bookmark. Orthogonal to the Lifecycle bucket, and never a
reason to reorder the list.
_Avoid_: starred, pinned, priority
+299
View File
@@ -0,0 +1,299 @@
# SQLite → Postgres cutover runbook
One-way, one-time. Moves the owner's reading history out of the retired SQLite
volume (`<compose project>_bookmarks-data`, holding `/data/bookmarks.db`) and into
the Postgres schema the migration runner builds. There is no dual-write period:
the old database is read once, at cutover, from a **fresh export** — anything
written to SQLite after the export is lost, so the old API must already be down.
Routine deploys are `REDEPLOY.md`; first-time setup is `DEPLOY.md`. This file is
run once and then only ever read for reference.
Proven end to end on 2026-08-08 against a copy of `bookmarks-20260807-213515.db`
into a scratch Postgres: 29 Bookmarks (18 reading, 11 archived, 7 favourites),
29 Series, all owned by the seeded Reader, and every field of every row matching
the source exactly. Production was not touched.
---
## 0. The generator is throwaway
It is written at cutover, run once, and deleted. It is deliberately **not** in
this repository and never will be:
- Its output is the owner's personal reading history. That does not enter
version control.
- It reads SQLite. The backend module dropped `modernc.org/sqlite` (ADR-0001);
a committed generator would drag the dependency back in through the side door.
So §3 specifies the transformation rather than shipping a script. It is a
twenty-line program against a sixteen-column table (fifteen after `key`, which
is dropped) — writing it from the spec below costs less than maintaining it
would.
Beyond `DEPLOY.md`'s prerequisites (Docker and Compose), this runbook needs
`python3`: its stdlib `sqlite3` module is the whole SQLite dependency, and §5's
read-path check uses it in place of `jq`, which the server does not have. It
does not have to run on the server — §3 only reads the snapshot copy, so it can
run on a laptop and the resulting `import.sql` be copied over.
---
## 1. Stop the old API and take a fresh export
**Order matters.** Export after the API stops, or you migrate a snapshot that is
already stale.
```bash
cd ~/mangaBookmark # wherever the checkout lives
COMPOSE="docker compose -f docker-compose.yml -f docker-compose.prod.yml"
BACKUP_DIR="$(cd .. && pwd)/$(basename "$PWD")-backups"; mkdir -p "$BACKUP_DIR"
STAMP=$(date -u +%Y%m%d-%H%M%S)
# The volume is <compose project>_bookmarks-data, and the project name defaults
# to the lowercased *directory* name, not the repo name — on this host the
# checkout is ~/mangaBookmark, so the volume is mangabookmark_bookmarks-data.
# Derive it exactly rather than with a `--filter name=` substring match, which
# would return every volume whose name merely contains the string.
VOL="$(basename "$PWD" | tr '[:upper:]' '[:lower:]')_bookmarks-data"
docker volume inspect "$VOL" >/dev/null && echo "$VOL"
$COMPOSE stop bookmark-api
# A clean SIGTERM closes the store, which checkpoints and unlinks the -wal, so
# bookmarks.db alone is then the whole database. But `compose stop` SIGKILLs
# after 10s, and a surviving -wal holds writes the main file does not — assert
# it is gone rather than assuming the shutdown was clean.
docker run --rm -v "$VOL":/d:ro alpine ls -l /d # -> bookmarks.db, alone
docker run --rm -v "$VOL":/from:ro -v "$BACKUP_DIR":/to \
alpine cp /from/bookmarks.db "/to/bookmarks-$STAMP.db"
ls -lh "$BACKUP_DIR/bookmarks-$STAMP.db"
```
If `-wal` and `-shm` are still there, the container was killed mid-write. Copy
all three under the same basename and let SQLite replay the log when §3 opens
it — copying only `bookmarks.db` silently drops whatever the log still holds.
Work on a **copy** of that file for the rest of this runbook. The export is the
last line of retreat; nothing below should be able to write to it.
```bash
mkdir -p /tmp/cutover && cp "$BACKUP_DIR/bookmarks-$STAMP.db" /tmp/cutover/snapshot.db
chmod 444 /tmp/cutover/snapshot.db
```
---
## 2. Bring up Postgres with the schema and the owner Reader
The new stack builds its own schema and seeds exactly one Reader from
`OWNER_DISCORD_ID` — do not hand-write either. Pull the Postgres-era commit
first: on a server that has only ever run the SQLite build, `--build` without a
pull silently rebuilds the old image and the checks below fail with
"relation readers does not exist".
```bash
git pull --ff-only
git log --oneline -1
# .env needs the new required vars (DATABASE_URL is built from
# POSTGRES_PASSWORD; TOKEN_KEY, OWNER_DISCORD_ID and the DISCORD_* set are
# required). Compose fails at start for a missing one.
git diff HEAD@{1} HEAD -- .env.example docker-compose.yml docker-compose.prod.yml
$COMPOSE up -d --build
docker logs bookmark-api --tail 20 # -> "listening on :8080"
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks -c '\dt'
# -> bookmarks, readers, schema_migrations, series, sessions
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks \
-c 'select id, discord_id from readers'
# -> exactly one row, and discord_id is the owner's
```
Two rows in `readers`, or zero, means `OWNER_DISCORD_ID` is wrong or the seed
failed. Stop here — the import attaches history to "the oldest reader row", and
that is only unambiguous while there is one.
`bookmarks` and `series` are empty at this point. That is what makes the import
a plain sequence of `INSERT`s with no conflict handling.
---
## 3. Generate the import SQL
Read `/tmp/cutover/snapshot.db` and emit plain SQL on stdout. The old table is
flat and its columns map one-for-one onto the split schema — no transformation
beyond the split itself:
| SQLite `bookmarks` column | lands in | notes |
|---|---|---|
| `site`, `series_id` | both tables | the Series key; the wire `key` column is dropped, it is re-derived as `site:series_id` on read |
| `title`, `series_url`, `cover`, `kind` | `series` | shared facts (ADR-0003) |
| `latest_chapter`, `latest_chapter_num`, `latest_checked_at` | `series` | `latest_chapter_num` is nullable on **both** sides and `NULL` is meaningful — never coerce it to `0` |
| `last_chapter`, `last_chapter_num`, `last_chapter_url` | `bookmarks` | Progress |
| `favorite`, `status`, `updated_at` | `bookmarks` | `favorite` is `0`/`1` in SQLite and a real `boolean` in Postgres — emit `true`/`false` |
| — | `bookmarks.reader_id` | the seeded owner |
**`latest_chapter_num` is the only column where `NULL` survives.** The SQLite
table declares `title`, `series_url`, `cover`, `last_chapter`,
`last_chapter_url` as bare `TEXT` and `last_chapter_num` as bare `REAL` — all
six nullable — while their Postgres targets are `NOT NULL DEFAULT ''` /
`NOT NULL DEFAULT 0`. One `NULL` in any of them aborts the whole import on a
not-null violation. Coalesce them in the `SELECT` (`ifnull(title,'')`,
`ifnull(last_chapter_num,0)`, …) rather than discovering it at §5. The
2026-08-07 export happened to have none; a fresh export is not promised the
same.
Rules the generator must follow:
- **Series first, Bookmarks second.** `bookmarks` has a foreign key onto
`series (site, series_id)`; the reverse order fails on the first row.
- **`SELECT DISTINCT` the Series.** The old key's uniqueness already makes
`(site, series_id)` unique, so this is belt and braces — but if it ever
collapses two rows, the count check in §5 catches it.
- **Never hardcode the reader id.** Emit
`INSERT INTO bookmarks (reader_id, …) SELECT id, … FROM owner`, where `owner`
is a temp table built once at the top:
`CREATE TEMP TABLE owner ON COMMIT DROP AS SELECT id FROM readers ORDER BY id LIMIT 1;`
A literal id is a number nobody verifies; this one cannot be wrong.
- **Wrap the whole file in `BEGIN; … COMMIT;`, temp table included.** Postgres
has transactional DDL and DML: a failure half way leaves an empty database
rather than half a library. The ordering is load-bearing —
`ON COMMIT DROP` outside the transaction means the temp table drops itself
the instant it is created (psql autocommits) and every
`SELECT … FROM owner` then fails.
- **Quote strings by doubling `'`.** Titles contain apostrophes and the URLs
contain `%5C%27` escapes. Emit standard SQL literals only — no `E''` strings,
no backslash escaping (`standard_conforming_strings` is on, so a backslash is
a literal backslash and the URLs survive verbatim).
```bash
python3 gen_import.py /tmp/cutover/snapshot.db > /tmp/cutover/import.sql
wc -l /tmp/cutover/import.sql # -> 2 header + 29 series + 29 bookmarks + framing
```
---
## 4. Review it by eye
29 rows is small enough to actually read, and this is the last point at which a
mistake is free:
```bash
less /tmp/cutover/import.sql
grep -c '^INSERT INTO series' /tmp/cutover/import.sql # -> 29
grep -c '^INSERT INTO bookmarks' /tmp/cutover/import.sql # -> 29
```
Look for: a title whose apostrophe is not doubled, a `favorite` that is still
`0`/`1`, a `latest_chapter_num` that turned into `0`, and any `reader_id`
written as a bare number.
---
## 5. Apply it
```bash
docker cp /tmp/cutover/import.sql "$($COMPOSE ps -q postgres)":/tmp/import.sql
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks -v ON_ERROR_STOP=1 \
-f /tmp/import.sql
```
`ON_ERROR_STOP=1` is not optional: without it `psql` reports the error, keeps
going, and exits `0` on a half-imported database.
Then the checklist. Every number here is asserted, not eyeballed:
```bash
OWNER=$(grep -E '^OWNER_DISCORD_ID=' .env | cut -d= -f2)
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks -x -c "
SELECT (SELECT count(*) FROM bookmarks) AS bookmarks_total,
(SELECT count(*) FROM bookmarks WHERE status='reading') AS reading,
(SELECT count(*) FROM bookmarks WHERE status='archived')AS archived,
(SELECT count(*) FROM series) AS series_total,
(SELECT count(*) FROM readers) AS readers_total,
(SELECT count(*) FROM bookmarks
WHERE reader_id <> (SELECT id FROM readers WHERE discord_id='$OWNER'))
AS not_owned_by_owner;"
```
`not_owned_by_owner` resolves the Reader by **Discord id**, not by
`ORDER BY id LIMIT 1`. The second form is the expression §3 tells the generator
to import with, so comparing against it is true by construction and could never
fail; resolving by Discord id is an independent check that the rows landed on
the identity the owner will actually log in as. If that subquery returns NULL
the whole count comes back `0` for the wrong reason — hence `readers_total`
beside it.
Expected, for the 2026-08-07 export: `29`, `18`, `11`, `29`, `1`, `0`. Against a
different export, the invariants rather than the literals are what hold:
- `bookmarks_total` equals the SQLite row count.
- `reading + archived` equals `bookmarks_total` (nothing was `finished`).
- `series_total` equals `SELECT count(*) FROM (SELECT DISTINCT site, series_id FROM bookmarks)`
in the source.
- `readers_total` is `1` and `not_owned_by_owner` is `0`.
Then spot-check the values themselves against the source — read position,
favourite flag and latest chapter. Take the sample from each bucket explicitly:
`ORDER BY updated_at DESC LIMIT 5` alone returns the most recently *progressed*
rows, which are the ones least likely to be archived.
```bash
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks -c "
SELECT s.title, b.last_chapter, b.last_chapter_num, b.favorite,
s.latest_chapter, b.status
FROM bookmarks b JOIN series s USING (site, series_id)
WHERE b.status='reading' ORDER BY b.updated_at DESC LIMIT 3;"
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks -c "
SELECT s.title, b.last_chapter, b.last_chapter_num, b.favorite,
s.latest_chapter, b.status
FROM bookmarks b JOIN series s USING (site, series_id)
WHERE b.status='archived' ORDER BY b.updated_at DESC LIMIT 2;"
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks -c "
SELECT s.title, b.last_chapter, b.last_chapter_num, b.favorite,
s.latest_chapter, b.status
FROM bookmarks b JOIN series s USING (site, series_id)
WHERE b.favorite ORDER BY b.updated_at DESC LIMIT 2;"
```
Compare each against the same row in the snapshot — the generator's own source
is the reference, so read it back with the same `python3` you used in §3.
Finally, prove the **read path**, not just the tables — this is the check that
would catch a correct import behind a broken join:
```bash
API=https://bookmark-api.violetcrown.my.id
# Your own Reader credential: sign in to the web UI and take it from the
# Userscripts panel's install link, or read the API_TOKEN constant out of an
# already-installed script. There is no credential in .env to grep.
TOKEN=<your Reader credential>
curl -s -H "Authorization: Bearer $TOKEN" $API/bookmarks |
python3 -c 'import json,sys; print(len(json.load(sys.stdin)))' # -> 29
```
---
## 6. Afterwards
- **Keep the old SQLite volume for a month.** It is already undeclared in
compose, so `docker compose down -v` cannot take it. Remove it by hand once
the Postgres data has been trusted for a while. That happens in a shell where
`$VOL` from §1 is long gone, so re-derive it:
`docker volume rm "$(basename ~/mangaBookmark | tr '[:upper:]' '[:lower:]')_bookmarks-data"`
(see `REDEPLOY.md` §1).
- **Delete the generator and the working copies:** `rm -rf /tmp/cutover`. The
timestamped export in `$BACKUP_DIR` is the copy that is kept.
- **Take the first Postgres dump immediately** — `REDEPLOY.md` §1. Until that
exists, the only backup of the migrated data is the SQLite file it came from.
If the import is wrong, there is nothing to unpick: drop the rows and start
again from §3 — `TRUNCATE bookmarks, series;` leaves the seeded Reader and the
schema in place.
+357 -64
View File
@@ -11,7 +11,7 @@ ACME/cert resolver, and control a domain.
- Docker + Docker Compose on the server.
- A Traefik instance watching a Docker network (default name assumed: `proxy`).
- DNS: an `A`/`AAAA` record for `bookmark-api.<yourdomain>` pointing at the server.
- The repo copied to the server, e.g. `/opt/bookmarkmanager/` (needs `backend/`,
- The repo copied to the server, e.g. `~/mangaBookmark/` (needs `backend/`,
`docker-compose.yml`, `docker-compose.prod.yml`, `.env.example`).
Confirm the Traefik network exists (create if not):
@@ -25,22 +25,46 @@ docker network ls | grep proxy || docker network create proxy
## 1. Configure `.env`
```bash
cd /opt/bookmarkmanager
cd ~/mangaBookmark
cp .env.example .env
```
Edit `.env`:
```ini
# Required — long random secret, also goes in the userscript.
API_TOKEN=<paste output of: openssl rand -hex 32>
# Required — secret every Reader's userscript credential is derived from.
# Only SHA-256 hashes of credentials are stored.
TOKEN_KEY=<paste output of: openssl rand -hex 32>
# Required — the owner's Discord user ID. Seeds the first Reader: the
# administrator, and the owner of every bookmark that predates registration.
# The value is the snowflake in your Discord profile (Settings →
# Advanced → Developer Mode → right-click your name → Copy User ID).
OWNER_DISCORD_ID=<discord user id>
# CORS allowlist — leave as-is unless a site changes hostname.
ALLOWED_ORIGINS=https://asuracomic.net,https://asurascans.com,https://demonicscans.org,https://comix.to,https://kagane.to
ALLOWED_ORIGINS=https://asurascans.com,https://demonicscans.org,https://comix.to,https://kagane.to,https://novelfull.com,https://lightnovelworld.net
# Required — password for the bundled Postgres container. Compose builds the
# backend's DATABASE_URL out of it and has no fallback for either.
POSTGRES_PASSWORD=<paste output of: openssl rand -hex 24>
# Leave unset. Only set this to point the backend at a Postgres compose does
# not run; it then replaces the URL built from POSTGRES_PASSWORD above.
# DATABASE_URL=postgres://user:pass@host:5432/bookmarks?sslmode=require
# Required path inside bookmark-api. Compose builds the image and mounts the
# named cover-data volume at this path.
COVER_DIR=/covers
# Required — the origin this deployment answers on, no trailing slash. Cover
# URLs on the wire are absolute, because the userscript renders them on a
# Site's own origin (ADR-0007). Same host as BOOKMARK_API_HOST below.
PUBLIC_BASE_URL=https://bookmark-api.violetcrown.my.id
# Required for the Traefik override. Both have no fallback — compose refuses
# to start without them. BOOKMARK_WEB_HOST is required even if you never set
# WEB_PASSWORD; see 1b.
# to start without them. BOOKMARK_WEB_HOST is required even if the web UI
# were unused; see 1b.
BOOKMARK_API_HOST=bookmark-api.violetcrown.my.id
BOOKMARK_WEB_HOST=bookmark.violetcrown.my.id
@@ -50,13 +74,33 @@ BOOKMARK_WEB_HOST=bookmark.violetcrown.my.id
# TRAEFIK_CERTRESOLVER=le
```
Generate + insert the token in one line:
Generate + insert the two secrets in three lines:
```bash
sed -i "s|^API_TOKEN=.*|API_TOKEN=$(openssl rand -hex 32)|" .env
grep -E '^API_TOKEN=' .env # copy this — the userscript needs the same value
sed -i "s|^TOKEN_KEY=.*|TOKEN_KEY=$(openssl rand -hex 32)|" .env
sed -i "s|^POSTGRES_PASSWORD=.*|POSTGRES_PASSWORD=$(openssl rand -hex 24)|" .env
grep -E '^TOKEN_KEY=' .env
```
`TOKEN_KEY` derives every Reader's userscript credential (issue #24); only
SHA-256 hashes of the credentials are stored, so this secret is what a
database leak alone cannot recover. Changing it invalidates every installed
script at once.
`POSTGRES_PASSWORD` is read **only while the `postgres-data` volume is empty**,
which in practice means at first boot. Changing it afterwards changes the URL
the backend dials but not the password the database expects, and `bookmark-api`
crash-loops on `password authentication failed`. Set it before §2 and leave it
alone.
An `.env` written before issue #100 carries the old poll-pace names
(`LATEST_CHAPTER_POLL_COOLDOWN`, `_BROWSER_COOLDOWN`, `_INTERVAL`, `_BATCH`,
`_STAGGER`). All five are dead configuration now — the pace lives in the Site
registry (`backend/internal/latest/sites.go`), so **delete those lines** and
keep only the kill switch `LATEST_CHAPTER_POLL_ENABLED`. Leaving them behind
is harmless (nothing reads them) but silently misleads the next person who
edits the file.
> Match `TRAEFIK_ENTRYPOINT` / `TRAEFIK_CERTRESOLVER` to your Traefik's actual
> names (check your Traefik static config — common alternatives: `https`,
> `myresolver`, `cloudflare`). Wrong names = no certificate issued.
@@ -65,44 +109,58 @@ grep -E '^API_TOKEN=' .env # copy this — the userscript needs the same value
## 1b. Web UI
The browser UI is served by the same container on a second hostname.
The browser UI is served by the same container on a second hostname. Sign-in
is a Discord authorization code grant (ADR-0002): the owner's Discord account,
gated by membership in one configured guild.
1. Add a DNS `A`/`AAAA` record for `bookmark.<yourdomain>` pointing at the server —
the same address as `bookmark-api.<yourdomain>`.
1. Add a DNS `A`/`AAAA` record for `bookmark.<yourdomain>` pointing at the
server — the same address as `bookmark-api.<yourdomain>`.
2. Set both variables in `.env`:
2. Create the Discord application at <https://discord.com/developers/applications>:
- **OAuth2 → Redirects:** add the exact callback URL
`https://bookmark.violetcrown.my.id/auth/discord/callback`. Discord
matches it verbatim — a trailing slash or different hostname breaks
sign-in.
- **OAuth2 → General:** note the Client ID, and generate a Client Secret.
- No scopes or bot setup are needed in the dashboard; the service requests
`identify` and `guilds.members.read` itself, and checks the *user's*
membership of the guild, not the application's.
3. Set the variables in `.env`:
```ini
BOOKMARK_WEB_HOST=bookmark.violetcrown.my.id
WEB_PASSWORD=<paste output of: openssl rand -base64 18>
DISCORD_CLIENT_ID=<client id>
DISCORD_CLIENT_SECRET=<client secret>
DISCORD_GUILD_ID=<guild snowflake>
DISCORD_REDIRECT_URI=https://bookmark.violetcrown.my.id/auth/discord/callback
# Optional: only members holding this role may sign in.
# DISCORD_REQUIRED_ROLE=<role snowflake>
```
Generate and insert in one line:
The guild id is in Discord's client with Developer Mode on: right-click the
server name → Copy Server ID. The four uncommented variables are required —
the backend refuses to start without them. Guild membership *is*
registration: any member of `DISCORD_GUILD_ID` becomes a Reader with their
own library on their first sign-in. `OWNER_DISCORD_ID` from §1 is only the
administrator — the Reader who can revoke another Reader's sessions.
```bash
sed -i "s|^WEB_PASSWORD=.*|WEB_PASSWORD=$(openssl rand -base64 18)|" .env
grep -E '^WEB_PASSWORD=' .env # this is what you type into the site
```
3. Redeploy and check:
4. Redeploy and check:
```bash
docker compose -f docker-compose.yml -f docker-compose.prod.yml up -d --build
curl -s -o /dev/null -w '%{http_code}\n' https://bookmark.violetcrown.my.id/
```
Expected `200`, serving the login page.
Expected `200`, serving the login page with the Discord button. Signing in
lands on the library; an account outside the guild is refused with a message
that names neither the guild nor its id.
Leaving `WEB_PASSWORD` unset is safe: the web routes are not registered and `/`
returns 404. The userscript's API on `BOOKMARK_API_HOST` is unaffected either way.
`BOOKMARK_WEB_HOST` itself is required by the prod override regardless — like
`BOOKMARK_API_HOST`, its Traefik label has no fallback, so `docker compose up`
refuses to start without it even if `WEB_PASSWORD` is unset and the web UI is
otherwise dormant.
Sessions are signed with a key derived from `API_TOKEN` and `WEB_PASSWORD`, so
rotating either one logs every browser out. The session cookie lasts 60 days.
Sessions are rows in the database: the cookie carries only an opaque id, and
every request looks the row up and checks its expiry. Deleting a session row —
or the whole `sessions` table — logs the browser out immediately; nothing is
signed, so rotating a credential does not affect browser sessions. Sessions
last 60 days.
---
@@ -116,16 +174,23 @@ This merges the base file (build/image/env/volume) with the prod override
(no host port, Traefik network + router labels). Always pass **both** `-f`
flags — the prod file is not standalone.
Two services come up: `bookmark-api` (the backend) and `headless-shell`, a CDP
sidecar the poller uses to fetch kagane (behind a Cloudflare JS challenge).
It has no published port — only `bookmark-api` can reach it, over
`BROWSER_WS_URL`. Missing or unreachable, the poller just skips kagane and
logs it; nothing else is affected.
Two services come up: `bookmark-api` (the backend) and `postgres` (its
database, `postgres:17-alpine`). Postgres publishes no port — it sits alone
with `bookmark-api` on an `internal: true` network — and stops everything if it
is missing: `bookmark-api` waits for `pg_isready` to pass, then applies its
embedded migrations, and only then listens. The schema is created that way;
there is nothing to import by hand.
There is deliberately no browser here. Kagane and novelfull need one, and it
runs on a **separate machine** over the tailnet — §7. Until you do that step,
`BROWSER_WS_URL` is unset, the poller logs and skips those two sites, and
everything else works normally.
Check it's up and healthy:
```bash
docker compose -f docker-compose.yml -f docker-compose.prod.yml ps
# bookmark-api Up; postgres Up (healthy)
docker logs bookmark-api --tail 20 # expect: "listening on :8080 ..."
```
@@ -143,7 +208,10 @@ curl -s https://bookmark-api.violetcrown.my.id/healthz # -> ok
curl -s -o /dev/null -w '%{http_code}\n' \
https://bookmark-api.violetcrown.my.id/bookmarks # -> 401
TOKEN=$(grep -E '^API_TOKEN=' .env | cut -d= -f2)
# A Reader's own credential. It is derived, never stored in .env — take it from
# the Userscripts panel's install link after signing in, or from an installed
# script's API_TOKEN constant.
TOKEN=<your Reader credential>
curl -s -H "Authorization: Bearer $TOKEN" \
https://bookmark-api.violetcrown.my.id/bookmarks # -> []
@@ -162,29 +230,30 @@ a bad cert makes the browser block the userscript's `fetch()` (mixed content).
## 4. Configure the userscript
Edit the config block at the top of `userscript/manga-bookmark.user.js`:
The bindmounted `userscript/*.user.js` files carry `__API_TOKEN__` placeholders
and the deployment's `@downloadURL`/`@updateURL` lines. Check the metadata
block — it ships hardcoded to this deployment's domain, so a deployer who
copies the repo to another domain must edit the two lines or the script
auto-updates from someone else's backend:
```js
const API_BASE = "https://bookmark-api.yourdomain.com"; // no trailing slash
const API_TOKEN = "<same token as .env>";
// @downloadURL https://bookmark-api.yourdomain.com/u/__API_TOKEN__/manga-bookmark.user.js
// @updateURL https://bookmark-api.yourdomain.com/u/__API_TOKEN__/manga-bookmark.user.js
```
The token sits in the userscript's isolated world — the manga sites' JS can't
read it.
Also edit the `@downloadURL`/`@updateURL` metadata lines near the top of the
file — they ship hardcoded to this deployment's domain and token, so a
deployer who skips them ends up auto-updating from someone else's backend.
See "Installing / updating the userscript" below for how those two lines are
used.
The backend substitutes `__API_TOKEN__` with the requesting Reader's derived
credential at serve time (issue #24), so no real credential ever sits in the
file. Only the `API_BASE` constant and the metadata hostname are deployer
edits; do not put a credential in this file.
---
## 5. Install on Bromite
1. Bromite → **Settings → User scripts** → enable (accept the permission prompt).
2. Put the edited `manga-bookmark.user.js` on the device (save the file, or open
its raw URL). Bromite detects `.user.js` and offers to install.
2. Sign in to the web UI, open the **Userscripts** panel, and open the install
link — Bromite detects `.user.js` and offers to install. The script already
carries your credential; you never see or type one.
3. Confirm install — the `@match` list covers both sites.
4. Open a series on asurascans.com or demonicscans.org → a 📑 button appears
bottom-right → tap → **+ Bookmark this**.
@@ -192,6 +261,9 @@ used.
Optional desktop test: the script is `GM_*`-free, so the same file installs in
Tampermonkey/Violentmonkey for quick checks before going mobile.
Rotating the credential in the same web-UI panel invalidates every installed
copy immediately — reinstall on all devices, or they silently stop syncing.
---
## 6. Smoke-test the full loop
@@ -206,6 +278,206 @@ Tampermonkey/Violentmonkey for quick checks before going mobile.
---
## 7. The browser, on the home machine
Kagane and novelfull sit behind a Cloudflare JavaScript challenge no TLS
fingerprint clears, so the poller reaches them through a real Chrome over CDP.
That browser does **not** run on the VPS: it held 471 MiB of a 1974 MiB box
with no swap, and a residential IP avoids the cloud-hosting-IP signature Bot
Fight Mode challenges anyway (ADR-0006). It
is its own compose unit, deployed and updated independently of everything
above.
Do this after §2, on the second machine. Both machines must already be on the
same tailnet.
First, on the VPS, record what you are reclaiming — this is the whole point of
the move and there is no way to measure it afterwards:
```bash
free -m | awk '/^Mem:/ {print "available before:", $NF, "MiB"}'
```
Take it again after §7 is finished and the old sidecar is gone. Expect roughly
the sidecar's former footprint back (measured at 471 MiB working set, 595 MiB
cgroup).
**On the home machine:**
```bash
git clone <this repo> ~/mangaBookmark && cd ~/mangaBookmark/chrome
tailscale ip -4 # -> 100.x.y.z, this machine's tailnet IP
cp .env.example .env
echo "BROWSER_BIND_ADDR=$(tailscale ip -4)" >> .env
docker compose up -d --build
```
The clone is only for `chrome/`; nothing else on this machine reads the rest of
the repo. The unit is its own compose project (`bookmark-browser`), so it shares
no volume, network or lifecycle with an API stack that happens to sit beside it.
`BROWSER_BIND_ADDR` has no default on purpose. CDP authenticates nothing —
whatever reaches port 9222 drives the browser and, through it, this host — so
the bind address *is* the access control, backed by Tailscale device identity.
On the VPS that job was done by Docker network membership; this machine has a
real LAN, so `0.0.0.0` would be a hole punched into your home network. Compose
refuses to start rather than guess.
**Narrow it to the one device that needs it.** The bind address keeps CDP off
your LAN; it still leaves port 9222 open to every device on the tailnet, and
CDP has no login — a compromised phone is enough to drive this host. A new
tailnet's policy is allow-all, so this is the step that makes "Tailscale
identity is the access control" true rather than aspirational.
Tailscale has no `deny`, so a restriction is expressed by removing the blanket
grant and enumerating what is left. That only works if the browser machine can
be *excluded* from a selector that still covers your own devices — which is
what tagging buys: a tagged device has no user, so `autogroup:member` and
`autogroup:self` stop matching it. Tagging is the mechanism, not decoration.
In the admin console, under **Access controls**, the shipped policy grants
`{"src": ["*"], "dst": ["*"], "ip": ["*"]}`. Replace it:
```jsonc
{
"tagOwners": {
// Empty list: implicitly owned by the tailnet Owner/Admins, which is you.
"tag:bookmark-api": [],
"tag:bookmark-browser": [],
},
"grants": [
// The only thing on the tailnet that may drive the browser.
{
"src": ["tag:bookmark-api"],
"dst": ["tag:bookmark-browser"],
"ip": ["tcp:9222"],
},
// Your own devices reach your own devices, and the VPS, in full.
{
"src": ["autogroup:member"],
"dst": ["autogroup:self", "tag:bookmark-api"],
"ip": ["*"],
},
// On the browser machine you get SSH and nothing else. Widen this to `*`
// and the restriction above is void; delete it and you are locked out.
{
"src": ["autogroup:member"],
"dst": ["tag:bookmark-browser"],
"ip": ["tcp:22"],
},
// Uncomment if you route traffic through an exit node — dropping the
// blanket grant takes exit-node access with it.
// {"src": ["autogroup:member"], "dst": ["autogroup:internet"], "ip": ["*"]},
],
// Tagged devices left `autogroup:self`, so Tailscale SSH needs them named.
// Irrelevant if you reach these boxes with ordinary sshd over the tailnet —
// that is the `tcp:22` grant above.
"ssh": [
{
"action": "check",
"src": ["autogroup:member"],
"dst": ["autogroup:self", "tag:bookmark-api", "tag:bookmark-browser"],
"users": ["autogroup:nonroot", "root"],
},
],
// Run on every save, so a later edit that reopens 9222 is rejected outright.
"tests": [
{ "src": "tag:bookmark-api", "accept": ["tag:bookmark-browser:9222"] },
{
"src": "you@example.com",
"accept": ["tag:bookmark-browser:22"],
"deny": ["tag:bookmark-browser:9222"],
},
],
}
```
Then apply the tags — on the VPS and the home machine respectively:
```bash
sudo tailscale up --advertise-tags=tag:bookmark-api
sudo tailscale up --advertise-tags=tag:bookmark-browser
```
Each re-authenticates in a browser and issues a new node key; the tailnet IP is
unchanged, so `BROWSER_WS_URL` and `BROWSER_BIND_ADDR` still hold. Key expiry is
disabled once a device is tagged, which is what you want for a server — an
expired key would otherwise take the poller down every few months.
**Tagging replaces the device's user identity**, so do this only to machines
that exist to run these services. If your "home machine" is also your daily
driver, tag it anyway and reach it through the `:22` rule above, or skip the
tag and accept that any device of yours can reach CDP.
Enforcement is by the destination's packet filter, so the check below is real,
not advisory.
Prove the bind is tight, from the home machine itself:
```bash
curl -s -m 3 http://$(tailscale ip -4):9222/json/version # -> JSON
curl -s -m 3 http://<this machine's LAN IP>:9222/json/version
# -> curl: (7) Failed to connect ... Connection refused
```
The first call is also what wakes Chrome: it is not running until something
connects, and it is reaped again after five idle minutes. A cold first response
takes a few seconds; that is the browser starting, not a fault.
That check proves the *bind*, not the ACL — traffic that starts on the node is
not filtered. Prove the ACL from somewhere else: on your laptop or phone the
same URL must now time out, and from the VPS it must answer.
```bash
# on any other device of yours -> hangs until timeout
curl -s -m 5 http://<home machine tailnet IP>:9222/json/version
# on the VPS -> JSON
curl -s -m 20 http://<home machine tailnet IP>:9222/json/version
```
**On the VPS:**
```bash
cd ~/mangaBookmark
echo 'BROWSER_WS_URL=ws://100.x.y.z:9222' >> .env # the home machine's tailnet IP
docker compose -f docker-compose.yml -f docker-compose.prod.yml up -d
```
It must be the tailnet **IP**. A MagicDNS hostname fails: Chrome's DevTools HTTP
handler answers `/json/version` with a 500 for any `Host` header that is not an
IP or `localhost`, and the failure looks like a broken site rather than a broken
hostname.
**Prove it end to end.** This is the only check that says the challenge actually
clears from that machine's egress — it fetches a real kagane cover and a real
chapter list:
```bash
cd backend
SMOKE_BROWSER_WS_URL=ws://100.x.y.z:9222 go test -run TestSmokeKagane ./internal/latest
```
A red run means "not clearing from this address right now", which is a live
fact to re-check before it is a defect — a Site's Cloudflare settings, and the
fingerprint this Chrome presents after an update, both move. Then, from
the web UI, open a bookmarked kagane series and confirm the cover renders. Once
a cover is stored it is served from Postgres forever after, so the browser being
asleep, unreachable, or mid-power-outage costs chapter freshness and nothing
visible.
Finally, take the VPS `free -m` reading again and compare it against the one
from the top of this section.
**Updating the browser** is independent of the API stack and has its own
runbook — `REDEPLOY.md` §8.
---
## Updating
Pull new code, then rebuild:
@@ -214,7 +486,14 @@ Pull new code, then rebuild:
docker compose -f docker-compose.yml -f docker-compose.prod.yml up -d --build
```
SQLite data persists in the named volume `bookmarks-data` across rebuilds.
Data persists in the named volume `postgres-data` across rebuilds. (If this
server predates the Postgres migration, the old SQLite volume `bookmarks-data`
is still on disk and deliberately undeclared in compose so `down -v` cannot take
it; see `REDEPLOY.md` §1 for when to remove it.)
The browser is a separate unit on a separate machine with its own update
command — §7. Nothing above touches it, and it needs no coordination: the API
picks up a restarted Chrome's new debugger UUID by itself.
---
@@ -225,9 +504,18 @@ SQLite data persists in the named volume `bookmarks-data` across rebuilds.
| No cert / TLS error at the domain | `TRAEFIK_ENTRYPOINT` or `TRAEFIK_CERTRESOLVER` name wrong; or DNS not resolving yet. Check `docker logs <traefik>`. |
| 404 from Traefik | Service not on the `proxy` network, or `BOOKMARK_API_HOST` mismatch. Confirm `docker network inspect proxy` lists `bookmark-api`. |
| `fetch` fails in the userscript, `curl` works | Origin missing from `ALLOWED_ORIGINS`, or mixed content (backend not HTTPS). |
| 401 with the right token | Trailing space/newline in `API_TOKEN`; regenerate and restart. |
| 401 with the right credential | The script's credential no longer matches the stored hash — most likely a rotation happened and the device was not reinstalled. Reinstall from the web UI. |
| 401 after rotation, even right after reinstalling | `TOKEN_KEY` changed between the rotation and the reinstall; credentials are derived from it, so changing it invalidates every credential. Keep it stable. |
| Panel button absent | URL didn't match an adapter, or user scripts disabled in Bromite. |
| `compose ... config` errors about `API_TOKEN` | Run compose from the dir with `.env`, or export the vars. |
| `compose ... config` errors about `TOKEN_KEY`, `OWNER_DISCORD_ID` or `POSTGRES_PASSWORD` | Run compose from the dir with `.env`, or export the vars. All three are required and none has a fallback. |
| `bookmark-api` restarts in a loop, `password authentication failed for user "bookmarks"` | `POSTGRES_PASSWORD` was changed after first boot; Postgres only applies it to an empty `postgres-data`. Restore the old value, or reset the role (`REDEPLOY.md` troubleshooting). |
| `bookmark-api` never logs `listening on :8080` | It is blocked on `postgres` passing `pg_isready`, or a migration failed. `docker compose -f docker-compose.yml -f docker-compose.prod.yml logs postgres`. |
| kagane rows never get a `latest_chapter`; log says `browser fetcher disabled` or nothing at all | `BROWSER_WS_URL` unset. Expected before §7 is done. |
| kagane polls all fail; log shows a 500 from `/json/version` | `BROWSER_WS_URL` names a MagicDNS hostname (or any name). Chrome's DevTools handler only accepts an IP or `localhost` — use the tailnet IP. |
| kagane polls fail with a connection error | Home machine off, off the tailnet, or the unit is down. `tailscale ping <machine>`, then `docker compose ps` in its `chrome/`. Costs freshness only; stored covers keep serving. |
| kagane cover is a placeholder for a newly bookmarked series | Its cover has never been fetched and the browser is unreachable. It fills in on the next successful poll of that series. |
| `compose` in `chrome/` errors `set BROWSER_BIND_ADDR to this machine's tailnet IP` | No `chrome/.env`, or the variable is empty. Deliberate — it has no default so an unset value cannot publish CDP to the LAN. |
| browser container restarts, or is OOM-killed | `docker inspect bookmark-browser --format '{{.RestartCount}} {{.State.OOMKilled}}'`. The 512 MiB cap is sized against a measured 645 MiB untuned peak; a real breach is a Chrome regression worth reading `docker logs` for, not a number to raise reflexively. |
Backend config reference and endpoint list: see `README.md`.
@@ -236,20 +524,25 @@ Backend config reference and endpoint list: see `README.md`.
## Installing / updating the userscript
The backend serves the script itself, so Violentmonkey can auto-update it.
Complements §4 above — that step points `API_BASE`/`API_TOKEN` at your
backend; this one points `@downloadURL`/`@updateURL` at the same place so
auto-updates come from it too.
Complements §4 above — the `@downloadURL`/`@updateURL` lines point at the
credential-bearing path, so auto-updates come from the same place as the
install.
Install once, on the phone (Cromite + Violentmonkey):
Install once, on the phone (Cromite + Violentmonkey): sign in to the web UI,
open the **Userscripts** panel, and open the install link for the library —
the script is served with your credential already inside it. Its
`@downloadURL`/`@updateURL` point at the same credential-bearing path for
updates:
```
https://bookmark-api.<your-domain>/u/<API_TOKEN>/manga-bookmark.user.js
https://bookmark-api.<your-domain>/u/<your credential>/manga-bookmark.user.js
```
Open that URL in Cromite; Violentmonkey offers to install it. The token is in
the path because Violentmonkey's update poll sends no `Authorization` header,
and the script embeds `API_TOKEN` in plain text — an open URL would leak it. A
wrong token answers 404.
Violentmonkey offers to install it. The credential is in the path because
Violentmonkey's update poll sends no `Authorization` header, and the script
embeds the credential in plain text — an open URL would leak it. A wrong
credential answers 404. The credential is derived from `TOKEN_KEY` and never
appears anywhere but this URL and the rendered script.
Updating, without a redeploy:
+20 -14
View File
@@ -8,47 +8,53 @@ web
## Users
Single user (self-hosted, no accounts, no multi-user planned). Reads manga on **asurascans.com** and **demonicscans.org** primarily via Bromite on mobile, also checks/updates from a desktop browser. The web UI is the cross-device view into progress captured by the userscript while reading.
Members of one private Discord guild, each with their own library. Accounts exist and are created by signing in — there is no signup form, no invite code and no approval step: any member of the configured guild becomes a Reader on their first Discord login. The person running the deployment is the owner, seeded at startup, and the only Reader with an administrative capability (revoking another Reader's sessions).
Reading happens on **asurascans.com**, **demonicscans.org**, **comix.to** and **kagane.to** for manga and **novelfull.com** and **lightnovelworld.net** for novels, primarily via Bromite on mobile, with checks and corrections from a desktop browser. The web UI is the cross-device view into progress the userscripts capture while reading.
## Product Purpose
Tracks read-progress ("last chapter read") per manga series across two otherwise-unrelated manga sites that each have their own separate `localStorage`. A Go backend unifies bookmarks into one store; the web UI is a password-gated browser view of that store for reviewing, favouriting, correcting, or removing bookmarks, and jumping back into a series to continue reading. A background poller also refreshes each series' latest-published-chapter so the list can flag "NEW" without the user visiting the site.
Tracks read-progress ("last chapter read") per series across sites that each have their own separate `localStorage`. A Go backend unifies bookmarks into one store; the web UI is a Discord-gated browser view of one Reader's own bookmarks, for reviewing, favouriting, correcting, shelving or removing them, and jumping back into a series to continue reading. A background poller refreshes each series' latest-published-chapter so the list can flag "NEW" without the Reader visiting the site.
## Positioning
Not a public reading tracker or social app — a private, self-hosted sync layer purpose-built for two specific scraped sites, with no server-side account system (single bearer token + one password-gated session).
Not a public reading tracker or social app — a private, self-hosted sync layer for one Discord community, purpose-built for a fixed set of scraped sites. Multi-Reader, not multi-tenant: libraries are isolated, but the deployment belongs to one group and its membership is the whole access model.
## Operating Context
- Primary reading device: Bromite (mobile Chromium), where a userscript captures progress automatically.
- Primary reading device: Bromite (mobile Chromium), where a userscript captures progress automatically. Each Reader installs their own copy, rendered with their own credential.
- Web UI is a secondary surface: checking list state, correcting a wrong chapter number, removing dead bookmarks, jumping to "continue reading."
- Manga cover art and titles come from the source sites' `og:image`/`og:title` — real content, not placeholders.
- List order is driven by `updated_at`, which moves only on real reading progress (not favouriting, not a newly detected chapter) — a UI constraint the redesign must not break.
- Cover art and titles come from the source sites' `og:image`/`og:title` — real content, not placeholders. They are facts about the series, so they are shared between Readers who track it; progress is not.
- List order is driven by `updated_at`, which moves only on real reading progress (not favouriting, not a newly detected chapter) — a UI constraint the design must not break.
## Capabilities and Constraints
- Two tabs: All / Favourites. Search-filter by title (client-side, `filter.js`).
- Card actions: continue (opens source site), toggle favourite, manual chapter override, delete (with confirm).
- "Continue reading" horizontal strip for recently-progressed series.
- htmx-driven partial updates (card re-render on favourite/chapter/delete), no client-side framework/build step — templates are Go `html/template`, `go:embed`-ed.
- Two libraries (manga, novels) with lifecycle tabs: All / Updated / Favourites / Archived / Finished. Search-filter by title (client-side, `filter.js`).
- Card actions: continue (opens source site), toggle favourite, manual chapter override, archive, finish, remove — each move out of the list confirm-gated.
- "Continue reading" horizontal strip for series with an unread chapter.
- A Reader with no bookmarks at all sees a deliberate empty library offering both userscript install links, not an error and not a blank page.
- Isolation is the load-bearing invariant: two Readers cannot see or change each other's bookmarks. A series both track is one shared row polled once, with independent progress on each side.
- The owner can revoke a specific Reader's sessions; nothing else in the UI differs by Reader.
- htmx-driven partial updates, no client-side framework or build step — templates are Go `html/template`, `go:embed`-ed.
- Mobile-first is a hard functional constraint (primary device is a phone), not just a starting breakpoint.
## Brand Commitments
- Name: **BookmarkManager**.
- **Dark-first is binding**: current dark-by-default / light-follows-system-preference behavior must be preserved as a design constraint, not just a starting default, because reading happens at night.
- **Dark-first is binding**: dark-by-default / light-follows-system-preference must be preserved as a design constraint, not just a starting default, because reading happens at night.
## Evidence on Hand
- Live templates/CSS at `backend/templates/*.html`, `backend/static/style.css` — current implemented UI, functional but not yet treated as an intentional design system.
- No logo, screenshots, or marketing copy exist; none should be fabricated.
- Live templates/CSS at `backend/internal/web/templates/*.html`, `backend/internal/web/static/style.css`, governed by the Cinder design system (`docs/design-system.md`).
- No logo beyond the wordmark, no screenshots, no marketing copy; none should be fabricated.
## Product Principles
- Dark-first, night-reading-optimized — never regress to a light-default or high-glare surface.
- Mobile is the primary target; desktop is an enhancement, not the design center.
- Progress data integrity over visual flourish: `updated_at`/list-ordering behavior is a correctness constraint the UI must respect, not decorate over.
- No accounts, no multi-tenant chrome — the whole product is for one reader.
- A leak between Readers fails silently and looks like working software — isolation is asserted from both directions, never inferred from counting one Reader's rows.
- No roles, no org chrome: the owner's Readers panel is one list with one button (revoke someone's sessions), not an admin console, and otherwise every Reader's view is the same.
- Prefer native platform affordances (system dark/light, native touch targets) over custom widgetry — this is a lean self-hosted tool, not a product to demo.
## Accessibility & Inclusion
+92 -28
View File
@@ -1,22 +1,35 @@
# Manga Bookmark
Track manga read-progress on **asurascans.com** (a.k.a. asuracomic.net),
Track manga read-progress on **asurascans.com**,
**demonicscans.org**, **comix.to**, and **kagane.to** from a phone (Bromite /
mobile Chromium), synced to a self-hosted Go backend so bookmarks unify across
all four sites and all devices.
Two parts:
- **`backend/`** — tiny Go (`net/http` + pure-Go SQLite) sync service. 4 routes,
- **`backend/`** — tiny Go (`net/http` + Postgres via pure-Go `pgx`) sync service. 4 routes,
static binary, distroless container.
- **`userscript/manga-bookmark.user.js`** — single Bromite-compatible userscript
(no `GM_*` APIs) that injects an on-page bookmark UI and syncs via `fetch()`.
```
Bromite userscript (isolated world, Shadow DOM UI, localStorage cache)
-- fetch() HTTPS --> reverse proxy (TLS + CORS) --> Go net/http --> SQLite (volume)
-- fetch() HTTPS --> reverse proxy (TLS + CORS) --> Go net/http --> Postgres (volume)
|
| CDP over the tailnet
v
headless Chrome, on-demand,
on a separate machine
(chrome/, ADR-0006)
```
Kagane and novelfull sit behind a Cloudflare JavaScript challenge no TLS
fingerprint clears, so the poller reaches those two through a real Chrome over
CDP. That browser is **not** part of the API stack: it is its own compose unit
on a second machine, spawned on the first connection and reaped when idle. The
API needs it only to discover new chapters and to fetch a kagane cover once —
covers are stored, so the library renders in full with the browser switched off.
---
## 1. Backend
@@ -25,21 +38,42 @@ Bromite userscript (isolated world, Shadow DOM UI, localStorage cache)
| Var | Default | Notes |
|-----|---------|-------|
| `API_TOKEN` | *(required)* | Bearer token shared with the userscript. |
| `TOKEN_KEY` | *(required)* | Secret every Reader's userscript credential is derived from (issue #24); only SHA-256 hashes of credentials are stored. |
| `OWNER_DISCORD_ID` | *(required)* | Discord user ID of the owner: seeded as the first Reader, owns every pre-registration bookmark, and is the only Reader who can revoke another's sessions. |
| `ALLOWED_ORIGINS` | Asura + Demonic + Comix + Kagane origins | Comma-separated CORS allowlist. |
| `DB_PATH` | `/data/bookmarks.db` | SQLite file location. |
| `DATABASE_URL` | *(required)* | Postgres connection URL, e.g. `postgres://bookmarks:…@postgres:5432/bookmarks?sslmode=disable`. Compose builds it from `POSTGRES_PASSWORD`. |
| `COVER_DIR` | *(required)* | Filesystem volume for immutable, content-addressed Cover bytes. Compose builds the image and mounts `cover-data` at this path; standalone runs may choose another writable durable path. |
| `PORT` | `8080` | Plain HTTP; TLS terminated by the proxy. |
| `BROWSER_WS_URL` | `ws://172.28.0.10:9222` | Headless-shell CDP endpoint used to poll Kagane past its JS challenge. Must be an IP or `localhost` — Chrome's DevTools handler 500s any other Host header. |
| `BROWSER_WS_URL` | empty | CDP endpoint of the remote browser (`ws://<tailnet IP>:9222`), used to poll Kagane/Novelfull past their JS challenge and to fetch uncached Kagane covers. Must be an IP or `localhost` — Chrome's DevTools handler 500s any other Host header, MagicDNS names included. Unset disables both; stored covers still serve. |
| `DISCORD_CLIENT_ID` | *(required)* | Discord application credentials for the browser sign-in (ADR-0002). |
| `DISCORD_CLIENT_SECRET` | *(required)* | As above. Never logged, never echoed in an error. |
| `DISCORD_GUILD_ID` | *(required)* | The one guild whose membership gates sign-in, checked at login only. Membership *is* registration: any member becomes a Reader on first login. |
| `DISCORD_REDIRECT_URI` | *(required)* | Exact callback URL; Discord matches it verbatim against the registered redirect. |
| `DISCORD_REQUIRED_ROLE` | empty | Role snowflake a member must additionally hold. Empty means guild membership alone suffices. |
| `DISCORD_API_BASE` | `https://discord.com/api/v10` | Test seam — tests point it at a local stub so the real token exchange runs. |
| `USERSCRIPT_PATH` | `/userscript/manga-bookmark.user.js` | Bindmounted file served at `/u/{token}/manga-bookmark.user.js`. |
| `NOVEL_USERSCRIPT_PATH` | `/userscript/novel-bookmark.user.js` | Same, for the novel library. |
| `LATEST_CHAPTER_POLL_ENABLED` | `1` | `0` turns the poller off entirely. Pace is per Site in the registry — one Poll Lane per Site, each with its own rest and gap (issue #100) — so no other knobs exist. |
Compose reads a few more from the same `.env` that the backend never sees:
`POSTGRES_PASSWORD` (required — `DATABASE_URL` is built from it, and Postgres
only applies it while `postgres-data` is empty), `BOOKMARK_API_HOST` and
`BOOKMARK_WEB_HOST` (required by the prod override), and the optional
`PROXY_NETWORK` / `TRAEFIK_ENTRYPOINT` / `TRAEFIK_CERTRESOLVER`. The browser
unit has its own `chrome/.env` on its own machine — `BROWSER_BIND_ADDR`
(required, the tailnet IP the CDP port is published on) and the optional
`BROWSER_TZ`. Full commentary is in `.env.example` and `chrome/.env.example`;
deployment order is `DEPLOY.md`.
### Endpoints
| Method | Path | Auth | Description |
|--------|------|------|-------------|
| `GET` | `/bookmarks` | Bearer | All bookmarks (single-user). |
| `GET` | `/bookmarks` | Bearer | All bookmarks of the acting Reader. |
| `PUT` | `/bookmarks/{key}` | Bearer | Upsert one series; returns the row as stored. |
| `DELETE` | `/bookmarks/{key}` | Bearer | Remove one. |
| `GET` | `/healthz` | none | `200 ok`. |
| `GET` | `/u/{token}/manga-bookmark.user.js` | token in path | Serves the userscript with an mtime-derived `@version`. |
| `GET` | `/u/{token}/manga-bookmark.user.js` | credential in path | Serves the userscript with the requesting Reader's credential substituted in and an mtime-derived `@version`. |
`key` is `<site>:<series_id>` — e.g. `asura:trash-of-the-counts-family-f886a8af`,
`demonic:Infinite-Level-Up-in-Murim`, `comix:12345`, or
@@ -60,19 +94,45 @@ go test ./... # unit + handler tests
CGO_ENABLED=0 go build # static binary
```
**`go test ./...` requires Docker.** The store talks to a real Postgres, so
each test package starts a throwaway `postgres:17-alpine` container and gives
every test its own database inside it (`internal/pgtest`). Nothing is stubbed
and nothing reaches the network beyond the local Docker daemon.
### Run the stack
```bash
cp .env.example .env
# edit .env: set API_TOKEN (openssl rand -hex 32)
# edit .env: set TOKEN_KEY (openssl rand -hex 32) and
# POSTGRES_PASSWORD (openssl rand -hex 24)
docker compose up -d --build # binds 127.0.0.1:8080
```
That brings up two services — the API and Postgres. The browser is deliberately
not one of them; without `BROWSER_WS_URL` the poller logs and skips kagane and
novelfull, and everything else works. To run one locally, publish it on the
Docker bridge gateway so the API container can name it by IP:
```bash
cd chrome
echo 'BROWSER_BIND_ADDR=172.17.0.1' > .env
docker compose up -d --build
# then in the repo's own .env: BROWSER_WS_URL=ws://172.17.0.1:9222
```
Bind it to `127.0.0.1` instead if you only want to drive it from the host, e.g.
`SMOKE_BROWSER_WS_URL=ws://127.0.0.1:9222 go test -run TestSmokeKagane ./internal/latest`.
In production that address is the home machine's tailnet IP and nothing else —
see `DEPLOY.md` §7 and ADR-0006.
Smoke test:
```bash
TOKEN=$(grep '^API_TOKEN=' .env | cut -d= -f2)
# The credential is per Reader and derived, so there is no token in .env to
# grep. Take yours from the Userscripts panel's install link after signing in,
# or read it out of an installed script's API_TOKEN constant.
TOKEN=<your Reader credential>
curl -s localhost:8080/healthz # ok
curl -s localhost:8080/bookmarks # 401
curl -s -H "Authorization: Bearer $TOKEN" localhost:8080/bookmarks # []
@@ -107,17 +167,18 @@ CORS headers.
## 2. Userscript
### Configure
### Install
Edit the config block at the top of `userscript/manga-bookmark.user.js`:
Sign in to the web UI and open the **Userscripts** panel: it offers one
install link per library. Each link serves a script rendered with your own
credential already inside it — you never see, type or copy a credential. The
served script carries `@downloadURL`/`@updateURL` pointing at its
credential-bearing path, so Violentmonkey keeps auto-updating it.
```js
const API_BASE = "https://bookmark-api.<domain>"; // no trailing slash
const API_TOKEN = "<same token as backend>";
```
The token lives in the userscript's **isolated world** — the manga sites' own
JS cannot read it.
The bindmounted files carry `__API_TOKEN__` placeholders; the backend
substitutes the requesting Reader's credential at serve time, so no real
credential is ever committed. Rotating the credential (same panel) invalidates
every installed copy immediately — reinstall on all devices.
### Install on Bromite (mobile)
@@ -125,8 +186,8 @@ Bromite runs Chromium's native userscript engine (no Tampermonkey needed):
1. Bromite → **Settings → User scripts** → enable user scripts (allow the
permission prompt).
2. Save the configured `manga-bookmark.user.js` to the device (or open its raw
URL). Bromite detects the `.user.js` and offers to install it.
2. Open the install link from the web UI — Bromite detects the `.user.js` and
offers to install it.
3. Confirm the install; the `@match` list covers both sites.
4. Open a series on either site — a 📑 button appears bottom-right.
@@ -195,8 +256,10 @@ an API.
## Adapter reference (verified live 2026-07-24)
The site adapters key everything off URL regex, with `title`/`cover` from
`og:title` / `og:image`. Confirmed against live pages via Playwright:
The site adapters key everything off URL regex, with `title` from `og:title`
(or the page's own heading where a site ships none). No adapter reads a cover:
the backend acquires, stores and serves every Cover from its own origin
(ADR-0007). Confirmed against live pages via Playwright:
| Site | Series URL | Chapter URL | `series_id` |
|------|-----------|-------------|-------------|
@@ -206,11 +269,12 @@ The site adapters key everything off URL regex, with `title`/`cover` from
| **Kagane** (`kagane.to`) | `/series/<uuid>` | `/series/<uuid>/reader/<bookUuid>` | `<uuid>` |
Notes:
- **`asuracomic.net` deep links are dead (re-checked 2026-07-25).** They 301 to
the `asurascans.com` **root**, discarding the path, at the edge — before the
userscript gets a document — so nothing client-side can rescue them. Reach
series through `asurascans.com`. The host stays matched in case the redirect
starts preserving paths again.
- **`asuracomic.net` is no longer matched (deep links dead, re-checked
2026-07-25).** They 301 to the `asurascans.com` **root**, discarding the path,
at the edge — before the userscript gets a document — so nothing client-side
can rescue them. The backend rejects stored addresses on that host too, since
the poller pins each Site to one hostname. Reach series through
`asurascans.com`.
- Asura `og:title` carries a `Chapter N - Read Online \| Asura Scans` suffix that
the adapter strips; Demonic chapter `og:title` is `<Title> Chapter N`.
- Demonic's `<slug>` is identical on `/manga/…` and the canonical `/title/…`
+218 -92
View File
@@ -9,16 +9,16 @@ Whole thing is ~5 minutes, most of it waiting on `docker build`. Order matters:
**back up before you pull.** A backup taken after a bad migration is a backup of
the damage.
Paths below assume the checkout is at `/opt/bookmarkmanager`; substitute your own. The
one absolute rule about paths: **backups live in `../bookmarkmanager-backups/`**, a
sibling of the project directory (`/opt/bookmarkmanager-backups`), never inside it. It
Paths below assume the checkout is at `~/mangaBookmark`, which is where it lives
on this deployment; substitute your own. The one absolute rule about paths:
**backups live in a `-backups` sibling of the checkout**, never inside it. It
sits outside the repo so `git pull`, `git clean -fd` and a bad `rm -rf` inside
the checkout cannot take the backups with them.
```
/opt/
├── bookmarkmanager/ <- the checkout (this repo)
└── bookmarkmanager-backups/ <- bookmarks-YYYYmmdd-HHMMSS.db
~/
├── mangaBookmark/ <- the checkout (this repo)
└── mangaBookmark-backups/ <- bookmarks-YYYYmmdd-HHMMSS.dump
```
---
@@ -26,7 +26,7 @@ the checkout cannot take the backups with them.
## 0. Preflight
```bash
cd /opt/bookmarkmanager
cd ~/mangaBookmark
# Both -f flags, every time. The prod override is not standalone.
COMPOSE="docker compose -f docker-compose.yml -f docker-compose.prod.yml"
@@ -44,87 +44,117 @@ dirty tree fails halfway and leaves you in a worse spot than either.
Create the backup directory once, and make sure it is a sibling, not a child:
```bash
mkdir -p ../bookmarkmanager-backups
BACKUP_DIR="$(cd .. && pwd)/bookmarkmanager-backups" # absolute — Docker needs it
echo "$BACKUP_DIR" # -> /opt/bookmarkmanager-backups
BACKUP_DIR="$(cd .. && pwd)/$(basename "$PWD")-backups" # absolute — Docker needs it
mkdir -p "$BACKUP_DIR"
echo "$BACKUP_DIR" # -> /home/sulthan/mangaBookmark-backups
```
---
## 1. Back up the database
The database is a single SQLite file in the named Docker volume, at
`/data/bookmarks.db` inside the container. Find the volume's real name — Compose
prefixes it with the project directory:
The database is Postgres, running as the `postgres` service on the named volume
`postgres-data`. It has **no published port** — nothing outside the internal `db`
network can reach it — so every command below goes in through the container:
```bash
docker volume ls --filter name=bookmarks-data
# -> local bookmarkmanager_bookmarks-data
VOL=$(docker volume ls --filter name=bookmarks-data -q | head -1)
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks -c '\dt'
# -> bookmarks, covers, readers, schema_migrations, series, sessions
```
### Preferred: hot backup, no downtime
The `covers` table is metadata only after the filesystem cutover: bytes live in
the separate `cover-data` volume. Back that volume up with the database dump;
restoring only Postgres leaves stored Cover addresses without files.
The store runs in **WAL mode**, so recent writes may still be sitting in
`bookmarks.db-wal`. Copying `bookmarks.db` alone while the container runs can
therefore silently drop the newest bookmarks. `VACUUM INTO` folds the WAL in and
writes one consistent file, safe to run against a live database:
Inside the container that connects over the local socket as the `bookmarks`
superuser, so no password is needed anywhere in this section. `-T` is not
optional: without it Compose allocates a TTY, which rewrites `\n` to `\r\n` and
silently corrupts any binary stream flowing back out — see the dump below.
### Preferred: hot dump, no downtime
`pg_dump` runs in a single repeatable-read transaction, so it writes one
point-in-time-consistent snapshot while the API keeps serving. No stopping, no
WAL to worry about — that is the server's problem, not yours.
```bash
STAMP=$(date -u +%Y%m%d-%H%M%S) # UTC, sorts chronologically as text
docker run --rm \
-v "$VOL":/data \
-v "$BACKUP_DIR":/backup \
alpine sh -c "apk add -q sqlite &&
sqlite3 /data/bookmarks.db \"VACUUM INTO '/backup/bookmarks-$STAMP.db'\""
$COMPOSE exec -T postgres pg_dump -U bookmarks -d bookmarks -Fc \
> "$BACKUP_DIR/bookmarks-$STAMP.dump"
ls -lh "$BACKUP_DIR"/bookmarks-$STAMP.db
ls -lh "$BACKUP_DIR"/bookmarks-$STAMP.dump
```
`$STAMP` is the "time in the name" — `bookmarks-20260730-014233.db`. UTC, so the
files sort in real order and never collide across a DST shift.
`-Fc` is the custom archive format rather than plain SQL: it is compressed, and
`pg_restore` can inspect and replay it selectively — list its table of contents,
restore one table, restore schema without data, reorder. A plain `.sql` dump can
only be piped into `psql` whole, and gives you no way to check what is in it
short of reading it.
Note the source volume is mounted **read-write**, which looks wrong for a backup
and is not. Opening a WAL database requires creating the `-shm` shared-memory
file; with `:ro` the command fails with `unable to open database file` and no
backup is produced. `VACUUM INTO` never writes to the source itself.
`$STAMP` is the "time in the name" — `bookmarks-20260730-014233.dump`. UTC, so
the files sort in real order and never collide across a DST shift.
Verify it before you trust it. An unreadable backup is worse than none, because
you will act as though you have one:
```bash
docker run --rm -v "$BACKUP_DIR":/backup alpine sh -c "apk add -q sqlite &&
sqlite3 /backup/bookmarks-$STAMP.db 'PRAGMA integrity_check;' &&
sqlite3 /backup/bookmarks-$STAMP.db 'SELECT count(*) FROM bookmarks;'"
# -> ok
# 1. The dump parses and contains the tables. Uses the same image compose
# already pulls, so nothing new to install.
docker run --rm -v "$BACKUP_DIR":/backup postgres:17-alpine \
pg_restore --list "/backup/bookmarks-$STAMP.dump" | grep 'TABLE DATA'
# -> 1234; 0 0 TABLE DATA public bookmarks bookmarks
# -> 1235; 0 0 TABLE DATA public covers bookmarks
# -> 1236; 0 0 TABLE DATA public readers bookmarks
# -> 1237; 0 0 TABLE DATA public schema_migrations bookmarks
# -> 1238; 0 0 TABLE DATA public series bookmarks
# -> 1239; 0 0 TABLE DATA public sessions bookmarks
# 2. Sanity-check the live row count you just captured.
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks \
-c 'select count(*) from bookmarks'
# -> 37
```
The count should match what the web UI shows. Zero rows on a server you know has
bookmarks means you backed up the wrong volume.
A custom-format archive stores row counts nowhere, so step 1 proves the file is
a readable archive with the right tables in it, not that the rows are there;
step 2 is the number those rows should be. It should match what the web UI
shows. Zero on a server you know has bookmarks means the API and your `psql`
are looking at different databases — check `DATABASE_URL`.
### Fallback: cold copy (no network for `apk add sqlite`)
### Fallback: cold volume archive
Stop the service first, then copy the database **and its sidecars** — the `-wal`
is not optional, it is where the newest writes are:
Use this when you want the whole data directory rather than a logical dump — a
like-for-like restore of the same Postgres major version onto the same host.
**The stack must be stopped first.** A running Postgres has dirty pages in
shared buffers and WAL that has not been replayed into the data files, and `tar`
walks the directory over several seconds while the server keeps writing to it.
The archive you get is torn: files from different instants, possibly a
half-written page. It may restore, start, and be quietly wrong. Online
filesystem-level backup is `pg_basebackup`'s job, not `tar`'s; with the
container stopped the shutdown checkpoint has already flushed everything and a
plain archive of the volume is consistent.
```bash
# Derived exactly, not with a `--filter name=` substring match plus `head -1`:
# that quietly picks the first of however many volumes happen to contain the
# string, and archiving the wrong data directory is not a visible failure.
VOL="$(basename "$PWD" | tr '[:upper:]' '[:lower:]')_postgres-data"
docker volume inspect "$VOL" >/dev/null && echo "$VOL" # -> mangabookmark_postgres-data
$COMPOSE stop
docker run --rm -v "$VOL":/data:ro -v "$BACKUP_DIR":/backup alpine sh -c "
cp /data/bookmarks.db /backup/bookmarks-$STAMP.db
[ -f /data/bookmarks.db-wal ] && cp /data/bookmarks.db-wal /backup/bookmarks-$STAMP.db-wal
[ -f /data/bookmarks.db-shm ] && cp /data/bookmarks.db-shm /backup/bookmarks-$STAMP.db-shm
ls -1 /backup"
docker run --rm -v "$VOL":/from:ro -v "$BACKUP_DIR":/to alpine \
tar czf "/to/postgres-data-$STAMP.tgz" -C /from .
$COMPOSE start
ls -lh "$BACKUP_DIR"/postgres-data-$STAMP.tgz
```
Costs ~10 seconds of downtime. A clean shutdown usually checkpoints the WAL away,
so seeing only the `.db` file is normal and fine — the `[ -f ]` guards exist for
the case where it did not. Restoring this variant means putting whichever files
you got back together, under their original names.
Read-only is safe here precisely because nothing opens the database: it is a file
copy, not a SQLite connection.
Costs ~15 seconds of downtime. Read-only on the source is safe here precisely
because nothing is running against it. Restoring this variant means untarring it
back into an *empty* `postgres-data` volume with the stack down — it is a whole
data directory, not a file you can drop next to the live one, and it will only
start under `postgres:17`.
### Retention
@@ -132,7 +162,21 @@ Keep a month, drop the rest — a bookmark database this small compresses the
decision to "disk is free, but not infinite":
```bash
ls -1t "$BACKUP_DIR"/bookmarks-*.db | tail -n +31 | xargs -r rm -v
ls -1t "$BACKUP_DIR"/bookmarks-*.dump | tail -n +31 | xargs -r rm -v
```
### A note on the old `bookmarks-data` volume
`bookmarks-data` is the **pre-migration SQLite volume**. It is deliberately not
declared in `docker-compose.yml` any more, which is what keeps `docker compose
down -v` from taking it with the rest of the stack. It is not the live database
and nothing reads it — the one-way move out of it is `CUTOVER.md`. Once the
Postgres data has been trusted for a while, remove it by hand — nothing else will.
Its full name is `<compose project>_bookmarks-data`, and the project name is the
lowercased directory name of the checkout:
```bash
docker volume rm "$(basename "$PWD" | tr '[:upper:]' '[:lower:]')_bookmarks-data"
```
---
@@ -171,11 +215,15 @@ rebuilt. The one exception is `userscript/manga-bookmark.user.js`, which is
bindmounted read-only and read fresh per request.
```bash
$COMPOSE ps # Up, and recently (re)created
$COMPOSE ps # bookmark-api Up; postgres Up (healthy)
docker logs bookmark-api --tail 20 # -> "listening on :8080 ..."
```
Nothing in the log about the database or the poller failing. The image is tagged
Nothing in the log about the database, the migrations or the poller failing.
`bookmark-api` waits on `postgres` reporting healthy before it starts and the
binary applies any pending migration before it listens, so an API that never
says "listening" is usually the database, not the code — `$COMPOSE logs
postgres` first. The image is tagged
`bookmarkmanager-backend:latest`, so the previous image is still on disk untagged —
that is what makes the rollback in §6 quick.
@@ -188,7 +236,10 @@ Same four API checks as `DEPLOY.md` §3, plus the web UI. Set the host names onc
```bash
API=https://bookmark-api.violetcrown.my.id
WEB=https://bookmark.violetcrown.my.id
TOKEN=$(grep -E '^API_TOKEN=' .env | cut -d= -f2)
# Your own Reader credential - derived, never stored in .env. Take it from the
# Userscripts panel's install link after signing in, or from an installed
# script's API_TOKEN constant.
TOKEN=<your Reader credential>
curl -s $API/healthz # -> ok
curl -s -o /dev/null -w '%{http_code}\n' $API/bookmarks # -> 401
@@ -199,9 +250,16 @@ curl -s -i -X OPTIONS -H 'Origin: https://asurascans.com' \
$API/bookmarks/x | grep -i access-control # -> allow-origin echoed
```
`[]` from the third call is the alarm that matters: the volume is not attached
and you are looking at an empty database. Stop and check `$COMPOSE config
--volumes` before touching anything else.
`[]` from the third call is the alarm that matters: you are talking to an empty
database, which means the API found a *different* Postgres than the one holding
your data — a renamed project directory, a fresh `postgres-data`, or a
`DATABASE_URL` override in `.env` pointing elsewhere. Stop and check, before
touching anything else:
```bash
$COMPOSE config --volumes # -> postgres-data
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks -c 'select count(*) from bookmarks'
```
Web UI and its assets:
@@ -267,33 +325,41 @@ git checkout <previous-hash>
$COMPOSE up -d --build
```
**Database damaged** — restore the backup from §1. Stop first: the running
process holds the WAL, and dropping a file under a live SQLite connection
corrupts what you were trying to save.
**Database damaged** — restore the dump from §1. Stop **only the API**, not the
whole stack: `pg_restore` needs the server up to restore into, and it needs
`bookmark-api`'s connection pool gone, because `--clean` cannot drop a table
other sessions are holding open.
```bash
$COMPOSE stop
$COMPOSE stop bookmark-api
docker run --rm -v "$VOL":/data -v "$BACKUP_DIR":/backup alpine sh -c '
rm -f /data/bookmarks.db /data/bookmarks.db-wal /data/bookmarks.db-shm &&
cp /backup/bookmarks-<STAMP>.db /data/bookmarks.db &&
chown 65532:65532 /data/bookmarks.db &&
ls -l /data'
$COMPOSE exec -T postgres pg_restore -U bookmarks -d bookmarks --clean --if-exists \
< "$BACKUP_DIR/bookmarks-<STAMP>.dump"
$COMPOSE start
$COMPOSE start bookmark-api
docker logs bookmark-api --tail 20
curl -s -H "Authorization: Bearer $TOKEN" $API/bookmarks | head -c 200
```
Two steps here are easy to skip and both bite:
Three things here are easy to skip and all three bite:
- **Delete the stale `-wal` and `-shm`.** Leaving them beside a restored database
mixes two different histories; SQLite will either refuse to open it or quietly
reapply writes you meant to discard.
- **`chown 65532:65532`.** The image is `distroless/static:nonroot` and runs as
that uid, while the helper container above writes as root. A root-owned
database opens read-only-ish: reads work, so `/bookmarks` looks fine, and then
every write fails. That is the worst possible failure mode — it looks restored.
- **`--clean --if-exists`.** Without `--clean` the dump's rows land *on top of*
what is already there and you get primary-key collisions half way through, a
partially restored database, and a non-zero exit you may not notice.
`--if-exists` only suppresses the "does not exist" noise when the target is
already empty; it is not the part doing the work.
- **`-T` again.** Feeding a custom-format archive into a TTY-allocated `exec`
corrupts it in flight and `pg_restore` fails with a garbled-header error on a
file that is perfectly fine on disk.
- **Stop the API, not Postgres.** `$COMPOSE stop` (everything) leaves you with
nothing to restore into; leaving `bookmark-api` running leaves connections
that block the drops *and* lets the poller write into a half-restored table.
No ownership fixing is needed any more — the Postgres image owns `postgres-data`
itself and `pg_restore` writes through the server, not the filesystem.
`schema_migrations` is inside the dump, so the database comes back at whatever
schema version the backup was taken at; the migration runner applies anything
newer the next time `bookmark-api` starts.
---
@@ -305,21 +371,73 @@ For a routine redeploy where nothing needs deciding:
cd /opt/bookmarkmanager
COMPOSE="docker compose -f docker-compose.yml -f docker-compose.prod.yml"
BACKUP_DIR="$(cd .. && pwd)/bookmarkmanager-backups"; mkdir -p "$BACKUP_DIR"
VOL=$(docker volume ls --filter name=bookmarks-data -q | head -1)
STAMP=$(date -u +%Y%m%d-%H%M%S)
docker run --rm -v "$VOL":/data -v "$BACKUP_DIR":/backup alpine sh -c \
"apk add -q sqlite && sqlite3 /data/bookmarks.db \"VACUUM INTO '/backup/bookmarks-$STAMP.db'\" &&
sqlite3 /backup/bookmarks-$STAMP.db 'PRAGMA integrity_check;'" &&
$COMPOSE exec -T postgres pg_dump -U bookmarks -d bookmarks -Fc \
> "$BACKUP_DIR/bookmarks-$STAMP.dump" &&
docker run --rm -v "$BACKUP_DIR":/backup postgres:17-alpine \
pg_restore --list "/backup/bookmarks-$STAMP.dump" > /dev/null &&
git pull --ff-only &&
$COMPOSE up -d --build &&
sleep 5 &&
curl -sf https://bookmark-api.violetcrown.my.id/healthz && echo " deploy ok"
```
The `&&` chain is deliberate: if the backup or its integrity check fails,
nothing is pulled and nothing is rebuilt. Then still do §5 by hand — no shell
command can tell you the panel works on the phone.
The `&&` chain is deliberate: if the dump or its `pg_restore --list` check
fails, nothing is pulled and nothing is rebuilt. A failed dump still leaves a
short or empty `.dump` behind — the shell creates the file before `pg_dump`
runs — so delete it rather than letting it sit in the backup directory looking
like a backup. Then still do §5 by hand — no shell command can tell you the
panel works on the phone.
---
## 8. The browser unit (separate machine, separate cadence)
Everything above is the API stack on the VPS. The headless browser is its own
compose unit on the home machine (ADR-0006, `DEPLOY.md` §7) and is redeployed
on its own schedule — it holds no data you can lose, so there is nothing to
back up and no ordering constraint against the API.
```bash
cd ~/mangaBookmark/chrome
git pull --ff-only
docker compose up -d --build
```
Then confirm it answers, and that a stopped-and-restarted Chrome is invisible
to the API:
```bash
curl -s -m 15 http://$(tailscale ip -4):9222/json/version | head -c 120
# -> {"Browser":"Chrome/1xx...","webSocketDebuggerUrl":"ws://...<new uuid>"}
```
The first call takes a few seconds: Chrome is not running until something
connects, and it is reaped again after five idle minutes. The debugger UUID
changes on every start and the API does not care — chromedp re-runs
`/json/version` discovery per fetch, which is exactly why `chromedp.NoModifyURL`
must never be added to `browser.go`.
**Rebuild is the Chrome upgrade path.** The image installs
`google-chrome-stable` unpinned on purpose: a stale browser is what Cloudflare
turns away, and the pinned Chrome 124 in `zenika/alpine-chrome` is the worked
example. The `chrome-profile` volume survives `--build`, so clearance cookies
are reused rather than re-solved.
Two things worth a glance after several days, both from the acceptance criteria
of the move:
```bash
docker inspect bookmark-browser --format '{{.RestartCount}} {{.State.OOMKilled}}'
# -> 0 false
free -m # the Gitea runner should still have its headroom
```
Nothing here needs doing during an API redeploy. The API stack does not
`depends_on` the browser, and an unreachable one degrades exactly as an unset
`BROWSER_WS_URL`: plain-TLS libraries unaffected, kagane and comix logged
and skipped, novelfull attempted over plain TLS, stored covers still served.
---
@@ -327,16 +445,24 @@ command can tell you the panel works on the phone.
| Symptom | Cause / fix |
|---|---|
| `/bookmarks` returns `[]` after redeploy | Volume not attached — check `$COMPOSE config --volumes` and that you passed both `-f` files. Do **not** re-bookmark; the data is still in the volume. |
| `/bookmarks` returns `[]` after redeploy | You are on an empty Postgres. Check `$COMPOSE config --volumes` lists `postgres-data`, that you passed both `-f` files, and that `.env` has no stray `DATABASE_URL` override. Do **not** re-bookmark; the data is still in the volume. |
| UI looks like plain Georgia / system sans | `static/fonts/` missing from the image, or the browser cached an old `style.css`. `/static/*` is served `max-age=3600`, so hard-reload or wait an hour. |
| CSS or template change did not appear | You restarted without `--build`. Assets are `//go:embed`ed. |
| Font answers `application/octet-stream` | Old binary — the `.woff2` MIME registration is in `web.go`. Rebuild. |
| Everyone logged out of the web UI | `API_TOKEN` or `WEB_PASSWORD` changed; sessions are derived from both. Expected, just log in again. |
| Everyone logged out of the web UI | The `sessions` table was wiped; sessions are database rows, not signed cookies. Expected after a deliberate revoke. |
| `compose` errors about `BOOKMARK_WEB_HOST` | Run from the directory holding `.env`. Both host vars are required even when the web UI is unused. |
| Userscript did not update on the phone | Violentmonkey polls on its own schedule; force a check. `@version` comes from the file's mtime, so confirm the pull actually touched it. |
| `apk add sqlite` fails (no network) | Use the cold-copy fallback in §1 — and copy `bookmarks.db-wal` too. |
| Reads work but every write fails after a restore | Restored file is root-owned; the container is uid 65532. `chown 65532:65532` it (§6). |
| Backup command: `unable to open database file` | Source volume mounted `:ro`. WAL needs to create `-shm`; mount it read-write (§1). |
| `bookmark-api` crash-loops, log says `password authentication failed for user "bookmarks"` | `POSTGRES_PASSWORD` in `.env` no longer matches the one burned into `postgres-data` at first init — Postgres reads that variable only when initialising an empty volume. Put the old value back, or reset the role: `$COMPOSE exec postgres psql -U bookmarks -d bookmarks -c '\password bookmarks'` (prompts, so nothing lands in shell history) and then match `.env` to it. |
| `compose` errors `set POSTGRES_PASSWORD in .env` | Unset. Compose builds the backend's `DATABASE_URL` out of it, so it is required even though you never write that URL yourself. Run from the directory holding `.env`. |
| `postgres` never leaves `starting`; `bookmark-api` never starts either | The healthcheck (`pg_isready`) is failing and `bookmark-api` waits on it. `$COMPOSE logs postgres` — usually `postgres-data` was initialised by a different major version ("database files are incompatible with server"), or the disk is full. |
| `pg_restore`: `cannot drop … other objects depend on it` / `being accessed by other users` | Live connections block `--clean`. `$COMPOSE stop bookmark-api` first (§6). If they persist: `$COMPOSE exec -T postgres psql -U bookmarks -d postgres -c "select pg_terminate_backend(pid) from pg_stat_activity where datname='bookmarks' and pid <> pg_backend_pid()"`. |
| Dump is 0 bytes, or `pg_restore`: `did not find magic string in file header` | You ran `exec` without `-T`. The allocated TTY rewrites newlines in the binary stream and corrupts the archive in flight (§1). |
| `git pull`: `could not read Username for 'https://…'` | The checkout's remote is the HTTPS clone URL and the server has no credential helper, so the pull prompts into a closed stdin. Switch it to SSH once — `git remote set-url origin ssh://git@gitea.violetcrown.my.id:2222/sulthan/mangaBookmark.git`. Gitea's SSH listens on **2222**, not 22; port 22 is the host's own sshd and answers `Permission denied (publickey)` no matter which key is registered. |
| kagane rows stopped updating after a redeploy | Check `BROWSER_WS_URL` survived the `.env` edit and still names the home machine's tailnet **IP**. A hostname 500s at `/json/version`; an empty value disables the browser silently. Plain-TLS sites keep working either way, which is why this is easy to miss. |
| kagane covers went blank in the web UI | Covers use the `cover-data` volume now. Restore/check that volume alongside Postgres; rows in `covers` are metadata only. If the database has rows but files are missing, the next browser-backed request refetches them; without a browser it remains a 404. |
| Browser unit will not start: `set BROWSER_BIND_ADDR to this machine's tailnet IP` | `chrome/.env` is missing or the variable is empty. It has no default on purpose — an unset value must fail the deploy rather than publish an unauthenticated CDP port to the LAN. |
| `bookmark-browser` shows `OOMKilled true` | The cap did its job. Read `docker logs bookmark-browser` before raising it — the sizing and what the cap protects are in ADR-0006. |
Full first-time setup: `DEPLOY.md`. Config reference and endpoints: `README.md`.
Full first-time setup: `DEPLOY.md`. The one-off SQLite→Postgres move:
`CUTOVER.md`. Config reference and endpoints: `README.md`.
UI conventions: `docs/design-system.md`.
+189 -36
View File
@@ -1,8 +1,8 @@
Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGENTS.md` for the project-wide architecture diagram, hard constraints, and design system.
- **Backend** (`backend/`): stdlib `net/http` (handful routes, no framework) + `modernc.org/sqlite` (pure Go, `CGO_ENABLED=0` -> static binary -> distroless/scratch image). Reverse proxy terminates TLS; Go service listens plain `:8080`.
- **Backend** (`backend/`): stdlib `net/http` (handful routes, no framework) + Postgres over `jackc/pgx/v5` (pure Go, `CGO_ENABLED=0` -> static binary -> distroless/scratch image). Reverse proxy terminates TLS; Go service listens plain `:8080`.
Single binary, split into packages under `backend/internal/`: `store`
(Bookmark type, SQLite persistence, migrations), `latest` (background
(Bookmark type, Postgres persistence, migration runner), `latest` (background
poller, site parsers, TLS fetcher), `session` (cookie signing, login
rate limiter), `httpmw` (Auth/Gzip/CORS middleware), `api` (JSON
bookmark handlers), `userscript` (userscript-serving handler), `web`
@@ -11,16 +11,52 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
packages together into `newRouter`. Root-level `*_test.go` hold
integration tests that exercise the full router; unit tests for a
package live beside it under `internal/`.
- **Single-user store.** One `bookmarks` table keyed `<site>:<series_id>` (`asura`|`demonic`|`comix`|`kagane`). Sync **last-write-wins**. Schema and endpoint list in plan.
- **Schema is migration-owned.** `internal/store/migrations/*.sql` is
`go:embed`-ed and applied on every start by `store.migrate`: one numbered
file per change, one transaction each, versions recorded in
`schema_migrations`. Files are **append-only** — editing an applied one
changes nothing on a database that already ran it. No column probing, no
data-fixup migrations: both were SQLite-era machinery and are gone.
- **Tests need Docker.** `internal/pgtest` starts one `postgres:17-alpine`
container per test binary (`TestMain` -> `pgtest.Main`) and hands each test
its own database (`pgtest.URL(t)`). A package whose tests touch the store
must have that `TestMain`.
- **Reader-owned store, four tables.** `readers` is keyed by Discord user ID
and carries the SHA-256 of the Reader's userscript credential plus a
`token_epoch` (issue #24). Credentials are derived, never stored: `token.Token(TOKEN_KEY, discord_id, epoch)` (HMAC, `internal/token`), and only its SHA-256 sits in `readers.token_sha256`, so install URLs can be rebuilt after any restart while a database leak yields nothing but hashes. The seed creates the **owner** row at startup; its epoch-0 hash is refreshed on every start **only while the row has never been rotated**, so a restart can never resurrect a rotated-away credential. Every other row is created by that Reader's own first login (`Store.EnsureReader`, idempotent on `discord_id`, and it never rewrites an existing row's hash). Rotation is `Store.RotateToken` (epoch bump + hash rewrite in one transaction), driven by the web UI.
`series` keyed `(site, series_id)`
(`asura`|`demonic`|`comix`|`kagane`|`novelfull`|`lightnovelworld`) owns the
shared facts — title, cover, canonical URL, `kind` (`manga`|`novel`),
Latest Chapter, `latest_checked_at` — and `bookmarks` holds only what
differs between readers: progress, favourite, lifecycle bucket,
`updated_at`. A bookmark is keyed `(reader_id, site, series_id)` — no
surrogate id; the wire `key` is derived as `site:series_id` on read — and
every store read/write is scoped to the reader it names. Auth resolves the
acting Reader from the presented credential (`httpmw.Auth`) and nothing
else — there is no unauthenticated-by-Reader route and no global token; the
reader id travels in the request context. Sync **last-write-wins**; the wire format
stays flat (ADR-0004). `Store.Upsert` decomposes one flat body across two
tables and enforces the ownership rule: client `title`/`series_url`/`cover`
are written only when the series row is new (ADR-0003).
- **Endpoints:** `GET /bookmarks`, `PUT /bookmarks/{key}` (upsert; see `updated_at` rule below), `DELETE /bookmarks/{key}`, `GET /healthz` (no auth).
- **Web UI:** same binary serve password-gated browser UI on second
- **Web UI:** same binary serve the browser UI on a second
hostname — `GET /` (list, or login page when no session),
`POST /login`, `POST /logout`, `GET /static/*`, htmx fragment endpoints
`GET /auth/discord` + `GET /auth/discord/callback` (Discord OAuth,
ADR-0002), `POST /logout`, `GET /static/*`, htmx fragment endpoints
under `/ui/*`. Templates + assets `go:embed`-ed under
`backend/internal/web/`, so `backend/Dockerfile` must copy the whole
`internal/` tree, not just `*.go`. Sessions stateless
HMAC cookies keyed off `API_TOKEN`; `WEB_PASSWORD` gates them, and when empty,
web routes not registered at all. UI mutations read-modify-write
`internal/` tree, not just `*.go`. Sessions are rows in the `sessions`
table: the cookie carries only an opaque id, looked up (and expiry-
checked) on every request, and deleting the row revokes the session.
Guild membership *is* registration (issue #27): `discordCallback` gates on
membership (and `DISCORD_REQUIRED_ROLE` when set) and then calls
`Store.EnsureReader`, so a refusal creates nothing and a returning Reader
reuses their row. The owner is the only Reader with administrative reach:
`POST /readers/{id}/revoke` (404 for anyone else) drops that Reader's
sessions, and the `readers` panel renders only on the owner's page.
A Reader with no bookmarks at all sees `listView.Fresh`, whose empty state
offers both install links instead of describing a filter.
UI mutations read-modify-write
through `Store.Get` + `Store.Upsert` so `updated_at` rule stays one
place. See `docs/superpowers/specs/2026-07-25-web-ui-design.md`.
**Design-tool caveat:** templates link `/static/style.css` root-absolutely
@@ -38,27 +74,73 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
Remove's row wear ember wash, two reversible ones wear `.calm` grey.
`--ember` stay reserved for new-chapter signal: busy bar and inline
error use `--mute`.
- **Latest-chapter poller:** ticker goroutine in same binary re-check
each bookmarked series' newest published chapter from backend's own
network access, so `latest_chapter` stay fresh when user not
browsing. Second, parallel signal — userscript keep own
- **Latest-chapter poller:** one goroutine per Site (a Poll Lane, issue #100),
each re-checking that Site's bookmarked series' newest published chapter from
backend's own network access, so `latest_chapter` stays fresh when the user
isn't browsing. Second, parallel signal — the userscript keeps its own
`maybeCaptureLatestOnSeriesPage`/`backgroundRefreshLatest` logic unchanged.
Two independent clocks: per-bookmark cooldown (`latest_checked_at` column,
enforced by `Store.DueForLatestCheck`'s WHERE clause) and wake interval.
Row stamped *before* fetch so broken series wait out full
cooldown instead of retrying every tick, and writes go through
`Store.Get` + `Store.Upsert` so new chapter never reorders list.
Two independent clocks: per-series rest (`series.latest_checked_at`,
enforced by `Store.DueForLatestCheck`'s WHERE clause — `now - Rest`) and
per-Lane gap (the Lane sleeping between fetches, `effectiveGap`). Both live
in the Site registry (`internal/latest/sites.go`), not config: the five env
knobs that used to size a shared pace are gone.
The poller walks **Series, not Bookmarks** — a series referenced by several
bookmarks is fetched once per cycle, and the due queue orders
`reader_count DESC, latest_checked_at ASC` (ADR-0003). Series row stamped
*before* fetch so broken series wait out the rest instead of retrying
every tick; found chapter written straight to the series row via
`Store.SetLatestChapter`, so a bookmark's `updated_at` — and the list
order — is never touched.
Refusals and browser loss are Lane-local: two `errChallengeHeld` in one pass
stop that Site for `refuseBackoff` (15m) while other Lanes continue; an
`errBrowserInterrupted` (remote Chrome restart) sets a shared Poller flag
that makes the other browser Lanes skip their passes for the same 15m, so a
restarting Chrome doesn't stamp one Series per Lane per pass — after the
window the flag decays and they probe again. Browser Lanes wake Chrome only
when 5+ Series are due or one has waited 15m (ADR-0005 on-demand browser),
and cover work (both healing a stored source URL and filling a blank from
the series page) runs in the background so a slow CDN can't consume a
Lane's gap.
Fetches use `bogdanfinn/tls-client` with Chrome profile as defence in depth
against fingerprint-based blocking; any failure log and skip. kagane and
novelfull sit behind Cloudflare JavaScript challenges the TLS client can't
clear, so they are browser-only: fetched over CDP via `BROWSER_WS_URL`, and
simply not polled when that's unset. See
against fingerprint-based blocking; any failure log and skip. kagane, comix
and novelfull sit behind Cloudflare JavaScript challenges the TLS client
can't clear, so they are fetched over CDP via `BROWSER_WS_URL`; kagane and
comix are simply not polled when that's unset, while novelfull falls back to
a plain-TLS attempt — its challenge is a live time-varying fact, and its
cover bytes never need the browser. comix's browser read is an in-tab
`fetch()` of the Series URL, not a DOM render: it is an SPA, so rendering
costs ~65 requests for the same server-rendered HTML one fetch returns
(measured 2026-08-12, issue #98). See
`docs/superpowers/specs/2026-07-26-server-latest-chapter-polling-design.md`.
Poller's `Store.Get` + `Store.Upsert` not wrapped in transaction, so
userscript `PUT` that commits between the two can get overwritten by
poller's stale re-read — reverting that read progress and, since stored
value now differs, moving `updated_at` and reordering list. Known,
accepted limitation for single-user deployment, not bug to fix.
The poller's series write is a single-column UPDATE
(`Store.SetLatestChapter`), not a read-modify-write of the whole bookmark:
it cannot revert read progress or move `updated_at`, so the old
stale-re-read race is gone with the Get+Upsert flow.
- **Covers are acquired at creation, then served from our own origin
(ADR-0007):** the first Bookmark of a Series fires `Store.OnSeriesCreated`,
which `latest.Acquirer` turns into one series-page fetch yielding both the
Latest Chapter and the cover URL; the bytes then go through
`latest.CoverBytesFetcher` into `Store.SetSeriesCover`. It runs in a
goroutine — the Reader's PUT must neither block on a Site nor fail with one
— and every failure is logged and dropped, leaving the Bookmark intact. The
wire's `cover` is the absolute `PUBLIC_BASE_URL + /covers/{sha256}` once
bytes exist and `""` before, never an address that 404s. `GET /covers/{addr}`
is public and uncredentialed: the userscript renders it on a Site's origin,
where no cookie or token of ours travels. A client-sent `cover` is decoded
and discarded, permanently (ADR-0004 compatibility).
Browser-backed Sites join the same pipeline (issue #62, extended to comix by
#98): kagane and comix pages *and* cover bytes go through the browser sidecar
(nothing falls back to a plain fetch, which would only retrieve a challenge
page), while novelfull needs the browser only for its HTML — the cover URL
comes out of the browser-fetched page and the bytes go over plain TLS. With
no browser configured, kagane and comix Covers are simply absent; novelfull
still gets one — at creation and on the poll — when its page body happens to
answer a plain request (the challenge is a live time-varying fact). comix
cover bytes must arrive by direct navigation, not an in-page fetch: its
Series page sets `cross-origin-embedder-policy: require-corp`, which fails a
page-context fetch of `static.comix.to`. The old kagane-only
serving path (`/img/kagane/{id}`, template rewrite, `CoverFetcher`) is gone
(issue #63): the one public route serves every Site.
- **`updated_at` drives list order, so moves only on real reading progress:** server apply its timestamp when row new or `last_chapter_num` changes, else keep stored value — favouriting series or recording newly published chapter must not reorder list. `PUT` therefore returns row **as stored**, clients must adopt that response rather than own payload. See `plans/2026-07-25-bookmark-list-favorites-design.md` §4.
- **Lifecycle buckets:** `status` on each bookmark is `reading` | `archived` |
`finished`, orthogonal to `favorite`. Archived and finished appear only in
@@ -70,13 +152,84 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
`excluded.*` is post-evaluation row and default applied there would
wipe bucket on every PUT from client that predates column. See
`docs/superpowers/specs/2026-07-27-status-buckets-design.md`.
- **Config via env:** `API_TOKEN`, `ALLOWED_ORIGINS` (comma list), `DB_PATH`
(default `/data/bookmarks.db`), `PORT` (default `8080`), `WEB_PASSWORD`
(gates browser UI; unset disable it),
`LATEST_CHAPTER_POLL_ENABLED`/`_COOLDOWN`/`_INTERVAL`/`_BATCH`/`_STAGGER`
(background latest-chapter poller; defaults on, `1h`/`10m`/`14`/`20s`).
`USERSCRIPT_PATH` (file served at `/u/{token}/manga-bookmark.user.js`,
default `/userscript/manga-bookmark.user.js`, supplied by bindmount).
`BROWSER_WS_URL` (headless-shell CDP endpoint for kagane and novelfull;
unset disables browser polling and leaves those sites to the userscript
alone).
- **Config via env:** `TOKEN_KEY` (derives every Reader's userscript credential;
required), `OWNER_DISCORD_ID` (seeds the owner Reader — the administrator and
the owner of every pre-registration bookmark; required),
`ALLOWED_ORIGINS` (comma list),
`DATABASE_URL` (Postgres connection URL, required — no default),
`COVER_DIR` (required filesystem volume for content-addressed Cover bytes),
`PUBLIC_BASE_URL` (required origin this deployment answers on, trailing
slash trimmed; every Cover URL on the wire is built from it, absolute
because the userscript renders on a Site's origin — ADR-0007),
`PORT` (default `8080`), `DISCORD_CLIENT_ID`/`_CLIENT_SECRET`/`_GUILD_ID`/
`_REDIRECT_URI` (required; Discord OAuth for the browser UI),
`DISCORD_REQUIRED_ROLE` (optional role gate, empty by default),
`DISCORD_API_BASE` (default `https://discord.com/api/v10`),
`LATEST_CHAPTER_POLL_ENABLED` (background latest-chapter poller kill
switch, default on). Pace is per Site in the registry (issue #100): every
Site rests an hour and gaps ten seconds, a Site with more eligible Series
than 360 tightens its own gap toward the 1s floor, and browser Lanes wake
Chrome only on demand (ADR-0005). The `_COOLDOWN`/`_BROWSER_COOLDOWN`/
`_INTERVAL`/`_BATCH`/`_STAGGER` knobs that used to size a shared pace are
gone. The 1h rest for browser Sites is safe on documented grounds: a
challenged page costs seconds of a serialized single-tab browser, free-plan
zones have no bot score and no published per-IP rate input, and
`cf_clearance` expires in 30 minutes so every cadence at or above 1h
re-solves anyway —
`docs/research/cloudflare-bot-scoring-and-poll-cadence.md`.
`USERSCRIPT_PATH` and `NOVEL_USERSCRIPT_PATH` (files served at
`/u/{token}/manga-bookmark.user.js` and `/u/{token}/novel-bookmark.user.js`,
defaults `/userscript/manga-bookmark.user.js` and
`/userscript/novel-bookmark.user.js`, both supplied by bindmount; the
`__API_TOKEN__` placeholder inside them is substituted with the requesting
Reader's credential at serve time).
`BROWSER_WS_URL` (CDP endpoint of the browser, which runs on a **separate
machine** and is reached over the tailnet — ADR-0006, `chrome/docker-compose.yml`.
Used by the poller for kagane, comix and novelfull page fetches and by the
cover pipeline for kagane's and comix's image bytes (the browser is the only
route that clears the challenge those two serve their covers behind); unset —
the default — disables browser polling and leaves kagane and comix Covers
blank until stored bytes
exist. Must be a tailnet IP, never a hostname: Chrome's DevTools handler 500s
`/json/version` for any Host that isn't an IP or `localhost`).
- **No per-Site cover path (issue #63):** every Cover — all six Sites — is
served by the one public `GET /covers/{addr}` route from content-addressed
bytes. There is no proxy, no per-Site rewrite, no second place that decides
a Cover's renderable address: the wire `cover` is it. The only place a Site
name still appears in cover code is the extraction module (`latest`), where
kagane's and comix's image URLs are claimed by `browserOnlyCoverURL` — kagane
answers a plain fetch with a challenge and
`cross-origin-resource-policy: same-origin`, and `static.comix.to` answers
one with the same Cloudflare challenge its pages serve;
every other Site's CDN answers plain TLS. Templates render `.Cover` — the
wire value — never anything else.
- **Web UI also owns:** session-gated `GET /install/{manga,novel}-bookmark.user.js`
(renders the bindmounted script with the acting Reader's derived credential
substituted in — the credential never appears in page markup, the address
bar, or a redirect; `?download=1` adds `Content-Disposition: attachment` for
mobile Violentmonkey, which ignores a `.user.js` navigation) and
`POST /rotate-token` (atomic epoch bump + hash
rewrite; invalidates every installed copy, so the panel warns to reinstall
on all devices).
- **Owner-only admin page (`internal/web/admin.go`, issue #102):** `GET /admin`
carries the Reader roster (sessions, Sighting counters, `POST
/readers/{id}/revoke` and `POST /readers/{id}/clear-marks`) and Poll Lane
status (`GET /ui/admin/lanes`, self-refreshing every 30s). Every route that
reaches past the acting Reader is listed in `adminRoutes()` and wrapped in
`requireOwner` at registration — add a route there, not a check inside a
handler; `web.AdminPatterns()` is what the gate test walks. A non-owner gets
404, never 403. Lane figures come from the running poller through the
`web.LaneReporter` seam (`latest.Poller.LaneStatus`), never from a table: a
nil reporter or a Lane that has not finished a pass renders "no data yet"
rather than zeroes. `main.newRouter` takes the reporter as an interface and
converts a nil `*Poller` to a nil interface — a typed nil would make the page
claim a poller exists.
The one owner comparison left outside `requireOwner` is in `index`
(`view.Owner = readerID == h.store.OwnerID()`): it gates a link, not an
endpoint, so it is a rendering decision a registration-time wrapper cannot
express — do not "unify" it into the gate.
A Lane pass that returns before computing its figures (refusal backoff,
sidecar down) carries the previous pass's due count and gap forward rather
than recording zeroes; a Lane that has never reached a pace renders no gap at
all. `Checked` next to `Due` is what separates a stopped Lane from a quiet
one, so neither figure may be dropped from the row.
-81
View File
@@ -1,81 +0,0 @@
Guidance for Claude Code working under `backend/`. See root `CLAUDE.md` for the project-wide architecture diagram, hard constraints, and design system.
- **Backend** (`backend/`): stdlib `net/http` (handful routes, no framework) + `modernc.org/sqlite` (pure Go, `CGO_ENABLED=0` -> static binary -> distroless/scratch image). Reverse proxy terminates TLS; Go service listens plain `:8080`.
Single binary, split into packages under `backend/internal/`: `store`
(Bookmark type, SQLite persistence, migrations), `latest` (background
poller, site parsers, TLS fetcher), `session` (cookie signing, login
rate limiter), `httpmw` (Auth/Gzip/CORS middleware), `api` (JSON
bookmark handlers), `userscript` (userscript-serving handler), `web`
(browser UI handler + `templates/` + `static/`, `go:embed`-ed).
`backend/main.go` is the composition root — the only place that wires
packages together into `newRouter`. Root-level `*_test.go` hold
integration tests that exercise the full router; unit tests for a
package live beside it under `internal/`.
- **Single-user store.** One `bookmarks` table keyed `<site>:<series_id>` (`asura`|`demonic`|`comix`|`kagane`). Sync **last-write-wins**. Schema and endpoint list in plan.
- **Endpoints:** `GET /bookmarks`, `PUT /bookmarks/{key}` (upsert; see `updated_at` rule below), `DELETE /bookmarks/{key}`, `GET /healthz` (no auth).
- **Web UI:** same binary serve password-gated browser UI on second
hostname — `GET /` (list, or login page when no session),
`POST /login`, `POST /logout`, `GET /static/*`, htmx fragment endpoints
under `/ui/*`. Templates + assets `go:embed`-ed under
`backend/internal/web/`, so `backend/Dockerfile` must copy the whole
`internal/` tree, not just `*.go`. Sessions stateless
HMAC cookies keyed off `API_TOKEN`; `WEB_PASSWORD` gates them, and when empty,
web routes not registered at all. UI mutations read-modify-write
through `Store.Get` + `Store.Upsert` so `updated_at` rule stays one
place. See `docs/superpowers/specs/2026-07-25-web-ui-design.md`.
**Design-tool caveat:** templates link `/static/style.css` root-absolutely
(correct — served from `/`), but impeccable detector resolves
stylesheet href with `path.resolve(fileDir, href)`, drops directory
on leading `/` and silently skip file. Relative href don't help
either: template's directory isn't its served path. So
`detect.mjs backend/internal/web/templates` reports **false clean** —
always pass `backend/internal/web/static` too. One finding there,
`overused-font` on "Instrument Serif", deliberate identity choice, not debt.
- **Every action that moves series out of list is confirm-gated.**
Archive, finish, remove each open own `.confirm-row` disclosure
(`toggleConfirmRow(key, kind)` in `filter.js`, `kind` ∈
`archive|finish|remove`); restore fire instantly since it's the reversal.
Remove's row wear ember wash, two reversible ones wear `.calm` grey.
`--ember` stay reserved for new-chapter signal: busy bar and inline
error use `--mute`.
- **Latest-chapter poller:** ticker goroutine in same binary re-check
each bookmarked series' newest published chapter from backend's own
network access, so `latest_chapter` stay fresh when user not
browsing. Second, parallel signal — userscript keep own
`maybeCaptureLatestOnSeriesPage`/`backgroundRefreshLatest` logic unchanged.
Two independent clocks: per-bookmark cooldown (`latest_checked_at` column,
enforced by `Store.DueForLatestCheck`'s WHERE clause) and wake interval.
Row stamped *before* fetch so broken series wait out full
cooldown instead of retrying every tick, and writes go through
`Store.Get` + `Store.Upsert` so new chapter never reorders list.
Fetches use `bogdanfinn/tls-client` with Chrome profile as defence in depth
against fingerprint-based blocking; any failure log and skip. kagane sits
behind a Cloudflare JavaScript challenge the TLS client can't clear, so it is
browser-only: fetched over CDP via `BROWSER_WS_URL`, and simply not polled
when that's unset. See
`docs/superpowers/specs/2026-07-26-server-latest-chapter-polling-design.md`.
Poller's `Store.Get` + `Store.Upsert` not wrapped in transaction, so
userscript `PUT` that commits between the two can get overwritten by
poller's stale re-read — reverting that read progress and, since stored
value now differs, moving `updated_at` and reordering list. Known,
accepted limitation for single-user deployment, not bug to fix.
- **`updated_at` drives list order, so moves only on real reading progress:** server apply its timestamp when row new or `last_chapter_num` changes, else keep stored value — favouriting series or recording newly published chapter must not reorder list. `PUT` therefore returns row **as stored**, clients must adopt that response rather than own payload. See `plans/2026-07-25-bookmark-list-favorites-design.md` §4.
- **Lifecycle buckets:** `status` on each bookmark is `reading` | `archived` |
`finished`, orthogonal to `favorite`. Archived and finished appear only in
own tab — not in All, Updated, Favourites, or recent strip. Poller keeps
checking archived series and skip finished ones. `finished` settable
only from web UI; `PUT /bookmarks/{key}` reject it with 400.
**Empty incoming status means "keep stored one"** — resolved on the
`VALUES` side of `Store.Upsert`, not conflict clause, since
`excluded.*` is post-evaluation row and default applied there would
wipe bucket on every PUT from client that predates column. See
`docs/superpowers/specs/2026-07-27-status-buckets-design.md`.
- **Config via env:** `API_TOKEN`, `ALLOWED_ORIGINS` (comma list), `DB_PATH`
(default `/data/bookmarks.db`), `PORT` (default `8080`), `WEB_PASSWORD`
(gates browser UI; unset disable it),
`LATEST_CHAPTER_POLL_ENABLED`/`_COOLDOWN`/`_INTERVAL`/`_BATCH`/`_STAGGER`
(background latest-chapter poller; defaults on, `1h`/`10m`/`14`/`20s`).
`USERSCRIPT_PATH` (file served at `/u/{token}/manga-bookmark.user.js`,
default `/userscript/manga-bookmark.user.js`, supplied by bindmount).
`BROWSER_WS_URL` (headless-shell CDP endpoint for kagane; unset disables
browser polling and leaves that site to the userscript alone).
+1
View File
@@ -0,0 +1 @@
AGENTS.md
+8 -7
View File
@@ -2,6 +2,7 @@
# --- build stage: compile a static, CGO-free binary ---
FROM golang:1.26-alpine AS build
ARG COVER_DIR=/covers
WORKDIR /src
# Dependencies first for layer caching (changes rarely).
@@ -14,21 +15,21 @@ RUN go mod download
COPY *.go ./
COPY internal/ ./internal/
# Static binary: pure-Go sqlite means CGO_ENABLED=0 -> no libc dependency.
# Static binary: the Postgres driver (jackc/pgx) is pure Go, so CGO_ENABLED=0
# leaves no libc dependency.
# -trimpath + -ldflags strip paths and debug info for a smaller image.
RUN CGO_ENABLED=0 GOOS=linux go build -trimpath -ldflags="-s -w" -o /out/server .
# Data dir with the runtime user's ownership so the mounted volume inherits it.
RUN mkdir -p /out/data
# Create the source directory; runtime COPY sets ownership for the named volume.
RUN mkdir -p "$COVER_DIR"
# --- runtime stage: distroless static, non-root ---
FROM gcr.io/distroless/static:nonroot
ARG COVER_DIR=/covers
WORKDIR /
COPY --from=build --chown=65532:65532 ${COVER_DIR} ${COVER_DIR}
COPY --from=build /out/server /server
COPY --from=build --chown=65532:65532 /out/data /data
VOLUME ["/data"]
EXPOSE 8080
USER nonroot:nonroot
ENV DB_PATH=/data/bookmarks.db PORT=8080
ENV PORT=8080
ENTRYPOINT ["/server"]
+228 -57
View File
@@ -12,57 +12,104 @@ import (
"testing"
"time"
"bookmarkmanager/backend/internal/pgtest"
"bookmarkmanager/backend/internal/store"
"bookmarkmanager/backend/internal/token"
)
const testToken = "s3cret-token"
// testTokenKey derives every test Reader's credential; it must match the key
// newTestStoreURL seeds the owner with, or derived credentials authenticate
// nothing.
const testTokenKey = "test-token-key"
// testDiscordID is the owner row's discord_id (newTestStoreURL); the derived
// credential is a function of it.
const testDiscordID = "test-owner"
// testCoverBaseURL is the public origin cover URLs are built from, standing in
// for PUBLIC_BASE_URL.
const testCoverBaseURL = "https://bookmarks.test"
func testConfig() Config {
return Config{
Token: testToken,
AllowedOrigins: []string{"https://asuracomic.net", "https://demonicscans.org"},
TokenKey: testTokenKey,
AllowedOrigins: []string{"https://asurascans.com", "https://demonicscans.org"},
Port: "8080",
}
}
// ownerCredential is the owner's epoch-0 derived credential: the string the
// install links carry and the userscript routes authenticate.
func ownerCredential() string {
return token.Token([]byte(testTokenKey), testDiscordID, 0)
}
func TestMain(m *testing.M) { os.Exit(pgtest.Main(m)) }
func newTestServer(t *testing.T) http.Handler {
t.Helper()
dbPath := filepath.Join(t.TempDir(), "test.db")
s, err := store.Open(dbPath)
return newRouter(newTestStore(t), testConfig(), nil)
}
func newTestStore(t *testing.T) *store.Store {
t.Helper()
s, _ := newTestStoreURL(t)
return s
}
// newTestStoreURL is newTestStore plus the database URL, for tests that need
// to reach the same database directly.
func newTestStoreURL(t *testing.T) (*store.Store, string) {
t.Helper()
url := pgtest.URL(t)
s, err := store.Open(url, store.Owner{
DiscordID: testDiscordID, TokenHash: token.Hash(ownerCredential()),
}, t.TempDir(), testCoverBaseURL)
if err != nil {
t.Fatalf("store.Open: %v", err)
}
t.Cleanup(func() { s.Close() })
return newRouter(s, testConfig())
return s, url
}
// auth authenticates a request as the owner Reader, whose derived credential
// is the only thing the API accepts.
func auth(req *http.Request) *http.Request {
req.Header.Set("Authorization", "Bearer "+testToken)
req.Header.Set("Authorization", "Bearer "+ownerCredential())
return req
}
func floatPtr(f float64) *float64 { return &f }
// seedForCheck inserts a bookmark and forces its latest_checked_at.
// seedForCheck inserts a bookmark (and with it its series) and forces the
// series' latest_checked_at.
func seedForCheck(t *testing.T, s *store.Store, key, seriesURL string, checkedAt int64) {
t.Helper()
if _, err := s.Upsert(store.Bookmark{
site, seriesID, ok := strings.Cut(key, ":")
if !ok {
t.Fatalf("key %q: no ':' separator", key)
}
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key,
Site: "asura",
SeriesID: key,
Site: site,
SeriesID: seriesID,
SeriesURL: seriesURL,
UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed %q: %v", key, err)
}
if err := s.MarkLatestChecked(key, checkedAt); err != nil {
if err := s.MarkLatestChecked(site, seriesID, checkedAt); err != nil {
t.Fatalf("seed mark %q: %v", key, err)
}
}
func readLatestCheckedAt(t *testing.T, s *store.Store, key string) int64 {
t.Helper()
ts, err := s.LatestCheckedAt(key)
site, seriesID, ok := strings.Cut(key, ":")
if !ok {
t.Fatalf("key %q: no ':' separator", key)
}
ts, err := s.LatestCheckedAt(site, seriesID)
if err != nil {
t.Fatalf("LatestCheckedAt %q: %v", key, err)
}
@@ -89,7 +136,7 @@ func TestAuthRequired(t *testing.T) {
}{
{"no header", ""},
{"bad token", "Bearer wrong"},
{"not bearer", "Basic " + testToken},
{"not bearer", "Basic " + ownerCredential()},
{"empty bearer", "Bearer "},
}
for _, tc := range cases {
@@ -122,7 +169,7 @@ func TestAuthAccepted(t *testing.T) {
func TestCORSPreflight(t *testing.T) {
srv := newTestServer(t)
req := httptest.NewRequest(http.MethodOptions, "/bookmarks/asura:foo-1", nil)
req.Header.Set("Origin", "https://asuracomic.net")
req.Header.Set("Origin", "https://asurascans.com")
req.Header.Set("Access-Control-Request-Method", "PUT")
rr := httptest.NewRecorder()
srv.ServeHTTP(rr, req)
@@ -130,7 +177,7 @@ func TestCORSPreflight(t *testing.T) {
if rr.Code != http.StatusNoContent {
t.Fatalf("preflight status = %d, want 204", rr.Code)
}
if got := rr.Header().Get("Access-Control-Allow-Origin"); got != "https://asuracomic.net" {
if got := rr.Header().Get("Access-Control-Allow-Origin"); got != "https://asurascans.com" {
t.Fatalf("Allow-Origin = %q, want reflected origin", got)
}
if got := rr.Header().Get("Access-Control-Allow-Methods"); got == "" {
@@ -157,11 +204,11 @@ func TestBookmarkRoundTrip(t *testing.T) {
key := "asura:solo-leveling-123"
in := store.Bookmark{
Title: "Solo Leveling",
SeriesURL: "https://asuracomic.net/series/solo-leveling-123",
Cover: "https://asuracomic.net/cover.jpg",
SeriesURL: "https://asurascans.com/series/solo-leveling-123",
Cover: "https://asurascans.com/cover.jpg",
LastChapter: "Chapter 10",
LastChapterNum: 10,
LastChapterURL: "https://asuracomic.net/series/solo-leveling-123/chapter/10",
LastChapterURL: "https://asurascans.com/series/solo-leveling-123/chapter/10",
}
body, _ := json.Marshal(in)
@@ -222,6 +269,127 @@ func TestBookmarkRoundTrip(t *testing.T) {
}
}
// The wire contract (ADR-0004): GET and PUT speak exactly the flat field set
// they always did, with the series-owned fields as siblings of the bookmark
// fields, not nested. Asserted as a key set, not by inspection.
func TestFlatWireFieldSet(t *testing.T) {
srv := newTestServer(t)
key := "comix:some-title"
in := store.Bookmark{
Key: key,
Site: "comix",
SeriesID: "some-title",
Title: "Some Title",
SeriesURL: "https://comix.to/title/some-title",
Cover: "https://comix.to/covers/some-title.jpg",
LastChapter: "Chapter 7",
LastChapterNum: 7,
LastChapterURL: "https://comix.to/title/some-title/ch/7",
Favorite: true,
LatestChapter: "Chapter 8",
LatestChapterNum: floatPtr(8),
Status: store.StatusArchived,
Kind: store.KindManga,
}
body, _ := json.Marshal(in)
wantKeys := map[string]bool{
"key": true, "site": true, "series_id": true, "title": true,
"series_url": true, "cover": true, "last_chapter": true,
"last_chapter_num": true, "last_chapter_url": true, "favorite": true,
"latest_chapter": true, "latest_chapter_num": true, "updated_at": true,
"status": true, "kind": true,
}
checkFlat := func(t *testing.T, payload []byte) map[string]json.RawMessage {
t.Helper()
var obj map[string]json.RawMessage
if err := json.Unmarshal(payload, &obj); err != nil {
t.Fatalf("decode: %v", err)
}
if len(obj) != len(wantKeys) {
t.Fatalf("field count = %d, want %d (%s)", len(obj), len(wantKeys), payload)
}
for k := range obj {
if !wantKeys[k] {
t.Fatalf("unexpected field %q", k)
}
}
return obj
}
// PUT
rr := httptest.NewRecorder()
srv.ServeHTTP(rr, auth(httptest.NewRequest(http.MethodPut, "/bookmarks/"+key, bytes.NewReader(body))))
if rr.Code != http.StatusOK {
t.Fatalf("PUT status = %d, want 200", rr.Code)
}
checkFlat(t, rr.Body.Bytes())
// Every field round-trips with its value, and updated_at is server-stamped.
var stored store.Bookmark
if err := json.Unmarshal(rr.Body.Bytes(), &stored); err != nil {
t.Fatalf("decode PUT response: %v", err)
}
latestNum := floatPtr(8)
want := store.Bookmark{
Key: key, Site: "comix", SeriesID: "some-title",
Title: in.Title, SeriesURL: in.SeriesURL,
LastChapter: in.LastChapter, LastChapterNum: in.LastChapterNum,
LastChapterURL: in.LastChapterURL, Favorite: true,
LatestChapter: in.LatestChapter, LatestChapterNum: latestNum,
Status: store.StatusArchived, Kind: store.KindManga,
}
// Cover is deliberately absent above: the client's cover is discarded, and
// this wiring acquires none, so the field is present and empty (ADR-0007).
if stored.Title != want.Title || stored.SeriesURL != want.SeriesURL || stored.Cover != want.Cover ||
stored.LastChapter != want.LastChapter || stored.LastChapterNum != want.LastChapterNum ||
stored.LastChapterURL != want.LastChapterURL || stored.Favorite != want.Favorite ||
stored.LatestChapter != want.LatestChapter ||
stored.LatestChapterNum == nil || *stored.LatestChapterNum != *want.LatestChapterNum ||
stored.Status != want.Status || stored.Kind != want.Kind {
t.Fatalf("PUT response = %+v, want %+v", stored, want)
}
if stored.UpdatedAt == 0 {
t.Fatal("updated_at not server-stamped")
}
// GET reports the same flat shape.
list := getBookmarks(t, srv)
if len(list) != 1 {
t.Fatalf("list = %d items, want 1", len(list))
}
body2, _ := json.Marshal(list[0])
checkFlat(t, body2)
}
// A PUT naming an existing series must ignore client-supplied title, cover and
// URL — the security boundary from ADR-0003, where a hostile site's scraped
// values could otherwise land on a shared row — while progress still lands.
func TestPutExistingSeriesIgnoresClientTitleCoverURL(t *testing.T) {
srv := newTestServer(t)
key := "asura:solo"
first := putBookmark(t, srv, key, store.Bookmark{
Title: "Solo Leveling",
SeriesURL: "https://asurascans.com/comics/solo",
Cover: "https://asurascans.com/covers/solo.jpg",
LastChapterNum: 10,
})
second := putBookmark(t, srv, key, store.Bookmark{
Title: "Scraped Rename",
SeriesURL: "https://evil.example/solo",
Cover: "https://evil.example/solo.jpg",
LastChapterNum: 11,
})
if second.Title != first.Title || second.SeriesURL != first.SeriesURL || second.Cover != first.Cover {
t.Fatalf("stored = %+v, want original title/url/cover kept", second)
}
if second.LastChapterNum != 11 {
t.Fatalf("LastChapterNum = %v, want 11 — progress must still land", second.LastChapterNum)
}
}
// putBookmark PUTs b at key and returns the bookmark the server echoes back,
// which is the row as actually stored (not the request payload).
func putBookmark(t *testing.T, srv http.Handler, key string, b store.Bookmark) store.Bookmark {
@@ -400,16 +568,29 @@ func TestLatestChapterNullable(t *testing.T) {
}
}
func TestLoadConfigWebPassword(t *testing.T) {
t.Setenv("API_TOKEN", "token-abc")
t.Setenv("WEB_PASSWORD", "hunter2")
if got := loadConfig().WebPassword; got != "hunter2" {
t.Fatalf("WebPassword = %q, want hunter2", got)
func TestLoadConfigDiscord(t *testing.T) {
t.Setenv("DISCORD_CLIENT_ID", "client-1")
t.Setenv("DISCORD_CLIENT_SECRET", "client-secret-1")
t.Setenv("DISCORD_GUILD_ID", "guild-1")
t.Setenv("DISCORD_REQUIRED_ROLE", "role-9")
t.Setenv("DISCORD_REDIRECT_URI", "https://bm.example.com/auth/discord/callback")
t.Setenv("DISCORD_API_BASE", "https://stub.example/api")
if got := loadConfig().Discord; got.ClientID != "client-1" || got.ClientSecret != "client-secret-1" ||
got.GuildID != "guild-1" || got.RequiredRole != "role-9" ||
got.RedirectURI != "https://bm.example.com/auth/discord/callback" ||
got.APIBase != "https://stub.example/api" {
t.Fatalf("Discord config = %+v, want every field set", got)
}
t.Setenv("WEB_PASSWORD", "")
if got := loadConfig().WebPassword; got != "" {
t.Fatalf("WebPassword = %q with the variable unset, want empty", got)
// API base falls back to the Discord default; the role is optional.
t.Setenv("DISCORD_REQUIRED_ROLE", "")
t.Setenv("DISCORD_API_BASE", "")
got := loadConfig().Discord
if got.RequiredRole != "" {
t.Fatalf("RequiredRole = %q, want empty by default", got.RequiredRole)
}
if got.APIBase != "https://discord.com/api/v10" {
t.Fatalf("APIBase = %q, want the Discord default", got.APIBase)
}
}
@@ -417,13 +598,8 @@ func TestLoadConfigWebPassword(t *testing.T) {
// moved into bookmarkColumns, this test catches it: the PUT would reset the
// cooldown and the poller would re-fetch that series on every single tick.
func TestPutDoesNotClobberLatestCheckedAt(t *testing.T) {
dbPath := filepath.Join(t.TempDir(), "test.db")
s, err := store.Open(dbPath)
if err != nil {
t.Fatalf("store.Open: %v", err)
}
t.Cleanup(func() { s.Close() })
srv := newRouter(s, testConfig())
s := newTestStore(t)
srv := newRouter(s, testConfig(), nil)
seedForCheck(t, s, "asura:x", "https://asurascans.com/comics/x", 777)
@@ -432,7 +608,7 @@ func TestPutDoesNotClobberLatestCheckedAt(t *testing.T) {
"series_url":"https://asurascans.com/comics/x",
"last_chapter":"Chapter 5","last_chapter_num":5}`
req := httptest.NewRequest(http.MethodPut, "/bookmarks/asura:x", strings.NewReader(body))
req.Header.Set("Authorization", "Bearer "+testToken)
req.Header.Set("Authorization", "Bearer "+ownerCredential())
req.Header.Set("Content-Type", "application/json")
rec := httptest.NewRecorder()
srv.ServeHTTP(rec, req)
@@ -445,34 +621,33 @@ func TestPutDoesNotClobberLatestCheckedAt(t *testing.T) {
}
}
// The userscript route is registered outside the `if cfg.WebPassword != ""`
// block in newRouter, so it must keep working on a deployment that never set
// WEB_PASSWORD — see internal/userscript for the handler's own behaviour.
// The userscript route is registered outside the web UI's Discord auth, so it
// must keep working whatever the web config — see internal/userscript for the
// handler's own behaviour. The credential in the path is the owner's derived
// one, and the served script carries it substituted in.
func TestUserscriptServedWithWebUIDisabled(t *testing.T) {
path := filepath.Join(t.TempDir(), "manga-bookmark.user.js")
if err := os.WriteFile(path, []byte("console.log(1);\n"), 0o644); err != nil {
if err := os.WriteFile(path, []byte("const API_TOKEN = \"__API_TOKEN__\";\n"), 0o644); err != nil {
t.Fatalf("write script: %v", err)
}
dbPath := filepath.Join(t.TempDir(), "nopass.db")
s, err := store.Open(dbPath)
if err != nil {
t.Fatalf("store.Open: %v", err)
}
t.Cleanup(func() { s.Close() })
cfg := testConfig() // WebPassword empty
s := newTestStore(t)
cfg := testConfig() // no Discord config needed for the userscript route
cfg.UserscriptPath = path
rr := httptest.NewRecorder()
req := httptest.NewRequest(http.MethodGet, "/u/"+testToken+"/manga-bookmark.user.js", nil)
newRouter(s, cfg).ServeHTTP(rr, req)
req := httptest.NewRequest(http.MethodGet, "/u/"+ownerCredential()+"/manga-bookmark.user.js", nil)
newRouter(s, cfg, nil).ServeHTTP(rr, req)
if rr.Code != http.StatusOK {
t.Fatalf("status = %d, want 200", rr.Code)
}
if got := rr.Body.String(); !strings.Contains(got, `API_TOKEN = "`+ownerCredential()+`"`) {
t.Fatalf("served script does not carry the requesting Reader's credential:\n%s", got)
}
}
// Both scripts are served from the same handler on the same token, outside the
// WEB_PASSWORD gate — a wrong token is a 404, never a 401.
// Both scripts are served from the same handler, outside the web UI's auth —
// a wrong credential is a 404, never a 401.
func TestNovelUserscriptServed(t *testing.T) {
dir := t.TempDir()
novelPath := filepath.Join(dir, "novel-bookmark.user.js")
@@ -480,19 +655,15 @@ func TestNovelUserscriptServed(t *testing.T) {
t.Fatalf("write script: %v", err)
}
s, err := store.Open(filepath.Join(dir, "test.db"))
if err != nil {
t.Fatalf("store.Open: %v", err)
}
t.Cleanup(func() { s.Close() })
s := newTestStore(t)
cfg := testConfig()
cfg.NovelUserscriptPath = novelPath
srv := newRouter(s, cfg)
srv := newRouter(s, cfg, nil)
rr := httptest.NewRecorder()
srv.ServeHTTP(rr, httptest.NewRequest(http.MethodGet,
"/u/"+testToken+"/novel-bookmark.user.js", nil))
"/u/"+ownerCredential()+"/novel-bookmark.user.js", nil))
if rr.Code != http.StatusOK {
t.Fatalf("status = %d, want 200", rr.Code)
}
+134
View File
@@ -0,0 +1,134 @@
package main
import (
"database/sql"
"net/http"
"net/http/httptest"
"strings"
"testing"
"bookmarkmanager/backend/internal/store"
)
func getCover(t *testing.T, srv http.Handler, path string, cookie *http.Cookie) *httptest.ResponseRecorder {
t.Helper()
req := httptest.NewRequest(http.MethodGet, path, nil)
if cookie != nil {
req.AddCookie(cookie)
}
rr := httptest.NewRecorder()
srv.ServeHTTP(rr, req)
return rr
}
// The acquired Cover is served from this deployment's own origin, to any
// browser rendering a third-party page — no session, no credential (ADR-0007).
func TestPublicCoverServesStoredBytesUnauthenticated(t *testing.T) {
const sourceURL = "https://cdn.asurascans.com/covers/solo.webp"
srv, st := newWebTestServer(t, testConfig())
if err := st.PutCover(sourceURL, []byte("\x00webp-bytes"), "image/webp"); err != nil {
t.Fatalf("PutCover: %v", err)
}
// The wire URL is what a client actually requests, so the path under test
// is taken from it rather than rebuilt by hand.
wire := st.CoverWireURL(store.CoverAddress(sourceURL))
path, ok := strings.CutPrefix(wire, testCoverBaseURL)
if !ok {
t.Fatalf("wire URL %q is not on the public origin %q", wire, testCoverBaseURL)
}
rr := getCover(t, srv, path, nil)
if rr.Code != http.StatusOK {
t.Fatalf("status = %d, want 200 without any credential", rr.Code)
}
if got := rr.Body.String(); got != "\x00webp-bytes" {
t.Fatalf("body = %q, want the stored bytes", got)
}
if got := rr.Header().Get("Content-Type"); got != "image/webp" {
t.Fatalf("Content-Type = %q, want the stored one", got)
}
// Content-addressed bytes never change, so a client that has them must
// never need to ask again.
if got := rr.Header().Get("Cache-Control"); !strings.Contains(got, "immutable") {
t.Fatalf("Cache-Control = %q, want an immutable cache directive", got)
}
}
func TestPublicCoverRejectsUnknownAddress(t *testing.T) {
srv, _ := newWebTestServer(t, testConfig())
cases := map[string]string{
"unknown": "/covers/" + store.CoverAddress("https://cdn.example/never-stored.jpg"),
"malformed": "/covers/not-an-address",
"traversal": "/covers/../../etc/passwd",
"empty": "/covers/",
}
for name, path := range cases {
t.Run(name, func(t *testing.T) {
if rr := getCover(t, srv, path, nil); rr.Code == http.StatusOK {
t.Fatalf("%s: status = 200, want anything but a served body", path)
}
})
}
}
// A content type outside the image set is never echoed back. The old kagane
// proxy could fetch text/html from a challenged fetch and had to refuse it;
// the general route's only input is the store, and the store refuses to
// record anything that is not an image — but the guarantee is pinned at the
// serving boundary, not the write gate, so a poisoned row (migrated data, a
// writer that skips the gate) is also never served.
func TestPublicCoverNeverEchoesNonImage(t *testing.T) {
const sourceURL = "https://cdn.example/cover"
st, dsn := newTestStoreURL(t)
// The write gate refuses non-image content types outright.
if err := st.PutCover(sourceURL, []byte("<script>"), "text/html"); err == nil {
t.Fatal("PutCover accepted a non-image content type")
}
// A legitimate row, then the content type flipped behind the store's back:
// the bytes exist at the address, so only the type is hostile.
address := store.CoverAddress(sourceURL)
if err := st.SetSeriesCover("asura", "solo", sourceURL, []byte("<script>"), "image/png"); err != nil {
t.Fatalf("seed row: %v", err)
}
db, err := sql.Open("pgx", dsn)
if err != nil {
t.Fatalf("open %s: %v", dsn, err)
}
defer db.Close()
if _, err := db.Exec(`UPDATE covers SET content_type = 'text/html' WHERE address = $1`, address); err != nil {
t.Fatalf("poison row: %v", err)
}
rr := getCover(t, newRouter(st, testConfig(), nil), "/covers/"+address, nil)
if rr.Code == http.StatusOK {
t.Fatalf("status = 200, want a refusal for a non-image row (body %q)", rr.Body.String())
}
}
// The whole point of acquiring bytes is that the UI shows them: the card's
// <img> must carry the public address, not a third-party URL and not a
// placeholder.
func TestListRendersAcquiredCover(t *testing.T) {
const sourceURL = "https://static.comix.to/039d/i/1/34/6a6742bf15736@280.jpg"
srv, st := newWebTestServer(t, testConfig())
if _, err := st.Upsert(st.OwnerID(), store.Bookmark{
Key: "comix:n8we", Site: "comix", SeriesID: "n8we", Title: "Dungeons and Crayons",
SeriesURL: "https://comix.to/title/n8we", UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
if err := st.SetSeriesCover("comix", "n8we", sourceURL, []byte("\xff\xd8jpeg"), "image/jpeg"); err != nil {
t.Fatalf("SetSeriesCover: %v", err)
}
req := httptest.NewRequest(http.MethodGet, "/ui/list", nil)
req.AddCookie(sessionCookie(t, st))
rr := httptest.NewRecorder()
srv.ServeHTTP(rr, req)
if rr.Code != http.StatusOK {
t.Fatalf("status = %d, want 200", rr.Code)
}
want := `src="` + testCoverBaseURL + "/covers/" + store.CoverAddress(sourceURL) + `"`
if !strings.Contains(rr.Body.String(), want) {
t.Fatalf("rendered list does not contain %s", want)
}
}
+5 -13
View File
@@ -7,7 +7,7 @@ require (
github.com/bogdanfinn/tls-client v1.15.1
github.com/chromedp/cdproto v0.0.0-20260714215040-dc233986426f
github.com/chromedp/chromedp v0.16.0
modernc.org/sqlite v1.34.4
github.com/jackc/pgx/v5 v5.10.0
)
require (
@@ -18,27 +18,19 @@ require (
github.com/bogdanfinn/utls v1.7.7-barnius // indirect
github.com/bogdanfinn/websocket v1.5.5-barnius // indirect
github.com/chromedp/sysutil v1.1.0 // indirect
github.com/dustin/go-humanize v1.0.1 // indirect
github.com/go-json-experiment/json v0.0.0-20260623181947-01eb4420fa68 // indirect
github.com/gobwas/httphead v0.1.0 // indirect
github.com/gobwas/pool v0.2.1 // indirect
github.com/gobwas/ws v1.4.0 // indirect
github.com/google/uuid v1.6.0 // indirect
github.com/hashicorp/golang-lru/v2 v2.0.7 // indirect
github.com/jackc/pgpassfile v1.0.0 // indirect
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 // indirect
github.com/jackc/puddle/v2 v2.2.2 // indirect
github.com/klauspost/compress v1.18.2 // indirect
github.com/mattn/go-isatty v0.0.20 // indirect
github.com/ncruces/go-strftime v0.1.9 // indirect
github.com/quic-go/qpack v0.6.0 // indirect
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
github.com/tam7t/hpkp v0.0.0-20160821193359-2b70b4024ed5 // indirect
golang.org/x/crypto v0.46.0 // indirect
golang.org/x/net v0.48.0 // indirect
golang.org/x/sync v0.19.0 // indirect
golang.org/x/sys v0.47.0 // indirect
golang.org/x/text v0.32.0 // indirect
modernc.org/gc/v3 v3.0.0-20240107210532-573471604cb6 // indirect
modernc.org/libc v1.55.3 // indirect
modernc.org/mathutil v1.6.0 // indirect
modernc.org/memory v1.8.0 // indirect
modernc.org/strutil v1.2.0 // indirect
modernc.org/token v1.1.0 // indirect
)
+14 -44
View File
@@ -20,10 +20,9 @@ github.com/chromedp/chromedp v0.16.0 h1:rOO4deOm4CbZgBCa8mD9g2rDyIoNs0BkgvNrlbp5
github.com/chromedp/chromedp v0.16.0/go.mod h1:rbuGKFT1vMcFcFqKfPIO1GpX/N+2s8onm2qMxZLbU5U=
github.com/chromedp/sysutil v1.1.0 h1:PUFNv5EcprjqXZD9nJb9b/c9ibAbxiYo4exNWZyipwM=
github.com/chromedp/sysutil v1.1.0/go.mod h1:WiThHUdltqCNKGc4gaU50XgYjwjYIhKWoHGPTUfWTJ8=
github.com/davecgh/go-spew v1.1.0/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/dustin/go-humanize v1.0.1 h1:GzkhY7T5VNhEkwH0PVJgjz+fX1rhBrR7pRT3mDkpeCY=
github.com/dustin/go-humanize v1.0.1/go.mod h1:Mu1zIs6XwVuF/gI1OepvI0qD18qycQx+mFykh5fBlto=
github.com/go-json-experiment/json v0.0.0-20260623181947-01eb4420fa68 h1:KZaTBSyshWX3MP5jukJcNSuXDQTO+rNpt0J564dX/eg=
github.com/go-json-experiment/json v0.0.0-20260623181947-01eb4420fa68/go.mod h1:tphK2c80bpPhMOI4v6bIc2xWywPfbqi1Z06+RcrMkDg=
github.com/gobwas/httphead v0.1.0 h1:exrUm0f4YX0L7EBwZHuCF4GDp8aJfVeBrlLQrs6NqWU=
@@ -32,28 +31,27 @@ github.com/gobwas/pool v0.2.1 h1:xfeeEhW7pwmX8nuLVlqbzVc7udMDrwetjEv+TZIz1og=
github.com/gobwas/pool v0.2.1/go.mod h1:q8bcK0KcYlCgd9e7WYLm9LpyS+YeLd8JVDW6WezmKEw=
github.com/gobwas/ws v1.4.0 h1:CTaoG1tojrh4ucGPcoJFiAQUAsEWekEWvLy7GsVNqGs=
github.com/gobwas/ws v1.4.0/go.mod h1:G3gNqMNtPppf5XUz7O4shetPpcZ1VJ7zt18dlUeakrc=
github.com/google/pprof v0.0.0-20240409012703-83162a5b38cd h1:gbpYu9NMq8jhDVbvlGkMFWCjLFlqqEZjEmObmhUy6Vo=
github.com/google/pprof v0.0.0-20240409012703-83162a5b38cd/go.mod h1:kf6iHlnVGwgKolg33glAes7Yg/8iWP8ukqeldJSO7jw=
github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0=
github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo=
github.com/hashicorp/golang-lru/v2 v2.0.7 h1:a+bsQ5rvGLjzHuww6tVxozPZFVghXaHOwFs4luLUK2k=
github.com/hashicorp/golang-lru/v2 v2.0.7/go.mod h1:QeFd9opnmA6QUJc5vARoKUSoFhyfM2/ZepoAG6RGpeM=
github.com/jackc/pgpassfile v1.0.0 h1:/6Hmqy13Ss2zCq62VdNG8tM1wchn8zjSGOBJ6icpsIM=
github.com/jackc/pgpassfile v1.0.0/go.mod h1:CEx0iS5ambNFdcRtxPj5JhEz+xB6uRky5eyVu/W2HEg=
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 h1:iCEnooe7UlwOQYpKFhBabPMi4aNAfoODPEFNiAnClxo=
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761/go.mod h1:5TJZWKEWniPve33vlWYSoGYefn3gLQRzjfDlhSJ9ZKM=
github.com/jackc/pgx/v5 v5.10.0 h1:VhSvgU2jSli8o3AqIEOTJr7rZwAEUVo4E4XhR94Zfr0=
github.com/jackc/pgx/v5 v5.10.0/go.mod h1:mal1tBGAFfLHvZzaYh77YS/eC6IX9OWbRV1QIIM0Jn4=
github.com/jackc/puddle/v2 v2.2.2 h1:PR8nw+E/1w0GLuRFSmiioY6UooMp6KJv0/61nB7icHo=
github.com/jackc/puddle/v2 v2.2.2/go.mod h1:vriiEXHvEE654aYKXXjOvZM39qJ0q+azkZFrfEOc3H4=
github.com/klauspost/compress v1.18.2 h1:iiPHWW0YrcFgpBYhsA6D1+fqHssJscY/Tm/y2Uqnapk=
github.com/klauspost/compress v1.18.2/go.mod h1:R0h/fSBs8DE4ENlcrlib3PsXS61voFxhIs2DeRhCvJ4=
github.com/ledongthuc/pdf v0.0.0-20220302134840-0c2507a12d80 h1:6Yzfa6GP0rIo/kULo2bwGEkFvCePZ3qHDDTC3/J9Swo=
github.com/ledongthuc/pdf v0.0.0-20220302134840-0c2507a12d80/go.mod h1:imJHygn/1yfhB7XSJJKlFZKl/J+dCPAknuiaGOshXAs=
github.com/mattn/go-isatty v0.0.20 h1:xfD0iDuEKnDkl03q4limB+vH+GxLEtL/jb4xVJSWWEY=
github.com/mattn/go-isatty v0.0.20/go.mod h1:W+V8PltTTMOvKvAeJH7IuucS94S2C6jfK/D7dTCTo3Y=
github.com/ncruces/go-strftime v0.1.9 h1:bY0MQC28UADQmHmaF5dgpLmImcShSi2kHU9XLdhx/f4=
github.com/ncruces/go-strftime v0.1.9/go.mod h1:Fwc5htZGVVkseilnfgOVb9mKy6w1naJmn9CehxcKcls=
github.com/orisano/pixelmatch v0.0.0-20220722002657-fb0b55479cde h1:x0TT0RDC7UhAVbbWWBzr41ElhJx5tXPWkIHA2HWPRuw=
github.com/orisano/pixelmatch v0.0.0-20220722002657-fb0b55479cde/go.mod h1:nZgzbfBr3hhjoZnS66nKrHmduYNpc34ny7RK4z5/HM0=
github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM=
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/quic-go/qpack v0.6.0 h1:g7W+BMYynC1LbYLSqRt8PBg5Tgwxn214ZZR34VIOjz8=
github.com/quic-go/qpack v0.6.0/go.mod h1:lUpLKChi8njB4ty2bFLX2x4gzDqXwUpaO1DP9qMDZII=
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec h1:W09IVJc94icq4NjY3clb7Lk8O1qJ8BdBEF8z0ibU0rE=
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec/go.mod h1:qqbHyh8v60DhA7CoWK5oRCqLrMHRGoxYCSS9EjAz6Eo=
github.com/stretchr/objx v0.1.0/go.mod h1:HFkY916IF+rwdDfMAkV7OtwuqBVzrE8GR6GFx+wExME=
github.com/stretchr/testify v1.3.0/go.mod h1:M5WIy9Dh21IEIfnGCwXGc5bZfKNJtfHm1UVUgZn+9EI=
github.com/stretchr/testify v1.7.0/go.mod h1:6Fq8oRcR53rry900zMqJjRRixrwX3KX962/h/Wwjteg=
github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U=
github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U=
github.com/tam7t/hpkp v0.0.0-20160821193359-2b70b4024ed5 h1:YqAladjX7xpA6BM04leXMWAEjS0mTZ5kUU9KRBriQJc=
@@ -64,8 +62,6 @@ go.uber.org/mock v0.5.2 h1:LbtPTcP8A5k9WPXj54PPPbjcI4Y6lhyOZXn+VS7wNko=
go.uber.org/mock v0.5.2/go.mod h1:wLlUxC2vVTPTaE3UD51E0BGOAElKrILxhVSDYQLld5o=
golang.org/x/crypto v0.46.0 h1:cKRW/pmt1pKAfetfu+RCEvjvZkA9RimPbh7bhFjGVBU=
golang.org/x/crypto v0.46.0/go.mod h1:Evb/oLKmMraqjZ2iQTwDwvCtJkczlDuTmdJXoZVzqU0=
golang.org/x/mod v0.30.0 h1:fDEXFVZ/fmCKProc/yAXXUijritrDzahmwwefnjoPFk=
golang.org/x/mod v0.30.0/go.mod h1:lAsf5O2EvJeSFMiBxXDki7sCgAxEUcZHXoXMKT4GJKc=
golang.org/x/net v0.0.0-20211104170005-ce137452f963/go.mod h1:9nx3DQGgdP8bBQD5qxJ1jj9UTztislL4KSBs9R2vV5Y=
golang.org/x/net v0.48.0 h1:zyQRTTrjc33Lhh0fBgT/H3oZq9WuvRR5gPC70xpDiQU=
golang.org/x/net v0.48.0/go.mod h1:+ndRgGjkh8FGtu1w1FGbEC31if4VrNVMuKTgcAAnQRY=
@@ -81,33 +77,7 @@ golang.org/x/text v0.3.6/go.mod h1:5Zoc/QRtKVWzQhOtBMvqHzDpF6irO9z98xDceosuGiQ=
golang.org/x/text v0.32.0 h1:ZD01bjUt1FQ9WJ0ClOL5vxgxOI/sVCNgX1YtKwcY0mU=
golang.org/x/text v0.32.0/go.mod h1:o/rUWzghvpD5TXrTIBuJU77MTaN0ljMWE47kxGJQ7jY=
golang.org/x/tools v0.0.0-20180917221912-90fa682c2a6e/go.mod h1:n7NCudcB/nEzxVGmLbDWY5pfWTLqBcC2KZ6jyYvM4mQ=
golang.org/x/tools v0.39.0 h1:ik4ho21kwuQln40uelmciQPp9SipgNDdrafrYA4TmQQ=
golang.org/x/tools v0.39.0/go.mod h1:JnefbkDPyD8UU2kI5fuf8ZX4/yUeh9W877ZeBONxUqQ=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/yaml.v3 v3.0.0-20200313102051-9f266ea9e77c/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
modernc.org/cc/v4 v4.21.4 h1:3Be/Rdo1fpr8GrQ7IVw9OHtplU4gWbb+wNgeoBMmGLQ=
modernc.org/cc/v4 v4.21.4/go.mod h1:HM7VJTZbUCR3rV8EYBi9wxnJ0ZBRiGE5OeGXNA0IsLQ=
modernc.org/ccgo/v4 v4.19.2 h1:lwQZgvboKD0jBwdaeVCTouxhxAyN6iawF3STraAal8Y=
modernc.org/ccgo/v4 v4.19.2/go.mod h1:ysS3mxiMV38XGRTTcgo0DQTeTmAO4oCmJl1nX9VFI3s=
modernc.org/fileutil v1.3.0 h1:gQ5SIzK3H9kdfai/5x41oQiKValumqNTDXMvKo62HvE=
modernc.org/fileutil v1.3.0/go.mod h1:XatxS8fZi3pS8/hKG2GH/ArUogfxjpEKs3Ku3aK4JyQ=
modernc.org/gc/v2 v2.4.1 h1:9cNzOqPyMJBvrUipmynX0ZohMhcxPtMccYgGOJdOiBw=
modernc.org/gc/v2 v2.4.1/go.mod h1:wzN5dK1AzVGoH6XOzc3YZ+ey/jPgYHLuVckd62P0GYU=
modernc.org/gc/v3 v3.0.0-20240107210532-573471604cb6 h1:5D53IMaUuA5InSeMu9eJtlQXS2NxAhyWQvkKEgXZhHI=
modernc.org/gc/v3 v3.0.0-20240107210532-573471604cb6/go.mod h1:Qz0X07sNOR1jWYCrJMEnbW/X55x206Q7Vt4mz6/wHp4=
modernc.org/libc v1.55.3 h1:AzcW1mhlPNrRtjS5sS+eW2ISCgSOLLNyFzRh/V3Qj/U=
modernc.org/libc v1.55.3/go.mod h1:qFXepLhz+JjFThQ4kzwzOjA/y/artDeg+pcYnY+Q83w=
modernc.org/mathutil v1.6.0 h1:fRe9+AmYlaej+64JsEEhoWuAYBkOtQiMEU7n/XgfYi4=
modernc.org/mathutil v1.6.0/go.mod h1:Ui5Q9q1TR2gFm0AQRqQUaBWFLAhQpCwNcuhBOSedWPo=
modernc.org/memory v1.8.0 h1:IqGTL6eFMaDZZhEWwcREgeMXYwmW83LYW8cROZYkg+E=
modernc.org/memory v1.8.0/go.mod h1:XPZ936zp5OMKGWPqbD3JShgd/ZoQ7899TUuQqxY+peU=
modernc.org/opt v0.1.3 h1:3XOZf2yznlhC+ibLltsDGzABUGVx8J6pnFMS3E4dcq4=
modernc.org/opt v0.1.3/go.mod h1:WdSiB5evDcignE70guQKxYUl14mgWtbClRi5wmkkTX0=
modernc.org/sortutil v1.2.0 h1:jQiD3PfS2REGJNzNCMMaLSp/wdMNieTbKX920Cqdgqc=
modernc.org/sortutil v1.2.0/go.mod h1:TKU2s7kJMf1AE84OoiGppNHJwvB753OYfNl2WRb++Ss=
modernc.org/sqlite v1.34.4 h1:sjdARozcL5KJBvYQvLlZEmctRgW9xqIZc2ncN7PU0P8=
modernc.org/sqlite v1.34.4/go.mod h1:3QQFCG2SEMtc2nv+Wq4cQCH7Hjcg+p/RMlS1XK+zwbk=
modernc.org/strutil v1.2.0 h1:agBi9dp1I+eOnxXeiZawM8F4LawKv4NzGWSaLfyeNZA=
modernc.org/strutil v1.2.0/go.mod h1:/mdcBmfOibveCTBxUl5B5l6W+TTH1FXPLHZE6bTosX0=
modernc.org/token v1.1.0 h1:Xl7Ap9dKaEs5kLoOQeQmPWevfnk/DM5qcLcYlA8ys6Y=
modernc.org/token v1.1.0/go.mod h1:UGzOrNV1mAFSEB63lOFHIpNRUVMvYTc6yu1SMY/XTDM=
+45 -4
View File
@@ -7,6 +7,7 @@ import (
"strings"
"time"
"bookmarkmanager/backend/internal/httpmw"
"bookmarkmanager/backend/internal/store"
)
@@ -25,9 +26,9 @@ func writeJSON(w http.ResponseWriter, status int, v any) {
}
}
// List returns all bookmarks. GET /bookmarks
// List returns all bookmarks of the acting Reader. GET /bookmarks
func (h *Handler) List(w http.ResponseWriter, r *http.Request) {
items, err := h.Store.List()
items, err := h.Store.List(httpmw.ReaderID(r))
if err != nil {
log.Printf("list: %v", err)
http.Error(w, "internal error", http.StatusInternalServerError)
@@ -49,6 +50,13 @@ func (h *Handler) Put(w http.ResponseWriter, r *http.Request) {
http.Error(w, "invalid JSON body", http.StatusBadRequest)
return
}
// A body may carry a cover, and it is discarded here rather than
// rejected: an older installed userscript may still send one, and
// ADR-0004's compatibility argument depends on those scripts continuing
// to work. The Cover is acquired server-side (ADR-0007), so the field is
// permanently inert - not pending removal, and not a value any later code
// should start reading.
b.Cover = ""
// Path key is authoritative; derive site/series_id from it when the body
// omits them so the stored row is always self-consistent.
@@ -91,7 +99,7 @@ func (h *Handler) Put(w http.ResponseWriter, r *http.Request) {
// reading progress actually moved. Any client value is ignored.
b.UpdatedAt = time.Now().UnixMilli()
stored, err := h.Store.Upsert(b)
stored, err := h.Store.Upsert(httpmw.ReaderID(r), b)
if err != nil {
log.Printf("upsert: %v", err)
http.Error(w, "internal error", http.StatusInternalServerError)
@@ -109,7 +117,7 @@ func (h *Handler) Delete(w http.ResponseWriter, r *http.Request) {
http.Error(w, "missing key", http.StatusBadRequest)
return
}
if err := h.Store.Delete(key); err != nil {
if err := h.Store.Delete(httpmw.ReaderID(r), key); err != nil {
log.Printf("delete: %v", err)
http.Error(w, "internal error", http.StatusInternalServerError)
return
@@ -123,3 +131,36 @@ func Healthz(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(http.StatusOK)
_, _ = w.Write([]byte("ok"))
}
// Cover serves stored cover bytes. GET /covers/{address}
//
// Public on purpose: the userscript renders these on Sites the deployment
// does not control, where no credential of ours may be sent, and the address
// is the SHA-256 of a URL the Site already publishes (ADR-0007). An unknown
// address is a 404 rather than an error - "no Cover yet" is a normal state,
// and the clients fall back to their placeholder.
func (h *Handler) Cover(w http.ResponseWriter, r *http.Request) {
body, contentType, ok, err := h.Store.CoverByAddress(r.PathValue("address"))
if err != nil {
log.Printf("cover: %v", err)
http.Error(w, "internal error", http.StatusInternalServerError)
return
}
if !ok {
http.NotFound(w, r)
return
}
// Refuse anything the write gate would not have recorded: a poisoned row
// (migrated data, a writer that skips the gate) must never be echoed back
// as bytes of a type no Cover may have.
if _, ok := store.CoverContentType(contentType); !ok {
log.Printf("cover %s: refusing non-image content type %q", r.PathValue("address"), contentType)
http.NotFound(w, r)
return
}
w.Header().Set("Content-Type", contentType)
// Content-addressed, so the bytes at this URL can never change. Public
// rather than private: no credential gates the route.
w.Header().Set("Cache-Control", "public, max-age=604800, immutable")
_, _ = w.Write(body)
}
+35 -7
View File
@@ -2,28 +2,56 @@ package httpmw
import (
"compress/gzip"
"crypto/subtle"
"context"
"log"
"net/http"
"strings"
"bookmarkmanager/backend/internal/store"
"bookmarkmanager/backend/internal/token"
)
const bearerPrefix = "Bearer "
// Auth guards a handler with a constant-time bearer-token check.
func Auth(token string, next http.Handler) http.Handler {
want := []byte(token)
type ctxKey int
// readerCtxKey is where Auth stashes the authenticated Reader id.
const readerCtxKey ctxKey = iota
// ReaderID returns the Reader id Auth authenticated, for handlers that take
// the acting Reader from the request rather than from a fixed field.
func ReaderID(r *http.Request) int64 { return r.Context().Value(readerCtxKey).(int64) }
// ResolveReader maps a presented credential to a Reader. The credential is
// hashed and matched against readers.token_sha256 — an equality on 32-byte
// values, never a comparison of the credential itself. The same resolution
// backs the API bearer header and the userscript download path, so a Reader
// has exactly one credential with one blast radius.
func ResolveReader(s *store.Store, cred string) (int64, bool) {
readerID, ok, err := s.ReaderIDForTokenHash(token.Hash(cred))
if err != nil {
log.Printf("auth: reader lookup: %v", err)
return 0, false
}
return readerID, ok
}
// Auth guards a handler with a per-Reader bearer credential. The acting
// Reader travels in the request context, so a handler scopes every store call
// to exactly the Reader that authenticated.
func Auth(s *store.Store, next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
h := r.Header.Get("Authorization")
if !strings.HasPrefix(h, bearerPrefix) {
http.Error(w, "unauthorized", http.StatusUnauthorized)
return
}
got := []byte(strings.TrimPrefix(h, bearerPrefix))
if subtle.ConstantTimeCompare(got, want) != 1 {
readerID, ok := ResolveReader(s, strings.TrimPrefix(h, bearerPrefix))
if !ok {
http.Error(w, "unauthorized", http.StatusUnauthorized)
return
}
next.ServeHTTP(w, r)
next.ServeHTTP(w, r.WithContext(context.WithValue(r.Context(), readerCtxKey, readerID)))
})
}
+146
View File
@@ -0,0 +1,146 @@
package latest
import (
"context"
"errors"
"log"
"sync"
"time"
"bookmarkmanager/backend/internal/store"
)
// acquireTimeout bounds one creation-time acquisition end to end: the series
// page plus the cover bytes. Nothing is waiting on it — the Reader's write has
// already returned — so this only stops a stalled Site from holding a
// goroutine and a connection open forever.
const acquireTimeout = 45 * time.Second
// Acquirer gives a Series its Latest Chapter and its Cover the moment the
// first Bookmark creates it, instead of leaving the Reader to wait out the
// poll queue — which is ordered by Reader count, so a Series with one Reader
// sits behind every popular one (ADR-0007).
//
// Both facts come from a single series-page fetch, which is also why no
// client-supplied cover hint is worth accepting: the page has to be fetched
// for the chapter signal regardless, so a hint would save no request while
// adding a client-controlled input to a server-side fetch.
//
// Every failure path is "log and move on". The Bookmark, its progress and its
// Latest Chapter are already committed; a Site that is down or a Cover that
// cannot be produced must not disturb any of them, and the Series is simply
// left blank until the poll's own cover pass (#61) fills it.
type Acquirer struct {
Store *store.Store
// Fetch retrieves the series page over plain TLS. Nil with a nil
// BrowserFetch disables acquisition entirely.
Fetch Fetcher
// BrowserFetch retrieves kagane and novelfull pages through the browser
// sidecar, the only thing that clears their Cloudflare challenge. The
// per-site fallback policy lives in fetcherFor. Nil leaves those Sites
// unacquired when no fallback applies.
BrowserFetch Fetcher
// Covers retrieves the cover bytes. Nil leaves the Cover blank and the
// chapter half working.
Covers CoverBytesFetcher
// BrowserCoverFetch retrieves browser-claimed cover bytes through the
// sidecar. Nil leaves those Covers blank; nothing falls back to a plain
// fetch, which would only ever retrieve a challenge page.
BrowserCoverFetch BrowserCoverFetcher
// Ctx cancels in-flight acquisitions at shutdown. A hook signature has
// nowhere to pass one, so it lives here; nil means context.Background.
Ctx context.Context
inflight sync.WaitGroup
}
// acquireSlots caps how many creation-time fetches run at once. A Reader whose
// userscript bulk-syncs creates many Series at once, and a burst of
// simultaneous requests from one server IP is the traffic shape most likely to
// move that IP's bot score — the same reason the poller staggers its batch.
var acquireSlots = make(chan struct{}, 2)
// Acquire starts one acquisition and returns immediately: a Reader's bookmark
// action may not block on a third-party Site's latency, nor fail with it. It
// is the store's OnSeriesCreated hook, so it only ever runs for a Series no
// Reader had bookmarked before.
func (a *Acquirer) Acquire(sr store.Series) {
a.inflight.Add(1)
go func() {
defer a.inflight.Done()
defer func() {
if r := recover(); r != nil {
log.Printf("acquire %q: recovered from panic: %v", sr.Key(), r)
}
}()
parent := a.Ctx
if parent == nil {
parent = context.Background()
}
select {
case acquireSlots <- struct{}{}:
defer func() { <-acquireSlots }()
case <-parent.Done():
return
}
ctx, cancel := context.WithTimeout(parent, acquireTimeout)
defer cancel()
a.acquire(ctx, sr)
}()
}
// Wait blocks until every started acquisition has finished. It exists for
// tests: an asynchronous side effect is otherwise unobservable without
// polling for it.
func (a *Acquirer) Wait() { a.inflight.Wait() }
func (a *Acquirer) acquire(ctx context.Context, sr store.Series) {
if a.Fetch == nil && a.BrowserFetch == nil {
return
}
facts, err := readSeriesPage(ctx, sr.Site, sr.SeriesURL, a.BrowserFetch, a.Fetch)
if err != nil {
switch {
case errors.Is(err, errNotFetchable):
// series_url arrives in a client-supplied PUT body, so without the
// gate a token-holder chooses what the server fetches from its own
// network position.
log.Printf("acquire %q: not fetchable: site=%q url=%q", sr.Key(), sr.Site, sr.SeriesURL)
case errors.Is(err, errNoFetcher):
log.Printf("acquire %q: no fetcher for site %q", sr.Key(), sr.Site)
default:
log.Printf("acquire %q: %v", sr.Key(), err)
}
return
}
// This page just served the same purpose a poll tick would have; without
// the stamp the row stays due and the poller refetches it immediately.
//
// Stamped after success — the reverse of the poller, which stamps before
// the fetch: the Reader is here, watching the Series they just created, so
// a failed acquisition must leave the row due for a fast retry rather than
// consuming the rest. The stamp happens even when the page read
// succeeded but produced no facts to persist.
if err := a.Store.MarkLatestChecked(sr.Site, sr.SeriesID, time.Now().UnixMilli()); err != nil {
log.Printf("acquire %q: mark checked: %v", sr.Key(), err)
}
if facts.HasLatest {
if err := a.Store.SetLatestChapter(sr.Site, sr.SeriesID, facts.Latest.Label, facts.Latest.Num); err != nil {
log.Printf("acquire %q: set latest chapter: %v", sr.Key(), err)
}
}
if !facts.HasCover {
return
}
bytes, contentType, err := fetchCoverBytes(ctx, facts.Cover, a.BrowserCoverFetch, a.Covers)
if err != nil {
log.Printf("acquire %q: fetch cover %s: %v", sr.Key(), facts.Cover, err)
return
}
if err := a.Store.SetSeriesCover(sr.Site, sr.SeriesID, facts.Cover, bytes, contentType); err != nil {
log.Printf("acquire %q: persist cover: %v", sr.Key(), err)
}
}
+430
View File
@@ -0,0 +1,430 @@
package latest
import (
"context"
"errors"
"testing"
"time"
"bookmarkmanager/backend/internal/store"
)
// The series page carries both facts, which is the whole argument for taking
// them from one fetch.
const asuraSeriesAndCoverFixture = asuraSeriesFixture + asuraCoverFixture
const (
acquireKey = "asura:chronicles-of-the-demon-faction-f886a8af"
acquireSeriesID = "chronicles-of-the-demon-faction-f886a8af"
acquireSeriesURL = "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af"
acquireCoverURL = "https://cdn.asurascans.com/asura-images/covers/chronicles-of-the-demon-faction.d4dcb8.webp"
)
const (
kaganeKey = "kagane:019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
kaganeSeriesID = "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
kaganeSeriesURL = "https://kagane.to/series/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
kaganeImageID = "019fe11a-84c3-7fc3-a84b-88787374b617"
kaganeCoverSrc = "https://kagane.to/api/v2/image/" + kaganeImageID + "/compressed"
)
// kagane's browser-fetched body is one JSON object carrying both the chapter
// list (series_books) and the cover image ids (series_covers), so the single
// acquisition fetch yields both facts.
const kaganeSeriesAndCoverFixture = `{"series_id":"019f84bc-9ba0-7ed9-86f5-8b905ec7c28b",` +
`"series_books":[{"book_id":"b","title":"Episode 41","chapter_no":"41","sort_no":41}],` +
`"series_covers":[{"cover_id":"019fe11a-84d1-714b-9cf4-2827f277f3c0","language":"en",` +
`"image_id":"019fe11a-84c3-7fc3-a84b-88787374b617"}]}`
const (
novelfullKey = "novelfull:reverend-insanity"
novelfullSeriesID = "reverend-insanity"
novelfullSeriesURI = "https://novelfull.com/reverend-insanity.html"
novelfullCoverURL = "https://novelfull.com/uploads/webp/novel/reverend-insanity-82661d911a.webp"
)
// newAcquirer wires an acquirer onto the store's creation hook, which is how
// main wires it: the write path is what starts an acquisition.
func newAcquirer(s *store.Store, page *fakeFetcher, covers *fakeBytesCoverFetcher) *Acquirer {
a := &Acquirer{Store: s, Fetch: page, Covers: covers}
s.OnSeriesCreated = a.Acquire
return a
}
func bookmarkNewSeries(t *testing.T, s *store.Store, seriesURL string) store.Bookmark {
t.Helper()
stored, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: acquireKey, Site: "asura", SeriesID: acquireSeriesID,
Title: "Chronicles of the Demon Faction", SeriesURL: seriesURL,
Cover: "https://evil.example/client-supplied.jpg", UpdatedAt: 1000,
})
if err != nil {
t.Fatalf("Upsert: %v", err)
}
return stored
}
func readBookmark(t *testing.T, s *store.Store, key string) store.Bookmark {
t.Helper()
b, ok, err := s.Get(s.OwnerID(), key)
if err != nil || !ok {
t.Fatalf("Get %q = %v, %v", key, ok, err)
}
return b
}
func bookmarkNewKaganeSeries(t *testing.T, s *store.Store) store.Bookmark {
t.Helper()
stored, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: kaganeKey, Site: "kagane", SeriesID: kaganeSeriesID,
Title: "Infinite Decryption", SeriesURL: kaganeSeriesURL, UpdatedAt: 1000,
})
if err != nil {
t.Fatalf("Upsert: %v", err)
}
return stored
}
func bookmarkNewNovelfullSeries(t *testing.T, s *store.Store) store.Bookmark {
t.Helper()
stored, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: novelfullKey, Site: "novelfull", SeriesID: novelfullSeriesID,
Title: "Reverend Insanity", SeriesURL: novelfullSeriesURI, UpdatedAt: 1000,
})
if err != nil {
t.Fatalf("Upsert: %v", err)
}
return stored
}
// The reported bug: a Reader bookmarks a Series nobody holds and expects the
// Cover, not a broken image. Both facts come from the one series-page fetch.
func TestAcquireFillsChapterAndCoverFromOneFetch(t *testing.T) {
s, _ := newTestStore(t)
page := &fakeFetcher{body: asuraSeriesAndCoverFixture, status: 200}
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"}
acq := newAcquirer(s, page, covers)
// The write itself must not carry the acquisition: it returns before the
// Cover exists, and the field is empty until the bytes land.
stored := bookmarkNewSeries(t, s, acquireSeriesURL)
if stored.Cover != "" {
t.Fatalf("Cover on the creating write = %q, want empty", stored.Cover)
}
acq.Wait()
if got := page.callCount(); got != 1 {
t.Fatalf("series page fetches = %d, want exactly 1", got)
}
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetches = %d, want 1", got)
}
got := readBookmark(t, s, acquireKey)
if got.LatestChapterNum == nil || *got.LatestChapterNum != 181 {
t.Fatalf("LatestChapterNum = %v, want 181", got.LatestChapterNum)
}
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(acquireCoverURL); got.Cover != want {
t.Fatalf("Cover = %q, want the absolute address %q", got.Cover, want)
}
body, contentType, ok, err := s.CoverByAddress(store.CoverAddress(acquireCoverURL))
if err != nil || !ok {
t.Fatalf("CoverByAddress = %v, %v", ok, err)
}
if string(body) != "cover-bytes" || contentType != "image/jpeg" {
t.Fatalf("stored cover = (%q, %q), want the fetched bytes", body, contentType)
}
}
// A Series that already exists is not re-acquired: no fetch, and the Cover it
// already has is left alone.
func TestAcquireSkipsAnExistingSeries(t *testing.T) {
s, _ := newTestStore(t)
page := &fakeFetcher{body: asuraSeriesAndCoverFixture, status: 200}
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"}
acq := newAcquirer(s, page, covers)
bookmarkNewSeries(t, s, acquireSeriesURL)
acq.Wait()
bookmarkNewSeries(t, s, acquireSeriesURL)
acq.Wait()
if got := page.callCount(); got != 1 {
t.Fatalf("series page fetches = %d, want 1 — an existing series is not re-acquired", got)
}
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetches = %d, want 1", got)
}
got := readBookmark(t, s, acquireKey)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(acquireCoverURL); got.Cover != want {
t.Fatalf("Cover = %q, want the acquired one %q", got.Cover, want)
}
}
// A Site that is down costs the Cover and nothing else.
func TestAcquireFailureLeavesTheBookmarkIntact(t *testing.T) {
cases := []struct {
name string
page *fakeFetcher
covers *fakeBytesCoverFetcher
// wantLatest is the chapter that still lands; 0 means none did.
wantLatest float64
}{
{
"the series page is unreachable",
&fakeFetcher{err: errors.New("connection reset")},
&fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"},
0,
},
{
"the series page answers with a challenge",
&fakeFetcher{body: challengeFixture, status: 200},
&fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"},
0,
},
{
"only the cover bytes fail",
&fakeFetcher{body: asuraSeriesAndCoverFixture, status: 200},
&fakeBytesCoverFetcher{err: errors.New("403")},
181,
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
s, _ := newTestStore(t)
acq := newAcquirer(s, tc.page, tc.covers)
stored := bookmarkNewSeries(t, s, acquireSeriesURL)
acq.Wait()
got := readBookmark(t, s, acquireKey)
if got.Cover != "" {
t.Fatalf("Cover = %q, want empty rather than an address that 404s", got.Cover)
}
if got.Title != stored.Title || got.UpdatedAt != stored.UpdatedAt {
t.Fatalf("bookmark = %+v, want it untouched by the failed acquisition", got)
}
if tc.wantLatest == 0 {
if got.LatestChapterNum != nil {
t.Fatalf("LatestChapterNum = %v, want none captured", *got.LatestChapterNum)
}
return
}
if got.LatestChapterNum == nil || *got.LatestChapterNum != tc.wantLatest {
t.Fatalf("LatestChapterNum = %v, want %v", got.LatestChapterNum, tc.wantLatest)
}
})
}
}
// series_url arrives in a client-supplied body, so the acquisition reuses the
// poller's gate rather than deriving a second one: a non-https scheme, a
// site the parsers do not know, or a host pinned to another site is refused
// before the server spends a request from its own network position.
func TestAcquireRefusesAnUnfetchableSeriesURL(t *testing.T) {
for _, seriesURL := range []string{
"http://asurascans.com/comics/x",
"file:///etc/passwd",
"",
} {
t.Run(seriesURL, func(t *testing.T) {
s, _ := newTestStore(t)
page := &fakeFetcher{body: asuraSeriesAndCoverFixture, status: 200}
acq := newAcquirer(s, page, &fakeBytesCoverFetcher{})
bookmarkNewSeries(t, s, seriesURL)
acq.Wait()
if got := page.callCount(); got != 0 {
t.Fatalf("fetches for %q = %d, want 0", seriesURL, got)
}
})
}
}
// blockingFetcher stands in for a Site that never answers, so a synchronous
// acquisition would be visible as a stalled write rather than a slow one.
type blockingFetcher struct {
release <-chan struct{}
body string
}
func (f *blockingFetcher) Get(ctx context.Context, _ string) (string, int, error) {
select {
case <-f.release:
return f.body, 200, nil
case <-ctx.Done():
return "", 0, ctx.Err()
}
}
// The Reader's write may not wait on a third-party Site: with the acquisition
// wedged on an unanswering page, the PUT still returns.
func TestAcquireDoesNotBlockTheWrite(t *testing.T) {
s, _ := newTestStore(t)
release := make(chan struct{})
acq := &Acquirer{Store: s, Fetch: &blockingFetcher{release: release, body: asuraSeriesAndCoverFixture}}
s.OnSeriesCreated = acq.Acquire
upserted := make(chan error, 1)
go func() {
_, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: acquireKey, Site: "asura", SeriesID: acquireSeriesID,
Title: "Chronicles of the Demon Faction", SeriesURL: acquireSeriesURL, UpdatedAt: 1000,
})
upserted <- err
}()
select {
case err := <-upserted:
if err != nil {
t.Fatalf("Upsert: %v", err)
}
case <-time.After(10 * time.Second):
t.Fatal("the creating write blocked on the acquisition")
}
close(release)
acq.Wait()
}
// The second symptom of #47: a kagane Series bookmarked from a chapter page
// gets its Cover at creation, with the bytes fetched through the browser
// sidecar — the only path that clears the challenge — into the
// content-addressed store.
func TestAcquireKaganeCoverThroughBrowser(t *testing.T) {
s, _ := newTestStore(t)
tlsPage := &fakeFetcher{body: "", status: 403}
browserPage := &fakeFetcher{body: kaganeSeriesAndCoverFixture, status: 200}
covers := &fakeCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
acq := &Acquirer{
Store: s, Fetch: tlsPage, BrowserFetch: browserPage,
BrowserCoverFetch: covers, Covers: &fakeBytesCoverFetcher{},
}
s.OnSeriesCreated = acq.Acquire
bookmarkNewKaganeSeries(t, s)
acq.Wait()
if got := tlsPage.callCount(); got != 0 {
t.Fatalf("plain-TLS page fetches = %d, want 0 — kagane pages are browser-only", got)
}
if got := browserPage.callCount(); got != 1 {
t.Fatalf("browser page fetches = %d, want 1", got)
}
if got := covers.callCount(); got != 1 {
t.Fatalf("browser cover fetches = %d, want 1", got)
}
if got := covers.calls[0]; got != kaganeCoverSrc {
t.Fatalf("browser cover fetched URL %q, want %q", got, kaganeCoverSrc)
}
got := readBookmark(t, s, kaganeKey)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(kaganeCoverSrc); got.Cover != want {
t.Fatalf("Cover = %q, want the content-addressed URL %q", got.Cover, want)
}
body, contentType, ok, err := s.CoverByAddress(store.CoverAddress(kaganeCoverSrc))
if err != nil || !ok {
t.Fatalf("CoverByAddress = %v, %v", ok, err)
}
if string(body) != "cover-bytes" || contentType != "image/webp" {
t.Fatalf("stored cover = (%q, %q), want the browser-fetched bytes", body, contentType)
}
}
// novelfull needs the browser only for its HTML: the cover URL comes out of
// the browser-fetched page, but the bytes go over plain TLS through the
// ordinary gated fetcher, never through the browser (issue #62).
func TestAcquireNovelfullCoverOverPlainTLS(t *testing.T) {
s, _ := newTestStore(t)
browserPage := &fakeFetcher{body: novelfullSeriesFixture + novelfullCoverFixture, status: 200}
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
acq := &Acquirer{
Store: s, Fetch: &fakeFetcher{body: "", status: 403},
BrowserFetch: browserPage, Covers: covers,
}
s.OnSeriesCreated = acq.Acquire
bookmarkNewNovelfullSeries(t, s)
acq.Wait()
if got := browserPage.callCount(); got != 1 {
t.Fatalf("browser page fetches = %d, want 1", got)
}
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetches = %d, want 1 — novelfull bytes never touch the browser", got)
}
if got := covers.calls[0]; got != novelfullCoverURL {
t.Fatalf("cover fetched from %q, want %q", got, novelfullCoverURL)
}
got := readBookmark(t, s, novelfullKey)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(novelfullCoverURL); got.Cover != want {
t.Fatalf("Cover = %q, want %q", got.Cover, want)
}
}
// With no browser sidecar configured, kagane is simply not acquired: no
// request is spent on a page that could only ever answer with a challenge,
// and nothing falls back to a plain fetch.
func TestAcquireKaganeSkippedWithoutBrowser(t *testing.T) {
s, _ := newTestStore(t)
tlsPage := &fakeFetcher{body: kaganeSeriesAndCoverFixture, status: 200}
acq := &Acquirer{
Store: s, Fetch: tlsPage,
Covers: &fakeBytesCoverFetcher{body: []byte("x"), contentType: "image/webp"},
}
s.OnSeriesCreated = acq.Acquire
bookmarkNewKaganeSeries(t, s)
acq.Wait()
if got := tlsPage.callCount(); got != 0 {
t.Fatalf("plain-TLS fetches for kagane = %d, want 0", got)
}
if got := readBookmark(t, s, kaganeKey); got.Cover != "" {
t.Fatalf("Cover = %q, want empty without a browser", got.Cover)
}
}
// The byte half of "nothing falls back to a plain fetch": with a browser for
// the page but none for the bytes, a kagane Cover stays absent and the TLS
// cover fetcher is never consulted.
func TestAcquireKaganeBytesNeverFallBackToPlainTLS(t *testing.T) {
s, _ := newTestStore(t)
browserPage := &fakeFetcher{body: kaganeSeriesAndCoverFixture, status: 200}
tlsCovers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
acq := &Acquirer{
Store: s, Fetch: &fakeFetcher{body: "", status: 403},
BrowserFetch: browserPage, Covers: tlsCovers,
}
s.OnSeriesCreated = acq.Acquire
bookmarkNewKaganeSeries(t, s)
acq.Wait()
if got := tlsCovers.callCount(); got != 0 {
t.Fatalf("plain-TLS cover fetches = %d, want 0 — kagane bytes are browser-only", got)
}
if got := readBookmark(t, s, kaganeKey); got.Cover != "" {
t.Fatalf("Cover = %q, want empty without a browser cover fetcher", got.Cover)
}
}
// novelfull's no-browser degradation differs from kagane's: only its HTML
// needs the sidecar, so when the page body is available — the challenge is a
// live time-varying fact that sometimes answers a plain request — the Cover
// still lands, bytes over plain TLS.
func TestAcquireNovelfullCoverWithoutBrowser(t *testing.T) {
s, _ := newTestStore(t)
tlsPage := &fakeFetcher{body: novelfullSeriesFixture + novelfullCoverFixture, status: 200}
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
acq := &Acquirer{Store: s, Fetch: tlsPage, Covers: covers}
s.OnSeriesCreated = acq.Acquire
bookmarkNewNovelfullSeries(t, s)
acq.Wait()
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetches = %d, want 1", got)
}
got := readBookmark(t, s, novelfullKey)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(novelfullCoverURL); got.Cover != want {
t.Fatalf("Cover = %q, want %q", got.Cover, want)
}
}
+246 -57
View File
@@ -2,7 +2,9 @@ package latest
import (
"context"
"encoding/base64"
"encoding/json"
"errors"
"fmt"
"net/url"
"regexp"
@@ -16,18 +18,22 @@ import (
// challengeTimeout bounds one navigate-and-solve. A Cloudflare managed
// challenge clears in a few seconds when it clears at all; anything longer is a
// challenge that is not going to pass, and the caller's cooldown was already
// challenge that is not going to pass, and the caller's rest was already
// stamped before this ran.
const challengeTimeout = 45 * time.Second
var kaganeSeriesRe = regexp.MustCompile(`^/series/([0-9a-f-]{36})/?$`)
// comixSeriesPathRe matches the one path shape comixRead will open: a Series
// page, "/title/<id>-<slug>". Verified live 2026-08-12.
var comixSeriesPathRe = regexp.MustCompile(`^/title/[^/?#]+/?$`)
// BrowserFetcher retrieves pages through a remote headless Chrome over the
// DevTools Protocol.
//
// It exists for one reason: kagane.to and novelfull.com sit behind a
// Cloudflare JavaScript challenge. Verified 2026-08-03 (kagane) and 2026-08-05
// (novelfull) from the deployment host, plain HTTP and bogdanfinn/tls-client
// It exists for one reason: kagane.to, novelfull.com and comix.to sit behind a
// Cloudflare JavaScript challenge. Verified 2026-08-03 (kagane), 2026-08-05
// (novelfull) and 2026-08-12 (comix), plain HTTP and bogdanfinn/tls-client
// with a Chrome_133 profile both get 403 with cf-mitigated: challenge on every
// path, including the API, robots.txt and images. Clearing it requires
// executing the challenge script, which only a real browser does.
@@ -38,26 +44,30 @@ var kaganeSeriesRe = regexp.MustCompile(`^/series/([0-9a-f-]{36})/?$`)
// sync that break silently and separately. The browser's own cookie jar
// persists across polls, so the challenge is solved once every few hours.
//
// The two sites differ in how the chapter list is read: kagane serves it from
// a JSON API that must be called from inside the page (so the request carries
// the clearance cookie), while novelfull renders it into the HTML so the
// cleared DOM is the payload.
// The three sites differ in what a cleared tab is asked for: kagane fetches a
// JSON API from inside the page (the list exists nowhere else), comix fetches
// its own Series URL from inside the page (the served HTML carries the facts,
// and rendering the SPA costs ~65 requests instead of one), and novelfull
// renders its list into the HTML so the cleared DOM is the payload.
type BrowserFetcher struct {
allocCtx context.Context
cancel context.CancelFunc
// One page at a time: caps the sidecar's memory and keeps series from
// sharing page state.
// One page at a time: caps the browser's memory — it runs under a hard
// cgroup cap on a shared machine — and keeps series from sharing page state.
mu sync.Mutex
}
var _ Fetcher = (*BrowserFetcher)(nil)
// NewBrowserFetcher connects to a headless-shell over CDP. wsURL must name the
// sidecar by IP, e.g. ws://172.28.0.10:9222 — not by Docker DNS name. Chrome's
// DevTools HTTP handler 500s any /json/version request whose Host header
// isn't an IP or "localhost" (confirmed 2026-08-03 against
// chromedp/headless-shell:stable), so the compose network pins the sidecar's
// address for this to resolve at all.
// NewBrowserFetcher connects to a Chrome over CDP. The browser is not a
// sidecar: it runs on a separate machine and is reached over the tailnet
// (ADR-0006), so wsURL is that machine's tailnet address, e.g.
// ws://100.64.0.5:9222.
//
// It must be an IP, never a hostname — not MagicDNS, not a Docker service
// name. Chrome's DevTools HTTP handler 500s any /json/version request whose
// Host header isn't an IP or "localhost" (confirmed 2026-08-03), so a name
// fails at discovery and surfaces as a dead site rather than a bad URL.
//
// Do not add chromedp.NoModifyURL here: that option skips the /json/version
// discovery request entirely and dials wsURL as if it were already the full
@@ -65,8 +75,10 @@ var _ Fetcher = (*BrowserFetcher)(nil)
// /devtools/browser/<uuid>, a path chosen fresh at every Chrome start — dialing
// the bare host:port 404s. The default (discovery) path works precisely
// because Chrome's /json/version response echoes back the Host header of the
// discovery request in webSocketDebuggerUrl, so as long as wsURL is a
// container-reachable IP, the URL chromedp gets back already points at it.
// discovery request in webSocketDebuggerUrl, so as long as wsURL is an IP this
// process can reach, the URL chromedp gets back already points at it. That is
// also why a Chrome restarted behind a stable endpoint needs no reconnect
// here: the fresh UUID arrives with the next discovery.
func NewBrowserFetcher(wsURL string) (*BrowserFetcher, error) {
if wsURL == "" {
return nil, fmt.Errorf("empty browser websocket url")
@@ -79,25 +91,189 @@ func (f *BrowserFetcher) Close() {
f.cancel()
}
// Get navigates to seriesURL, lets any challenge resolve, then reads either the
// site's JSON API (kagane) from inside the page so the request carries the
// clearance cookie, or the served HTML itself (novelfull) — see
// novelfullSeriesURL for the latter case. The returned body is whatever the
// site's chapter list lives in, which is what latestChapterFrom's per-site
// switch expects.
// Get navigates to seriesURL, lets any challenge resolve, then reads the
// payload the Site's registry entry describes (the shapes are listed on
// BrowserFetcher). The returned body is whatever the Site's chapter list lives
// in, which is what the entry's LatestChapter parse expects.
func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int, error) {
apiURL, isKagane := kaganeAPIURL(seriesURL)
if !isKagane && !novelfullSeriesURL(seriesURL) {
return "", 0, fmt.Errorf("not a fetchable browser series url: %q", seriesURL)
var body string
// Sorted order (browserBackedSites sorts) makes dispatch deterministic:
// entries' Read funcs are expected to refuse any address owned by another
// Site, and the loop must not depend on that staying true.
for _, name := range browserBackedSites() {
s := sites[name]
read, ok := s.Browser.Read(seriesURL, &body)
if !ok {
continue
}
if err := f.run(ctx, seriesURL, read,
func() bool { return s.Browser.Done(body) }); err != nil {
// Challenge never cleared, or the payload was refused.
// Indistinguishable from here and handled identically by the caller.
if errors.Is(err, errChallengeHeld) {
return "", 403, nil
}
return "", 0, fmt.Errorf("browser fetch %q: %w", seriesURL, err)
}
return body, 200, nil
}
return "", 0, fmt.Errorf("not a fetchable browser series url: %q", seriesURL)
}
// kaganeRead builds the in-tab fetch of kagane's chapter-list API: the
// request must be made from inside the page so it carries the clearance
// cookie, and the API is the only place the list exists. Refusing any other
// address is the per-Site half of the SSRF gate, kept deliberately behind
// fetchableSeriesURL (see browserRead.Read).
func kaganeRead(seriesURL string, out *string) (chromedp.Action, bool) {
apiURL, ok := kaganeAPIURL(seriesURL)
if !ok {
return nil, false
}
return chromedp.Evaluate(
`fetch(`+jsString(apiURL)+`).then(r => r.ok ? r.text() : "")`,
out, awaitPromise), true
}
// novelfullRead reads the cleared DOM. novelfull renders its chapter list
// into the served HTML, so there is no API to call from inside the page — the
// challenge-cleared DOM is the payload.
func novelfullRead(seriesURL string, out *string) (chromedp.Action, bool) {
if !novelfullSeriesURL(seriesURL) {
return nil, false
}
return chromedp.OuterHTML("html", out, chromedp.ByQuery), true
}
// comixRead fetches the Series page from inside the cleared tab. comix is an
// SPA: rendering the page costs ~65 requests, while one same-origin fetch of
// the same address returns the server-rendered HTML — 24.5 KB, ~480 ms,
// carrying both parser anchors (measured 2026-08-12, issue #98). So this is
// kaganeRead's shape, not novelfullRead's, even though the payload is HTML.
// Refusing any other address is the per-Site half of the SSRF gate.
func comixRead(seriesURL string, out *string) (chromedp.Action, bool) {
pageURL, ok := comixSeriesPageURL(seriesURL)
if !ok {
return nil, false
}
return chromedp.Evaluate(
`fetch(`+jsString(pageURL)+`).then(r => r.ok ? r.text() : "")`,
out, awaitPromise), true
}
// Image retrieves one cover's bytes through the browser sidecar, and its
// content type.
//
// It exists because kagane and comix serve covers behind the same challenge as
// their pages — kagane additionally with
// `cross-origin-resource-policy: same-origin` — so the bytes are only
// reachable from inside a browser that already holds the clearance cookie
// (verified 2026-08-08 for kagane, 2026-08-12 for comix). Acquisition through
// the sidecar is the only route.
//
// The image URL is navigated to rather than fetched from another page of the
// Site: the challenge only runs on a top-level navigation, and once it clears
// the document *is* the image, so a same-origin fetch of location.href reads
// it straight back out of the cache. For comix the navigation is also the only
// route that works at all — its Series page sets
// `cross-origin-embedder-policy: require-corp`, which fails a page-context
// fetch of the cover host.
//
// The challenge is not solved by the first read: WaitReady("body") is satisfied
// by the interstitial too. run holds the tab open until the in-page fetch
// succeeds, which is what gives the challenge script the seconds it needs.
func (f *BrowserFetcher) Image(ctx context.Context, imageURL string) ([]byte, string, error) {
if !browserOnlyCoverURL(imageURL) {
return nil, "", fmt.Errorf("not a browser-fetchable cover url: %q", imageURL)
}
var dataURL string
err := f.run(ctx, imageURL,
chromedp.Evaluate(`fetch(location.href).then(r => r.ok
? r.blob().then(b => new Promise(res => {
const fr = new FileReader();
fr.onload = () => res(fr.result);
fr.readAsDataURL(b);
}))
: "")`, &dataURL, awaitPromise),
func() bool { return dataURL != "" })
if err != nil {
return nil, "", fmt.Errorf("browser image %s: %w", imageURL, err)
}
// "data:image/webp;base64,<payload>".
head, payload, ok := strings.Cut(dataURL, ";base64,")
if !ok {
return nil, "", fmt.Errorf("browser image %s: not a data url", imageURL)
}
raw, err := base64.StdEncoding.DecodeString(payload)
if err != nil {
return nil, "", fmt.Errorf("browser image %s: %w", imageURL, err)
}
return raw, strings.TrimPrefix(head, "data:"), nil
}
// errChallengeHeld reports that the budget ran out with the interstitial still
// up. Distinct from a transport failure: it means "this site said no", which
// the poller answers with a refusal backoff for that Site's Lane (issue #100).
var errChallengeHeld = errors.New("challenge held")
// errBrowserInterrupted distinguishes a remote Chrome restart from the
// caller's own deadline. chromedp reports both as context.Canceled.
var errBrowserInterrupted = errors.New("browser interrupted")
func classifyBrowserError(ctx context.Context, browserLost bool, err error) error {
if err == nil || ctx.Err() != nil {
return err
}
if !browserLost {
return err
}
if !errors.Is(err, context.Canceled) {
return err
}
return fmt.Errorf("%w: %w", errBrowserInterrupted, err)
}
func browserConnectionLost(ctx context.Context) bool {
c := chromedp.FromContext(ctx)
if c == nil || c.Browser == nil {
return true
}
select {
case <-c.Browser.LostConnection:
return true
default:
return false
}
}
// challengePollInterval paces re-reads while a challenge solves itself.
const challengePollInterval = 2 * time.Second
// isInterstitial reports whether html is Cloudflare's challenge page rather
// than the site's own. Matched on the challenge runtime's script path, which is
// stable across the interstitial's wording and locale — the visible "Just a
// moment..." title is neither.
func isInterstitial(html string) bool {
return strings.Contains(html, "/cdn-cgi/challenge-platform/")
}
// run navigates to target and re-reads until done reports an answer, bounded by
// challengeTimeout and by the caller's own deadline, in a tab that is closed on
// return so one wedged page cannot poison later calls.
//
// Holding the tab open across re-reads is the whole point. A Cloudflare
// interstitial needs several seconds of a live page to solve itself and write
// clearance into the browser's shared cookie jar; reading once and closing the
// tab — which is what this did before 2026-08-08 — never gives it that window,
// so every fetch lands on the interstitial and the clearance that would have
// unblocked all the later ones is never obtained.
func (f *BrowserFetcher) run(ctx context.Context, target string, read chromedp.Action, done func() bool) error {
f.mu.Lock()
defer f.mu.Unlock()
callerCtx := ctx
ctx, cancel := context.WithTimeout(ctx, challengeTimeout)
defer cancel()
// A fresh tab per fetch, closed on return, so one wedged page cannot
// poison later polls.
tabCtx, cancelTab := chromedp.NewContext(f.allocCtx)
defer cancelTab()
// Bind the caller's deadline to the tab.
@@ -108,37 +284,37 @@ func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int
cancelDeadline()
}()
var body string
// kagane's chapter list is only in its JSON API, which must be called from
// inside the page so the request carries the clearance cookie. novelfull
// renders its chapters into the HTML, so the cleared DOM is the answer.
// chromedp.OuterHTML returns a QueryAction and chromedp.Evaluate an
// EvaluateAction, so the variable has to be the interface both implement.
var read chromedp.Action = chromedp.OuterHTML("html", &body, chromedp.ByQuery)
if isKagane {
read = chromedp.Evaluate(
`fetch(`+jsString(apiURL)+`).then(r => r.ok ? r.text() : "")`,
&body,
awaitPromise,
)
}
err := chromedp.Run(tabCtx,
chromedp.Navigate(seriesURL),
// The challenge reloads the page itself when it passes; waiting for the
// site's own root element is what tells us we are through it.
if err := chromedp.Run(tabCtx,
chromedp.Navigate(target),
chromedp.WaitReady("body", chromedp.ByQuery),
read,
)
if err != nil {
return "", 0, fmt.Errorf("browser fetch %q: %w", seriesURL, err)
); err != nil {
return classifyBrowserError(callerCtx, browserConnectionLost(tabCtx), err)
}
if body == "" {
// Challenge still up, or the API refused. Indistinguishable from here
// and handled identically by the caller.
return "", 403, nil
var lastErr error
for {
// The challenge reloads the page when it passes, which tears down the
// execution context mid-read. That is a retry, not a failure.
if err := chromedp.Run(tabCtx, read); err != nil {
err = classifyBrowserError(callerCtx, browserConnectionLost(tabCtx), err)
if errors.Is(err, errBrowserInterrupted) {
return err
}
lastErr = err
} else if done() {
return nil
}
select {
case <-ctx.Done():
if err := callerCtx.Err(); err != nil {
return err
}
if lastErr != nil {
return fmt.Errorf("%w (last read: %v)", errChallengeHeld, lastErr)
}
return errChallengeHeld
case <-time.After(challengePollInterval):
}
}
return body, 200, nil
}
// kaganeAPIURL maps a stored series_url to the JSON endpoint carrying its
@@ -169,6 +345,19 @@ func novelfullSeriesURL(seriesURL string) bool {
strings.HasSuffix(u.Path, ".html")
}
// comixSeriesPageURL returns the address comixRead fetches inside the tab: the
// Series page itself, rebuilt from the pinned host and path so nothing else
// travels. Host-pinned here for the same reason kagane's is — series_url is
// client-supplied and a headless browser is a strong SSRF primitive.
func comixSeriesPageURL(seriesURL string) (string, bool) {
u, err := url.Parse(seriesURL)
if err != nil || u.Scheme != "https" || u.Hostname() != "comix.to" ||
!comixSeriesPathRe.MatchString(u.Path) {
return "", false
}
return "https://comix.to" + u.Path, true
}
// awaitPromise makes Evaluate resolve the promise rather than returning a
// serialised Promise object.
func awaitPromise(p *runtime.EvaluateParams) *runtime.EvaluateParams {
+75 -1
View File
@@ -1,6 +1,10 @@
package latest
import "testing"
import (
"context"
"errors"
"testing"
)
func TestKaganeAPIURL(t *testing.T) {
const uuid = "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
@@ -57,3 +61,73 @@ func TestNovelfullSeriesURL(t *testing.T) {
})
}
}
func TestComixSeriesPageURL(t *testing.T) {
const series = "https://comix.to/title/n8we-dungeons-and-crayons"
cases := []struct {
name string
url string
want string
}{
{"series page", series, series},
{"trailing slash kept", series + "/", series + "/"},
// Query and fragment are dropped: only the pinned path travels.
{"query dropped", series + "?tab=chapters", series},
{"foreign host", "https://evil.example/title/x", ""},
{"lookalike host", "https://comix.to.evil.example/title/x", ""},
{"not https", "http://comix.to/title/x", ""},
{"not a series path", "https://comix.to/search", ""},
{"chapter page", series + "/11139891-chapter-80", ""},
{"garbage", "://nope", ""},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
got, ok := comixSeriesPageURL(tc.url)
if ok != (tc.want != "") || got != tc.want {
t.Fatalf("comixSeriesPageURL(%q) = %q, %v; want %q", tc.url, got, ok, tc.want)
}
})
}
}
// The browser is an SSRF primitive and a cover address can originate in a
// client-supplied PUT body, so this gate decides what it may navigate to.
func TestBrowserOnlyCoverURL(t *testing.T) {
cases := []struct {
url string
want bool
}{
{"https://static.comix.to/039d/i/1/34/6a6742bf15736@280.jpg", true},
{"https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed", true},
// Every other Site's CDN answers plain TLS.
{"https://gg.asuracomic.net/covers/x.webp", false},
{"http://static.comix.to/039d/x.jpg", false},
{"https://static.comix.to.evil.example/039d/x.jpg", false},
{"https://evil.example/static.comix.to/x.jpg", false},
{"https://static.comix.to/039d/x.jpg?next=http://169.254.169.254/", false},
{"https://static.comix.to/039d/x.svg", false},
{"https://static.comix.to/../etc/passwd.jpg", false},
{"https://static.comix.to/", false},
}
for _, tc := range cases {
t.Run(tc.url, func(t *testing.T) {
if got := browserOnlyCoverURL(tc.url); got != tc.want {
t.Fatalf("browserOnlyCoverURL(%q) = %v, want %v", tc.url, got, tc.want)
}
})
}
}
func TestClassifyBrowserInterruption(t *testing.T) {
if err := classifyBrowserError(context.Background(), true, context.Canceled); !errors.Is(err, errBrowserInterrupted) {
t.Fatalf("classifyBrowserError(context.Canceled) = %v, want browser interruption", err)
}
if err := classifyBrowserError(context.Background(), false, context.Canceled); errors.Is(err, errBrowserInterrupted) {
t.Fatalf("ordinary cancellation misclassified as browser interruption: %v", err)
}
caller, cancel := context.WithCancel(context.Background())
cancel()
if err := classifyBrowserError(caller, true, context.Canceled); errors.Is(err, errBrowserInterrupted) {
t.Fatalf("caller cancellation misclassified as browser interruption: %v", err)
}
}
+205
View File
@@ -0,0 +1,205 @@
package latest
import (
"context"
"errors"
"fmt"
"io"
"mime"
"net"
"net/http"
"net/netip"
"net/url"
"strings"
"time"
"bookmarkmanager/backend/internal/store"
)
// CoverBytesFetcher retrieves one cover from its source URL. The caller owns
// persistence; this seam keeps network policy independent from the store.
type CoverBytesFetcher interface {
Fetch(ctx context.Context, sourceURL string) (body []byte, contentType string, err error)
}
// fetchCoverBytes routes a cover's byte retrieval by URL shape, not by Site
// name: the browser fetcher's module claims the addresses only it can fetch
// (kagane's image route answers a plain fetch with a challenge and
// `cross-origin-resource-policy: same-origin`, static.comix.to answers one with
// the same challenge its pages serve), and everything else goes over
// plain TLS. Missing fetchers degrade to an error the caller logs, never a
// fallback onto a path that cannot succeed. One routing rule for the poll and
// the acquirer, so the two cannot drift apart.
func fetchCoverBytes(ctx context.Context, cover string, browser BrowserCoverFetcher, tls CoverBytesFetcher) ([]byte, string, error) {
if browserOnlyCoverURL(cover) {
if browser == nil {
return nil, "", errors.New("no cover fetcher")
}
return browser.Image(ctx, cover)
}
if tls == nil {
return nil, "", errors.New("no cover fetcher")
}
return tls.Fetch(ctx, cover)
}
// CoverResolver resolves a host before any connection is attempted. Tests
// inject it to exercise hostile DNS results without touching the live network.
type CoverResolver func(context.Context, string) ([]netip.Addr, error)
// TLSCoverFetcher retrieves image bytes with the standard HTTPS client. Unlike
// TLSFetcher, it does not need a browser fingerprint: cover hosts are public
// CDNs and the response is accepted only after the destination gate passes.
type TLSCoverFetcher struct {
client *http.Client
resolve CoverResolver
}
var _ CoverBytesFetcher = (*TLSCoverFetcher)(nil)
const coverRequestTimeout = 30 * time.Second
var carrierGradeNAT = netip.MustParsePrefix("100.64.0.0/10")
// NewCoverFetcher builds the production cover client with the real resolver.
func NewCoverFetcher() *TLSCoverFetcher {
return NewCoverFetcherWithResolver(nil)
}
// NewCoverFetcherWithResolver builds a cover client using resolve, or the real
// system resolver when resolve is nil.
func NewCoverFetcherWithResolver(resolve CoverResolver) *TLSCoverFetcher {
if resolve == nil {
resolve = defaultCoverResolver
}
return newCoverFetcher(newCoverHTTPClient(resolve), resolve)
}
func newCoverFetcher(client *http.Client, resolve CoverResolver) *TLSCoverFetcher {
f := &TLSCoverFetcher{client: client, resolve: resolve}
client.CheckRedirect = func(req *http.Request, _ []*http.Request) error {
if err := f.validateURL(req.Context(), req.URL); err != nil {
return fmt.Errorf("redirect destination: %w", err)
}
return nil
}
return f
}
func defaultCoverResolver(ctx context.Context, host string) ([]netip.Addr, error) {
return net.DefaultResolver.LookupNetIP(ctx, "ip", host)
}
func newCoverHTTPClient(resolve CoverResolver) *http.Client {
base, ok := http.DefaultTransport.(*http.Transport)
if !ok {
base = &http.Transport{}
}
transport := base.Clone()
// A proxy would make the dial target the proxy rather than the cover host,
// defeating destination classification. Cover fetching is direct by design.
transport.Proxy = nil
dialer := &net.Dialer{}
transport.DialContext = func(ctx context.Context, network, address string) (net.Conn, error) {
host, port, err := net.SplitHostPort(address)
if err != nil {
return nil, fmt.Errorf("split cover address %q: %w", address, err)
}
addrs, err := resolveCoverHost(ctx, host, resolve)
if err != nil {
return nil, err
}
for _, addr := range addrs {
if !publicCoverAddress(addr) {
return nil, fmt.Errorf("cover host resolves to refused address %s", addr)
}
conn, err := dialer.DialContext(ctx, network, net.JoinHostPort(addr.String(), port))
if err == nil {
return conn, nil
}
}
return nil, fmt.Errorf("cover host %q has no reachable address", host)
}
return &http.Client{Transport: transport, Timeout: coverRequestTimeout}
}
func (f *TLSCoverFetcher) Fetch(ctx context.Context, sourceURL string) ([]byte, string, error) {
u, err := url.Parse(sourceURL)
if err != nil {
return nil, "", fmt.Errorf("parse cover URL: %w", err)
}
if err := f.validateURL(ctx, u); err != nil {
return nil, "", err
}
req, err := http.NewRequestWithContext(ctx, http.MethodGet, u.String(), nil)
if err != nil {
return nil, "", fmt.Errorf("build cover request: %w", err)
}
resp, err := f.client.Do(req)
if err != nil {
return nil, "", fmt.Errorf("fetch cover: %w", err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return nil, "", fmt.Errorf("fetch cover: status %d", resp.StatusCode)
}
raw, _, err := mime.ParseMediaType(resp.Header.Get("Content-Type"))
if err != nil {
return nil, "", fmt.Errorf("fetch cover: unsupported content type %q", resp.Header.Get("Content-Type"))
}
contentType, ok := store.CoverContentType(raw)
if !ok {
return nil, "", fmt.Errorf("fetch cover: unsupported content type %q", raw)
}
if resp.ContentLength > maxBodyBytes {
return nil, "", fmt.Errorf("fetch cover: response exceeds %d bytes", maxBodyBytes)
}
body, err := io.ReadAll(io.LimitReader(resp.Body, maxBodyBytes+1))
if err != nil {
return nil, "", fmt.Errorf("read cover: %w", err)
}
if len(body) > maxBodyBytes {
return nil, "", fmt.Errorf("fetch cover: response exceeds %d bytes", maxBodyBytes)
}
return body, contentType, nil
}
// This gate deliberately differs from fetchableSeriesURL: cover hosts are
// site-independent CDNs, so a Site host allowlist would reject valid covers.
func (f *TLSCoverFetcher) validateURL(ctx context.Context, u *url.URL) error {
if u == nil || u.Scheme != "https" || u.Host == "" || u.User != nil {
return errors.New("cover URL must use HTTPS without credentials")
}
host := u.Hostname()
if host == "" {
return errors.New("cover URL has no host")
}
addrs, err := resolveCoverHost(ctx, host, f.resolve)
if err != nil {
return fmt.Errorf("resolve cover host %q: %w", host, err)
}
if len(addrs) == 0 {
return fmt.Errorf("resolve cover host %q: no addresses", host)
}
for _, addr := range addrs {
if !publicCoverAddress(addr) {
return fmt.Errorf("cover host %q resolves to refused address %s", host, addr)
}
}
return nil
}
func resolveCoverHost(ctx context.Context, host string, resolve CoverResolver) ([]netip.Addr, error) {
if literal, err := netip.ParseAddr(host); err == nil {
return []netip.Addr{literal.Unmap()}, nil
}
return resolve(ctx, strings.TrimSuffix(host, "."))
}
func publicCoverAddress(addr netip.Addr) bool {
addr = addr.Unmap()
return addr.IsValid() && addr.IsGlobalUnicast() &&
!addr.IsLoopback() && !addr.IsPrivate() && !addr.IsLinkLocalUnicast() &&
!carrierGradeNAT.Contains(addr)
}
+226
View File
@@ -0,0 +1,226 @@
package latest
import (
"bytes"
"context"
"crypto/tls"
"io"
"net"
"net/http"
"net/http/httptest"
"net/netip"
"testing"
)
func TestCoverFetcherFetchesPublicHTTPSImage(t *testing.T) {
server := httptest.NewTLSServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.TLS == nil {
t.Fatal("cover request was not made over TLS")
}
w.Header().Set("Content-Type", "image/jpeg")
io.WriteString(w, "cover-bytes")
}))
defer server.Close()
transport := server.Client().Transport.(*http.Transport).Clone()
transport.TLSClientConfig = &tls.Config{InsecureSkipVerify: true} // test server certificate
transport.DialContext = func(ctx context.Context, network, _ string) (net.Conn, error) {
return (&net.Dialer{}).DialContext(ctx, network, server.Listener.Addr().String())
}
client := &http.Client{Transport: transport}
fetcher := newCoverFetcher(client, func(context.Context, string) ([]netip.Addr, error) {
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
})
body, contentType, err := fetcher.Fetch(context.Background(), "https://cdn.example/cover.jpg")
if err != nil {
t.Fatalf("Fetch: %v", err)
}
if string(body) != "cover-bytes" || contentType != "image/jpeg" {
t.Fatalf("Fetch = (%q, %q), want (cover-bytes, image/jpeg)", body, contentType)
}
}
func TestNewCoverFetcherRechecksResolverBeforeConnection(t *testing.T) {
var requests int
server := httptest.NewTLSServer(http.HandlerFunc(func(http.ResponseWriter, *http.Request) {
requests++
}))
defer server.Close()
_, port, err := net.SplitHostPort(server.Listener.Addr().String())
if err != nil {
t.Fatalf("server address: %v", err)
}
resolves := 0
fetcher := NewCoverFetcherWithResolver(func(context.Context, string) ([]netip.Addr, error) {
resolves++
if resolves == 1 {
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
}
return []netip.Addr{netip.MustParseAddr("127.0.0.1")}, nil
})
_, _, err = fetcher.Fetch(context.Background(), "https://cdn.example:"+port+"/cover.jpg")
if err == nil {
t.Fatal("Fetch accepted a destination that became private")
}
if resolves != 2 {
t.Fatalf("resolver calls = %d, want preflight and dial checks", resolves)
}
if requests != 0 {
t.Fatalf("requests = %d, want 0", requests)
}
}
type roundTripFunc func(*http.Request) (*http.Response, error)
func (f roundTripFunc) RoundTrip(r *http.Request) (*http.Response, error) { return f(r) }
func coverResponse(status int, contentType, location string, body []byte) *http.Response {
header := make(http.Header)
if contentType != "" {
header.Set("Content-Type", contentType)
}
if location != "" {
header.Set("Location", location)
}
return &http.Response{
StatusCode: status,
Status: http.StatusText(status),
Header: header,
Body: io.NopCloser(bytes.NewReader(body)),
ContentLength: int64(len(body)),
}
}
func TestCoverFetcherRefusesUnsafeDestinationsBeforeRequest(t *testing.T) {
var calls int
client := &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) {
calls++
return coverResponse(http.StatusOK, "image/jpeg", "", []byte("must not reach network")), nil
})}
resolve := func(_ context.Context, host string) ([]netip.Addr, error) {
switch host {
case "loopback.example":
return []netip.Addr{netip.MustParseAddr("127.0.0.1")}, nil
case "private.example":
return []netip.Addr{netip.MustParseAddr("10.0.0.1")}, nil
case "linklocal.example":
return []netip.Addr{netip.MustParseAddr("169.254.1.1")}, nil
case "unique-local.example":
return []netip.Addr{netip.MustParseAddr("fc00::1")}, nil
case "cgnat.example":
return []netip.Addr{netip.MustParseAddr("100.64.0.1")}, nil
default:
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
}
}
fetcher := newCoverFetcher(client, resolve)
tests := []string{
"http://public.example/cover.jpg",
"https://127.0.0.1/cover.jpg",
"https://10.0.0.1/cover.jpg",
"https://169.254.1.1/cover.jpg",
"https://[fc00::1]/cover.jpg",
"https://100.64.0.1/cover.jpg",
"https://loopback.example/cover.jpg",
"https://private.example/cover.jpg",
"https://linklocal.example/cover.jpg",
"https://unique-local.example/cover.jpg",
"https://cgnat.example/cover.jpg",
}
for _, sourceURL := range tests {
t.Run(sourceURL, func(t *testing.T) {
calls = 0
if _, _, err := fetcher.Fetch(context.Background(), sourceURL); err == nil {
t.Fatal("Fetch accepted refused destination")
}
if calls != 0 {
t.Fatalf("network calls = %d, want 0", calls)
}
})
}
}
func TestCoverFetcherStopsRedirectIntoPrivateAddress(t *testing.T) {
var calls int
client := &http.Client{Transport: roundTripFunc(func(req *http.Request) (*http.Response, error) {
calls++
if req.URL.Hostname() != "cdn.example" {
t.Fatalf("redirect reached %s", req.URL)
}
return coverResponse(http.StatusFound, "", "https://internal.example/cover.jpg", nil), nil
})}
fetcher := newCoverFetcher(client, func(_ context.Context, host string) ([]netip.Addr, error) {
if host == "internal.example" {
return []netip.Addr{netip.MustParseAddr("192.168.1.1")}, nil
}
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
})
if _, _, err := fetcher.Fetch(context.Background(), "https://cdn.example/cover.jpg"); err == nil {
t.Fatal("Fetch followed redirect into private address")
}
if calls != 1 {
t.Fatalf("network calls = %d, want only public first hop", calls)
}
}
func TestCoverFetcherRejectsOversizedBody(t *testing.T) {
var calls int
client := &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) {
calls++
response := coverResponse(http.StatusOK, "image/webp", "", bytes.Repeat([]byte("x"), maxBodyBytes+1))
response.ContentLength = -1
return response, nil
})}
fetcher := newCoverFetcher(client, func(context.Context, string) ([]netip.Addr, error) {
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
})
if _, _, err := fetcher.Fetch(context.Background(), "https://cdn.example/large.webp"); err == nil {
t.Fatal("Fetch accepted oversized body")
}
if calls != 1 {
t.Fatalf("network calls = %d, want 1", calls)
}
}
func TestCoverFetcherRejectsNonImage(t *testing.T) {
var calls int
client := &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) {
calls++
return coverResponse(http.StatusOK, "text/html", "", []byte("challenge")), nil
})}
fetcher := newCoverFetcher(client, func(context.Context, string) ([]netip.Addr, error) {
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
})
if _, _, err := fetcher.Fetch(context.Background(), "https://cdn.example/challenge"); err == nil {
t.Fatal("Fetch accepted non-image response")
}
if calls != 1 {
t.Fatalf("network calls = %d, want 1", calls)
}
}
// comix labels its covers "image/jpg", which is not a registered type but is
// what the Site actually answers with; the bytes are stored under the real
// name so one image cannot land under two spellings.
func TestCoverFetcherCanonicalisesJpgAlias(t *testing.T) {
client := &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) {
return coverResponse(http.StatusOK, "image/jpg", "", []byte("cover-bytes")), nil
})}
fetcher := newCoverFetcher(client, func(context.Context, string) ([]netip.Addr, error) {
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
})
body, contentType, err := fetcher.Fetch(context.Background(), "https://static.comix.to/cover.jpg")
if err != nil {
t.Fatalf("Fetch: %v", err)
}
if string(body) != "cover-bytes" || contentType != "image/jpeg" {
t.Fatalf("Fetch = (%q, %q), want (cover-bytes, image/jpeg)", body, contentType)
}
}
+3 -2
View File
@@ -11,8 +11,9 @@ import (
)
// maxBodyBytes caps what a single series page can cost in memory. Real pages
// measured 100-400 KB on 2026-07-26, so this is roughly 10x headroom and mostly
// guards against a proxy handing back something enormous.
// measured 100-400 KB on 2026-07-26; lightnovelworld runs larger — 685 KB and
// 1.18 MB measured 2026-08-11 — so the headroom there is roughly 3.5x, and the
// cap mostly guards against a proxy handing back something enormous.
const maxBodyBytes = 4 << 20
// chromeUA matches the client profile below. A Chrome fingerprint paired with a
+443 -172
View File
@@ -2,8 +2,11 @@ package latest
import (
"context"
"errors"
"log"
"net/url"
"sort"
"sync"
"time"
"bookmarkmanager/backend/internal/store"
@@ -15,237 +18,505 @@ type Fetcher interface {
Get(ctx context.Context, url string) (body string, status int, err error)
}
// BrowserCoverFetcher retrieves one cover's bytes through the browser-backed
// path — the only route that clears the challenge kagane's and comix's image
// URLs answer a plain fetch with. Satisfied by BrowserFetcher.
type BrowserCoverFetcher interface {
Image(ctx context.Context, imageURL string) (body []byte, contentType string, err error)
}
// Poller re-checks each bookmarked series' newest published chapter on a
// schedule, independent of the userscript's own in-browser checks. The two run
// in parallel and report the same observable fact, so whichever writes last wins
// and neither needs to know about the other.
//
// Two clocks, deliberately independent:
//
// - Interval is how often this goroutine wakes up and looks.
// - Cooldown is how long one bookmark rests since its own last check.
//
// Only the cooldown is per bookmark, and it is enforced by the WHERE clause in
// DueForLatestCheck rather than by any timer. Shortening Interval therefore
// cannot shorten anyone's cooldown; it only makes the poller wake up and find
// nothing due more often.
// Every Site gets its own Poll Lane: one independent stream of Polls with its
// own pace, running concurrently with every other Site's (issue #100). Rest
// time and gap live in the Site registry, not here — see sites.go. Rest is
// enforced by the WHERE clause in DueForLatestCheck rather than by any timer;
// the gap is enforced by the Lane sleeping between fetches.
type Poller struct {
Store *store.Store
Fetch Fetcher
Store *store.Store
Fetch Fetcher
// BrowserFetch handles sites behind a JavaScript challenge that Fetch
// cannot clear. Nil disables those sites entirely rather than falling back
// to Fetch, which would only ever retrieve a challenge page.
BrowserFetch Fetcher
Now func() time.Time // injected so tests can freeze it
Cooldown time.Duration
Interval time.Duration
Stagger time.Duration
Batch int
// CoverFetch is optional; failures are logged and never affect the chapter poll.
CoverFetch BrowserCoverFetcher
// CoverBytesFetch is optional; it handles plain-TLS sources through the
// same failure-isolated prefetch path.
CoverBytesFetch CoverBytesFetcher
Now func() time.Time // injected so tests can freeze it
// refuseUntil gates a Site's Lane after it refused twice in one run: no
// Series of that Site is attempted again before this time (issue #100).
// browserDownAt is when a browser Lane last lost the sidecar; the other
// browser Lanes skip their passes for the next refuseBackoff, so a
// restarting Chrome does not stamp one Series per pass per Lane (story 20).
mu sync.Mutex
refuseUntil map[string]time.Time
browserDownAt time.Time
// laneStates is the owner's page snapshot of each Lane's last pass
// (issue #102), keyed by Site. Guarded by mu; a Site appears only after
// its first pass, so a restart renders "no data yet" rather than zeroes.
laneStates map[string]LaneState
// coverWG tracks in-flight cover work. Covers heal in the background so a
// slow cover host cannot delay the next Series-page Poll; tests join it
// before asserting on cover fetches.
coverWG sync.WaitGroup
}
// fetcherFor returns the fetcher a site needs, or nil when the site cannot be
// fetched at all right now. kagane and novelfull both sit behind a Cloudflare
// JavaScript challenge that no TLS fingerprint clears — kagane verified
// 2026-08-03, novelfull verified 2026-08-05, both against the same Chrome_133
// profile TLSFetcher uses — so they are browser-only or nothing.
func (p *Poller) fetcherFor(site string) Fetcher {
switch site {
case "kagane", "novelfull":
return p.BrowserFetch
}
return p.Fetch
}
// Run polls until ctx is cancelled.
// fillBlankCover gives a Series its Cover when it has none. The blank state is
// what "no Cover yet" means on the wire (ADR-0007): permanently-blank rows
// created before acquisition existed, and rows whose creation-time fetch
// failed, both heal here. A non-blank CoverAddress is left alone — refetching
// would add a request per Series per cycle and change artwork under the Reader
// for no visible reason. A row that already carries a source URL is owned by
// prefetchCover instead; this path only records a Cover address already
// extracted from the series page.
//
// runOnce is called synchronously, so a batch that overruns the tick delays the
// next one instead of stacking a second batch on top of it. That is the intended
// failure mode for a misconfigured batch x stagger: a slower cadence, never
// concurrent fetch storms.
// Failures are logged against the Series and never returned: the chapter poll
// must not notice. A failed fill is retried the next time this Series is due;
// there is no separate retry queue.
func (p *Poller) fillBlankCover(ctx context.Context, sr store.Series, cover string) {
if sr.CoverAddress != "" || sr.Cover != "" {
return
}
if cover == "" {
return
}
// Like healCover, the fill runs in the background: a large import of
// blanks would otherwise pay one og:image fetch per Series against the
// Lane's gap (issue #100, story 12).
p.coverWG.Add(1)
go func() {
defer p.coverWG.Done()
p.storeCover(ctx, sr, cover)
}()
}
// prefetchCover heals Series that already carry a third-party source URL but
// no stored address — the state left by client-supplied covers before
// acquisition moved server-side. Every Site takes the same path; fetchCoverBytes
// routes by URL shape, so browser-claimed URLs still need the sidecar. New
// blanks have no source URL and go through fillBlankCover from the series page
// instead.
func (p *Poller) prefetchCover(ctx context.Context, sr store.Series) {
if sr.Cover == "" || sr.CoverAddress != "" {
return
}
body, contentType, found, err := p.Store.GetCover(sr.Cover)
if err != nil {
log.Printf("latest poll %q: read cover: %v", sr.Key(), err)
return
}
if found {
if err := p.Store.SetSeriesCover(sr.Site, sr.SeriesID, sr.Cover, body, contentType); err != nil {
log.Printf("latest poll %q: persist cover: %v", sr.Key(), err)
}
return
}
p.storeCover(ctx, sr, sr.Cover)
}
// storeCover fetches bytes for sourceURL and points the Series at them. Every
// failure is logged against the Series and swallowed so the chapter poll
// cannot see it.
func (p *Poller) storeCover(ctx context.Context, sr store.Series, sourceURL string) {
bytes, contentType, err := fetchCoverBytes(ctx, sourceURL, p.CoverFetch, p.CoverBytesFetch)
if err != nil {
log.Printf("latest poll %q: fetch cover %s: %v", sr.Key(), sourceURL, err)
return
}
if err := p.Store.SetSeriesCover(sr.Site, sr.SeriesID, sourceURL, bytes, contentType); err != nil {
log.Printf("latest poll %q: persist cover: %v", sr.Key(), err)
}
}
// fetcherFor returns the fetcher a site's page needs, or nil when the site
// cannot be fetched at all right now. A Site whose registry entry carries a
// Browser read — kagane, comix and novelfull, all behind a Cloudflare
// JavaScript challenge no TLS fingerprint clears — prefers the browser; when it
// is absent, the entry's Fallback decides whether plain TLS may take over. One
// routing rule for the poll and the acquirer, so the two cannot drift apart.
func fetcherFor(site string, browser, tls Fetcher) Fetcher {
s, known := sites[site]
if !known {
// No registry entry means nothing to fetch or parse; fail closed even
// though the only caller gates first, so a future caller that skips
// the gate cannot hand an arbitrary https URL to the TLS fetcher.
return nil
}
if s.Browser == nil {
return tls
}
if browser != nil {
return browser
}
if s.Browser.Fallback {
return tls
}
return nil
}
// Run polls until ctx is cancelled: one goroutine per Site Lane, each pacing
// itself by the Site's effective gap. Lanes share nothing but the store and
// the browser fetcher's single tab (BrowserFetcher serializes itself), so one
// hostile Site burns only its own budget.
// laneNames returns every registry Site in the deterministic order both Run
// and runOnce iterate: sorted, so lane behaviour and its tests agree on who
// runs first.
func laneNames() []string {
names := make([]string, 0, len(sites))
for name := range sites {
names = append(names, name)
}
sort.Strings(names)
return names
}
func (p *Poller) Run(ctx context.Context) {
log.Printf("latest-chapter poller: interval=%s cooldown=%s batch=%d stagger=%s",
p.Interval, p.Cooldown, p.Batch, p.Stagger)
t := time.NewTicker(p.Interval)
defer t.Stop()
names := laneNames()
log.Printf("latest-chapter poller: %d lanes, rest=%s gap=%s", len(names), defaultRest, defaultGap)
for _, name := range names {
go p.lane(ctx, name)
}
<-ctx.Done()
log.Println("latest-chapter poller: stopped")
}
// lane is one Site's Poll Lane: one pass, then sleep the pace the pass
// reported, then another pass, until ctx is cancelled. The sleep is the whole
// pace discipline — a pass that fetched nothing still reports its gap so the
// Lane wakes often enough to notice Series as they become due. The pass shares
// the Poller's browser-down state, so a sidecar loss is noticed once and the
// other browser Lanes skip passes until the backoff window decays.
func (p *Poller) lane(ctx context.Context, name string) {
for {
pace := p.runLanePass(ctx, name, true)
if ctx.Err() != nil {
return
}
select {
case <-ctx.Done():
log.Println("latest-chapter poller: stopped")
return
case <-t.C:
p.runOnce(ctx)
case <-time.After(pace):
}
}
}
// runOnce processes one batch of due bookmarks.
// runOnce processes one round: one pass of every Lane, back to back, no real
// time passing. This is the deterministic entry point the test suite drives a
// round at a time. The production Run loop does the same work paced by its own
// sleeps; pacing is the only difference.
func (p *Poller) runOnce(ctx context.Context) {
cutoff := p.Now().Add(-p.Cooldown).UnixMilli()
due, err := p.Store.DueForLatestCheck(cutoff, p.Batch)
if err != nil {
log.Printf("latest poll: due query: %v", err)
return
for _, name := range laneNames() {
p.runLanePass(ctx, name, false)
}
}
// runLanePass processes one pass of one Site's Lane: select the due Series,
// pace through them, and report how long the Lane should wait before its next
// pass. paced spaces consecutive fetches by the Site's effective gap — the
// production Lane's rate limit; the deterministic test entry runs back to back.
func (p *Poller) runLanePass(ctx context.Context, name string, paced bool) time.Duration {
now := p.Now()
// Snapshot this pass for the owner's page (issue #102). Recorded on every
// return path, with the figures filled in where the pass computes them.
st := LaneState{Site: name, LastRun: now, Browser: isBrowserSite(name)}
defer func() { p.recordLaneState(st) }()
if until := p.refusalBackoff(name); now.Before(until) {
// Cooling down after a refusal: do not attempt this Site at all.
return until.Sub(now)
}
if isBrowserSite(name) {
if downFor, down := p.browserDownFor(now); down && downFor < refuseBackoff {
// A sibling browser Lane lost the sidecar within the backoff
// window: skip this pass, so a restarting Chrome does not stamp
// this Site's Series one pass at a time. After refuseBackoff the
// flag decays and the Lane probes again (issue #100, story 20).
log.Printf("latest poll %s: browser lane skipping pass (sidecar down %s ago)", name, downFor)
return refuseBackoff - downFor
}
}
s := sites[name]
f := fetcherFor(name, p.BrowserFetch, p.Fetch)
if f == nil {
// No fetcher at all right now (browser absent, no fallback): every
// Series stays unstamped and due, so a browser that appears after a
// restart finds its full queue waiting (issue #100).
st.Gap = defaultGap
return defaultGap
}
checked := 0
for i, b := range due {
due, err := p.Store.DueForLatestCheck(name, now.Add(-s.Rest).UnixMilli())
if err != nil {
log.Printf("latest poll %s: due query: %v", name, err)
st.Gap = defaultGap
return defaultGap
}
st.Due = len(due)
if s.Browser != nil && f == p.BrowserFetch && !browserWakeDue(due, now, s.Rest) {
// Below both thresholds Chrome stays asleep (ADR-0005 on-demand
// browser): waking it for a single Poll would cost a challenge solve
// per request. The Lane still paces at the default gap, which is what
// the owner's page must show rather than a zero.
st.Gap = defaultGap
return defaultGap
}
if s.Browser != nil {
// Browser Lanes share one tab, so their combined ceiling is about 360
// Polls an hour. When they cannot keep up, the wait past the rest time
// grows — log by how much, every pass, so the decision to give them
// more pages is made from a measurement rather than a guess.
if behind := maxSeriesWait(due, now, s.Rest) - s.Rest; behind > 0 {
log.Printf("latest poll %s: browser lane behind by %s (browser Sites cannot keep up with the hour)", name, behind)
}
}
eligible, err := p.Store.EligibleSeriesCount(name)
if err != nil {
log.Printf("latest poll %s: eligible count: %v", name, err)
st.Gap = defaultGap
return defaultGap
}
gap, clamped := effectiveGap(s, eligible)
st.Gap, st.Clamped = gap, clamped
if clamped {
log.Printf("latest poll %s: gap clamped to %s floor (eligible series=%d)", name, minGap, eligible)
}
if eligible == 0 {
// Nothing to poll for the foreseeable future; sleep a full rest instead
// of re-querying every gap.
return s.Rest
}
refusals := 0
for i, sr := range due {
if ctx.Err() != nil {
break
}
// Staggered rather than fired together: a burst of simultaneous requests
// from one server IP is the traffic shape most likely to move that IP's
// bot score. This is the server-side analogue of the userscript's "one
// series per navigation ... indistinguishable from browsing" (L455-456).
stopped := false
if i > 0 && p.Stagger > 0 {
select {
case <-ctx.Done():
stopped = true
case <-time.After(p.Stagger):
}
}
if stopped {
if refusals >= 2 {
// This Site refused twice in a row: the remaining Series are left
// unstamped and due, and the Lane waits refuseBackoff before
// trying it again.
break
}
p.checkOne(ctx, b)
checked++
if paced && i > 0 {
select {
case <-ctx.Done():
break
case <-time.After(gap):
}
if ctx.Err() != nil {
break
}
}
if err := p.checkOne(ctx, sr); err != nil {
switch {
case errors.Is(err, errChallengeHeld):
refusals++
case errors.Is(err, errBrowserInterrupted):
p.setBrowserDown(now)
log.Printf("latest poll %s: browser unreachable, browser lanes skipping passes for %s", name, refuseBackoff)
return gap
default:
refusals = 0
}
} else {
refusals = 0
}
st.Checked++
}
// due vs checked is how you tell which constraint is binding: ticks that
// report due=0 mean the cooldown is the limit, ticks that report due==batch
// every time mean throughput is.
log.Printf("latest poll: due=%d checked=%d", len(due), checked)
if st.Checked > 0 {
log.Printf("latest poll %s: due=%d checked=%d", name, len(due), st.Checked)
}
if refusals >= 2 {
p.setRefusalBackoff(name, now.Add(refuseBackoff))
log.Printf("latest poll %s: refused twice this run, waiting %s", name, refuseBackoff)
return refuseBackoff
}
return gap
}
func (p *Poller) refusalBackoff(name string) time.Time {
p.mu.Lock()
defer p.mu.Unlock()
return p.refuseUntil[name]
}
func (p *Poller) setRefusalBackoff(name string, until time.Time) {
p.mu.Lock()
defer p.mu.Unlock()
if p.refuseUntil == nil {
p.refuseUntil = make(map[string]time.Time)
}
p.refuseUntil[name] = until
}
// setBrowserDown records when a browser Lane lost the sidecar. It is Poller
// state rather than pass state so the other browser Lanes see it too.
func (p *Poller) setBrowserDown(now time.Time) {
p.mu.Lock()
p.browserDownAt = now
p.mu.Unlock()
}
// browserDownFor reports how long the sidecar has been down and that it is
// down at all — the zero time means never down, which must not read as a
// zero-duration loss. The window decays: once refuseBackoff passes without a
// fresh loss, Lanes probe again.
func (p *Poller) browserDownFor(now time.Time) (time.Duration, bool) {
p.mu.Lock()
defer p.mu.Unlock()
if p.browserDownAt.IsZero() {
return 0, false
}
return now.Sub(p.browserDownAt), true
}
// isBrowserSite reports whether the registry routes this Site's page through
// the browser sidecar.
func isBrowserSite(name string) bool {
return sites[name].Browser != nil
}
// browserWakeDue reports whether a browser Lane may start a run: five or more
// of its Series are due, or any one of them has been due for browserWakeAge.
// Below both thresholds the Lane leaves Chrome asleep — Series Polled together
// become due together, so the group naturally stays clustered, and the age
// rule exists to stop a Series that drifted out of the group from starving.
func browserWakeDue(due []store.Series, now time.Time, rest time.Duration) bool {
if len(due) >= browserWakeCount {
return true
}
return maxSeriesWait(due, now, rest) >= browserWakeAge
}
// maxSeriesWait returns how long the most-overdue of the due Series has been
// waiting past its due moment (0 when due is empty).
func maxSeriesWait(due []store.Series, now time.Time, rest time.Duration) time.Duration {
var oldest time.Duration
for _, sr := range due {
if w := now.Sub(time.UnixMilli(sr.LatestCheckedAt).Add(rest)); w > oldest {
oldest = w
}
}
return oldest
}
// checkOne re-checks one series. Every failure path here is "log and move on":
// the poller is a best-effort enhancement, and no single bad series may stall a
// batch or take down the process.
func (p *Poller) checkOne(ctx context.Context, b store.Bookmark) {
// Lane or take down the process. The returned error is the page read's
// classified outcome so the Lane can tell a refusal from a loss of the
// browser; non-classified failures still return nil-equivalent behaviour.
func (p *Poller) checkOne(ctx context.Context, sr store.Series) error {
defer func() {
if r := recover(); r != nil {
log.Printf("latest poll %q: recovered from panic: %v", b.Key, r)
log.Printf("latest poll %q: recovered from panic: %v", sr.Key(), r)
}
}()
// Stamped before the fetch, not after, so an error, a timeout, or a shutdown
// mid-request still consumes the cooldown. Otherwise a renamed or deleted
// series would be retried on every single tick forever. The userscript
// stamps in the same order and for the same reason (L471-473).
if err := p.Store.MarkLatestChecked(b.Key, p.Now().UnixMilli()); err != nil {
log.Printf("latest poll %q: mark checked: %v", b.Key, err)
return
// mid-request still consumes the rest. Otherwise a renamed or deleted
// series would be retried on every single pass forever. The userscript
// stamps in the same order and for the same reason (L471-473). A Series
// never reaches checkOne without a fetcher — runLanePass skips those — so
// the stamp means "attempted", and an untried Series stays due.
if err := p.Store.MarkLatestChecked(sr.Site, sr.SeriesID, p.Now().UnixMilli()); err != nil {
log.Printf("latest poll %q: mark checked: %v", sr.Key(), err)
return nil
}
// series_url is client-supplied (PUT /bookmarks/{key} accepts any string),
// so this is not just an optimisation against burning a request on an
// unknown site: without it, the server would issue a GET from its own
// network position to whatever URL a token-holder writes, including
// link-local/internal addresses or non-https schemes. The cooldown above
// is already consumed, so a row that never passes this check is retried at
// cooldown pace rather than hot-looping.
if !fetchableSeriesURL(b.Site, b.SeriesURL) {
log.Printf("latest poll %q: not fetchable: site=%q url=%q", b.Key, b.Site, b.SeriesURL)
return
}
f := p.fetcherFor(b.Site)
if f == nil {
log.Printf("latest poll %q: no fetcher for site %q", b.Key, b.Site)
return
}
body, status, err := f.Get(ctx, b.SeriesURL)
facts, err := readSeriesPage(ctx, sr.Site, sr.SeriesURL, p.BrowserFetch, p.Fetch)
if err != nil {
log.Printf("latest poll %q: fetch %s: %v", b.Key, b.SeriesURL, err)
return
switch {
case errors.Is(err, errNotFetchable):
// The rest above is already consumed, so a row that never
// passes the gate is retried at rest pace rather than
// hot-looping.
log.Printf("latest poll %q: not fetchable: site=%q url=%q", sr.Key(), sr.Site, sr.SeriesURL)
return err
case errors.Is(err, errNoFetcher):
log.Printf("latest poll %q: no fetcher for site %q", sr.Key(), sr.Site)
return err
}
// A legacy cover heals independently of the page read: its source may
// answer — a CDN — while the origin does not, so a fetch failure does
// not skip the heal, matching the order the shared read replaced.
p.healCover(ctx, sr)
log.Printf("latest poll %q: %v", sr.Key(), err)
return err
}
if status != 200 {
log.Printf("latest poll %q: fetch %s: status %d", b.Key, b.SeriesURL, status)
return
}
latest, ok := latestChapterFrom(b.Site, b.SeriesURL, body)
if !ok {
// A legacy cover source is healed independently of the page read.
p.healCover(ctx, sr)
// Cover fill is independent of the chapter signal: a page that lost its
// chapter list may keep its og:image, and a blank Series heals either way.
p.fillBlankCover(ctx, sr, facts.Cover)
if !facts.HasLatest {
// Most likely a challenge page or a layout change. Either way the row is
// already stamped, so this waits out a cooldown instead of hot-looping.
log.Printf("latest poll %q: no chapter links in %d bytes", b.Key, len(body))
return
// already stamped, so this waits out a rest instead of hot-looping.
log.Printf("latest poll %q: no chapter links in %d bytes", sr.Key(), facts.BodyLen)
return nil
}
// Re-read: the row may have been updated or deleted while the fetch was in
// flight, and writing b back wholesale would undo that.
//
// ponytail: non-transactional read-modify-write, wrap Get+Upsert in a tx if
// this ever runs for more than one user. A client PUT that commits between
// these two statements is lost to the stale re-read — reverting read
// progress or a status change, and moving updated_at because the stored
// value now differs. Accepted for a single-user deployment: the window is
// milliseconds and the loser is one poll cycle.
cur, found, err := p.Store.Get(b.Key)
if err != nil {
log.Printf("latest poll %q: reread: %v", b.Key, err)
return
}
if !found {
return
}
// Equality, not >, mirroring the userscript (L427): a site that retracts a
// chapter should correct the stored number downward.
if cur.LatestChapterNum != nil && *cur.LatestChapterNum == latest.Num {
return
// chapter should correct the stored number downward. The comparison is
// against the due-query snapshot; a concurrent write in between only costs
// one redundant UPDATE of the same absolute value, never a wrong one.
if sr.LatestChapterNum != nil && *sr.LatestChapterNum == facts.Latest.Num {
return nil
}
num := latest.Num
cur.LatestChapter = latest.Label
cur.LatestChapterNum = &num
// A candidate only. last_chapter_num is untouched, so the CASE in Upsert
// keeps the stored updated_at and the bookmark list does not reorder.
cur.UpdatedAt = p.Now().UnixMilli()
if _, err := p.Store.Upsert(cur); err != nil {
log.Printf("latest poll %q: upsert: %v", b.Key, err)
return
// Series-level write: the row is shared, so one update refreshes every
// bookmark joining to it, and the bookmark's updated_at is never touched —
// a newly published chapter is not reading progress and must not reorder
// the list.
if err := p.Store.SetLatestChapter(sr.Site, sr.SeriesID, facts.Latest.Label, facts.Latest.Num); err != nil {
log.Printf("latest poll %q: set latest chapter: %v", sr.Key(), err)
return nil
}
log.Printf("latest poll %q: latest is now %s", b.Key, latest.Label)
log.Printf("latest poll %q: latest is now %s", sr.Key(), facts.Latest.Label)
return nil
}
// fetchableSeriesURL reports whether site is a site latestChapterFrom knows how
// to parse and seriesURL is safe to hand to a fetcher: an https URL with a
// non-empty host. series_url comes from client-supplied PUT bodies, so this is
// a defence against the poller being used to probe arbitrary hosts from the
// server's own network position, not just a check against wasted requests.
//
// Three sites are held to a stricter rule, each for a different reason:
//
// - kagane and novelfull are fetched by a headless browser, which executes
// JavaScript and carries cookies, and is therefore a far stronger SSRF
// primitive than an HTTP GET. Their hosts must match exactly, not merely
// be non-empty.
// - lightnovelworld's parser regex hardcodes its host, so a URL anywhere
// else could never yield a match — reject it here rather than burn the
// request.
// healCover runs prefetchCover in the background. Cover bytes come from a
// different host — often a CDN — and heal once in a Series's life, so they
// must not consume a Lane's gap: a large import with many blanks would
// otherwise make every Latest Chapter go stale behind a slow image host
// (issue #100).
func (p *Poller) healCover(ctx context.Context, sr store.Series) {
p.coverWG.Add(1)
go func() {
defer p.coverWG.Done()
p.prefetchCover(ctx, sr)
}()
}
// waitCovers blocks until every in-flight cover heal finishes. Tests call it
// after a round before asserting on cover fetches.
func (p *Poller) waitCovers() {
p.coverWG.Wait()
}
// fetchableSeriesURL reports whether site is a Site the registry knows and
// seriesURL is safe to hand to a fetcher: an https URL whose host matches the
// Site's pinned hostname exactly. series_url comes from client-supplied PUT
// bodies, so this is a defence against the poller being used to probe
// arbitrary hosts from the server's own network position, not just a check
// against wasted requests. The pin guards different things per Site — a
// browser Site guards a control that executes JavaScript and carries cookies,
// a parser Site guards a wasted request — but the rule is one rule, from the
// registry.
func fetchableSeriesURL(site, seriesURL string) bool {
switch site {
case "asura", "demonic", "comix", "kagane", "novelfull", "lightnovelworld":
default:
s, known := sites[site]
if !known {
return false
}
u, err := url.Parse(seriesURL)
if err != nil {
return false
}
if u.Scheme != "https" || u.Host == "" {
return false
}
switch site {
case "kagane":
return u.Hostname() == "kagane.to"
case "novelfull":
// Fetched by a real browser, same as kagane, so the host is pinned
// rather than merely non-empty.
return u.Hostname() == "novelfull.com"
case "lightnovelworld":
// Its parser regex hardcodes this host, so a URL anywhere else could
// never yield a match — reject it here rather than burn the request.
return u.Hostname() == "lightnovelworld.net"
}
return true
return u.Scheme == "https" && u.Hostname() == s.Host
}
File diff suppressed because it is too large Load Diff
+73
View File
@@ -0,0 +1,73 @@
package latest
import (
"context"
"errors"
"fmt"
)
// seriesRead carries the two facts the poll and the acquirer both extract
// from a series page. Persistence, stamps and scheduling stay with the
// callers, so the policies that keep the two flows distinct (stamp order,
// rests) are not swallowed by the module.
type seriesRead struct {
Latest latestChapter
HasLatest bool
Cover string
HasCover bool
// BodyLen is the fetched body's length, surfaced because the no-chapter
// log uses it to tell a markup change from a body the size cap cut short.
BodyLen int
}
// errNotFetchable and errNoFetcher separate the gate and the route from fetch
// failures so each caller keeps its own distinct log line for all three.
// errChallengeHeld (browser.go) is the outcome of a Site that answered with
// its interstitial — status 403 (cf-mitigated) or a challenge page body — and
// is how a Lane tells a refusal from an ordinary failure (issue #100).
var (
errNotFetchable = errors.New("series url not fetchable")
errNoFetcher = errors.New("no fetcher for site")
)
// readSeriesPage performs the series-page read the poll and the acquirer have
// in common: gate the address, choose the route, fetch the page, extract the
// Latest Chapter and the Cover address. It persists nothing and stamps
// nothing.
//
// series_url arrives in a client-supplied PUT body (PUT /bookmarks/{key}
// accepts any string), so the gate is not an optimisation against burning a
// request on an unknown site: without it, the server would issue a GET from
// its own network position to whatever URL a token-holder writes, including
// link-local/internal addresses or non-https schemes.
func readSeriesPage(ctx context.Context, site, seriesURL string, browser, tls Fetcher) (seriesRead, error) {
if !fetchableSeriesURL(site, seriesURL) {
return seriesRead{}, fmt.Errorf("%w: site=%q url=%q", errNotFetchable, site, seriesURL)
}
f := fetcherFor(site, browser, tls)
if f == nil {
return seriesRead{}, fmt.Errorf("%w: site %q", errNoFetcher, site)
}
body, status, err := f.Get(ctx, seriesURL)
if err != nil {
return seriesRead{}, fmt.Errorf("fetch %s: %w", seriesURL, err)
}
if status == 403 {
// Cloudflare's challenge response for these Sites (cf-mitigated). The
// browser fetcher returns exactly this on a held interstitial, and a
// plain-TLS 403 means the same: the Site is refusing.
return seriesRead{}, fmt.Errorf("%w: fetch %s: status %d", errChallengeHeld, seriesURL, status)
}
if status != 200 {
return seriesRead{}, fmt.Errorf("fetch %s: status %d", seriesURL, status)
}
if isInterstitial(body) {
// A 200 that is the challenge page, not the payload: the TLS route can
// receive this where the browser would have kept re-reading. Same
// refusal as the 403.
return seriesRead{}, fmt.Errorf("%w: fetch %s: interstitial body", errChallengeHeld, seriesURL)
}
latest, hasLatest := latestChapterFrom(site, seriesURL, body)
cover, hasCover := coverFrom(site, seriesURL, body)
return seriesRead{Latest: latest, HasLatest: hasLatest, Cover: cover, HasCover: hasCover, BodyLen: len(body)}, nil
}
+457 -82
View File
@@ -1,12 +1,17 @@
package latest
import (
"encoding/json"
"html"
"log"
"net/url"
"regexp"
"sort"
"strconv"
"strings"
"time"
"bookmarkmanager/backend/internal/store"
"github.com/chromedp/chromedp"
)
// latestChapter is the newest chapter a series page advertises.
@@ -15,6 +20,45 @@ type latestChapter struct {
Label string
}
// site answers the fixed questions every series-page read asks of its Site
// (ADR-0009): the host its addresses must carry, how to find the Latest
// Chapter and the Cover address in a body, and — for a Site behind a
// JavaScript challenge — how to read its payload from a cleared tab. One
// entry describes everything about one Site, and nowhere else gets to compare
// the site string.
type site struct {
// Host is the exact hostname a series_url for this Site must carry.
Host string
// LatestChapter finds the newest chapter in a fetched body.
LatestChapter func(seriesURL, body string) (latestChapter, bool)
// Cover finds the Cover address in a fetched body.
Cover func(seriesURL, body string) (string, bool)
// Rest is how long a Series of this Site rests between Polls.
Rest time.Duration
// Gap is the Lane's strictest pace: at least one second must pass between
// two consecutive Series-page Polls of this Site (issue #100).
Gap time.Duration
// Browser reads this Site's payload from a cleared browser tab; nil
// means the page is fetched over plain TLS.
Browser *browserRead
}
type browserRead struct {
// Read builds the tab read for seriesURL, refusing (false) an address
// this Site will not open in a browser — the per-Site half of the SSRF
// gate, kept deliberately behind fetchableSeriesURL: a headless browser
// executes JavaScript and carries cookies, and series_url is
// client-supplied.
Read func(seriesURL string, out *string) (chromedp.Action, bool)
// Done reports whether the payload arrived.
Done func(body string) bool
// Fallback allows the plain-TLS fetcher when no browser is configured.
// False skips the Site instead. kagane and comix are false — a plain fetch
// would only ever retrieve a challenge page — and novelfull is true,
// because its challenge is a live time-varying fact (AGENTS.md).
Fallback bool
}
// asuraSlugRe pulls the series slug out of a stored series_url.
// Shape verified live 2026-07-26: https://asurascans.com/comics/<slug>, where
// the slug carries a trailing build-hash suffix (e.g. "-f886a8af") that
@@ -22,6 +66,12 @@ type latestChapter struct {
// before using the slug to scope anything.
var asuraSlugRe = regexp.MustCompile(`/comics/([^/?#]+)`)
// asuraBuildHash matches the trailing "-xxxxxxxx" site-wide build ID Asura
// appends to every series slug. It rotates on each site redeploy, so it is
// never part of a stable series_id. Must stay in sync with stripBuildHash in
// userscript/manga-bookmark.user.js.
var asuraBuildHash = regexp.MustCompile(`-[0-9a-f]{8}$`)
// demonicChapterRe matches the pre-redirect anchors demonic series pages link
// through. Both the raw "&" and the HTML-escaped "&amp;" forms occur.
var demonicChapterRe = regexp.MustCompile(`chaptered\.php\?manga=\d+&(?:amp;)?chapter=([0-9.]+)`)
@@ -30,6 +80,18 @@ var demonicChapterRe = regexp.MustCompile(`chaptered\.php\?manga=\d+&(?:amp;)?ch
// Only the id prefix is stable; the slug tail follows the title.
var comixSlugRe = regexp.MustCompile(`/title/([^/?#]+)`)
func comixSeriesID(seriesURL string) (string, bool) {
m := comixSlugRe.FindStringSubmatch(seriesURL)
if m == nil {
return "", false
}
id := m[1]
if i := strings.Index(id, "-"); i != -1 {
id = id[:i]
}
return id, true
}
// kaganeChapterRe matches the chapter numbers in a kagane API response. This
// branch is fed by the browser fetcher, so the body is JSON rather than HTML —
// there are no anchors to scan.
@@ -40,91 +102,37 @@ var kaganeChapterRe = regexp.MustCompile(`"chapter_no":"([0-9.]+)"`)
// "/<slug>/chapter-<n>[-<title-slug>].html". Verified live 2026-08-05.
var novelfullSlugRe = regexp.MustCompile(`^/([^/?#]+)\.html$`)
// lnwSlugRe does the same for lightnovelworld, whose series pages live under
// /novel/<slug>/ while its chapter URLs are flat at the site root:
// "/<slug>-chapter-<n>/", absolute in the page's own anchors. Verified live
// 2026-08-05.
var lnwSlugRe = regexp.MustCompile(`^/novel/([^/?#]+)/?$`)
// lnwChapterRe matches any chapter-shaped address on lightnovelworld. Unlike
// asura, novelfull and comix — which scope to their stored series slug so a
// foreign chapter link cannot contribute — this Site's chapter addresses carry
// the Chapter Slug, which is not the Series identity: one Series may publish
// under several Chapter Slugs (measured 2026-08-11: a sampled novel serves
// 1-99 under one slug and 100-423 under another), so no stored-slug pattern can
// cover a Series' whole list. An unscoped match is safe because
// lnwLatestChapter truncates the body at the comment thread before scanning
// (lnwCommentMarker); without that, a visitor's comment could set the Latest
// Chapter on the shared Series row.
var lnwChapterRe = regexp.MustCompile(`lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/`)
// latestChapterFrom returns the highest chapter number body advertises for this
// series. ok is false when the body yields nothing usable — an unknown site, an
// empty body, a Cloudflare challenge page, and a site redesign all land here,
// and the caller treats all four identically.
//
// Ported from the userscript's latestChapterFromAnchors (asura L123-133,
// demonic L183-193), including its reason for taking a maximum rather than a
// first or last: neither site lists chapters in a dependable order.
// lnwCommentMarker is the boundary of lightnovelworld's server-rendered
// wpdiscuz comment thread. It occurs exactly once per page and follows every
// chapter anchor (measured 2026-08-11,
// docs/research/lightnovelworld-chapter-vs-series-slug.md §6), so cutting the
// body at its first occurrence keeps the whole chapter list while excluding a
// region any visitor can write to. Absent means the page shape changed: the
// body is skipped, never scanned whole.
const lnwCommentMarker = "wpd-threads"
// maxChapter returns the highest chapter number the regex finds in body. A
// maximum rather than a first or last, ported from the userscript's
// latestChapterFromAnchors (asura L123-133, demonic L183-193): neither site
// lists chapters in a dependable order.
//
// The userscript's asura rule additionally requires the anchor text to match
// /Chapter\s+[\d.]+/i. That check exists only to skip the "First Chapter"
// shortcut, which points at chapter/1 and therefore can never win a maximum, so
// it is redundant here. For asura, scoping the pattern to this series' own slug
// replaces it with a stronger guarantee: a chapter link belonging to some other
// series cannot contribute even if the page starts carrying them. demonic has no
// such guarantee — demonicChapterRe matches any chaptered.php?manga=<id> anchor
// with no per-series scoping, because the stored series_id for demonic is a
// slug, not the numeric id the URL carries, so it cannot easily be scoped.
func latestChapterFrom(site, seriesURL, body string) (latestChapter, bool) {
var re *regexp.Regexp
switch site {
case "asura":
m := asuraSlugRe.FindStringSubmatch(seriesURL)
if m == nil {
return latestChapter{}, false
}
// Stored URLs predating a redeploy may carry a stale build hash;
// chapter hrefs in the fetched body carry the current one. Strip to
// the stable ID (same rule as migrateAsuraKeys) and make the hash
// optional in the pattern, so scoping survives rotations.
slug := store.AsuraBuildHash.ReplaceAllString(m[1], "")
// Compiled per call rather than cached: this runs once per fetch, which
// is at most a few times a minute, and the slug varies per series.
re = regexp.MustCompile(`/comics/` + regexp.QuoteMeta(slug) + `(?:-[0-9a-f]{8})?/chapter/([0-9.]+)`)
case "demonic":
re = demonicChapterRe
case "comix":
m := comixSlugRe.FindStringSubmatch(seriesURL)
if m == nil {
return latestChapter{}, false
}
// comix ships an SPA: the served HTML carries a JSON state blob instead
// of chapter anchors, and latestChapterUrl is the only place the newest
// chapter appears. Scoping to this series' id prefix keeps a
// "recommended" strip's entries from winning the maximum.
id := m[1]
if i := strings.Index(id, "-"); i != -1 {
id = id[:i]
}
re = regexp.MustCompile(`"latestChapterUrl":"/title/` + regexp.QuoteMeta(id) + `-[^"]*-chapter-([0-9.]+)"`)
case "kagane":
re = kaganeChapterRe
case "novelfull":
u, err := url.Parse(seriesURL)
if err != nil {
return latestChapter{}, false
}
m := novelfullSlugRe.FindStringSubmatch(u.Path)
if m == nil {
return latestChapter{}, false
}
// Scoped to this series' slug for the same reason asura is: page 1
// carries a "latest chapters" widget and a "you may also like" strip,
// and neither may contribute to the maximum.
re = regexp.MustCompile(`/` + regexp.QuoteMeta(m[1]) + `/chapter-([0-9.]+)`)
case "lightnovelworld":
u, err := url.Parse(seriesURL)
if err != nil {
return latestChapter{}, false
}
m := lnwSlugRe.FindStringSubmatch(u.Path)
if m == nil {
return latestChapter{}, false
}
re = regexp.MustCompile(`lightnovelworld\.net/` + regexp.QuoteMeta(m[1]) + `-chapter-([0-9.]+)/`)
default:
return latestChapter{}, false
}
// shortcut, which points at chapter/1 and therefore can never win a maximum,
// so it is redundant once a maximum is taken.
func maxChapter(re *regexp.Regexp, body string) (latestChapter, bool) {
var best latestChapter
found := false
for _, m := range re.FindAllStringSubmatch(body, -1) {
@@ -142,3 +150,370 @@ func latestChapterFrom(site, seriesURL, body string) (latestChapter, bool) {
}
return best, found
}
// asuraLatestChapter scopes chapter links to this series' own slug, which
// replaces the userscript's anchor-text check with a stronger guarantee: a
// chapter link belonging to some other series cannot contribute even if the
// page starts carrying them.
func asuraLatestChapter(seriesURL, body string) (latestChapter, bool) {
m := asuraSlugRe.FindStringSubmatch(seriesURL)
if m == nil {
return latestChapter{}, false
}
// Stored URLs predating a redeploy may carry a stale build hash; chapter
// hrefs in the fetched body carry the current one. Strip to the stable ID
// and make the hash optional in the pattern, so scoping survives
// rotations.
slug := asuraBuildHash.ReplaceAllString(m[1], "")
// Compiled per call rather than cached: this runs once per fetch, which is
// at most a few times a minute, and the slug varies per series.
re := regexp.MustCompile(`/comics/` + regexp.QuoteMeta(slug) + `(?:-[0-9a-f]{8})?/chapter/([0-9.]+)`)
return maxChapter(re, body)
}
// demonicLatestChapter is not scoped: demonicChapterRe matches any
// chaptered.php?manga=<id> anchor, because the stored series_id is a slug,
// not the numeric id the URL carries, so it cannot be scoped.
func demonicLatestChapter(_, body string) (latestChapter, bool) {
return maxChapter(demonicChapterRe, body)
}
// comixLatestChapter reads comix's SPA: the served HTML carries a JSON state
// blob instead of chapter anchors, and latestChapterUrl is the only place the
// newest chapter appears. Scoping to this series' id prefix keeps a
// "recommended" strip's entries from winning the maximum.
func comixLatestChapter(seriesURL, body string) (latestChapter, bool) {
id, ok := comixSeriesID(seriesURL)
if !ok {
return latestChapter{}, false
}
re := regexp.MustCompile(`"latestChapterUrl":"/title/` + regexp.QuoteMeta(id) + `-[^"]*-chapter-([0-9.]+)"`)
return maxChapter(re, body)
}
// kaganeLatestChapter scans the kagane series API JSON that the browser read
// fetched from inside the page; the match rides on the property name,
// regardless of the surrounding JSON shape.
func kaganeLatestChapter(_, body string) (latestChapter, bool) {
return maxChapter(kaganeChapterRe, body)
}
// novelfullLatestChapter is scoped to this series' slug for the same reason
// asura is: page 1 carries a "latest chapters" widget and a "you may also
// like" strip, and neither may contribute to the maximum.
func novelfullLatestChapter(seriesURL, body string) (latestChapter, bool) {
u, err := url.Parse(seriesURL)
if err != nil {
return latestChapter{}, false
}
m := novelfullSlugRe.FindStringSubmatch(u.Path)
if m == nil {
return latestChapter{}, false
}
re := regexp.MustCompile(`/` + regexp.QuoteMeta(m[1]) + `/chapter-([0-9.]+)`)
return maxChapter(re, body)
}
// lnwLatestChapter truncates the body at the comment thread before scanning:
// it is the one region of the page any visitor can write to (see lnwChapterRe).
// A body without the marker is skipped, never scanned whole — a redesign must
// degrade into staleness, not into a wrong shared value; the logged body length
// tells a markup change from a body the size cap cut short.
func lnwLatestChapter(seriesURL, body string) (latestChapter, bool) {
i := strings.Index(body, lnwCommentMarker)
if i < 0 {
log.Printf("latest poll %q: no %s marker in %d bytes", seriesURL, lnwCommentMarker, len(body))
return latestChapter{}, false
}
return maxChapter(lnwChapterRe, body[:i])
}
// latestChapterFrom returns the highest chapter number body advertises for this
// series, via the Site's registry entry. ok is false when the body yields
// nothing usable — an unknown site, an empty body, a Cloudflare challenge page,
// and a site redesign all land here, and the caller treats all four identically.
func latestChapterFrom(site, seriesURL, body string) (latestChapter, bool) {
if fn := sites[site].LatestChapter; fn != nil {
return fn(seriesURL, body)
}
return latestChapter{}, false
}
var metaTagRe = regexp.MustCompile(`(?is)<meta\b[^>]*>`)
var doubleQuotedMetaAttrRe = regexp.MustCompile(`(?is)([a-z][a-z0-9:_-]*)\s*=\s*"([^"]*)"`)
var singleQuotedMetaAttrRe = regexp.MustCompile(`(?is)([a-z][a-z0-9:_-]*)\s*=\s*'([^']*)'`)
// comix's server-rendered page embeds query data in this JSON script; parsing
// the target detail entry avoids matching posters from recommended results.
var comixInitialDataRe = regexp.MustCompile(`(?is)<script\b[^>]*\bid\s*=\s*["']initial-data["'][^>]*>(.*?)</script>`)
// kaganeImageURLRe matches the canonical compressed image route kagane's API
// publishes — the only cover URL form the extractor emits and the browser
// fetcher accepts. The URL is matched in full (scheme, host, id shape) rather
// than trusted: the value a fetcher is pointed at may have been client-
// supplied, and a headless browser is a strong SSRF primitive.
var kaganeImageURLRe = regexp.MustCompile(`^https://kagane\.to/api/v2/image/([0-9a-f-]{36})/compressed$`)
// comixImageURLRe matches comix's cover host and path shape. Pinned in full
// (scheme, host, path characters, image extension) for the same reason
// kaganeImageURLRe is: the address reaches a headless browser, and it can
// originate in a client-supplied PUT body. No dot is allowed inside the path,
// so no traversal or second extension can hide in it. Shape from a live page,
// 2026-08-10: /039d/i/1/34/6a6742bf15736@280.jpg.
var comixImageURLRe = regexp.MustCompile(`^https://static\.comix\.to/[A-Za-z0-9@/_-]+\.(?:jpg|jpeg|png|webp)$`)
// browserOnlyCoverURL reports whether the browser sidecar is the only fetcher
// for cover bytes at imageURL. kagane's image route answers a plain fetch with
// a challenge and `cross-origin-resource-policy: same-origin`, and
// static.comix.to answers one with the same Cloudflare challenge its pages
// serve (measured 2026-08-12, issue #98), so a TLS fetch would only ever
// retrieve a challenge page and must not be attempted (ADR-0007). This is the
// byte-fetch router's per-Site knowledge; it lives in the extraction module,
// which owns those URL shapes.
func browserOnlyCoverURL(imageURL string) bool {
return kaganeImageURLRe.MatchString(imageURL) ||
comixImageURLRe.MatchString(imageURL)
}
// kagane's browser-fetched series response publishes cover image IDs under
// series_covers. The API's canonical compressed image route is the only URL
// form accepted by the store and browser fetcher; no rendition is guessed.
func kaganeCoverURL(body string) string {
var response struct {
SeriesCovers []struct {
ImageID string `json:"image_id"`
} `json:"series_covers"`
}
if err := json.Unmarshal([]byte(body), &response); err != nil {
return ""
}
for _, cover := range response.SeriesCovers {
// Validate the assembled URL against the same regex the browser
// fetcher enforces, so the extractor can never emit an address the
// fetch would refuse.
imageURL := "https://kagane.to/api/v2/image/" + cover.ImageID + "/compressed"
if kaganeImageURLRe.MatchString(imageURL) {
return imageURL
}
}
return ""
}
func comixCoverURL(seriesURL, body string) string {
id, ok := comixSeriesID(seriesURL)
if !ok {
return ""
}
data := comixInitialDataRe.FindStringSubmatch(body)
if data == nil {
return ""
}
var state struct {
Queries map[string]json.RawMessage `json:"queries"`
}
if err := json.Unmarshal([]byte(data[1]), &state); err != nil {
return ""
}
raw := state.Queries[`["manga","detail","`+id+`"]`]
if len(raw) == 0 {
return ""
}
var detail struct {
Poster struct {
Medium string `json:"medium"`
} `json:"poster"`
}
if err := json.Unmarshal(raw, &detail); err != nil {
return ""
}
return publishedCoverURL(detail.Poster.Medium)
}
// ogImageCover reads the og:image metadata shared by asura, demonic and
// lightnovelworld.
func ogImageCover(_, body string) (string, bool) {
cover := metaContent(body, "property", "og:image")
return cover, cover != ""
}
func novelfullCoverEntry(_, body string) (string, bool) {
cover := metaContent(body, "name", "image")
return cover, cover != ""
}
func comixCoverEntry(seriesURL, body string) (string, bool) {
cover := comixCoverURL(seriesURL, body)
return cover, cover != ""
}
func kaganeCoverEntry(_, body string) (string, bool) {
cover := kaganeCoverURL(body)
return cover, cover != ""
}
// coverFrom reports false for unknown sites, challenge bodies, and pages with
// no usable cover, via the Site's registry entry.
func coverFrom(site, seriesURL, body string) (string, bool) {
if fn := sites[site].Cover; fn != nil {
return fn(seriesURL, body)
}
return "", false
}
// metaContent returns the content of the first <meta> whose attrName is
// attrValue. It keeps scanning after an empty match so a later published cover
// is not hidden by an empty tag.
func metaContent(body, attrName, attrValue string) string {
for _, tag := range metaTagRe.FindAllString(body, -1) {
attrs := make(map[string]string)
for _, m := range doubleQuotedMetaAttrRe.FindAllStringSubmatch(tag, -1) {
attrs[strings.ToLower(m[1])] = m[2]
}
for _, m := range singleQuotedMetaAttrRe.FindAllStringSubmatch(tag, -1) {
attrs[strings.ToLower(m[1])] = m[2]
}
if strings.EqualFold(attrs[strings.ToLower(attrName)], attrValue) {
if cover := publishedCoverURL(attrs["content"]); cover != "" {
return cover
}
}
}
return ""
}
func publishedCoverURL(value string) string {
value = strings.TrimSpace(html.UnescapeString(value))
return strings.ReplaceAll(value, " ", "%20")
}
// Poll Lane constants (issue #100). The per-Site structure is deliberately
// uniform at first — every Site rests an hour and gaps ten seconds — but it
// exists so a single Site can be slowed if it turns hostile, and the numbers
// stay in the registry so the structure has a place to differ.
const (
// defaultRest is how long every Series rests between Polls.
defaultRest = time.Hour
// defaultGap is the strictest pace of every Lane unless the eligible
// Series count forces it tighter.
defaultGap = 10 * time.Second
// minGap floors the effective gap. One request per second is already an
// order of magnitude past the strictest rate rule a free-plan Site can
// express (docs/research/cloudflare-bot-scoring-and-poll-cadence.md);
// below it the Lane is outrunning its own plan and says so loudly.
minGap = time.Second
// refuseBackoff is how long a Lane waits after its Site refused twice in
// one run before attempting it again.
refuseBackoff = 15 * time.Minute
// browserWakeCount and browserWakeAge gate a browser Lane's run: five or
// more due Series, or any one of them waiting this long, or Chrome stays
// asleep (ADR-0005 on-demand browser).
browserWakeCount = 5
browserWakeAge = 15 * time.Minute
)
// effectiveGap is a Site's pace: the registry gap, or one rest divided by the
// eligible Series count when that is smaller, never below one second. The
// denominator follows defaultRest rather than a literal hour so a Site whose
// rest is ever changed keeps its per-Series pace in step. The second return is
// true when the one-second floor engaged (and the Lane logs a warning naming
// the Site, every round it does).
func effectiveGap(s site, eligible int) (time.Duration, bool) {
gap := s.Gap
if eligible > 0 {
if perSeries := defaultRest / time.Duration(eligible); perSeries < gap {
gap = perSeries
}
}
if gap < minGap {
return minGap, true
}
return gap, false
}
// sites is the registry: one entry per Site, keyed by the stored site string.
// Adding a Site means adding an entry here and nowhere else — the dispatch
// functions above and the poller's route list are lookups into this map. An
// unknown site string resolves to the zero entry, which fails the existing
// not-fetchable and no-fetcher paths unchanged.
var sites = map[string]site{
"asura": {
Host: "asurascans.com",
LatestChapter: asuraLatestChapter,
Cover: ogImageCover,
Rest: defaultRest,
Gap: defaultGap,
},
"demonic": {
Host: "demonicscans.org",
LatestChapter: demonicLatestChapter,
Cover: ogImageCover,
Rest: defaultRest,
Gap: defaultGap,
},
"comix": {
Host: "comix.to",
LatestChapter: comixLatestChapter,
Cover: comixCoverEntry,
Rest: defaultRest,
Gap: defaultGap,
Browser: &browserRead{
Read: comixRead,
// The interstitial is served in place of the page, so "arrived"
// has to exclude it explicitly, as novelfull's does.
Done: func(body string) bool { return body != "" && !isInterstitial(body) },
// Never falls back: a plain fetch of a comix page or cover
// retrieves only a challenge page (measured 2026-08-12).
Fallback: false,
},
},
"kagane": {
Host: "kagane.to",
LatestChapter: kaganeLatestChapter,
Cover: kaganeCoverEntry,
Rest: defaultRest,
Gap: defaultGap,
Browser: &browserRead{
Read: kaganeRead,
Done: func(body string) bool { return body != "" },
// Never falls back: a plain fetch of a kagane page or cover would
// only ever retrieve a challenge page (verified 2026-08-03).
Fallback: false,
},
},
"novelfull": {
Host: "novelfull.com",
LatestChapter: novelfullLatestChapter,
Cover: novelfullCoverEntry,
Rest: defaultRest,
Gap: defaultGap,
Browser: &browserRead{
Read: novelfullRead,
// The interstitial has a DOM too, so "the payload arrived" has to
// exclude it explicitly.
Done: func(body string) bool { return body != "" && !isInterstitial(body) },
Fallback: true,
},
},
"lightnovelworld": {
Host: "lightnovelworld.net",
LatestChapter: lnwLatestChapter,
Cover: ogImageCover,
Rest: defaultRest,
Gap: defaultGap,
},
}
// browserBackedSites is derived from the registry: the Sites whose pages are
// read through the browser sidecar. Sorted so callers that range it (the
// browser fetcher's dispatch) see a stable order instead of map-iteration
// noise.
func browserBackedSites() []string {
out := make([]string, 0, len(sites))
for name, s := range sites {
if s.Browser != nil {
out = append(out, name)
}
}
sort.Strings(out)
return out
}
+254 -16
View File
@@ -1,6 +1,9 @@
package latest
import "testing"
import (
"strings"
"testing"
)
// Trimmed from https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af
// fetched 2026-07-26. The first anchor is the "First Chapter" shortcut: it is a
@@ -38,6 +41,13 @@ const challengeFixture = `<!DOCTYPE html><html><head><title>Just a moment...</ti
// https://comix.to/title/n8we-dungeons-and-crayons fetched 2026-08-03. comix is
// an SPA: the page ships a JSON state blob rather than a list of chapter
// anchors, and latestChapterUrl is where the newest chapter actually lives.
//
// Still the right fixture after comix moved behind the challenge (#98): the
// browser read is an in-tab fetch of the Series URL, so the body a poll parses
// is this same server-rendered HTML, not a rendered DOM. Confirmed against a
// live cleared tab 2026-08-16 (TestSmokeComix): the in-tab fetch returned
// 24793 bytes of server-rendered HTML that these same parses read a chapter
// and a cover out of.
const comixSeriesFixture = `
{"firstChapterUrl":"/title/n8we-dungeons-and-crayons/5038739-chapter-1","latestChapterUrl":"/title/n8we-dungeons-and-crayons/11139891-chapter-80"},
{""manga","recommended","n8we",1]":{"items":[{"latestChapterUrl":"/title/qqwrm-full-time-awakening/99999999-chapter-999"}]}
@@ -69,16 +79,230 @@ const novelfullSeriesFixture = `
<a href="/release-that-witch/chapter-9999.html">Chapter 9999</a>
`
// Trimmed from https://lightnovelworld.net/novel/a-will-eternal/ fetched
// 2026-08-05. Its chapter anchors are absolute and flat — /<slug>-chapter-<n>/
// at the site root, not under /novel/. The last anchor is another series'.
// Trimmed from
// https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/
// fetched 2026-08-11, whole document (~307 KB decoded; the wire body is ~32 KB
// zstd-compressed). This
// novel publishes its chapters under two Chapter Slugs: 1–99 at
// …-not-chapter-<n>/ and 100–423 at …-not-them-all-chapter-<n>/, and the page
// lists them newest-first, so the …-not-them-all anchors precede the …-not
// anchors. Every anchor through the comment-thread marker is verbatim page
// text (the site renders this novel's chapter titles as "[ ... words ]"). The
// comment block after the marker is the real wpdiscuz comment #wpd-comm-358_0
// from https://lightnovelworld.net/novel/the-sword-illuminates-the-great-wilderness/
// — the pinned page serves zero comments — with its share/link/vote/reply
// boilerplate trimmed. The comment's body carried no link, so the bare <a
// href> to https://lightnovelworld.net/overgeared-chapter-2059/ inside
// wpd-comment-text is the one composed element; that URL is a real chapter of
// a real different novel (overgeared; fetched, HTTP 200).
const lnwSeriesFixture = `
<a href="https://lightnovelworld.net/a-will-eternal-chapter-1/">Chapter 1</a>
<a href="https://lightnovelworld.net/a-will-eternal-chapter-1317/">Chapter 1317</a>
<a href="https://lightnovelworld.net/a-will-eternal-chapter-1298/">Chapter 1298</a>
<a href="https://lightnovelworld.net/overgeared-chapter-9999/">Chapter 9999</a>
<li data-ID="102741">
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-404/">
<div class="epl-num">Vol. 1 Ch. 404</div>
<div class="epl-title">[ ... words ]</div>
<div class="epl-date">April 12, 2026</div>
</a>
</li>
<li data-ID="102780">
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-423/">
<div class="epl-num">Vol. 1 Ch. 423</div>
<div class="epl-title">[ ... words ]</div>
<div class="epl-date">April 7, 2026</div>
</a>
</li>
<li data-ID="102527">
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-300/">
<div class="epl-num">Vol. 1 Ch. 300</div>
<div class="epl-title">[ ... words ]</div>
<div class="epl-date">April 4, 2026</div>
</a>
</li>
<li data-ID="102325">
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-200/">
<div class="epl-num">Vol. 1 Ch. 200</div>
<div class="epl-title">[ ... words ]</div>
<div class="epl-date">March 29, 2026</div>
</a>
</li>
<li data-ID="102121">
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-100/">
<div class="epl-num">Vol. 1 Ch. 100</div>
<div class="epl-title">[ ... words ]</div>
<div class="epl-date">March 22, 2026</div>
</a>
</li>
<li data-ID="26014">
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-chapter-99/">
<div class="epl-num">Vol. 1 Ch. 99</div>
<div class="epl-title">Chapter 99</div>
<div class="epl-date">November 5, 2025</div>
</a>
</li>
<li data-ID="25916">
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-chapter-50/">
<div class="epl-num">Vol. 1 Ch. 50</div>
<div class="epl-title">Chapter 50</div>
<div class="epl-date">October 29, 2025</div>
</a>
</li>
<li class='tseplsfrst' data-ID="25818">
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-chapter-1/">
<div class="epl-num">Vol. 1 Ch. 1</div>
<div class="epl-title">Chapter 01</div>
<div class="epl-date">October 11, 2025</div>
</a>
</li>
<div id="wpd-threads" class="wpd-thread-wrapper">
<div class="wpd-thread-list">
<div id='wpd-comm-358_0' class='comment byuser comment-author-jimbear even thread-even depth-1 wpd-comment wpd_comment_level-1'><div class="wpd-comment-wrap wpd-blog-user wpd-blog-subscriber">
<div class="wpd-comment-left ">
<div class="wpd-avatar ">
<img alt='hasbi asy' src='https://secure.gravatar.com/avatar/3c1792cb31cab842f90e7c463f0948e98536cf938aebbd2f678f937ca60fb799?s=64&#038;d=mm&#038;r=g' srcset='https://secure.gravatar.com/avatar/3c1792cb31cab842f90e7c463f0948e98536cf938aebbd2f678f937ca60fb799?s=128&#038;d=mm&#038;r=g 2x' class='avatar avatar-64 photo' height='64' width='64' decoding='async'/>
</div>
<div class="wpd-comment-label" wpd-tooltip="Member" wpd-tooltip-position="right">
<span>Member</span>
</div>
</div>
<div id="comment-358" class="wpd-comment-right">
<div class="wpd-comment-header">
<div class="wpd-comment-author ">
hasbi asy
</div>
<div class="wpd-comment-date" title="July 9, 2026 1:44 am">
<i class='far fa-clock' aria-hidden='true'></i>
1 month ago
</div>
</div>
<div class="wpd-comment-text">
<p>where&#8217;s everyone</p>
<a href="https://lightnovelworld.net/overgeared-chapter-2059/">https://lightnovelworld.net/overgeared-chapter-2059/</a>
</div>
</div>
</div>
<div id='wpdiscuz_form_anchor-358_0'></div>
</div>
</div>
`
// Trimmed from https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af
// (redirected to ...-00dcbf97) on 2026-08-10.
const asuraCoverFixture = `<meta property="og:image" content="https://cdn.asurascans.com/asura-images/covers/chronicles-of-the-demon-faction.d4dcb8.webp">`
// Trimmed from https://demonicscans.org/manga/Catastrophic-Necromancer on 2026-08-10.
// The source publishes the raw space in this URL.
const demonicCoverFixture = `<meta property="og:image" content="https://readermc.org/images/thumbnails/Catastrophic Necromancer.webp">`
// Trimmed from https://comix.to/title/n8we-dungeons-and-crayons on 2026-08-10.
// The state includes a recommended poster before the target detail object and
// nested IDs inside that object; no og:image is present.
const comixCoverFixture = `<script type="application/json" id="initial-data">{"queries":{"[\"manga\",\"recommended\",\"n8we\",1]":{"poster":{"medium":"https://static.comix.to/recommended@280.jpg","large":"https://static.comix.to/recommended.jpg"}},"[\"manga\",\"detail\",\"n8we\"]":{"chapters":[{"hid":"nested"}],"poster":{"medium":"https://static.comix.to/039d/i/1/34/6a6742bf15736@280.jpg","large":"https://static.comix.to/039d/i/1/34/6a6742bf15736.jpg"}}}}</script>`
// Trimmed from GET https://kagane.to/api/v2/series/019fe11a-8670-7cf3-8343-0b02057d3787 on 2026-08-10.
const kaganeCoverFixture = `{"series_covers":[{"cover_id":"019fe11a-84d1-714b-9cf4-2827f277f3c0","language":"en","volume_number":"1","chapter_number":null,"note":null,"image_id":"019fe11a-84c3-7fc3-a84b-88787374b617"}]}`
// Trimmed from https://novelfull.com/reverend-insanity.html on 2026-08-10.
const novelfullCoverFixture = `<meta name="image" content="https://novelfull.com/uploads/webp/novel/reverend-insanity-82661d911a.webp">`
// Trimmed from
// https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/
// on 2026-08-11.
const lnwCoverFixture = `<meta property="og:image" content="https://i1.wp.com/lightnovelworld.net/wp-content/uploads/2025/10/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all.jpg" />`
func TestCoverFrom(t *testing.T) {
const comixURL = "https://comix.to/title/n8we-dungeons-and-crayons"
tests := []struct {
name string
site string
seriesURL string
body string
wantOK bool
wantCover string
}{
{
name: "asura uses published metadata URL",
site: "asura", body: asuraCoverFixture, wantOK: true,
wantCover: "https://cdn.asurascans.com/asura-images/covers/chronicles-of-the-demon-faction.d4dcb8.webp",
},
{
name: "demonic escapes raw spaces",
site: "demonic", body: demonicCoverFixture, wantOK: true,
wantCover: "https://readermc.org/images/thumbnails/Catastrophic%20Necromancer.webp",
},
{
name: "comix takes target medium poster",
site: "comix", seriesURL: comixURL, body: comixCoverFixture, wantOK: true,
wantCover: "https://static.comix.to/039d/i/1/34/6a6742bf15736@280.jpg",
},
{
name: "kagane reads API cover image ID",
site: "kagane", body: kaganeCoverFixture, wantOK: true,
wantCover: "https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
},
{
name: "novelfull reads image metadata",
site: "novelfull", body: novelfullCoverFixture, wantOK: true,
wantCover: "https://novelfull.com/uploads/webp/novel/reverend-insanity-82661d911a.webp",
},
{
name: "lightnovelworld reads og image",
site: "lightnovelworld", body: lnwCoverFixture, wantOK: true,
wantCover: "https://i1.wp.com/lightnovelworld.net/wp-content/uploads/2025/10/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all.jpg",
},
{
name: "later metadata cover survives empty match",
site: "asura",
body: `<meta property="og:image" content="">` + asuraCoverFixture,
wantOK: true,
wantCover: "https://cdn.asurascans.com/asura-images/covers/chronicles-of-the-demon-faction.d4dcb8.webp",
},
{
name: "page without cover is empty",
site: "asura", body: `<meta property="og:title" content="No Cover">`,
},
{
name: "unknown site is empty",
site: "unknown", body: asuraCoverFixture,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
got, ok := coverFrom(tt.site, tt.seriesURL, tt.body)
if ok != tt.wantOK {
t.Fatalf("ok = %v, want %v (got %q)", ok, tt.wantOK, got)
}
if got != tt.wantCover {
t.Errorf("cover = %q, want %q", got, tt.wantCover)
}
})
}
}
func TestCoverFromChallenge(t *testing.T) {
tests := []struct {
site string
seriesURL string
}{
{"asura", "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af"},
{"demonic", "https://demonicscans.org/manga/Catastrophic-Necromancer"},
{"comix", "https://comix.to/title/n8we-dungeons-and-crayons"},
{"kagane", "https://kagane.to/series/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"},
{"novelfull", "https://novelfull.com/reverend-insanity.html"},
{"lightnovelworld", "https://lightnovelworld.net/novel/a-will-eternal/"},
}
for _, tt := range tests {
t.Run(tt.site, func(t *testing.T) {
if got, ok := coverFrom(tt.site, tt.seriesURL, challengeFixture); ok || got != "" {
t.Fatalf("cover = %q, ok = %v, want empty", got, ok)
}
})
}
}
func TestLatestChapterFrom(t *testing.T) {
const asuraURL = "https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af"
const demonicURL = "https://demonicscans.org/manga/Catastrophic-Necromancer"
@@ -194,24 +418,38 @@ func TestLatestChapterFrom(t *testing.T) {
body: novelfullSeriesFixture,
wantOK: false,
},
// Stored before the slug split, so the address carries the ...-not
// Chapter Slug; the 100-423 block under the other slug must still win.
// The body is the chapter-list portion of lnwSeriesFixture with the
// comment block omitted; the marker is kept, because a body without it
// is skipped, not scanned.
{
name: "lightnovelworld takes the max and ignores another series",
name: "lightnovelworld max spans both chapter slugs",
site: "lightnovelworld",
seriesURL: "https://lightnovelworld.net/novel/a-will-eternal/",
body: lnwSeriesFixture,
wantOK: true, wantNum: 1317, wantLabel: "Chapter 1317",
seriesURL: "https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not/",
body: strings.SplitN(lnwSeriesFixture, lnwCommentMarker, 2)[0] + lnwCommentMarker,
wantOK: true, wantNum: 423, wantLabel: "Chapter 423",
},
{
name: "lightnovelworld tolerates a series url with no trailing slash",
name: "lightnovelworld comment anchor cannot set the latest chapter",
site: "lightnovelworld",
seriesURL: "https://lightnovelworld.net/novel/a-will-eternal",
seriesURL: "https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/",
body: lnwSeriesFixture,
wantOK: true, wantNum: 1317, wantLabel: "Chapter 1317",
wantOK: true, wantNum: 423, wantLabel: "Chapter 423",
},
// Marker removed from the fixture, comment block still present: a
// redesign must degrade into a skip, never into the comment's number.
{
name: "lightnovelworld body without the comment marker is skipped",
site: "lightnovelworld",
seriesURL: "https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/",
body: strings.ReplaceAll(lnwSeriesFixture, lnwCommentMarker, ""),
wantOK: false,
},
{
name: "lightnovelworld yields nothing on a challenge page",
site: "lightnovelworld",
seriesURL: "https://lightnovelworld.net/novel/a-will-eternal/",
seriesURL: "https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/",
body: challengeFixture,
wantOK: false,
},
@@ -0,0 +1,72 @@
package latest
import (
"context"
"os"
"testing"
"time"
"bookmarkmanager/backend/internal/store"
)
// TestSmokeComix answers "is comix's challenge clearing from this browser right
// now" — a live, time-varying fact, so a red run is something to re-check
// before it is a defect. Needs the real browser unit with outbound network:
//
// cd chrome && BROWSER_BIND_ADDR=127.0.0.1 docker compose up -d --build
// SMOKE_BROWSER_WS_URL=ws://127.0.0.1:9222 go test -run TestSmokeComix ./internal/latest
//
// It walks the whole read: the in-tab page fetch, both parses, and the Cover
// bytes by direct navigation to static.comix.to. The Cover address comes out of
// the page rather than being pinned in the test, because a stored one rots.
func TestSmokeComix(t *testing.T) {
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
if ws == "" {
t.Skip("SMOKE_BROWSER_WS_URL unset")
}
const seriesURL = "https://comix.to/title/m12d-classmate"
f, err := NewBrowserFetcher(ws)
if err != nil {
t.Fatalf("NewBrowserFetcher: %v", err)
}
defer f.Close()
ctx, cancel := context.WithTimeout(context.Background(), 120*time.Second)
defer cancel()
body, status, err := f.Get(ctx, seriesURL)
if err != nil {
t.Fatalf("Get: %v", err)
}
t.Logf("status=%d bytes=%d", status, len(body))
if status != 200 {
t.Fatalf("status = %d, want 200 — the sidecar is not clearing the challenge", status)
}
chapter, ok := latestChapterFrom("comix", seriesURL, body)
if !ok {
t.Fatalf("no latest chapter in %d bytes — page shape changed", len(body))
}
t.Logf("latest chapter: %v %q", chapter.Num, chapter.Label)
cover, ok := coverFrom("comix", seriesURL, body)
if !ok {
t.Fatalf("no cover address in %d bytes — page shape changed", len(body))
}
t.Logf("cover: %s", cover)
if !browserOnlyCoverURL(cover) {
t.Fatalf("cover %q is not claimed by the browser gate: the pin and the live URL shape disagree", cover)
}
bytes, contentType, err := f.Image(ctx, cover)
if err != nil {
t.Fatalf("Image: %v", err)
}
if len(bytes) < 1000 {
t.Fatalf("cover is %d bytes, want a real image", len(bytes))
}
t.Logf("fetched %d bytes of %s", len(bytes), contentType)
if _, ok := store.CoverContentType(contentType); !ok {
t.Fatalf("content type %q is not storable", contentType)
}
}
+157
View File
@@ -0,0 +1,157 @@
package latest
import (
"context"
"net/http"
"os"
"testing"
"time"
"bookmarkmanager/backend/internal/store"
)
// TestSmokeKaganeImage is the live proof that the acquisition path's browser
// fetch actually clears Cloudflare and returns image bytes. It needs the real
// browser unit with outbound network, so it runs only when SMOKE_BROWSER_WS_URL
// is set:
//
// cd chrome && BROWSER_BIND_ADDR=127.0.0.1 docker compose up -d --build
// SMOKE_BROWSER_WS_URL=ws://127.0.0.1:9222 go test -run TestSmokeKaganeImage ./internal/latest
//
// Not chromedp/headless-shell: its challenge never clears (see chrome/Dockerfile),
// so a red run there proves nothing about kagane.
func TestSmokeKaganeImage(t *testing.T) {
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
if ws == "" {
t.Skip("SMOKE_BROWSER_WS_URL unset")
}
const imageURL = "https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed" // SP Baby's cover
// The same URL through a plain client is what any other fetcher would get.
// Asserting on it keeps the test honest about why the browser is needed.
req, err := http.NewRequest(http.MethodGet, imageURL, nil)
if err != nil {
t.Fatal(err)
}
if res, err := (&http.Client{Timeout: 15 * time.Second}).Do(req); err == nil {
res.Body.Close()
if res.StatusCode == http.StatusOK {
t.Log("note: kagane answered a plain request 200 — the challenge is not up right now")
}
}
f, err := NewBrowserFetcher(ws)
if err != nil {
t.Fatalf("NewBrowserFetcher: %v", err)
}
defer f.Close()
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
defer cancel()
body, contentType, err := f.Image(ctx, imageURL)
if err != nil {
t.Fatalf("Image: %v", err)
}
if len(body) < 1000 {
t.Fatalf("body is %d bytes, want a real image", len(body))
}
if contentType != "image/webp" {
t.Fatalf("content type = %q, want image/webp", contentType)
}
// WebP files start with "RIFF....WEBP".
if string(body[:4]) != "RIFF" || string(body[8:12]) != "WEBP" {
t.Fatalf("body is not a WebP: % x", body[:12])
}
t.Logf("fetched %d bytes of %s", len(body), contentType)
// The browser module claims only the cover URL shape it can clear a
// challenge for; anything else must be refused before any navigation.
if _, _, err := f.Image(ctx, "https://cdn.example/cover.jpg"); err == nil {
t.Fatal("Image accepted a cover URL the browser module does not claim")
}
}
// Control for the test above: the poller's own kagane path, same sidecar. If
// this fails too, the sidecar is not clearing the challenge at all and the
// image result says nothing about Image itself.
func TestSmokeKaganeGet(t *testing.T) {
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
if ws == "" {
t.Skip("SMOKE_BROWSER_WS_URL unset")
}
f, err := NewBrowserFetcher(ws)
if err != nil {
t.Fatalf("NewBrowserFetcher: %v", err)
}
defer f.Close()
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
defer cancel()
body, status, err := f.Get(ctx, "https://kagane.to/series/019fe11a-8670-7cf3-8343-0b02057d3787")
if err != nil {
t.Fatalf("Get: %v", err)
}
t.Logf("status=%d bytes=%d head=%.80q", status, len(body), body)
if status != 200 {
t.Fatalf("status = %d, want 200 — the sidecar is not clearing the challenge", status)
}
}
// TestSmokeAcquireKaganeCover proves the #62 acquisition path end to end
// against the real browser: a kagane Series bookmarked at creation gets its
// Cover, bytes fetched through the sidecar into the content-addressed store.
// Same SMOKE_BROWSER_WS_URL gate as the tests above; a red run means the
// challenge is not clearing from this IP (a live fact to re-check), not
// necessarily a defect in the pipeline.
func TestSmokeAcquireKaganeCover(t *testing.T) {
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
if ws == "" {
t.Skip("SMOKE_BROWSER_WS_URL unset")
}
const (
seriesID = "019fe11a-8670-7cf3-8343-0b02057d3787"
coverURL = "https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed"
)
s, _ := newTestStore(t)
bf, err := NewBrowserFetcher(ws)
if err != nil {
t.Fatalf("NewBrowserFetcher: %v", err)
}
defer bf.Close()
tlsF, err := NewTLSFetcher()
if err != nil {
t.Fatalf("NewTLSFetcher: %v", err)
}
acq := &Acquirer{
Store: s, Fetch: tlsF, BrowserFetch: bf,
BrowserCoverFetch: bf, Covers: NewCoverFetcher(),
}
s.OnSeriesCreated = acq.Acquire
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: "kagane:" + seriesID, Site: "kagane", SeriesID: seriesID,
Title: "smoke", SeriesURL: "https://kagane.to/series/" + seriesID, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("Upsert: %v", err)
}
acq.Wait()
got, found, err := s.Get(s.OwnerID(), "kagane:"+seriesID)
if err != nil || !found {
t.Fatalf("Get: %v found=%v", err, found)
}
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(coverURL); got.Cover != want {
t.Fatalf("Cover = %q, want %q — the acquire path did not store the browser-fetched bytes", got.Cover, want)
}
body, contentType, ok, err := s.CoverByAddress(store.CoverAddress(coverURL))
if err != nil || !ok {
t.Fatalf("CoverByAddress: %v found=%v", err, ok)
}
if len(body) < 1000 {
t.Fatalf("stored cover is %d bytes, want a real image", len(body))
}
if contentType != "image/webp" {
t.Fatalf("content type = %q, want image/webp", contentType)
}
t.Logf("stored %d bytes of %s", len(body), contentType)
}
+101
View File
@@ -0,0 +1,101 @@
package latest
import (
"context"
"fmt"
"net/http"
"os"
"strings"
"testing"
"time"
)
// lnwSeriesPageFloor is the smallest body that can still be a whole
// lightnovelworld Series page. Whole pages measured 685 KB..1.18 MB on
// 2026-08-11 and carry the marker at ~94% of the document, so a body under
// 100 KB is a challenge, a notice, or a truncated read — asserting on it
// would report the marker missing when it was never fetched.
const lnwSeriesPageFloor = 100 << 10
// TestSmokeLnwCommentBoundary is the live proof that the comment-thread marker
// the lightnovelworld chapter scan truncates at (lnwCommentMarker,
// "wpd-threads") still holds on the Site. The scan depends on it: when the
// marker vanishes every Series is skipped and logged — correct, but silent
// until a Reader notices their Latest Chapter has stopped moving. It runs only
// when SMOKE_LNW_SERIES_URL is set — the URL of the live Series page to check.
// The immortality-simulator page measured 2026-08-11
// (docs/research/lightnovelworld-chapter-vs-series-slug.md) is the default to
// point it at:
//
// SMOKE_LNW_SERIES_URL=https://lightnovelworld.net/novel/immortality-simulator/ go test -v -run TestSmokeLnwCommentBoundary ./internal/latest
//
// A red run means the Site's markup has moved — the marker is gone, occurs
// more than once, or no longer follows the last chapter anchor — and the scan
// in sites.go is now skipping this Site. Revisit sites.go before anything
// else; the test is not flaky. A Cloudflare challenge or a non-200 is
// distinguished from a marker failure by the "not a marker failure" messages
// below, which carry the observed status and body length.
func TestSmokeLnwCommentBoundary(t *testing.T) {
seriesURL := os.Getenv("SMOKE_LNW_SERIES_URL")
if seriesURL == "" {
t.Skip("SMOKE_LNW_SERIES_URL unset")
}
if !fetchableSeriesURL("lightnovelworld", seriesURL) {
t.Fatalf("%q is not a fetchable lightnovelworld series URL", seriesURL)
}
f, err := NewTLSFetcher()
if err != nil {
t.Fatalf("NewTLSFetcher: %v", err)
}
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
body, status, err := f.Get(ctx, seriesURL)
if err != nil {
t.Fatalf("Get: %v", err)
}
if status != http.StatusOK {
t.Fatalf("status = %d, body %d bytes — not a marker failure; the Site did not answer this IP with a Series page", status, len(body))
}
if len(body) < lnwSeriesPageFloor {
t.Fatalf("body %d bytes — not a whole Series page (measured 685 KB..1.18 MB); not a marker failure, likely a Cloudflare challenge or a non-Series response", len(body))
}
if !lnwChapterRe.MatchString(body) {
t.Fatalf("no chapter anchor in %d bytes — not a lightnovelworld Series page; not a marker failure, likely a Cloudflare challenge or a non-Series response", len(body))
}
markerIdx := strings.Index(body, lnwCommentMarker)
if failures := checkLnwCommentBoundary(body); len(failures) > 0 {
t.Fatalf("%s (body %d bytes)", strings.Join(failures, "; "), len(body))
}
t.Logf("ok: %q once at byte %d, body %d bytes", lnwCommentMarker, markerIdx, len(body))
}
// checkLnwCommentBoundary verifies the three marker assertions against a
// fetched Series body: the marker occurs exactly once, every chapter anchor
// precedes it, and at least one anchor precedes it at all. It returns one
// human-readable failure per broken assertion — with observed offsets and body
// length — and empty when the page is healthy.
func checkLnwCommentBoundary(body string) []string {
markerIdx := strings.Index(body, lnwCommentMarker)
switch n := strings.Count(body, lnwCommentMarker); {
case n == 0:
return []string{fmt.Sprintf("%q occurs 0 times in %d bytes, want exactly 1", lnwCommentMarker, len(body))}
case n != 1:
return []string{fmt.Sprintf("%q occurs %d times in %d bytes (first at byte %d), want exactly 1", lnwCommentMarker, n, len(body), markerIdx)}
}
lastAnchor, anchorsBefore := -1, 0
for _, m := range lnwChapterRe.FindAllStringIndex(body, -1) {
if m[0] < markerIdx {
anchorsBefore++
}
lastAnchor = m[0]
}
var failures []string
if lastAnchor >= markerIdx {
failures = append(failures, fmt.Sprintf("last chapter anchor at byte %d does not precede the marker at byte %d", lastAnchor, markerIdx))
}
if anchorsBefore == 0 {
failures = append(failures, fmt.Sprintf("no chapter anchor before the marker at byte %d — the truncated prefix the scan sees yields nothing", markerIdx))
}
return failures
}
+74
View File
@@ -0,0 +1,74 @@
package latest
import "time"
// LaneState is the administrative page's view of one Poll Lane (issue #102):
// what the Lane's last pass saw. Due, Gap and Checked are filled in as the
// pass computes them; a pass that returned before reaching a figure (refusal
// backoff, sidecar down) carries the previous pass's figures forward rather
// than overwriting them with zeroes the page would state as fact.
type LaneState struct {
Site string
Due int
LastRun time.Time
Gap time.Duration
// Checked is how many Series this pass actually read. A Lane with Series
// due and nothing checked has stopped working; one with nothing due is
// merely quiet, and the page must not draw the two the same (story 13).
Checked int
Clamped bool
Refusing bool
Browser bool
}
// Status is the owner's page snapshot of the whole poller (issue #102).
type Status struct {
Lanes []LaneState
BrowserConfigured bool
BrowserReachable bool
}
// LaneStatus returns a copy of the poller's Lane state for the owner's page.
// Only Sites that have completed a pass appear — a restart therefore renders
// "no data yet" instead of confident zeroes — in the same order Run iterates.
// Refusing is derived at snapshot time from the refusal backoff, not stored,
// so a Lane that cooled down between passes reports false without a new pass.
// BrowserReachable mirrors the Lanes' own gate: the sidecar is down only
// within the refuseBackoff window since its last loss.
func (p *Poller) LaneStatus() Status {
p.mu.Lock()
defer p.mu.Unlock()
lanes := make([]LaneState, 0, len(p.laneStates))
now := p.Now()
for _, name := range laneNames() {
st, ok := p.laneStates[name]
if !ok {
continue
}
st.Refusing = now.Before(p.refuseUntil[name])
lanes = append(lanes, st)
}
configured := p.BrowserFetch != nil
reachable := configured
if reachable && !p.browserDownAt.IsZero() && now.Sub(p.browserDownAt) < refuseBackoff {
reachable = false
}
return Status{Lanes: lanes, BrowserConfigured: configured, BrowserReachable: reachable}
}
// recordLaneState stores one Lane's last pass for LaneStatus. Called deferred
// from runLanePass so every return path records, even a pass that refused.
// A pass that never reached the pace (Gap zero) keeps the last pass's figures:
// the Lane's due count and gap did not become zero because this pass declined
// to look, and the row's own marks say why it declined.
func (p *Poller) recordLaneState(st LaneState) {
p.mu.Lock()
defer p.mu.Unlock()
if p.laneStates == nil {
p.laneStates = make(map[string]LaneState)
}
if prev, ok := p.laneStates[st.Site]; ok && st.Gap == 0 {
st.Due, st.Gap, st.Clamped, st.Checked = prev.Due, prev.Gap, prev.Clamped, prev.Checked
}
p.laneStates[st.Site] = st
}
+120
View File
@@ -0,0 +1,120 @@
// Package pgtest runs the Postgres the test suite needs: one throwaway
// container per test binary, one fresh database per test. Docker is therefore
// a hard prerequisite for `go test ./...`.
//
// Rolled by hand rather than pulled in as a dependency — it is one `docker
// run`, one `docker port` and a ping loop, against a module list that is
// otherwise stdlib plus what the poller genuinely needs.
package pgtest
import (
"database/sql"
"fmt"
"os/exec"
"strconv"
"strings"
"sync/atomic"
"testing"
"time"
_ "github.com/jackc/pgx/v5/stdlib"
)
const (
image = "postgres:17-alpine"
readyLimit = 60 * time.Second
)
var (
adminURL string
dbSeq atomic.Int64
)
// Main starts the container, runs the package's tests and tears the container
// down. Every test package that touches the store calls it from TestMain:
//
// func TestMain(m *testing.M) { os.Exit(pgtest.Main(m)) }
func Main(m *testing.M) int {
id, url, err := start()
if err != nil {
fmt.Println("pgtest:", err)
return 1
}
defer exec.Command("docker", "rm", "-f", id).Run()
adminURL = url
return m.Run()
}
// URL creates a database of its own for t and returns a connection URL for it.
// Nothing drops it again: the container goes away wholesale when Main returns.
func URL(t testing.TB) string {
t.Helper()
if adminURL == "" {
t.Fatal("pgtest: no container; this package needs TestMain to call pgtest.Main")
}
// Generated, never derived from the test name, so it needs no quoting and
// cannot collide when tests run in parallel.
name := "test_" + strconv.FormatInt(dbSeq.Add(1), 10)
admin, err := sql.Open("pgx", adminURL)
if err != nil {
t.Fatalf("pgtest: open admin connection: %v", err)
}
defer admin.Close()
if _, err := admin.Exec(`CREATE DATABASE ` + name); err != nil {
t.Fatalf("pgtest: create database %s: %v", name, err)
}
return strings.Replace(adminURL, "/postgres?", "/"+name+"?", 1)
}
// start launches the container and waits for it to accept queries, returning
// its id and a connection URL for the default database.
func start() (id, url string, err error) {
out, err := exec.Command("docker", "run", "-d", "--rm",
"-e", "POSTGRES_PASSWORD=pgtest",
"-P", image,
// Durability buys nothing for a database that dies with the test
// binary, and turning it off is most of the container's start-up cost.
"-c", "fsync=off", "-c", "full_page_writes=off",
).Output()
if err != nil {
return "", "", fmt.Errorf("docker run %s: %w", image, err)
}
id = strings.TrimSpace(string(out))
port, err := exec.Command("docker", "port", id, "5432/tcp").Output()
if err != nil {
exec.Command("docker", "rm", "-f", id).Run()
return "", "", fmt.Errorf("docker port: %w", err)
}
// "0.0.0.0:32768" (and possibly a second, IPv6 line); the port is all we want.
first, _, _ := strings.Cut(strings.TrimSpace(string(port)), "\n")
url = fmt.Sprintf("postgres://postgres:pgtest@127.0.0.1:%s/postgres?sslmode=disable",
first[strings.LastIndex(first, ":")+1:])
if err := waitReady(url); err != nil {
exec.Command("docker", "rm", "-f", id).Run()
return "", "", err
}
return id, url, nil
}
func waitReady(url string) error {
db, err := sql.Open("pgx", url)
if err != nil {
return err
}
defer db.Close()
deadline := time.Now().Add(readyLimit)
for {
if err = db.Ping(); err == nil {
return nil
}
if time.Now().After(deadline) {
return fmt.Errorf("postgres not ready after %s: %w", readyLimit, err)
}
time.Sleep(200 * time.Millisecond)
}
}
+19 -51
View File
@@ -1,13 +1,10 @@
package session
import (
"crypto/hmac"
"crypto/sha256"
"crypto/subtle"
"encoding/base64"
"crypto/rand"
"encoding/hex"
"net"
"net/http"
"strconv"
"strings"
"sync"
"time"
@@ -16,49 +13,18 @@ import (
const (
CookieName = "bmgr_session"
// 60 days: long enough that a phone stays logged in between reading spells.
sessionTTL = 60 * 24 * time.Hour
// Domain separation, so the session key can never collide with any other
// use of the secrets it is derived from. Changing this string logs
// everyone out.
sessionKeyPurpose = "bmgr-web-session-v1"
SessionTTL = 60 * 24 * time.Hour
)
// Key derives the cookie-signing key from both secrets. Sessions are
// stateless — there is no session table — so rotating either API_TOKEN or
// WEB_PASSWORD invalidates every outstanding cookie at once. The \x00
// separator prevents the concatenation ambiguity a bare apiToken+webPassword
// would have (e.g. "ab"+"c" colliding with "a"+"bc").
func Key(apiToken, webPassword string) []byte {
sum := sha256.Sum256([]byte(apiToken + "\x00" + webPassword + sessionKeyPurpose))
return sum[:]
}
// Sign encodes "<expiryMs>.<base64url HMAC(expiryMs)>".
func Sign(key []byte, expiryMs int64) string {
payload := strconv.FormatInt(expiryMs, 10)
return payload + "." + sessionMAC(key, payload)
}
func sessionMAC(key []byte, payload string) string {
mac := hmac.New(sha256.New, key)
mac.Write([]byte(payload))
return base64.RawURLEncoding.EncodeToString(mac.Sum(nil))
}
// Verify checks shape, then expiry, then the signature — in that order.
// The signature comparison is constant-time; the checks before it only look at
// data the holder already supplied, so their timing leaks nothing.
func Verify(key []byte, value string, nowMs int64) bool {
payload, sig, ok := strings.Cut(value, ".")
if !ok {
return false
// NewID returns an opaque session id: 32 random bytes, hex-encoded. The id is
// all the cookie carries and all the sessions table keys on, so its entropy is
// what stops a guessed id from being someone else's session.
func NewID() string {
var b [32]byte
if _, err := rand.Read(b[:]); err != nil {
panic("session id: " + err.Error())
}
expiry, err := strconv.ParseInt(payload, 10, 64)
if err != nil || expiry <= nowMs {
return false
}
want := sessionMAC(key, payload)
return subtle.ConstantTimeCompare([]byte(sig), []byte(want)) == 1
return hex.EncodeToString(b[:])
}
// isHTTPS reports whether the browser's connection is encrypted. Behind Traefik
@@ -69,12 +35,14 @@ func isHTTPS(r *http.Request) bool {
return r.TLS != nil || r.Header.Get("X-Forwarded-Proto") == "https"
}
func SetCookie(w http.ResponseWriter, r *http.Request, key []byte) {
// SetCookie writes the session cookie. The value is the session id and nothing
// else; the row behind it is looked up on every request.
func SetCookie(w http.ResponseWriter, r *http.Request, id string) {
http.SetCookie(w, &http.Cookie{
Name: CookieName,
Value: Sign(key, time.Now().Add(sessionTTL).UnixMilli()),
Value: id,
Path: "/",
MaxAge: int(sessionTTL / time.Second),
MaxAge: int(SessionTTL / time.Second),
HttpOnly: true,
Secure: isHTTPS(r),
SameSite: http.SameSiteLaxMode,
@@ -120,14 +88,14 @@ func ClientIP(r *http.Request) string {
return host
}
// LoginLimiter throttles password guessing: MaxFailures failures inside a
// rolling Window blocks further attempts from that IP until the oldest one
// LoginLimiter throttles failed sign-in attempts: MaxFailures failures inside
// a rolling Window blocks further attempts from that IP until the oldest one
// ages out. There is no permanent ban and no unlock step.
//
// Behind carrier-grade NAT this budget is shared with every other subscriber on
// the same public address, so a stranger can lock the owner out for up to one
// window. That is accepted: the block self-heals, and ten attempts is generous
// for a mistyped password.
// for the occasional fumbled sign-in.
//
// State is in memory and per-process, so a restart clears it. Entries are
// pruned lazily on access; for a single-user deployment the map cannot grow
+19 -63
View File
@@ -9,66 +9,19 @@ import (
"time"
)
func TestSessionRoundTrip(t *testing.T) {
key := Key("token-abc", "pw-abc")
now := time.Now().UnixMilli()
value := Sign(key, now+60_000)
if !Verify(key, value, now) {
t.Fatal("Verify = false for a freshly signed cookie, want true")
func TestNewID(t *testing.T) {
a := NewID()
b := NewID()
if a == b {
t.Fatal("NewID returned the same value twice")
}
}
func TestSessionRejects(t *testing.T) {
key := Key("token-abc", "pw-abc")
now := time.Now().UnixMilli()
valid := Sign(key, now+60_000)
payload, sig, _ := strings.Cut(valid, ".")
cases := []struct {
name string
value string
}{
{"empty", ""},
{"no separator", payload + sig},
{"unparseable expiry", "notanumber." + sig},
{"expired", Sign(key, now-1)},
{"tampered signature", payload + "." + flipLastChar(sig)},
{"tampered expiry", "99999999999999." + sig},
{"signed with another key", Sign(Key("other-token", "pw-abc"), now+60_000)},
if len(a) != 64 { // 32 random bytes, hex
t.Fatalf("NewID() length = %d, want 64", len(a))
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
if Verify(key, tc.value, now) {
t.Fatalf("Verify(%q) = true, want false", tc.value)
}
})
}
}
func flipLastChar(s string) string {
if s == "" {
return "x"
}
last := s[len(s)-1]
if last == 'A' {
return s[:len(s)-1] + "B"
}
return s[:len(s)-1] + "A"
}
func TestSessionKeyDependsOnToken(t *testing.T) {
a := Key("token-a", "pw-abc")
b := Key("token-b", "pw-abc")
if string(a) == string(b) {
t.Fatal("Key collided for different API tokens")
}
}
func TestSessionKeyDependsOnWebPassword(t *testing.T) {
a := Key("token-abc", "pw-a")
b := Key("token-abc", "pw-b")
if string(a) == string(b) {
t.Fatal("Key collided for different web passwords with the same API token")
for _, r := range a {
if !strings.ContainsRune("0123456789abcdef", r) {
t.Fatalf("NewID() = %q, want hex", a)
}
}
}
@@ -86,7 +39,7 @@ func TestSetSessionCookieAttributes(t *testing.T) {
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
r := httptest.NewRequest(http.MethodPost, "/login", nil)
r := httptest.NewRequest(http.MethodPost, "/", nil)
if tc.tls {
r.TLS = &tls.ConnectionState{}
}
@@ -94,7 +47,7 @@ func TestSetSessionCookieAttributes(t *testing.T) {
r.Header.Set("X-Forwarded-Proto", tc.forwarded)
}
rr := httptest.NewRecorder()
SetCookie(rr, r, Key("token-abc", "pw-abc"))
SetCookie(rr, r, "abc123")
cookies := rr.Result().Cookies()
if len(cookies) != 1 {
@@ -104,6 +57,9 @@ func TestSetSessionCookieAttributes(t *testing.T) {
if c.Name != CookieName {
t.Fatalf("cookie name = %q, want %q", c.Name, CookieName)
}
if c.Value != "abc123" {
t.Fatalf("cookie value = %q, want the session id verbatim", c.Value)
}
if !c.HttpOnly {
t.Fatal("cookie HttpOnly = false, want true")
}
@@ -116,8 +72,8 @@ func TestSetSessionCookieAttributes(t *testing.T) {
if c.Secure != tc.wantSecure {
t.Fatalf("cookie Secure = %v, want %v", c.Secure, tc.wantSecure)
}
if c.MaxAge != int(sessionTTL/time.Second) {
t.Fatalf("cookie MaxAge = %d, want %d", c.MaxAge, int(sessionTTL/time.Second))
if c.MaxAge != int(SessionTTL/time.Second) {
t.Fatalf("cookie MaxAge = %d, want %d", c.MaxAge, int(SessionTTL/time.Second))
}
})
}
@@ -163,7 +119,7 @@ func TestClientIP(t *testing.T) {
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
r := httptest.NewRequest(http.MethodPost, "/login", nil)
r := httptest.NewRequest(http.MethodPost, "/", nil)
r.RemoteAddr = tc.remoteAddr
for _, v := range tc.xff {
r.Header.Add("X-Forwarded-For", v)
@@ -0,0 +1,28 @@
-- One row per tracked series, keyed "<site>:<series_id>".
--
-- Everything is NOT NULL with a default except latest_chapter_num, where NULL
-- is a distinct state: nothing has been captured yet, which is not the same as
-- chapter zero.
--
-- Timestamps are unix milliseconds as bigint, not timestamptz: the userscripts
-- send Date.now() over the wire and the ordering rule compares them directly.
CREATE TABLE bookmarks (
key text PRIMARY KEY,
site text NOT NULL,
series_id text NOT NULL,
title text NOT NULL DEFAULT '',
series_url text NOT NULL DEFAULT '',
cover text NOT NULL DEFAULT '',
last_chapter text NOT NULL DEFAULT '',
last_chapter_num double precision NOT NULL DEFAULT 0,
last_chapter_url text NOT NULL DEFAULT '',
favorite boolean NOT NULL DEFAULT false,
latest_chapter text NOT NULL DEFAULT '',
latest_chapter_num double precision,
-- When the server last polled this series, unix ms; 0 means never, and sorts
-- first so a new bookmark is picked up on the next tick with no special case.
latest_checked_at bigint NOT NULL DEFAULT 0,
status text NOT NULL DEFAULT 'reading',
kind text NOT NULL DEFAULT 'manga',
updated_at bigint NOT NULL
);
@@ -0,0 +1,45 @@
-- One row per distinct work, shared by every bookmark that tracks it
-- (ADR-0003). Keyed (site, series_id), the pair a bookmark key decomposes
-- into. title/series_url/cover are written once, at creation, and never
-- again: client-supplied values are ignored once the row exists and the
-- poller is the only party that may change them. kind and the latest-chapter
-- fields are last-write-wins like the bookmark's own fields.
CREATE TABLE series (
site text NOT NULL,
series_id text NOT NULL,
title text NOT NULL DEFAULT '',
series_url text NOT NULL DEFAULT '',
cover text NOT NULL DEFAULT '',
kind text NOT NULL DEFAULT 'manga',
latest_chapter text NOT NULL DEFAULT '',
latest_chapter_num double precision,
-- When the server last polled this series, unix ms; 0 means never, and sorts
-- first so a new bookmark is picked up on the next tick with no special case.
latest_checked_at bigint NOT NULL DEFAULT 0,
PRIMARY KEY (site, series_id)
);
-- Backfill from today's rows. The bookmark key's uniqueness makes
-- (site, series_id) unique in practice; DISTINCT is belt and braces.
INSERT INTO series (site, series_id, title, series_url, cover, kind,
latest_chapter, latest_chapter_num, latest_checked_at)
SELECT DISTINCT site, series_id, title, series_url, cover, kind,
latest_chapter, latest_chapter_num, latest_checked_at
FROM bookmarks;
-- The bookmark keeps only what differs between readers (ADR-0003): progress,
-- favourite, lifecycle bucket. The dropped columns now live on series.
ALTER TABLE bookmarks
DROP COLUMN title,
DROP COLUMN series_url,
DROP COLUMN cover,
DROP COLUMN kind,
DROP COLUMN latest_chapter,
DROP COLUMN latest_chapter_num,
DROP COLUMN latest_checked_at;
-- A bookmark may not point at a series that does not exist. No cascade: a
-- series outlives its last bookmark, and deleting one is not a store operation.
ALTER TABLE bookmarks
ADD CONSTRAINT bookmarks_series_fk
FOREIGN KEY (site, series_id) REFERENCES series (site, series_id);
@@ -0,0 +1,17 @@
-- One row per person. Keyed by their Discord user ID; carries the SHA-256 of
-- their userscript token and when they were created. Hashed because a token
-- in the database is a token anyone with the database can replay; SHA-256 is
-- enough because the tokens are high-entropy random values with nothing to
-- brute-force. No one can register yet, so this table holds exactly the one
-- owner row the seed creates at startup (see Store.Open).
CREATE TABLE readers (
id bigserial PRIMARY KEY,
discord_id text NOT NULL UNIQUE,
token_sha256 bytea NOT NULL UNIQUE,
created_at timestamptz NOT NULL DEFAULT now()
);
-- Every bookmark now belongs to a reader. Added nullable: rows created before
-- this migration have no owner yet — 0004 attaches them to the seeded owner
-- before NOT NULL and the composite key land.
ALTER TABLE bookmarks ADD COLUMN reader_id bigint;
@@ -0,0 +1,19 @@
-- Attach every pre-existing bookmark to the owner reader, seeded between the
-- two migrate passes (Store.Open). The oldest reader is the owner by
-- construction: only the seed creates readers, and it runs once per database.
-- Run-once via the version table, like every migration.
UPDATE bookmarks SET reader_id = (SELECT id FROM readers ORDER BY id LIMIT 1);
-- Ownership lands structurally: reader_id becomes part of the key, so a
-- bookmark is one Reader's progress on one Series and a duplicate for the
-- same pair is impossible at the database level. Deleting a Reader takes
-- their bookmarks with them. The old text key is gone — the wire "key" is
-- derived as site:series_id on read, and nothing references the column.
-- Dropping it drops the primary key it carried; the composite key replaces
-- it, and the FK index the series constraint needs is created automatically.
ALTER TABLE bookmarks
ALTER COLUMN reader_id SET NOT NULL,
DROP COLUMN key,
ADD PRIMARY KEY (reader_id, site, series_id),
ADD CONSTRAINT bookmarks_reader_fk
FOREIGN KEY (reader_id) REFERENCES readers (id) ON DELETE CASCADE;
@@ -0,0 +1,11 @@
-- One row per browser session. The id is an opaque random value the cookie
-- carries verbatim; a request is authenticated by looking the row up, and
-- deleting the row is how a session is revoked. Expired rows are removed
-- lazily on lookup and swept by the next login, so nothing runs a background
-- cleanup.
CREATE TABLE sessions (
id text PRIMARY KEY,
reader_id bigint NOT NULL REFERENCES readers (id) ON DELETE CASCADE,
created_at timestamptz NOT NULL DEFAULT now(),
expires_at timestamptz NOT NULL
);
@@ -0,0 +1,7 @@
-- Rotation is an epoch bump: a Reader's credential is derived from the
-- deployment secret, their Discord id and this epoch, so bumping it issues a
-- new credential and the rewritten token_sha256 invalidates the old one the
-- moment the transaction commits. The seed's ON CONFLICT refresh (Store.Open)
-- is gated on this being 0, so a restart can never undo a rotation by
-- restoring the epoch-0 hash.
ALTER TABLE readers ADD COLUMN token_epoch bigint NOT NULL DEFAULT 0;
@@ -0,0 +1,8 @@
-- Kagane cover bytes belong in their own table so image blobs never enter the
-- series queries that drive the latest-chapter poller.
CREATE TABLE covers (
image_id text PRIMARY KEY,
body bytea NOT NULL,
content_type text NOT NULL,
fetched_at timestamptz NOT NULL DEFAULT now()
);
@@ -0,0 +1,9 @@
-- Cover bytes move out of Postgres. Existing rows are intentionally dropped:
-- the old kagane path already refetches missing Covers on demand.
DROP TABLE covers;
CREATE TABLE covers (
address text PRIMARY KEY,
path text NOT NULL,
content_type text NOT NULL
);
@@ -0,0 +1,8 @@
-- The Cover splits into two facts. `cover` keeps the third-party address the
-- bytes come from, which is what the acquisition path refetches and dedupes
-- on; `cover_address` is the content address of the bytes once they are
-- actually stored, and is what the wire's absolute URL is built from.
--
-- Empty `cover_address` therefore means "no Cover yet" rather than "a Cover
-- that 404s", which is the distinction the API and the UI both depend on.
ALTER TABLE series ADD COLUMN cover_address text NOT NULL DEFAULT '';
@@ -0,0 +1,6 @@
-- The owner's administrative page (issue #102) renders these counters and
-- offers a control to clear them, deliberately shipped before the Sighting
-- feature (issue #103) that fills them, so a false mark never needs SQL
-- against production. Zero counters mean a trusted Reader.
ALTER TABLE readers ADD COLUMN sighting_agreements integer NOT NULL DEFAULT 0;
ALTER TABLE readers ADD COLUMN sighting_disagreements integer NOT NULL DEFAULT 0;
+80
View File
@@ -0,0 +1,80 @@
package store
import (
"database/sql"
"fmt"
"time"
)
// Session is one browser login: an opaque id the cookie carries verbatim,
// the Reader it belongs to, and when it stops being valid.
type Session struct {
ID string
ReaderID int64
ExpiresAt time.Time
}
// CreateSession stores a new session row for reader. The id is generated by
// the caller (session.NewID) — the store only persists it. Expired rows that
// were never looked up are swept in the same transaction: this is the one
// write every login makes, so the table stays bounded without a background
// job.
func (s *Store) CreateSession(id string, readerID int64, ttl time.Duration) (Session, error) {
tx, err := s.db.Begin()
if err != nil {
return Session{}, err
}
defer tx.Rollback()
expires := time.Now().Add(ttl)
if _, err := tx.Exec(`INSERT INTO sessions (id, reader_id, expires_at) VALUES ($1, $2, $3)`,
id, readerID, expires); err != nil {
return Session{}, err
}
if _, err := tx.Exec(`DELETE FROM sessions WHERE expires_at < now()`); err != nil {
return Session{}, err
}
if err := tx.Commit(); err != nil {
return Session{}, err
}
return Session{ID: id, ReaderID: readerID, ExpiresAt: expires}, nil
}
// GetSession returns the live session row for id, or ok=false when the id is
// unknown or expired. An expired row is deleted on the way out, so the table
// never grows past sessions that are still valid.
func (s *Store) GetSession(id string, now time.Time) (Session, bool, error) {
var sess Session
err := s.db.QueryRow(
`SELECT id, reader_id, expires_at FROM sessions WHERE id = $1`, id,
).Scan(&sess.ID, &sess.ReaderID, &sess.ExpiresAt)
if err == sql.ErrNoRows {
return Session{}, false, nil
}
if err != nil {
return Session{}, false, err
}
if !sess.ExpiresAt.After(now) {
// Best-effort: the row is dead either way; failing the request over a
// cleanup delete would only hide the real error. CreateSession's
// sweep catches anything this misses.
_, _ = s.db.Exec(`DELETE FROM sessions WHERE id = $1`, id)
return Session{}, false, nil
}
return sess, true, nil
}
// DeleteSession revokes one session. Deleting an unknown id is not an error.
func (s *Store) DeleteSession(id string) error {
_, err := s.db.Exec(`DELETE FROM sessions WHERE id = $1`, id)
return err
}
// DeleteReaderSessions revokes every session one Reader holds — the owner's
// remedy when a Reader's browser must be logged out everywhere at once. The
// next request carrying any of those cookies finds no row and is rejected.
func (s *Store) DeleteReaderSessions(readerID int64) error {
if _, err := s.db.Exec(`DELETE FROM sessions WHERE reader_id = $1`, readerID); err != nil {
return fmt.Errorf("delete sessions for reader %d: %w", readerID, err)
}
return nil
}
+84
View File
@@ -0,0 +1,84 @@
package store
import (
"testing"
"time"
)
func TestCreateAndGetSession(t *testing.T) {
s := newTestStore(t)
owner := s.OwnerID()
sess, err := s.CreateSession("sess-1", owner, time.Hour)
if err != nil {
t.Fatalf("CreateSession: %v", err)
}
if sess.ID != "sess-1" || sess.ReaderID != owner {
t.Fatalf("CreateSession returned %+v, want id sess-1 reader %d", sess, owner)
}
got, ok, err := s.GetSession("sess-1", time.Now())
if err != nil || !ok {
t.Fatalf("GetSession: ok=%v err=%v, want ok", ok, err)
}
if got.ReaderID != owner {
t.Fatalf("session reader = %d, want %d", got.ReaderID, owner)
}
}
func TestGetSessionUnknownID(t *testing.T) {
s := newTestStore(t)
if _, ok, err := s.GetSession("nope", time.Now()); err != nil || ok {
t.Fatalf("GetSession(unknown) = ok=%v err=%v, want ok=false", ok, err)
}
}
func TestExpiredSessionIsGone(t *testing.T) {
s := newTestStore(t)
owner := s.OwnerID()
if _, err := s.CreateSession("sess-exp", owner, -time.Minute); err != nil {
t.Fatalf("CreateSession: %v", err)
}
now := time.Now()
if _, ok, err := s.GetSession("sess-exp", now); err != nil || ok {
t.Fatalf("GetSession(expired) = ok=%v err=%v, want ok=false", ok, err)
}
// The expired row is deleted on lookup, so the next call cannot revive it.
if _, ok, err := s.GetSession("sess-exp", now.Add(-time.Hour)); err != nil || ok {
t.Fatalf("GetSession(expired again) = ok=%v err=%v, want ok=false", ok, err)
}
}
func TestDeleteSessionRevokes(t *testing.T) {
s := newTestStore(t)
owner := s.OwnerID()
if _, err := s.CreateSession("sess-del", owner, time.Hour); err != nil {
t.Fatalf("CreateSession: %v", err)
}
if err := s.DeleteSession("sess-del"); err != nil {
t.Fatalf("DeleteSession: %v", err)
}
if _, ok, err := s.GetSession("sess-del", time.Now()); err != nil || ok {
t.Fatalf("GetSession after delete = ok=%v err=%v, want ok=false", ok, err)
}
// Deleting twice is not an error.
if err := s.DeleteSession("sess-del"); err != nil {
t.Fatalf("DeleteSession twice: %v", err)
}
}
func TestDeleteSessionIsPerReader(t *testing.T) {
s := newTestStore(t)
other := secondReader(t, s)
if _, err := s.CreateSession("sess-other", other, time.Hour); err != nil {
t.Fatalf("CreateSession: %v", err)
}
got, ok, err := s.GetSession("sess-other", time.Now())
if err != nil || !ok {
t.Fatalf("GetSession: ok=%v err=%v, want ok", ok, err)
}
if got.ReaderID != other {
t.Fatalf("session reader = %d, want %d", got.ReaderID, other)
}
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+37
View File
@@ -0,0 +1,37 @@
package token
import (
"crypto/hmac"
"crypto/sha256"
"encoding/hex"
"strconv"
)
// Token derives one Reader's userscript credential from the deployment
// secret, the Reader's Discord id and their token epoch.
//
// The credential is deterministic rather than stored random because the
// server must be able to rebuild the install URL after a restart while the
// database holds only hashes: a random token with no plaintext copy anywhere
// would be unreconstructible, and keeping plaintext in memory would break
// every install link on restart. HMAC output is high-entropy, indistinguishable
// from random to anyone without the secret, and changes whenever the epoch
// does — which is what rotation is. The stored form is Hash of this value,
// so a database leak yields nothing but hashes of unguessable strings.
func Token(key []byte, discordID string, epoch int64) string {
mac := hmac.New(sha256.New, key)
// The separator is unambiguous: discord ids are decimal snowflakes and
// epochs are plain integers, so no two (id, epoch) pairs can collide.
mac.Write([]byte(discordID))
mac.Write([]byte{0})
mac.Write([]byte(strconv.FormatInt(epoch, 10)))
return hex.EncodeToString(mac.Sum(nil))
}
// Hash is the SHA-256 of a credential — the only form that ever touches the
// database (readers.token_sha256). SHA-256 rather than a password hash is
// deliberate: these are unguessable values with nothing to brute-force, so a
// slow hash would only add per-request cost.
func Hash(cred string) [32]byte {
return sha256.Sum256([]byte(cred))
}
+53
View File
@@ -0,0 +1,53 @@
package token
import (
"bytes"
"crypto/sha256"
"testing"
)
func TestTokenDeterministicPerReaderAndEpoch(t *testing.T) {
key := []byte("deployment-secret")
a := Token(key, "reader-1", 0)
b := Token(key, "reader-1", 0)
if a != b {
t.Fatal("same (reader, epoch) derived different credentials")
}
if a == Token(key, "reader-2", 0) {
t.Fatal("different readers derived the same credential")
}
if a == Token(key, "reader-1", 1) {
t.Fatal("rotation epoch derived the same credential")
}
}
func TestTokenChangesWithSecret(t *testing.T) {
a := Token([]byte("key-1"), "reader-1", 0)
b := Token([]byte("key-2"), "reader-1", 0)
if a == b {
t.Fatal("different secrets derived the same credential")
}
}
func TestTokenFormat(t *testing.T) {
cred := Token([]byte("key"), "reader-1", 0)
// 32 bytes of HMAC-SHA256, hex-encoded: the length the install URL and
// the committed placeholder both assume.
if len(cred) != 64 {
t.Fatalf("credential length = %d, want 64", len(cred))
}
for _, c := range cred {
if !(c >= '0' && c <= '9' || c >= 'a' && c <= 'f') {
t.Fatalf("credential contains non-hex byte %q", c)
}
}
}
func TestHashIsSha256OfCredential(t *testing.T) {
cred := Token([]byte("key"), "reader-1", 0)
got := Hash(cred)
want := sha256.Sum256([]byte(cred))
if !bytes.Equal(got[:], want[:]) {
t.Fatal("Hash is not the SHA-256 of the credential")
}
}
+65 -25
View File
@@ -1,14 +1,24 @@
package userscript
import (
"crypto/subtle"
"bytes"
"log"
"net/http"
"os"
"regexp"
"time"
"bookmarkmanager/backend/internal/httpmw"
"bookmarkmanager/backend/internal/store"
)
// tokenPlaceholder is what the bindmounted userscript carries where the
// Reader's credential goes: in the API_TOKEN constant and in the @downloadURL
// and @updateURL metadata lines. The handler substitutes the requesting
// Reader's credential for it at serve time, so no credential literal is ever
// committed or deployed, and each Reader's copy carries exactly their own.
var tokenPlaceholder = []byte("__API_TOKEN__")
// versionLine matches the userscript metadata block's @version directive.
var versionLine = regexp.MustCompile(`(?m)^// @version[ \t]+.*$`)
@@ -26,35 +36,65 @@ func stampVersion(src []byte, mod time.Time) []byte {
return versionLine.ReplaceAll(src, []byte("// @version "+mod.UTC().Format("2006.01.02.1504")))
}
// userscriptHandler serves the userscript to Violentmonkey's updater.
// substituteToken replaces every tokenPlaceholder with the Reader's
// credential. A file without the placeholder is returned unchanged so Render
// can warn about it rather than silently serving a credential-less script.
func substituteToken(src []byte, credential string) []byte {
return bytes.ReplaceAll(src, tokenPlaceholder, []byte(credential))
}
// Render writes one userscript file with the credential substituted and the
// mtime-derived version stamped. Shared by the download path (Handler) and
// the web UI's install endpoints, so both serve byte-identical scripts.
//
// The token lives in the path because the update poll sends no Authorization
// header, and the file embeds API_TOKEN in plain text, so an open path would
// hand that token to anyone who guessed the URL. A mismatch answers 404 rather
// than 401: a prober learns nothing about whether the route exists.
// The file is read per request — that is what lets a bindmounted copy be
// edited on the host without a restart. It is ~50 KB and polled about once a
// day.
func Render(w http.ResponseWriter, r *http.Request, path, credential string) {
info, err := os.Stat(path)
if err != nil {
log.Printf("userscript: stat %s: %v", path, err)
http.NotFound(w, r)
return
}
src, err := os.ReadFile(path)
if err != nil {
log.Printf("userscript: read %s: %v", path, err)
http.NotFound(w, r)
return
}
rendered := substituteToken(src, credential)
if bytes.Equal(rendered, src) {
// The bindmounted file was not built for per-Reader rendering. Serving
// it as written is the operator's freedom, but a credential-less copy
// is a deployment bug worth one log line — the symptom (silent 401s on
// every device) is otherwise indistinguishable from a network fault.
log.Printf("userscript: %s has no %s placeholder; serving as written", path, tokenPlaceholder)
}
w.Header().Set("Content-Type", "text/javascript; charset=utf-8")
w.Header().Set("Cache-Control", "no-cache")
w.Write(stampVersion(rendered, info.ModTime()))
}
// Handler serves the userscript to Violentmonkey's updater, rendered for the
// Reader whose credential is in the path.
//
// The file is read per request — that is what lets a bindmounted copy be edited
// on the host without a restart. It is ~50 KB and polled about once a day.
func Handler(token, path string) http.HandlerFunc {
// The credential lives in the path because the update poll sends no
// Authorization header, and the rendered file embeds the credential in
// plaintext, so an open path would hand it to anyone who guessed the URL. A
// mismatch answers 404 rather than 401: a prober learns nothing about whether
// the route exists. The same credential authenticates the API bearer header,
// so the two are one secret with one blast radius.
//
// The path segment is the credential itself, so once it resolves it is also
// exactly what the served copy must carry — no re-derivation needed.
func Handler(s *store.Store, path string) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
if subtle.ConstantTimeCompare([]byte(r.PathValue("token")), []byte(token)) != 1 {
cred := r.PathValue("token")
if _, ok := httpmw.ResolveReader(s, cred); !ok {
http.NotFound(w, r)
return
}
info, err := os.Stat(path)
if err != nil {
log.Printf("userscript: stat %s: %v", path, err)
http.NotFound(w, r)
return
}
src, err := os.ReadFile(path)
if err != nil {
log.Printf("userscript: read %s: %v", path, err)
http.NotFound(w, r)
return
}
w.Header().Set("Content-Type", "text/javascript; charset=utf-8")
w.Header().Set("Cache-Control", "no-cache")
w.Write(stampVersion(src, info.ModTime()))
Render(w, r, path, cred)
}
}
+41 -89
View File
@@ -1,115 +1,67 @@
package userscript
import (
"net/http"
"net/http/httptest"
"os"
"path/filepath"
"strings"
"testing"
"time"
)
const testToken = "s3cret-token"
// sampleScript is a stand-in for the real userscript: a metadata block with a
// @version line, plus a body that must survive the rewrite untouched.
// @version line, the credential placeholder in its metadata and body, plus
// content that must survive the rewrites untouched.
const sampleScript = `// ==UserScript==
// @name Manga Bookmark Sync
// @version 1.5.0
// @downloadURL https://api.example/u/__API_TOKEN__/manga-bookmark.user.js
// @match https://asurascans.com/*
// ==/UserScript==
(function () { "use strict"; })();
(function () { "use strict";
const API_TOKEN = "__API_TOKEN__";
})();
`
// writeScript drops a userscript in a temp dir with a known mtime and returns
// its path plus the version string the handler is expected to stamp.
func writeScript(t *testing.T, body string) (path, wantVersion string) {
t.Helper()
path = filepath.Join(t.TempDir(), "manga-bookmark.user.js")
if err := os.WriteFile(path, []byte(body), 0o644); err != nil {
t.Fatalf("write script: %v", err)
}
func TestStampVersionReplacesVersionLineOnly(t *testing.T) {
mod := time.Date(2026, 7, 28, 16, 42, 0, 0, time.UTC)
if err := os.Chtimes(path, mod, mod); err != nil {
t.Fatalf("chtimes: %v", err)
}
return path, "2026.07.28.1642"
}
got := string(stampVersion([]byte(sampleScript), mod))
// newTestMux registers Handler the same way main.go's router does, without
// pulling in the store or the rest of the app.
func newTestMux(token, path string) http.Handler {
mux := http.NewServeMux()
mux.HandleFunc("GET /u/{token}/manga-bookmark.user.js", Handler(token, path))
return mux
}
func getScript(t *testing.T, srv http.Handler, token string) *httptest.ResponseRecorder {
t.Helper()
rr := httptest.NewRecorder()
srv.ServeHTTP(rr, httptest.NewRequest(http.MethodGet, "/u/"+token+"/manga-bookmark.user.js", nil))
return rr
}
func TestUserscriptServedWithStampedVersion(t *testing.T) {
path, wantVersion := writeScript(t, sampleScript)
rr := getScript(t, newTestMux(testToken, path), testToken)
if rr.Code != http.StatusOK {
t.Fatalf("status = %d, want 200", rr.Code)
if !strings.Contains(got, "// @version "+mod.UTC().Format("2006.01.02.1504")) {
t.Errorf("body has no stamped version:\n%s", got)
}
if ct := rr.Header().Get("Content-Type"); !strings.HasPrefix(ct, "text/javascript") {
t.Errorf("Content-Type = %q, want text/javascript", ct)
if strings.Contains(got, "1.5.0") {
t.Errorf("body still carries the file's own version:\n%s", got)
}
if cc := rr.Header().Get("Cache-Control"); cc != "no-cache" {
t.Errorf("Cache-Control = %q, want no-cache", cc)
}
body := rr.Body.String()
if !strings.Contains(body, "// @version "+wantVersion) {
t.Errorf("body has no stamped version %q:\n%s", wantVersion, body)
}
if strings.Contains(body, "1.5.0") {
t.Errorf("body still carries the file's own version:\n%s", body)
}
// Everything outside the @version line is served verbatim.
if !strings.Contains(body, `(function () { "use strict"; })();`) {
t.Errorf("body was altered beyond the version line:\n%s", body)
}
if !strings.Contains(body, "// @name Manga Bookmark Sync") {
t.Errorf("metadata block was altered:\n%s", body)
// Everything outside the @version line is served verbatim, including the
// placeholder — stamping must not do the substitution's job.
if !strings.Contains(got, `const API_TOKEN = "__API_TOKEN__";`) {
t.Errorf("body was altered beyond the version line:\n%s", got)
}
}
// The empty-token case ("/u//manga-bookmark.user.js") is covered at the
// router level (see backend's guardEmptyUserscriptToken): ServeMux 307s it to
// "/u/manga-bookmark.user.js" before this handler's own token check ever runs.
func TestUserscriptWrongTokenIs404(t *testing.T) {
path, _ := writeScript(t, sampleScript)
srv := newTestMux(testToken, path)
for _, tok := range []string{"wrong", testToken + "x", testToken[:3]} {
if got := getScript(t, srv, tok).Code; got != http.StatusNotFound {
t.Errorf("token %q: status = %d, want 404", tok, got)
}
}
}
func TestUserscriptMissingFileIs404(t *testing.T) {
srv := newTestMux(testToken, filepath.Join(t.TempDir(), "absent.user.js"))
if got := getScript(t, srv, testToken).Code; got != http.StatusNotFound {
t.Fatalf("status = %d, want 404", got)
}
}
func TestUserscriptWithoutVersionLineServedUnmodified(t *testing.T) {
func TestStampVersionWithoutVersionLineServedUnmodified(t *testing.T) {
const noVersion = "// ==UserScript==\n// @name x\n// ==/UserScript==\nconsole.log(1);\n"
path, _ := writeScript(t, noVersion)
rr := getScript(t, newTestMux(testToken, path), testToken)
if rr.Code != http.StatusOK {
t.Fatalf("status = %d, want 200", rr.Code)
}
if rr.Body.String() != noVersion {
t.Fatalf("body = %q, want it unmodified", rr.Body.String())
if got := string(stampVersion([]byte(noVersion), time.Now())); got != noVersion {
t.Errorf("stampVersion altered a file with no @version line:\n%s", got)
}
}
func TestSubstituteTokenReplacesEveryPlaceholder(t *testing.T) {
got := string(substituteToken([]byte(sampleScript), "abc123"))
if strings.Contains(got, "__API_TOKEN__") {
t.Errorf("placeholder survived substitution:\n%s", got)
}
// The credential lands in the constant and in both metadata lines.
if want := `const API_TOKEN = "abc123";`; !strings.Contains(got, want) {
t.Errorf("no substituted constant %q:\n%s", want, got)
}
if want := "https://api.example/u/abc123/manga-bookmark.user.js"; !strings.Contains(got, want) {
t.Errorf("no substituted download URL %q:\n%s", want, got)
}
}
func TestSubstituteTokenWithoutPlaceholderServedUnmodified(t *testing.T) {
const noPlaceholder = "// ==UserScript==\n// @name x\n// ==/UserScript==\n"
if got := string(substituteToken([]byte(noPlaceholder), "abc123")); got != noPlaceholder {
t.Errorf("substituteToken altered a file without the placeholder:\n%s", got)
}
}
+251
View File
@@ -0,0 +1,251 @@
package web
import (
"log"
"net/http"
"strconv"
"time"
"bookmarkmanager/backend/internal/latest"
"bookmarkmanager/backend/internal/store"
)
// LaneReporter is the administrative page's whole window onto the running
// poller: one snapshot of Poll Lane state, copied out of memory on request.
// The Poller satisfies it in production and a fake with fixed values satisfies
// it in tests, so the page's tests need neither a poller nor a Site.
type LaneReporter interface {
LaneStatus() latest.Status
}
// adminView is what the administrative page and the roster fragment receive.
type adminView struct {
Readers []store.ReaderSummary
// OwnerID travels with the roster so it can tell the owner's own row from
// the Readers they may act on.
OwnerID int64
Lanes lanesView
}
// lanesView is the Lane status block: one row per Site that has run, plus the
// browser fact, which is shared by the three browser Sites rather than held
// once per Site.
type lanesView struct {
Rows []laneRow
// PollerOff means no poller is running at all (disabled by config, or its
// client could not be built). The browser line must not answer "not
// configured" then: the sidecar is not the reason nothing is polled.
PollerOff bool
BrowserConfigured bool
BrowserReachable bool
}
// laneRow is one Lane formatted for reading rather than for arithmetic: the
// template renders strings and flags, and every judgement about what they mean
// is made here.
type laneRow struct {
Site string
Due int
Ran string
// Checked is how many Series the last pass read. Due without Checked is a
// Lane that has stopped working; the two figures side by side are what
// separate that from a Lane with nothing to do.
Checked int
// Gap is empty when no pass has reached the pace yet, so the row omits the
// figure instead of stating a zero.
Gap string
Clamped bool
Refusing bool
// BrowserLost marks a Lane whose pages can only be read through the
// sidecar while the sidecar is unreachable — including the case where none
// is configured, which stops those Series just as completely.
BrowserLost bool
// Stalled marks a Lane with Series waiting that its last pass did not read
// — the difference between a stopped Lane and a quiet one (story 13).
Stalled bool
// Attention is the one flag the template colours on, so an unhealthy Lane
// is found at a glance rather than read for.
Attention bool
}
// adminRoute pairs a route pattern with its handler so the route list and the
// gate cannot drift apart.
type adminRoute struct {
pattern string
handler http.HandlerFunc
}
// adminRoutes is every route that reaches past the acting Reader. Register
// wraps each one in requireOwner, so a new administrative route is gated by
// being listed here rather than by remembering to write a check inside it.
func (h *Handler) adminRoutes() []adminRoute {
return []adminRoute{
{"GET /admin", h.admin},
{"GET /ui/admin/lanes", h.uiLanes},
{"POST /readers/{id}/revoke", h.revokeReaderSessions},
{"POST /readers/{id}/clear-marks", h.clearReaderMarks},
}
}
// AdminPatterns names every administrative route, so one test can prove the
// owner gate covers all of them rather than one test per route. The receiver is
// nil because only the patterns are read; the bound handlers are never called.
func AdminPatterns() []string {
routes := (*Handler)(nil).adminRoutes()
out := make([]string, 0, len(routes))
for _, rt := range routes {
out = append(out, rt.pattern)
}
return out
}
// requireOwner is the owner test, in one place, layered on the session gate: no
// session is still 401, and a signed-in Reader who is not the owner gets 404
// rather than 403 — a refusal that confirms the address exists is a refusal
// that helps whoever is probing for it.
func (h *Handler) requireOwner(next http.HandlerFunc) http.HandlerFunc {
return h.requireSession(func(w http.ResponseWriter, r *http.Request) {
if readerOf(r) != h.store.OwnerID() {
http.NotFound(w, r)
return
}
next(w, r)
})
}
// admin renders the owner's page: the Reader roster and Poll Lane status.
func (h *Handler) admin(w http.ResponseWriter, r *http.Request) {
readers, err := h.store.Readers()
if err != nil {
log.Printf("admin: %v", err)
http.Error(w, "internal error", http.StatusInternalServerError)
return
}
h.render(w, http.StatusOK, "admin", adminView{
Readers: readers,
OwnerID: h.store.OwnerID(),
Lanes: h.lanesView(),
})
}
// uiLanes answers the status block's own refresh. Only the block refreshes on a
// timer; the roster re-renders after an action, as it always has.
func (h *Handler) uiLanes(w http.ResponseWriter, r *http.Request) {
h.render(w, http.StatusOK, "lanes", h.lanesView())
}
// lanesView copies the poller's snapshot into display form. A nil reporter (no
// poller running) and a poller no Lane has reported to yet are the same thing
// to the page: no data, which it must say rather than draw as confident zeroes
// — an empty page a few seconds after a restart must not read as a stopped one.
func (h *Handler) lanesView() lanesView {
if h.lanes == nil {
return lanesView{PollerOff: true}
}
snap := h.lanes.LaneStatus()
v := lanesView{
Rows: make([]laneRow, 0, len(snap.Lanes)),
BrowserConfigured: snap.BrowserConfigured,
BrowserReachable: snap.BrowserReachable,
}
now := time.Now()
for _, l := range snap.Lanes {
lost := l.Browser && !snap.BrowserReachable
// Series waiting and none read is the shape of a Lane that has stopped
// working, as distinct from one that is quiet for want of work.
stalled := l.Due > 0 && l.Checked == 0
gap := ""
if l.Gap > 0 {
gap = l.Gap.Truncate(time.Second).String()
}
v.Rows = append(v.Rows, laneRow{
Site: l.Site,
Due: l.Due,
Ran: since(now, l.LastRun),
Checked: l.Checked,
Gap: gap,
Clamped: l.Clamped,
Refusing: l.Refusing,
BrowserLost: lost,
Stalled: stalled,
Attention: l.Clamped || l.Refusing || lost || stalled,
})
}
return v
}
// since formats how long ago a Lane last ran, at second resolution: the block
// refreshes every thirty seconds, so anything finer is noise the owner would
// have to ignore.
func since(now, then time.Time) string {
d := now.Sub(then).Truncate(time.Second)
if d < time.Second {
return "just now"
}
return d.String() + " ago"
}
// revokeReaderSessions logs one Reader out of every browser they are signed in
// on. The owner gate is the route's, not this handler's.
func (h *Handler) revokeReaderSessions(w http.ResponseWriter, r *http.Request) {
target, ok := readerPathID(w, r)
if !ok {
return
}
// The owner is not one of the Readers this endpoint reaches: revoking
// themselves would sign out the browser making the request, which is what
// logout is for. The roster hides the button; this refuses the hand-rolled
// POST behind it.
if target == h.store.OwnerID() {
http.NotFound(w, r)
return
}
if err := h.store.DeleteReaderSessions(target); err != nil {
log.Printf("revoke sessions: %v", err)
http.Error(w, "internal error", http.StatusInternalServerError)
return
}
h.renderRoster(w, "revoke sessions")
}
// clearReaderMarks zeroes one Reader's Sighting counters. The guard those
// counters feed has one known false positive — a Site changing its page shape
// makes a correct adapter read a wrong high number and marks every honest
// Reader of that Site at once (issue #103) — and this is its remedy. It
// restores a privilege rather than destroying anything, so the control is
// confirmed but never wears the destruction accent.
func (h *Handler) clearReaderMarks(w http.ResponseWriter, r *http.Request) {
target, ok := readerPathID(w, r)
if !ok {
return
}
if err := h.store.ClearReaderMarks(target); err != nil {
log.Printf("clear marks: %v", err)
http.Error(w, "internal error", http.StatusInternalServerError)
return
}
h.renderRoster(w, "clear marks")
}
// readerPathID reads the Reader a route names, answering the request itself
// when there is nobody to act on.
func readerPathID(w http.ResponseWriter, r *http.Request) (int64, bool) {
id, err := strconv.ParseInt(r.PathValue("id"), 10, 64)
if err != nil {
http.Error(w, "bad reader id", http.StatusBadRequest)
return 0, false
}
return id, true
}
// renderRoster answers an action with the whole roster, so the counts and marks
// it shows cannot describe the state before the tap.
func (h *Handler) renderRoster(w http.ResponseWriter, what string) {
readers, err := h.store.Readers()
if err != nil {
log.Printf("%s: %v", what, err)
http.Error(w, "internal error", http.StatusInternalServerError)
return
}
h.render(w, http.StatusOK, "readers", adminView{Readers: readers, OwnerID: h.store.OwnerID()})
}
+320
View File
@@ -0,0 +1,320 @@
package web
import (
"context"
"encoding/json"
"errors"
"fmt"
"io"
"log"
"net/http"
"net/url"
"slices"
"strconv"
"strings"
"sync"
"time"
"bookmarkmanager/backend/internal/session"
"bookmarkmanager/backend/internal/token"
)
const (
// oauthStateTTL bounds how long a started sign-in stays valid. Ten
// minutes is generous for Discord's round trip and short enough that a
// captured state is stale before it is worth replaying.
oauthStateTTL = 10 * time.Minute
// maxStates caps the state map so a flood of /auth/discord hits cannot
// grow memory; past the cap the oldest state is evicted, which at worst
// invalidates an in-flight sign-in.
maxStates = 256
// maxResponseBytes caps Discord API bodies; they are small, and an
// unbounded read is an OOM handed to Discord's CDN.
maxResponseBytes = 1 << 20
// discordTimeout keeps a hung Discord request from hanging the login
// callback forever.
discordTimeout = 15 * time.Second
)
// DiscordConfig is the OAuth application this service registers as, plus the
// guild that gates access.
type DiscordConfig struct {
ClientID string
ClientSecret string
GuildID string
// RequiredRole, when non-empty, is a role ID a member must hold on top of
// guild membership. Empty by default: membership alone suffices.
RequiredRole string
// APIBase is the Discord API root; configurable so tests run the whole
// flow against a local stub.
APIBase string
// RedirectURI is the full public URL of the callback — Discord requires
// the exact string, so it is configured, never derived from headers.
RedirectURI string
}
// oauthStates stores one-time sign-in states. A state is generated at
// /auth/discord, echoed back by Discord at the callback, and consumed there.
type oauthStates struct {
mu sync.Mutex
expiry map[string]time.Time
}
func newOAuthStates() *oauthStates {
return &oauthStates{expiry: make(map[string]time.Time)}
}
func (s *oauthStates) put(state string, expires time.Time) {
s.mu.Lock()
defer s.mu.Unlock()
now := time.Now()
for k, at := range s.expiry {
if !at.After(now) {
delete(s.expiry, k)
}
}
// Evict the state closest to expiring when full, so a flood of starts
// cannot grow memory; at worst it invalidates an in-flight sign-in.
if len(s.expiry) >= maxStates {
var oldest string
var oldestAt time.Time
for k, at := range s.expiry {
if oldest == "" || at.Before(oldestAt) {
oldest, oldestAt = k, at
}
}
delete(s.expiry, oldest)
}
s.expiry[state] = expires
}
// take validates and consumes a state in one step: a state works exactly
// once, which is what makes a replayed callback useless.
func (s *oauthStates) take(state string) bool {
s.mu.Lock()
defer s.mu.Unlock()
expires, ok := s.expiry[state]
if !ok || !expires.After(time.Now()) {
return false
}
delete(s.expiry, state)
return true
}
// discordStart begins the authorization code grant: a fresh state, then a
// redirect to Discord's authorize page.
func (h *Handler) discordStart(w http.ResponseWriter, r *http.Request) {
state := session.NewID()
h.states.put(state, time.Now().Add(oauthStateTTL))
u := h.discord.APIBase + "/oauth2/authorize?" + url.Values{
"client_id": {h.discord.ClientID},
"redirect_uri": {h.discord.RedirectURI},
"response_type": {"code"},
"scope": {"identify guilds.members.read"},
"state": {state},
}.Encode()
http.Redirect(w, r, u, http.StatusSeeOther)
}
// discordCallback completes the grant: exchange the code, verify identity,
// membership and role, then mint a session. Every failure path renders the
// login page with an author-written message — nothing Discord supplied is
// ever interpolated into a page, and no secret reaches a log line.
func (h *Handler) discordCallback(w http.ResponseWriter, r *http.Request) {
ip := session.ClientIP(r)
if wait := h.limiter.RetryAfter(ip, time.Now()); wait > 0 {
secs := int(wait.Seconds()) + 1
w.Header().Set("Retry-After", strconv.Itoa(secs))
h.renderLogin(w, http.StatusTooManyRequests,
"Too many attempts. Try again in "+strconv.Itoa((secs+59)/60)+" min.")
return
}
// Discord refuses the grant (the reader hit cancel, or the application
// was misconfigured). The state is consumed so the flow is cleanly over;
// this makes no Discord calls, so it is not a failure the limiter counts.
if oerr := r.URL.Query().Get("error"); oerr != "" {
h.states.take(r.URL.Query().Get("state"))
h.renderLogin(w, http.StatusBadRequest, "Sign-in was cancelled.")
return
}
code := r.URL.Query().Get("code")
if code == "" || !h.states.take(r.URL.Query().Get("state")) {
h.limiter.Fail(ip, time.Now())
h.renderLogin(w, http.StatusBadRequest,
"This sign-in link was invalid or already used. Start again.")
return
}
tok, err := h.exchangeToken(r.Context(), code)
if err != nil {
h.limiter.Fail(ip, time.Now())
log.Printf("discord token exchange: %v", err)
h.renderLogin(w, http.StatusBadGateway,
"Discord sign-in is unavailable right now. Try again in a moment.")
return
}
userID, err := h.discordUserID(r.Context(), tok.AccessToken)
if err != nil {
h.limiter.Fail(ip, time.Now())
log.Printf("discord users/@me: %v", err)
h.renderLogin(w, http.StatusBadGateway,
"Discord sign-in is unavailable right now. Try again in a moment.")
return
}
member, isMember, err := h.discordMember(r.Context(), tok.AccessToken)
if err != nil {
h.limiter.Fail(ip, time.Now())
log.Printf("discord member check: %v", err)
h.renderLogin(w, http.StatusBadGateway,
"Discord sign-in is unavailable right now. Try again in a moment.")
return
}
// The refusal is the same for a non-member and a member without the
// required role, and it names neither the guild nor its id: an outsider
// cannot tell whether the guild exists, let alone which one gates.
//
// It also returns before EnsureReader, so a refused sign-in leaves no
// Reader row behind — the gate is the only thing standing between guild
// membership and a library.
if !isMember || (h.discord.RequiredRole != "" && !slices.Contains(member.Roles, h.discord.RequiredRole)) {
h.limiter.Fail(ip, time.Now())
h.renderLogin(w, http.StatusForbidden,
"This Discord account is not a member of this community.")
return
}
// Registration is the login (issue #27): first sight of a guild member
// creates their Reader, every later sight returns the same one. Their
// userscript credential is derived at epoch 0 the way the owner's is, so
// the install links work before they have read anything.
readerID, err := h.store.EnsureReader(userID, token.Hash(token.Token(h.tokenKey, userID, 0)))
if err != nil {
log.Printf("register reader: %v", err)
http.Error(w, "internal error", http.StatusInternalServerError)
return
}
h.limiter.Reset(ip)
sess, err := h.store.CreateSession(session.NewID(), readerID, session.SessionTTL)
if err != nil {
log.Printf("create session: %v", err)
http.Error(w, "internal error", http.StatusInternalServerError)
return
}
session.SetCookie(w, r, sess.ID)
http.Redirect(w, r, "/", http.StatusSeeOther)
}
// exchangeToken trades an authorization code for an access token. The body is
// form-encoded because that is what Discord accepts — it rejects a JSON
// payload — so the wire format is fixed here, not in a client library.
func (h *Handler) exchangeToken(ctx context.Context, code string) (discordToken, error) {
form := url.Values{
"client_id": {h.discord.ClientID},
"client_secret": {h.discord.ClientSecret},
"grant_type": {"authorization_code"},
"code": {code},
"redirect_uri": {h.discord.RedirectURI},
}
req, err := http.NewRequestWithContext(ctx, http.MethodPost,
h.discord.APIBase+"/oauth2/token", strings.NewReader(form.Encode()))
if err != nil {
return discordToken{}, err
}
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
req.Header.Set("Accept", "application/json")
resp, err := h.httpClient.Do(req)
if err != nil {
return discordToken{}, err
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return discordToken{}, fmt.Errorf("status %d", resp.StatusCode)
}
var tok discordToken
if err := json.NewDecoder(io.LimitReader(resp.Body, maxResponseBytes)).Decode(&tok); err != nil {
return discordToken{}, err
}
if tok.AccessToken == "" {
return discordToken{}, errors.New("empty access token")
}
return tok, nil
}
// discordUserID fetches the signed-in user's id via the identify scope.
func (h *Handler) discordUserID(ctx context.Context, accessToken string) (string, error) {
req, err := http.NewRequestWithContext(ctx, http.MethodGet,
h.discord.APIBase+"/users/@me", nil)
if err != nil {
return "", err
}
req.Header.Set("Authorization", "Bearer "+accessToken)
req.Header.Set("Accept", "application/json")
resp, err := h.httpClient.Do(req)
if err != nil {
return "", err
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return "", fmt.Errorf("status %d", resp.StatusCode)
}
var u struct {
ID string `json:"id"`
}
if err := json.NewDecoder(io.LimitReader(resp.Body, maxResponseBytes)).Decode(&u); err != nil {
return "", err
}
if u.ID == "" {
return "", errors.New("empty user id")
}
return u.ID, nil
}
type discordMember struct {
Roles []string `json:"roles"`
}
// discordMember fetches the current user's membership in the configured guild.
//
// This is the OAuth endpoint (Get Current User Guild Member), the one the
// guilds.members.read scope grants. Its bot-side twin, GET /guilds/{id}/
// members/{user}, reads almost identically and is the wrong one: it wants a
// Bot token and the application present in the guild, and answers a user
// Bearer token with 401 — which fails as an outage rather than a refusal, so
// nobody could sign in at all.
//
// A 404 or 403 (not in the guild, or the token lacks the scope) is a
// non-member, not an error.
func (h *Handler) discordMember(ctx context.Context, accessToken string) (discordMember, bool, error) {
u := h.discord.APIBase + "/users/@me/guilds/" +
url.PathEscape(h.discord.GuildID) + "/member"
req, err := http.NewRequestWithContext(ctx, http.MethodGet, u, nil)
if err != nil {
return discordMember{}, false, err
}
req.Header.Set("Authorization", "Bearer "+accessToken)
req.Header.Set("Accept", "application/json")
resp, err := h.httpClient.Do(req)
if err != nil {
return discordMember{}, false, err
}
defer resp.Body.Close()
if resp.StatusCode == http.StatusNotFound || resp.StatusCode == http.StatusForbidden {
return discordMember{}, false, nil
}
if resp.StatusCode != http.StatusOK {
return discordMember{}, false, fmt.Errorf("status %d", resp.StatusCode)
}
var m discordMember
if err := json.NewDecoder(io.LimitReader(resp.Body, maxResponseBytes)).Decode(&m); err != nil {
return discordMember{}, false, err
}
return m, true, nil
}
type discordToken struct {
AccessToken string `json:"access_token"`
}
+55
View File
@@ -0,0 +1,55 @@
package web
import (
"testing"
"time"
)
func TestOAuthStateSingleUse(t *testing.T) {
s := newOAuthStates()
s.put("st", time.Now().Add(time.Minute))
if !s.take("st") {
t.Fatal("take of a fresh state = false, want true")
}
if s.take("st") {
t.Fatal("take of a consumed state = true, want false")
}
}
func TestOAuthStateUnknownOrExpired(t *testing.T) {
s := newOAuthStates()
if s.take("never-seen") {
t.Fatal("take of an unknown state = true, want false")
}
s.put("stale", time.Now().Add(-time.Minute))
if s.take("stale") {
t.Fatal("take of an expired state = true, want false")
}
}
// The map is capped: a flood of starts evicts older states instead of growing,
// and consumed states must not change that.
func TestOAuthStateEviction(t *testing.T) {
s := newOAuthStates()
key := func(i, salt int) string {
return string(rune('a'+i%26)) + string(rune('0'+i/26+salt*16))
}
now := time.Now().Add(time.Hour)
for i := 0; i < maxStates*2; i++ {
s.put(key(i, 0), now)
}
if got := len(s.expiry); got != maxStates {
t.Fatalf("states after a flood = %d, want %d", got, maxStates)
}
// Consume everything, then flood again: the map stays bounded.
for state := range s.expiry {
s.take(state)
}
for i := 0; i < maxStates; i++ {
s.put(key(i, 1), now)
}
if got := len(s.expiry); got != maxStates {
t.Fatalf("states after consume+flood = %d, want %d", got, maxStates)
}
}
+130 -21
View File
@@ -87,6 +87,11 @@
--moss: #7fae86; /* finished */
--clay: #b5906f; /* set chapter */
--trash: #977671; /* remove, resting — icons need 3:1, not 4.5:1 */
/* A Lane needing attention: the admin page's only accent. Verdigris — cool,
the far side of the wheel from ember's crimson, and clear of the archive
blue. Neither ember (new chapter) nor danger (destruction) may say
"system unhealthy". */
--patina: #5fb3a6;
/* Desktop cell borders for the two coloured action states. */
--play-hot-line: #3a1d18;
@@ -146,6 +151,7 @@
--moss: #3d6c46;
--clay: #7c5533;
--trash: #8c6558;
--patina: #1f6f66;
--play-hot-line: #f0cfc6;
--fav-line: #e3d3a4;
@@ -243,6 +249,117 @@ button { cursor: pointer; }
/* The label is 15px tall by design; the thumb gets 44 without moving it. */
.ghost::after { content: ""; position: absolute; inset: -15px -12px; }
/* ---- userscript setup: collapsed by default, one hairline, no card ---- */
.setup {
margin: 0 20px;
padding: 12px 0 0;
border-bottom: 1px solid var(--rule);
color: var(--mute);
}
.setup summary {
display: flex;
align-items: center;
min-height: 44px;
padding: 0;
font: 500 10px/1 var(--font-mono);
letter-spacing: .2em;
text-transform: uppercase;
color: var(--mute-2);
cursor: pointer;
list-style: none;
}
.setup summary::-webkit-details-marker { display: none; }
.setup summary:hover { color: var(--paper); }
.setup[open] { padding-bottom: 16px; }
.setup-copy {
margin: 0;
padding: 4px 0 12px;
font: 14px/1.55 var(--font-body);
color: var(--mute);
}
.setup-links {
display: flex;
flex-wrap: wrap;
gap: 8px 20px;
margin: 0 0 14px;
}
.setup-links .ghost { font-size: 11px; }
.setup-rotate { margin: 0; }
/* Rotation confirmation: the one hot state the panel wears, and it is
destruction, not new-chapter signal — danger, never ember. */
.setup-warn {
margin: 0;
padding: 10px 12px;
border: 1px solid var(--danger);
color: var(--danger);
font: 500 12px/1.5 var(--font-mono);
letter-spacing: .04em;
}
/* ---- admin page: two sections on the same measured sheet, no cards ----
The reading page is a list of series; this is a list of facts. Both are
sheets of hairline-separated rows, so the roster keeps the shape it had as
a fold-out and the Lane block copies it. */
.readers, .lanes { margin: 0 20px; padding: 12px 0 16px; border-bottom: 1px solid var(--rule); }
.readers h2, .lanes h2 {
margin: 0;
padding: 8px 0;
font: 500 10px/1 var(--font-mono);
letter-spacing: .2em;
text-transform: uppercase;
color: var(--mute-2);
}
.readerlist, .lanelist { margin: 0; padding: 0; list-style: none; }
.readerlist li, .lanelist li {
display: flex;
align-items: center;
flex-wrap: wrap;
gap: 4px 16px;
min-height: 44px;
border-top: 1px solid var(--rule);
}
.reader-actions { display: flex; gap: 18px; margin-left: auto; }
.readerlist form { margin: 0; }
.reader-id {
font: 500 13px/1.4 var(--font-mono);
letter-spacing: .04em;
color: var(--paper);
}
.reader-sessions {
font: 500 10px/1 var(--font-mono);
letter-spacing: .14em;
text-transform: uppercase;
color: var(--mute);
}
.reader-sightings, .lane-fact {
font: 500 10px/1 var(--font-mono);
letter-spacing: .14em;
text-transform: uppercase;
color: var(--mute-2);
}
/* Two states the owner is meant to find rather than read for: a Reader whose
reports no longer defer a Poll, and a Lane that is not keeping its promise.
Both wear --patina — never ember, which means one thing, and never danger,
which is destruction. */
.reader-blocked, .lane-mark {
font: 500 10px/1 var(--font-mono);
letter-spacing: .14em;
text-transform: uppercase;
color: var(--patina);
}
.lane-site {
font: 400 19px/1.2 var(--font-display);
color: var(--paper-dim);
}
/* The whole row leans patina when the Lane needs attention, so the scan is one
pass down the left edge rather than a read of every mark. */
.lanelist li.attention .lane-site { color: var(--patina); }
.lane-browser { padding: 12px 0 0; }
/* Revocation cuts someone off, so it wears --danger. Ember stays reserved for
the new-chapter signal. */
.ghost.danger { color: var(--danger); }
.ghost.danger:hover { color: var(--danger); border-bottom-color: var(--danger); }
.chrome { display: flex; flex-direction: column; }
.searchbar {
@@ -526,6 +643,9 @@ button { cursor: pointer; }
box-shadow: inset 0 -2px 0 var(--ember);
}
.topbar form { margin-left: 18px; }
/* The admin page's topbar has no switch to fill the middle, so its back link
keeps company with Log out at the right edge instead of floating centre. */
.topbar .back { margin-left: auto; }
/* At phone width brand + switch + Log out do not fit on one line, so the
switch takes its own row under the wordmark rather than pushing Log out
off-screen. */
@@ -535,6 +655,9 @@ button { cursor: pointer; }
.libswitch { order: 3; margin-left: 0; }
.libswitch a { flex: 1; text-align: center; padding: 8px 14px; }
.topbar form { margin-left: 12px; }
/* The admin page has no switch to take the second row, so its brand claims
the first outright and the back link keeps Log out company below. */
.topbar:has(.back) .brand { flex: 1 1 100%; }
}
/* ---- action strip: full-width on a phone, hairline-divided cells ---- */
@@ -774,26 +897,6 @@ button { cursor: pointer; }
filter: drop-shadow(0 0 34px var(--ember-wash)) drop-shadow(0 18px 24px rgba(0,0,0,.5));
}
.login-card form { display: flex; flex-direction: column; gap: 18px; }
.login-card label {
font: 500 10px/1 var(--font-mono);
letter-spacing: .16em;
text-transform: uppercase;
color: var(--mute);
}
.login-card input {
width: 100%;
height: 54px;
margin-top: 9px;
padding: 0 2px;
border: none;
border-bottom: 1px solid var(--field-line);
background: transparent;
color: var(--paper);
font: 500 20px var(--font-mono);
letter-spacing: .16em;
outline: none;
}
.login-card input:focus { border-bottom-color: var(--paper); }
.login-card .error {
margin: 0;
min-height: 20px;
@@ -810,7 +913,13 @@ button { cursor: pointer; }
.login-card button:hover {
background: var(--ember);
border-color: var(--ember);
color: #fff;
color: var(--ember-ink);
}
.login-card .login-note {
margin: 14px 0 0;
text-align: center;
font: 400 12px/1.4 var(--font-body);
color: var(--mute);
}
/* ---- laptop and up: the whole sheet is drawn 20% larger, which is what
+39
View File
@@ -0,0 +1,39 @@
{{/* The owner's administrative page: everything that reaches past one Reader,
at its own address so it can be bookmarked rather than hunted for inside
the reading page. Owner-only at route registration (requireOwner), which
is why nothing in here re-tests who is asking. */}}
{{define "admin"}}
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1, viewport-fit=cover">
<meta name="color-scheme" content="dark light">
<title>BookmarkManager — Admin</title>
<link rel="icon" href="/static/logo.svg" type="image/svg+xml">
<link rel="stylesheet" href="/static/style.css">
<link rel="preload" href="/static/fonts/instrument-serif-400-latin.woff2" as="font" type="font/woff2" crossorigin>
<script src="/static/htmx.min.js" defer></script>
</head>
<body>
<div class="sheet">
<header class="topbar">
<h1 class="brand">{{template "mark" .}}<span>Bookmark<em>Manager</em></span></h1>
{{/* Back to the library, no switch: this page belongs to neither library,
and the ember-lit switch says which library you are reading. */}}
<a class="ghost back" href="/">Library</a>
<form method="post" action="/logout">
<button type="submit" class="ghost">Log out</button>
</form>
</header>
{{/* The live region wraps the swapped block rather than being it: the
refresh replaces the section wholesale, and a region recreated on every
update is never announced. */}}
<div aria-live="polite">{{template "lanes" .Lanes}}</div>
{{template "readers" .}}
</div>
</body>
</html>
{{end}}
+6
View File
@@ -28,6 +28,10 @@
<a href="/?lib=novel&amp;tab=all" class="{{if eq .Lib "novel"}}active{{end}}"
{{if eq .Lib "novel"}}aria-current="page"{{end}}>Novels</a>
</nav>
{{/* The owner's only difference on this page: a link out to the
administrative one. It sits beside Log out rather than in the library
switch — that switch says which library, not which page. */}}
{{if .Owner}}<a class="ghost" href="/admin">Admin</a>{{end}}
<form method="post" action="/logout">
<button type="submit" class="ghost">Log out</button>
</form>
@@ -74,6 +78,8 @@
</nav>
</div>
{{template "setup" .}}
{{template "keyrow" .}}
{{template "recent" .}}
+41
View File
@@ -0,0 +1,41 @@
{{/* Poll Lane status: one row per Site, refreshing itself so a run can be
watched rather than sampled by reloading. The refresh is one attribute on
the fragment root and the endpoint answers with this same fragment, so the
swap replaces the element that asked for it.
Every figure here is read out of the running poller, never out of a table:
a Site absent from Rows has not completed a pass since the last restart,
which the empty state must say — zeroes would read as a stopped Lane. */}}
{{define "lanes"}}
<section class="lanes" id="lanes"
hx-get="/ui/admin/lanes" hx-trigger="every 30s" hx-swap="outerHTML">
<h2>Poll Lanes</h2>
{{if .Rows}}
<ul class="lanelist">
{{range .Rows}}
<li{{if .Attention}} class="attention"{{end}}>
<span class="lane-site">{{.Site}}</span>
<span class="lane-fact">{{.Due}} due</span>
<span class="lane-fact">{{.Checked}} checked</span>
<span class="lane-fact">ran {{.Ran}}</span>
{{if .Gap}}<span class="lane-fact">gap {{.Gap}}</span>{{end}}
{{if .Clamped}}<span class="lane-mark">gap at floor</span>{{end}}
{{if .Refusing}}<span class="lane-mark">refusing</span>{{end}}
{{if .BrowserLost}}<span class="lane-mark">no browser</span>{{end}}
{{if .Stalled}}<span class="lane-mark">not checking</span>{{end}}
</li>
{{end}}
</ul>
{{else}}
<p class="setup-copy">No data yet — no Lane has completed a pass since the
backend started.</p>
{{end}}
<p class="setup-copy lane-browser">
{{if .PollerOff}}Polling is switched off in this deployment: no Lane runs,
and Latest Chapter comes from the userscripts alone.
{{else}}Browser sidecar:
{{if not .BrowserConfigured}}not configured — comix, kagane and novelfull
pages are not fetched through it{{else if .BrowserReachable}}reachable
{{else}}unreachable{{end}}.{{end}}</p>
</section>
{{end}}
+11
View File
@@ -16,6 +16,17 @@
<div class="empty"><strong>Nothing archived.</strong><p>Shelve a series to park it here — it keeps getting checked for new chapters.</p></div>
{{else if eq .Tab "finished"}}
<div class="empty"><strong>Nothing finished yet.</strong><p>Mark a series finished and it moves out of your reading list.</p></div>
{{else if .EmptyLibrary}}
{{/* Nothing in either library, so the links are the only thing this page can
usefully say. Both scripts: the two libraries are separate installs. */}}
<div class="empty">
<strong>Nothing here yet.</strong>
<p>Install the userscripts, then open a series and read a chapter — bookmarks arrive on their own.</p>
<p class="setup-links">
<a class="ghost" href="/install/manga-bookmark.user.js">Install Manga script</a>
<a class="ghost" href="/install/novel-bookmark.user.js">Install Novels script</a>
</p>
</div>
{{else}}
<div class="empty"><strong>Nothing here yet.</strong><p>Bookmarks appear once the userscript records a chapter.</p></div>
{{end}}
+3 -7
View File
@@ -19,17 +19,13 @@
<figure class="login-art" aria-hidden="true">
<img src="/static/login-art.png" alt="">
</figure>
<form method="post" action="/login">
<div>
<label for="password">Password</label>
<input id="password" name="password" type="password"
autocomplete="current-password" autofocus required>
</div>
<form method="get" action="/auth/discord">
{{/* The page reloads on a failed sign-in, so the message is present from
the start; role=alert is what gets it announced anyway. */}}
<p class="error" role="alert">{{.Error}}</p>
<button type="submit">Sign in</button>
<button type="submit">Continue with Discord</button>
</form>
<p class="login-note">Guild membership is required to sign in.</p>
</main>
</body>
</html>
@@ -0,0 +1,48 @@
{{/* The Reader roster, on the owner's administrative page. Re-rendered whole
as the response to an action so the counts and marks it shows cannot
describe the state before the tap. Both actions are confirm-gated: one
signs a Reader out of every device at once, the other wipes a record.
Owner-only at route registration, so nothing here re-tests who is asking. */}}
{{define "readers"}}
<section class="readers" id="readers">
<h2>Readers</h2>
<p class="setup-copy">Everyone who has signed in through Discord. Revoking
signs a Reader out of every device; their library and bookmarks are
untouched, and they can sign in again. The Sighting counters record how
often a later Poll confirmed or contradicted what that Reader's browser
reported; enough contradictions stop their reports deferring a Poll, and
clearing the marks gives that back.</p>
<ul class="readerlist">
{{range .Readers}}
<li>
<span class="reader-id">{{.DiscordID}}</span>
<span class="reader-sessions">{{.Sessions}} session{{if ne .Sessions 1}}s{{end}}</span>
<span class="reader-sightings">{{.Agreements}} confirmed / {{.Disagreements}} contradicted</span>
{{/* Blocked is spelled out rather than left to be worked out from two
numbers and a threshold. */}}
{{if .Blocked}}<span class="reader-blocked">deferral blocked</span>{{end}}
<span class="reader-actions">
{{/* Clearing restores a privilege, so it is a plain ghost button —
the destruction accent belongs to revocation alone. It is offered
on every row, including one reading zero: the remedy must be
findable before the counters climb, not after. */}}
<form hx-post="/readers/{{.ID}}/clear-marks" hx-target="#readers" hx-swap="outerHTML"
hx-confirm="Clearing wipes this Reader's whole Sighting record, confirmations included. Clear?">
<button type="submit" class="ghost">Clear marks</button>
</form>
{{/* The owner's own row never offers Revoke: it is the one row where the
button would sign the tapping browser out, and the endpoint refuses
it anyway. Logout is the deliberate way to do that. */}}
{{if and .Sessions (ne .ID $.OwnerID)}}
<form hx-post="/readers/{{.ID}}/revoke" hx-target="#readers" hx-swap="outerHTML"
hx-confirm="Revoking signs this Reader out on every device immediately. Revoke?">
<button type="submit" class="ghost danger">Revoke sessions</button>
</form>
{{end}}
</span>
</li>
{{end}}
</ul>
</section>
{{end}}
+34
View File
@@ -0,0 +1,34 @@
{{/* The userscript install panel. Each link serves the script rendered
with the acting Reader's credential inside it, so the credential never
appears in this page's markup, the address bar, or a redirect. Rotation
is confirm-gated because it invalidates every installed copy at once;
the response swaps this same panel open with the reinstall warning. */}}
{{define "setup"}}
<details class="setup" id="setup"{{if .Rotated}} open{{end}}>
<summary>Userscripts</summary>
<p class="setup-copy">Install each script once per device. They keep your
bookmarks in sync across every site and update themselves from here.</p>
<p class="setup-links">
<a class="ghost" href="/install/manga-bookmark.user.js">Install Manga script</a>
<a class="ghost" href="/install/novel-bookmark.user.js">Install Novels script</a>
</p>
<p class="setup-copy">On mobile, Violentmonkey does not pick up the install
links — the script opens as text. Download the file instead, then add it
from Violentmonkey's own menu.</p>
<p class="setup-links">
<a class="ghost" href="/install/manga-bookmark.user.js?download=1">Download Manga script</a>
<a class="ghost" href="/install/novel-bookmark.user.js?download=1">Download Novels script</a>
</p>
{{if .Rotated}}
<p class="setup-warn" role="status">Credential rotated — the old one no
longer works. Reinstall both scripts on every device now, or they will
silently stop syncing.</p>
{{else}}
<form class="setup-rotate" hx-post="/rotate-token" hx-target="#setup"
hx-swap="outerHTML"
hx-confirm="Rotation invalidates the current credential on every device immediately. You will have to reinstall both scripts everywhere. Rotate?">
<button type="submit" class="ghost">Rotate credential</button>
</form>
{{end}}
</details>
{{end}}
+174 -57
View File
@@ -1,7 +1,7 @@
package web
import (
"crypto/subtle"
"context"
"embed"
"html/template"
"io/fs"
@@ -16,6 +16,8 @@ import (
"bookmarkmanager/backend/internal/session"
"bookmarkmanager/backend/internal/store"
"bookmarkmanager/backend/internal/token"
"bookmarkmanager/backend/internal/userscript"
)
//go:embed templates
@@ -31,11 +33,26 @@ const RecentCount = 5
// It is a separate handler from api.Handler because the two speak different
// representations (HTML versus JSON) to different clients under different auth.
type Handler struct {
store *store.Store
tmpl *template.Template
key []byte
password string
limiter *session.LoginLimiter
store *store.Store
// tokenKey derives Readers' userscript credentials (internal/token): the
// install endpoints render the scripts with the credential inside, which
// is the one place the UI needs the secret.
tokenKey []byte
// mangaUserscriptPath / novelUserscriptPath are the bindmounted script
// files the install endpoints render — the same files the /u/ download
// paths serve.
mangaUserscriptPath string
novelUserscriptPath string
tmpl *template.Template
discord DiscordConfig
states *oauthStates
limiter *session.LoginLimiter
// httpClient is the plain stdlib client that talks to Discord. It is not
// an injected interface: tests point APIBase at a stub server instead.
httpClient *http.Client
// lanes is the Poll Lane snapshot source the administrative page reads.
// Nil is a running deployment with no poller, not a bug.
lanes LaneReporter
}
// listView is what every list-rendering template receives.
@@ -43,7 +60,7 @@ type listView struct {
// Lib is the library this view renders: store.KindManga or store.KindNovel.
// Manga is the default and carries no query parameter, so every pre-novel
// URL keeps meaning exactly what it did.
Lib string
Lib string
Tab string // "all", "fav", or "new"
Recent []store.Bookmark
Items []store.Bookmark
@@ -55,6 +72,17 @@ type listView struct {
// OOB marks a render of the chrome partials as an out-of-band swap rather
// than the inline copy app.html lays out.
OOB bool
// Rotated marks the setup panel as having just rotated the credential:
// it swaps the reinstall warning in over the button row.
Rotated bool
// EmptyLibrary means this Reader holds no bookmarks in either library, so
// the empty state can offer the installs instead of reporting on a filter.
// It is not "newly registered": a Reader who deletes their last bookmark is
// in the same position and needs the same links.
EmptyLibrary bool
// Owner marks the acting Reader as the deployment's owner, which offers
// the link to the administrative page. Nothing else in the UI differs.
Owner bool
}
// PageURL and ListURL are the two link shapes every tab needs. Building them
@@ -81,23 +109,32 @@ type loginView struct {
// New parses every template up front so a broken one kills the process at
// startup rather than the first request that touches it.
func New(s *store.Store, apiToken, webPassword string) (*Handler, error) {
//
// lanes is the administrative page's window onto the running Poller; nil means
// nothing is polling, which the page reports rather than hides.
func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, novelPath string, lanes LaneReporter) (*Handler, error) {
tmpl, err := template.ParseFS(templateFS, "templates/*.html")
if err != nil {
return nil, err
}
return &Handler{
store: s,
tmpl: tmpl,
key: session.Key(apiToken, webPassword),
password: webPassword,
limiter: session.NewLoginLimiter(),
store: s,
tokenKey: tokenKey,
mangaUserscriptPath: mangaPath,
novelUserscriptPath: novelPath,
tmpl: tmpl,
discord: discord,
states: newOAuthStates(),
limiter: session.NewLoginLimiter(),
httpClient: &http.Client{Timeout: discordTimeout},
lanes: lanes,
}, nil
}
func (h *Handler) Register(mux *http.ServeMux) {
mux.HandleFunc("GET /{$}", h.index)
mux.HandleFunc("POST /login", h.login)
mux.HandleFunc("GET /auth/discord", h.discordStart)
mux.HandleFunc("GET /auth/discord/callback", h.discordCallback)
mux.HandleFunc("POST /logout", h.logout)
mux.Handle("GET /static/", staticHandler())
@@ -106,6 +143,19 @@ func (h *Handler) Register(mux *http.ServeMux) {
mux.HandleFunc("POST /ui/bookmarks/{key}/status", h.requireSession(h.uiStatus))
mux.HandleFunc("POST /ui/bookmarks/{key}/chapter", h.requireSession(h.uiChapter))
mux.HandleFunc("DELETE /ui/bookmarks/{key}", h.requireSession(h.uiDelete))
// Install endpoints render the script directly under the session: the
// credential travels inside the served bytes, never in the address bar or
// the page markup. Updates after install use the credential-bearing /u/
// path the script embeds, which needs no session.
mux.HandleFunc("GET /install/manga-bookmark.user.js", h.requireSession(h.installUserscript("manga-bookmark.user.js")))
mux.HandleFunc("GET /install/novel-bookmark.user.js", h.requireSession(h.installUserscript("novel-bookmark.user.js")))
mux.HandleFunc("POST /rotate-token", h.requireSession(h.rotateToken))
// Owner-only: every route that reaches past the acting Reader is gated in
// one place, so a missing gate is visible in the route list.
for _, rt := range h.adminRoutes() {
mux.HandleFunc(rt.pattern, h.requireOwner(rt.handler))
}
}
// staticHandler serves the embedded assets. An hour, not longer: assets are
@@ -131,10 +181,26 @@ func staticHandler() http.Handler {
}))
}
// authed reports whether the request carries a valid session cookie.
func (h *Handler) authed(r *http.Request) bool {
type ctxKey int
// readerCtxKey is where requireSession stashes the authenticated Reader id.
const readerCtxKey ctxKey = iota
// sessionReader reports whether the request carries a live session, and for
// whom. The cookie holds only the session id; the row behind it is looked up
// on every request, so deleting a session takes effect immediately. Expiry is
// enforced here, in the store, which also removes rows that have lapsed.
func (h *Handler) sessionReader(r *http.Request) (int64, bool) {
c, err := r.Cookie(session.CookieName)
return err == nil && session.Verify(h.key, c.Value, time.Now().UnixMilli())
if err != nil {
return 0, false
}
sess, ok, err := h.store.GetSession(c.Value, time.Now())
if err != nil {
log.Printf("session lookup: %v", err)
return 0, false
}
return sess.ReaderID, ok
}
// requireSession guards the fragment endpoints. It answers 401 rather than
@@ -142,14 +208,18 @@ func (h *Handler) authed(r *http.Request) bool {
// redirected login page would be spliced into the card list.
func (h *Handler) requireSession(next http.HandlerFunc) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
if !h.authed(r) {
readerID, ok := h.sessionReader(r)
if !ok {
http.Error(w, "unauthorized", http.StatusUnauthorized)
return
}
next(w, r)
next(w, r.WithContext(context.WithValue(r.Context(), readerCtxKey, readerID)))
}
}
// readerOf returns the authenticated Reader id requireSession stashed.
func readerOf(r *http.Request) int64 { return r.Context().Value(readerCtxKey).(int64) }
func (h *Handler) render(w http.ResponseWriter, status int, name string, data any) {
w.Header().Set("Content-Type", "text/html; charset=utf-8")
w.WriteHeader(status)
@@ -163,16 +233,20 @@ func (h *Handler) render(w http.ResponseWriter, status int, name string, data an
// page is served at / with status 200 rather than as a redirect to a separate
// URL: one page, no redirect loop to reason about.
func (h *Handler) index(w http.ResponseWriter, r *http.Request) {
if !h.authed(r) {
readerID, ok := h.sessionReader(r)
if !ok {
h.render(w, http.StatusOK, "login", loginView{})
return
}
view, err := h.buildListView(libOf(r.URL.Query().Get("lib")), r.URL.Query().Get("tab"))
view, err := h.buildListView(readerID, libOf(r.URL.Query().Get("lib")), r.URL.Query().Get("tab"))
if err != nil {
log.Printf("index: %v", err)
http.Error(w, "internal error", http.StatusInternalServerError)
return
}
// The owner's page differs only by the link to the administrative page:
// the roster lives there now, so the page read every day is only reading.
view.Owner = readerID == h.store.OwnerID()
h.render(w, http.StatusOK, "app", view)
}
@@ -208,18 +282,22 @@ func libOf(q string) string {
return store.KindManga
}
// buildListView loads the list once and derives both the tab-filtered items and
// the recent strip from it.
// buildListView loads one reader's list once and derives both the tab-filtered
// items and the recent strip from it.
//
// Archived and finished series appear in their own tab and nowhere else — not
// in All, not in Updated, not in Favourites, and not in the recent strip. An
// archived favourite therefore shows only under Archived: Favourites means
// "favourites I am currently reading".
func (h *Handler) buildListView(lib, tab string) (listView, error) {
all, err := h.store.List() // already ordered updated_at DESC
func (h *Handler) buildListView(readerID int64, lib, tab string) (listView, error) {
all, err := h.store.List(readerID) // already ordered updated_at DESC
if err != nil {
return listView{}, err
}
// Taken before the filter narrows the slice: a Reader with novels but no
// manga has a working install already, and does not need to be told to go
// and get one.
emptyLibrary := len(all) == 0
// Narrow to one library first: reading, withNew and recent all derive from
// this slice, so doing it later would let the other library's rows into the
// strip and the Updated badge.
@@ -263,11 +341,12 @@ func (h *Handler) buildListView(lib, tab string) (listView, error) {
recent = recent[:RecentCount]
}
}
return listView{Lib: lib, Tab: tab, Recent: recent, Items: items, NewCount: len(withNew)}, nil
return listView{Lib: lib, Tab: tab, Recent: recent, Items: items,
NewCount: len(withNew), EmptyLibrary: emptyLibrary}, nil
}
func (h *Handler) uiList(w http.ResponseWriter, r *http.Request) {
view, err := h.buildListView(libOf(r.URL.Query().Get("lib")), r.URL.Query().Get("tab"))
view, err := h.buildListView(readerOf(r), libOf(r.URL.Query().Get("lib")), r.URL.Query().Get("tab"))
if err != nil {
log.Printf("ui list: %v", err)
http.Error(w, "internal error", http.StatusInternalServerError)
@@ -325,7 +404,7 @@ func (h *Handler) writeChromeOOB(w http.ResponseWriter, view listView) {
// refreshChrome rebuilds the chrome for the reader's current tab after a
// mutation and appends it to the response.
func (h *Handler) refreshChrome(w http.ResponseWriter, r *http.Request) {
view, err := h.buildListView(currentLib(r), currentTab(r))
view, err := h.buildListView(readerOf(r), currentLib(r), currentTab(r))
if err != nil {
log.Printf("ui chrome: %v", err)
return
@@ -333,35 +412,21 @@ func (h *Handler) refreshChrome(w http.ResponseWriter, r *http.Request) {
h.writeChromeOOB(w, view)
}
func (h *Handler) login(w http.ResponseWriter, r *http.Request) {
ip := session.ClientIP(r)
if wait := h.limiter.RetryAfter(ip, time.Now()); wait > 0 {
secs := int(wait.Seconds()) + 1
w.Header().Set("Retry-After", strconv.Itoa(secs))
h.render(w, http.StatusTooManyRequests, "login", loginView{
Error: "Too many attempts. Try again in " +
strconv.Itoa((secs+59)/60) + " min.",
})
return
}
if err := r.ParseForm(); err != nil {
http.Error(w, "invalid form", http.StatusBadRequest)
return
}
got := r.PostFormValue("password")
if subtle.ConstantTimeCompare([]byte(got), []byte(h.password)) != 1 {
h.limiter.Fail(ip, time.Now())
h.render(w, http.StatusUnauthorized, "login", loginView{Error: "Wrong password."})
return
}
h.limiter.Reset(ip)
session.SetCookie(w, r, h.key)
http.Redirect(w, r, "/", http.StatusSeeOther)
// renderLogin renders the login page with an error message, for refused or
// failed sign-ins. Every message is author-written text — nothing Discord
// supplied is ever interpolated into a page.
func (h *Handler) renderLogin(w http.ResponseWriter, status int, msg string) {
h.render(w, status, "login", loginView{Error: msg})
}
// logout revokes the session row and clears the cookie in one step: the next
// request finds no row and is rejected.
func (h *Handler) logout(w http.ResponseWriter, r *http.Request) {
if c, err := r.Cookie(session.CookieName); err == nil {
if err := h.store.DeleteSession(c.Value); err != nil {
log.Printf("delete session: %v", err)
}
}
session.ClearCookie(w, r)
http.Redirect(w, r, "/", http.StatusSeeOther)
}
@@ -374,7 +439,7 @@ func (h *Handler) loadForMutation(w http.ResponseWriter, r *http.Request) (store
http.Error(w, "missing key", http.StatusBadRequest)
return store.Bookmark{}, false
}
b, ok, err := h.store.Get(key)
b, ok, err := h.store.Get(readerOf(r), key)
if err != nil {
log.Printf("ui get %q: %v", key, err)
http.Error(w, "internal error", http.StatusInternalServerError)
@@ -397,7 +462,7 @@ func (h *Handler) loadForMutation(w http.ResponseWriter, r *http.Request) (store
// describe the whole library, so they are rebuilt out of band on every
// mutation, at the cost of one extra list read per toggle.
func (h *Handler) saveAndRenderCard(w http.ResponseWriter, r *http.Request, b store.Bookmark) {
stored, err := h.store.Upsert(b)
stored, err := h.store.Upsert(readerOf(r), b)
if err != nil {
log.Printf("ui upsert %q: %v", b.Key, err)
http.Error(w, "internal error", http.StatusInternalServerError)
@@ -489,7 +554,7 @@ func (h *Handler) uiDelete(w http.ResponseWriter, r *http.Request) {
http.Error(w, "missing key", http.StatusBadRequest)
return
}
if err := h.store.Delete(key); err != nil {
if err := h.store.Delete(readerOf(r), key); err != nil {
log.Printf("ui delete %q: %v", key, err)
http.Error(w, "internal error", http.StatusInternalServerError)
return
@@ -500,3 +565,55 @@ func (h *Handler) uiDelete(w http.ResponseWriter, r *http.Request) {
// the library got smaller.
h.refreshChrome(w, r)
}
// installUserscript renders the bindmounted script with the acting Reader's
// derived credential substituted in. The credential is derived, not stored,
// so installs work after any restart; the Reader never types or copies it —
// clicking Install is the whole setup.
//
// ?download=1 forces a save instead. Mobile Violentmonkey (Chromium) does not
// intercept navigation to a .user.js URL, so the Install link only renders the
// source as text there; the Reader needs the file on disk to add it by hand.
func (h *Handler) installUserscript(name string) http.HandlerFunc {
path := h.mangaUserscriptPath
if name == "novel-bookmark.user.js" {
path = h.novelUserscriptPath
}
return func(w http.ResponseWriter, r *http.Request) {
discordID, epoch, err := h.store.ReaderTokenInfo(readerOf(r))
if err != nil {
log.Printf("install %s: %v", name, err)
http.Error(w, "internal error", http.StatusInternalServerError)
return
}
if r.URL.Query().Has("download") {
w.Header().Set("Content-Disposition", `attachment; filename="`+name+`"`)
}
userscript.Render(w, r, path, token.Token(h.tokenKey, discordID, epoch))
}
}
// rotateToken issues the acting Reader a new credential: the epoch bumps and
// the stored hash is rewritten, so the old credential stops authenticating
// the moment the statement commits. Every device must reinstall, or its
// script keeps failing silently — the setup panel states that warning next
// to the button, and the response repeats it as confirmation.
func (h *Handler) rotateToken(w http.ResponseWriter, r *http.Request) {
readerID := readerOf(r)
discordID, epoch, err := h.store.ReaderTokenInfo(readerID)
if err != nil {
log.Printf("rotate token: %v", err)
http.Error(w, "internal error", http.StatusInternalServerError)
return
}
// The hash is computed for epoch+1 and guarded by it in the store, so a
// concurrent rotation cannot leave the stored hash describing another
// epoch.
if err := h.store.RotateToken(readerID, epoch, token.Hash(token.Token(h.tokenKey, discordID, epoch+1))); err != nil {
log.Printf("rotate token: %v", err)
http.Error(w, "internal error", http.StatusInternalServerError)
return
}
view := listView{Lib: store.KindManga, Rotated: true}
h.render(w, http.StatusOK, "setup", view)
}
+187 -129
View File
@@ -7,7 +7,6 @@ import (
"net/http"
"os"
"os/signal"
"strconv"
"strings"
"syscall"
"time"
@@ -16,18 +15,38 @@ import (
"bookmarkmanager/backend/internal/httpmw"
"bookmarkmanager/backend/internal/latest"
"bookmarkmanager/backend/internal/store"
"bookmarkmanager/backend/internal/token"
"bookmarkmanager/backend/internal/userscript"
"bookmarkmanager/backend/internal/web"
)
// Config holds all runtime settings, sourced from environment variables.
type Config struct {
Token string
// TokenKey derives every Reader's userscript credential (internal/token).
// Required: without it no install URL can ever be built.
TokenKey string
AllowedOrigins []string
DBPath string
Port string
// WebPassword gates the browser UI. Empty disables the web routes entirely.
WebPassword string
// DatabaseURL is the Postgres connection URL; required, no default,
// because a wrong guess would silently start on an empty database.
DatabaseURL string
// CoverDir is the filesystem volume for immutable cover bytes. Required:
// serving a stored address without durable bytes would be worse than a
// startup failure.
CoverDir string
// PublicBaseURL is the origin this deployment answers on, e.g.
// "https://bookmarks.example.com". Required: cover URLs go out absolute
// because the userscript renders them on third-party origins, where a
// relative path would resolve against the Site (ADR-0007), and there is
// no way to guess it from a request the poller never sees.
PublicBaseURL string
Port string
// OwnerDiscordID identifies the seeded owner Reader (issue #22). Required:
// bookmarks are scoped to a Reader, and a fresh deployment needs one
// before anybody logs in. The owner is also the only Reader who can revoke
// another Reader's sessions.
OwnerDiscordID string
// Discord is the OAuth application the browser UI signs in with.
Discord web.DiscordConfig
// UserscriptPath is the file served at /u/{token}/manga-bookmark.user.js.
// Supplied by a bindmount so the script can be edited without a rebuild.
UserscriptPath string
@@ -41,28 +60,15 @@ type Config struct {
// LatestPoll configures the background latest-chapter poller.
//
// Sizing: batch x (cooldown / interval) is how many series hold a true cooldown
// cadence — 14 x (1h / 10m) = 84 with these defaults, which covers this
// deployment. Past that nothing breaks; the effective cadence stretches to
// N x interval / batch and the oldest-checked-first ordering keeps it uniform.
// Only the kill switch lives here. Pace is per Site — rest time and gap are
// registry properties (internal/latest/sites.go, issue #100), because each
// Lane has to be able to differ from the others. The five environment
// settings that used to size a shared pace (cooldown, browser cooldown,
// interval, stagger, batch) are gone with it: no deployed .env may carry them.
type LatestPoll struct {
Enabled bool
Cooldown time.Duration
Interval time.Duration
Stagger time.Duration
Batch int
Enabled bool
}
const (
defaultPollCooldown = time.Hour
defaultPollInterval = 10 * time.Minute
defaultPollStagger = 20 * time.Second
defaultPollBatch = 14
// minPollCooldown keeps a typo from turning a polite background check into
// a hammer against sites that are already bot-scoring us.
minPollCooldown = 15 * time.Minute
)
func envOr(key, def string) string {
if v := os.Getenv(key); v != "" {
return v
@@ -85,71 +91,33 @@ func envBool(key string, def bool) bool {
}
}
// envDuration reads a duration env var. An unparseable or non-positive value
// falls back to def and logs rather than failing startup: the poller is an
// enhancement, and a typo in one of its knobs must not stop bookmark sync.
func envDuration(key string, def time.Duration) time.Duration {
raw := strings.TrimSpace(os.Getenv(key))
if raw == "" {
return def
}
d, err := time.ParseDuration(raw)
if err != nil || d <= 0 {
log.Printf("config: %s=%q is not a positive duration, using %s", key, raw, def)
return def
}
return d
}
// envInt reads a positive integer env var, with the same fallback policy.
func envInt(key string, def int) int {
raw := strings.TrimSpace(os.Getenv(key))
if raw == "" {
return def
}
n, err := strconv.Atoi(raw)
if err != nil || n <= 0 {
log.Printf("config: %s=%q is not a positive integer, using %d", key, raw, def)
return def
}
return n
}
// loadLatestPoll reads the poller's settings, clamping anything that would make
// it antisocial.
// loadLatestPoll reads the poller's settings. The pace knobs that used to be
// clamped here are registry properties now (issue #100), so there is nothing
// left to clamp.
func loadLatestPoll() LatestPoll {
p := LatestPoll{
Enabled: envBool("LATEST_CHAPTER_POLL_ENABLED", true),
Cooldown: envDuration("LATEST_CHAPTER_POLL_COOLDOWN", defaultPollCooldown),
Interval: envDuration("LATEST_CHAPTER_POLL_INTERVAL", defaultPollInterval),
Stagger: envDuration("LATEST_CHAPTER_POLL_STAGGER", defaultPollStagger),
Batch: envInt("LATEST_CHAPTER_POLL_BATCH", defaultPollBatch),
}
if p.Cooldown < minPollCooldown {
log.Printf("config: cooldown %s is below the %s floor, clamping", p.Cooldown, minPollCooldown)
p.Cooldown = minPollCooldown
}
// batch x stagger has to fit inside one tick or a batch is still running
// when the next one is due. Run() serialises them, so this degrades to a
// slower cadence rather than to overlapping fetches — worth a warning, not
// a failure.
if span := time.Duration(p.Batch) * p.Stagger; span > p.Interval {
log.Printf("config: batch(%d) x stagger(%s) = %s exceeds interval %s; batches will overrun their tick",
p.Batch, p.Stagger, span, p.Interval)
}
return p
return LatestPoll{Enabled: envBool("LATEST_CHAPTER_POLL_ENABLED", true)}
}
func loadConfig() Config {
c := Config{
Token: os.Getenv("API_TOKEN"),
DBPath: envOr("DB_PATH", "/data/bookmarks.db"),
TokenKey: os.Getenv("TOKEN_KEY"),
DatabaseURL: os.Getenv("DATABASE_URL"),
CoverDir: os.Getenv("COVER_DIR"),
PublicBaseURL: os.Getenv("PUBLIC_BASE_URL"),
Port: envOr("PORT", "8080"),
WebPassword: os.Getenv("WEB_PASSWORD"),
OwnerDiscordID: os.Getenv("OWNER_DISCORD_ID"),
UserscriptPath: envOr("USERSCRIPT_PATH", "/userscript/manga-bookmark.user.js"),
NovelUserscriptPath: envOr("NOVEL_USERSCRIPT_PATH", "/userscript/novel-bookmark.user.js"),
LatestPoll: loadLatestPoll(),
}
c.Discord = web.DiscordConfig{
ClientID: os.Getenv("DISCORD_CLIENT_ID"),
ClientSecret: os.Getenv("DISCORD_CLIENT_SECRET"),
GuildID: os.Getenv("DISCORD_GUILD_ID"),
RequiredRole: os.Getenv("DISCORD_REQUIRED_ROLE"),
APIBase: envOr("DISCORD_API_BASE", "https://discord.com/api/v10"),
RedirectURI: os.Getenv("DISCORD_REDIRECT_URI"),
}
for _, o := range strings.Split(os.Getenv("ALLOWED_ORIGINS"), ",") {
if o = strings.TrimSpace(o); o != "" {
c.AllowedOrigins = append(c.AllowedOrigins, o)
@@ -161,36 +129,45 @@ func loadConfig() Config {
// newRouter wires routes and middleware. CORS is the outermost layer so
// preflight OPTIONS short-circuits before auth; /bookmarks* is auth-protected,
// /healthz is public.
func newRouter(s *store.Store, cfg Config) http.Handler {
//
// lanes may be nil — polling disabled, or its client could not be built. The
// admin page reports that rather than pretending Lanes exist.
func newRouter(s *store.Store, cfg Config, lanes web.LaneReporter) http.Handler {
mux := http.NewServeMux()
h := &api.Handler{Store: s}
mux.HandleFunc("GET /healthz", api.Healthz)
// Public: cover bytes are rendered by the userscript on origins that may
// not send our credentials, and the address is the hash of a URL the Site
// already publishes (ADR-0007).
mux.HandleFunc("GET /covers/{address}", h.Cover)
// Outside httpmw.Auth (the updater sends no Authorization header) and
// outside the WEB_PASSWORD gate (the script must be installable either
// way). The path segment carries the token instead.
mux.HandleFunc("GET /u/{token}/manga-bookmark.user.js", userscript.Handler(cfg.Token, cfg.UserscriptPath))
mux.HandleFunc("GET /u/{token}/novel-bookmark.user.js", userscript.Handler(cfg.Token, cfg.NovelUserscriptPath))
// outside the web UI's Discord auth (the script must be installable
// without a browser session). The path segment carries the credential
// instead, and the script is rendered with the resolved Reader's
// credential substituted in.
mux.HandleFunc("GET /u/{token}/manga-bookmark.user.js",
userscript.Handler(s, cfg.UserscriptPath))
mux.HandleFunc("GET /u/{token}/novel-bookmark.user.js",
userscript.Handler(s, cfg.NovelUserscriptPath))
h := &api.Handler{Store: s}
protected := http.NewServeMux()
protected.HandleFunc("GET /bookmarks", h.List)
protected.HandleFunc("PUT /bookmarks/{key}", h.Put)
protected.HandleFunc("DELETE /bookmarks/{key}", h.Delete)
auth := httpmw.Auth(cfg.Token, protected)
auth := httpmw.Auth(s, protected)
mux.Handle("/bookmarks", auth)
mux.Handle("/bookmarks/", auth)
// The browser UI is registered only when a password is configured, so a
// deployment that forgets WEB_PASSWORD exposes nothing rather than
// exposing an unprotected list.
if cfg.WebPassword != "" {
wh, err := web.New(s, cfg.Token, cfg.WebPassword)
if err != nil {
log.Fatalf("web handler: %v", err)
}
wh.Register(mux)
// The browser UI is always registered; signing in is Discord OAuth, so
// there is no password to forget and no gate to leave unset.
wh, err := web.New(s, cfg.Discord, []byte(cfg.TokenKey),
cfg.UserscriptPath, cfg.NovelUserscriptPath, lanes)
if err != nil {
log.Fatalf("web handler: %v", err)
}
wh.Register(mux)
return httpmw.CORS(cfg.AllowedOrigins, httpmw.Gzip(guardEmptyUserscriptToken(mux)))
}
@@ -212,11 +189,42 @@ func guardEmptyUserscriptToken(next http.Handler) http.Handler {
func main() {
cfg := loadConfig()
if cfg.Token == "" {
log.Fatal("API_TOKEN is required")
if cfg.TokenKey == "" {
log.Fatal("TOKEN_KEY is required")
}
if cfg.OwnerDiscordID == "" {
log.Fatal("OWNER_DISCORD_ID is required")
}
if cfg.DatabaseURL == "" {
log.Fatal("DATABASE_URL is required")
}
if cfg.CoverDir == "" {
log.Fatal("COVER_DIR is required")
}
if cfg.PublicBaseURL == "" {
log.Fatal("PUBLIC_BASE_URL is required")
}
// The web UI signs in through Discord, so a deployment without the OAuth
// application is misconfigured rather than passwordless.
for key, v := range map[string]string{
"DISCORD_CLIENT_ID": cfg.Discord.ClientID,
"DISCORD_CLIENT_SECRET": cfg.Discord.ClientSecret,
"DISCORD_GUILD_ID": cfg.Discord.GuildID,
"DISCORD_REDIRECT_URI": cfg.Discord.RedirectURI,
} {
if v == "" {
log.Fatalf("%s is required", key)
}
}
// The owner's userscript credential is derived from TOKEN_KEY at epoch 0
// (internal/token); the readers row carries its SHA-256, not the
// credential itself.
owner := store.Owner{
DiscordID: cfg.OwnerDiscordID,
TokenHash: token.Hash(token.Token([]byte(cfg.TokenKey), cfg.OwnerDiscordID, 0)),
}
s, err := store.Open(cfg.DBPath)
s, err := store.Open(cfg.DatabaseURL, owner, cfg.CoverDir, cfg.PublicBaseURL)
if err != nil {
log.Fatalf("open store: %v", err)
}
@@ -225,18 +233,68 @@ func main() {
// The poller is off the request path entirely: if it cannot start, the
// service still serves bookmarks and the userscript still captures latest
// chapters on its own.
//
// One headless browser serves both consumers that need a Cloudflare
// challenge cleared: the poller's kagane/novelfull page fetches and
// kagane's cover bytes. Optional — unset leaves kagane unpolled and its
// Covers blank until the bytes exist.
var browser latest.Fetcher
pollCtx, stopPoll := context.WithCancel(context.Background())
defer stopPoll()
startLatestPoller(pollCtx, s, cfg.LatestPoll)
if ws := strings.TrimSpace(os.Getenv("BROWSER_WS_URL")); ws != "" {
bf, err := latest.NewBrowserFetcher(ws)
if err != nil {
log.Printf("browser fetcher disabled: %v", err)
} else {
browser = bf
context.AfterFunc(pollCtx, bf.Close)
log.Printf("browser fetcher at %s", ws)
}
}
// A Series nobody had bookmarked before gets its Latest Chapter and its
// Cover from one fetch, at creation, instead of waiting out a poll queue
// ordered by Reader count. Off the write path: the hook returns as soon
// as the goroutine is started.
var tlsFetch latest.Fetcher
if f, err := latest.NewTLSFetcher(); err != nil {
log.Printf("creation-time acquisition: plain-TLS Sites disabled, cannot build client: %v", err)
} else {
tlsFetch = f
}
var browserCover latest.BrowserCoverFetcher
if b, ok := browser.(latest.BrowserCoverFetcher); ok {
browserCover = b
}
// The Acquirer must survive a TLS client failure: kagane needs only the
// sidecar, and novelfull degrades to whatever is left.
if tlsFetch != nil || browser != nil {
acq := &latest.Acquirer{
Store: s,
Fetch: tlsFetch,
BrowserFetch: browser,
BrowserCoverFetch: browserCover,
Covers: latest.NewCoverFetcher(),
Ctx: pollCtx,
}
s.OnSeriesCreated = acq.Acquire
}
// A nil *Poller must not become a non-nil interface holding a nil pointer:
// the admin page tests the reporter for nil to decide whether anything is
// polling at all.
var lanes web.LaneReporter
if poller := startLatestPoller(pollCtx, s, cfg.LatestPoll, browser); poller != nil {
lanes = poller
}
srv := &http.Server{
Addr: ":" + cfg.Port,
Handler: newRouter(s, cfg),
Handler: newRouter(s, cfg, lanes),
ReadHeaderTimeout: 10 * time.Second,
}
go func() {
log.Printf("listening on :%s (db=%s, origins=%v)", cfg.Port, cfg.DBPath, cfg.AllowedOrigins)
// The connection URL carries a password, so it stays out of the log.
log.Printf("listening on :%s (origins=%v)", cfg.Port, cfg.AllowedOrigins)
if err := srv.ListenAndServe(); err != nil && !errors.Is(err, http.ErrServerClosed) {
log.Fatalf("serve: %v", err)
}
@@ -257,43 +315,43 @@ func main() {
}
}
// newLatestPoller wires the fetcher seams into the poller. Pace is registry
// property, not config (issue #100), so there are no knobs to pass through.
func newLatestPoller(s *store.Store, cfg LatestPoll, fetch, browser latest.Fetcher) *latest.Poller {
var covers latest.BrowserCoverFetcher
if f, ok := browser.(latest.BrowserCoverFetcher); ok {
covers = f
}
return &latest.Poller{
Store: s,
Fetch: fetch,
BrowserFetch: browser,
CoverFetch: covers,
CoverBytesFetch: latest.NewCoverFetcher(),
Now: time.Now,
}
}
// startLatestPoller launches the background poller unless it is disabled or its
// HTTP client cannot be built. Any problem here is logged and skipped: this
// feature going missing degrades the service to userscript-only latest-chapter
// tracking, which is exactly how it behaved before.
func startLatestPoller(ctx context.Context, s *store.Store, cfg LatestPoll) {
// tracking, which is exactly how it behaved before. It returns the running
// Poller, or nil when there is none — the admin page's Lane status reads it.
func startLatestPoller(ctx context.Context, s *store.Store, cfg LatestPoll, browser latest.Fetcher) *latest.Poller {
if !cfg.Enabled {
log.Println("latest-chapter poller: disabled by config")
return
return nil
}
f, err := latest.NewTLSFetcher()
if err != nil {
log.Printf("latest-chapter poller: disabled, cannot build client: %v", err)
return
}
p := &latest.Poller{
Store: s,
Fetch: f,
Now: time.Now,
Cooldown: cfg.Cooldown,
Interval: cfg.Interval,
Stagger: cfg.Stagger,
Batch: cfg.Batch,
}
// Optional: without it, sites behind a JavaScript challenge are simply not
// polled, and their latest_chapter comes from the userscript alone — which
// is how the service behaved before the sidecar existed.
if ws := strings.TrimSpace(os.Getenv("BROWSER_WS_URL")); ws != "" {
bf, err := latest.NewBrowserFetcher(ws)
if err != nil {
log.Printf("latest-chapter poller: browser fetcher disabled: %v", err)
} else {
p.BrowserFetch = bf
context.AfterFunc(ctx, bf.Close)
log.Printf("latest-chapter poller: browser fetcher at %s", ws)
}
return nil
}
// Nil browser: sites behind a JavaScript challenge are simply not polled,
// and their latest_chapter comes from the userscript alone — which is how
// the service behaved before the sidecar existed.
p := newLatestPoller(s, cfg, f, browser)
go p.Run(ctx)
return p
}
+28 -82
View File
@@ -9,30 +9,22 @@ import (
"net/http/httptest"
"strings"
"testing"
"time"
"bookmarkmanager/backend/internal/latest"
"bookmarkmanager/backend/internal/store"
)
func TestLoadLatestPollDefaults(t *testing.T) {
for _, k := range []string{
"LATEST_CHAPTER_POLL_ENABLED", "LATEST_CHAPTER_POLL_COOLDOWN",
"LATEST_CHAPTER_POLL_INTERVAL", "LATEST_CHAPTER_POLL_STAGGER",
"LATEST_CHAPTER_POLL_BATCH",
} {
t.Setenv(k, "")
t.Setenv("LATEST_CHAPTER_POLL_ENABLED", "")
if got := loadLatestPoll(); got != (LatestPoll{Enabled: true}) {
t.Fatalf("loadLatestPoll() = %+v, want %+v", got, LatestPoll{Enabled: true})
}
}
got := loadLatestPoll()
want := LatestPoll{
Enabled: true,
Cooldown: time.Hour,
Interval: 10 * time.Minute,
Stagger: 20 * time.Second,
Batch: 14,
}
if got != want {
t.Fatalf("loadLatestPoll() = %+v, want %+v", got, want)
func TestLoadConfigReadsCoverDirectory(t *testing.T) {
t.Setenv("COVER_DIR", "/covers")
if got := loadConfig().CoverDir; got != "/covers" {
t.Fatalf("CoverDir = %q, want /covers", got)
}
}
@@ -55,71 +47,25 @@ func TestLoadLatestPollEnabledParsing(t *testing.T) {
}
}
func TestLoadLatestPollClampsAndFallsBack(t *testing.T) {
tests := []struct {
name string
env map[string]string
wantFrom func(LatestPoll) any
want any
}{
{
name: "cooldown below the floor is clamped up",
env: map[string]string{"LATEST_CHAPTER_POLL_COOLDOWN": "1m"},
wantFrom: func(p LatestPoll) any { return p.Cooldown },
want: 15 * time.Minute,
},
{
name: "cooldown at the floor is kept",
env: map[string]string{"LATEST_CHAPTER_POLL_COOLDOWN": "15m"},
wantFrom: func(p LatestPoll) any { return p.Cooldown },
want: 15 * time.Minute,
},
{
name: "a valid override is honoured",
env: map[string]string{"LATEST_CHAPTER_POLL_INTERVAL": "5m"},
wantFrom: func(p LatestPoll) any { return p.Interval },
want: 5 * time.Minute,
},
{
name: "an unparseable duration falls back",
env: map[string]string{"LATEST_CHAPTER_POLL_INTERVAL": "ten minutes"},
wantFrom: func(p LatestPoll) any { return p.Interval },
want: 10 * time.Minute,
},
{
name: "a zero duration falls back",
env: map[string]string{"LATEST_CHAPTER_POLL_STAGGER": "0s"},
wantFrom: func(p LatestPoll) any { return p.Stagger },
want: 20 * time.Second,
},
{
name: "a valid batch is honoured",
env: map[string]string{"LATEST_CHAPTER_POLL_BATCH": "30"},
wantFrom: func(p LatestPoll) any { return p.Batch },
want: 30,
},
{
name: "a negative batch falls back",
env: map[string]string{"LATEST_CHAPTER_POLL_BATCH": "-5"},
wantFrom: func(p LatestPoll) any { return p.Batch },
want: 14,
},
{
name: "a non-numeric batch falls back",
env: map[string]string{"LATEST_CHAPTER_POLL_BATCH": "lots"},
wantFrom: func(p LatestPoll) any { return p.Batch },
want: 14,
},
// newLatestPoller wires the fetcher seams; pace lives in the registry, so
// nothing here sizes a cooldown any more.
func TestNewLatestPollerWiresFetchers(t *testing.T) {
tls := &latest.TLSFetcher{}
p := newLatestPoller(nil, LatestPoll{Enabled: true}, tls, nil)
if p.Fetch != tls {
t.Fatalf("Fetch not wired")
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
for k, v := range tt.env {
t.Setenv(k, v)
}
if got := tt.wantFrom(loadLatestPoll()); got != tt.want {
t.Fatalf("got %v, want %v", got, tt.want)
}
})
if p.BrowserFetch != nil {
t.Fatalf("BrowserFetch = %v, want nil for a browser-less deployment", p.BrowserFetch)
}
if p.CoverFetch != nil {
t.Fatalf("CoverFetch = %v, want nil when the browser is absent", p.CoverFetch)
}
if p.CoverBytesFetch == nil {
t.Fatalf("CoverBytesFetch = nil, want the TLS cover fetcher")
}
if p.Now == nil {
t.Fatalf("Now = nil, want the live clock")
}
}
@@ -202,7 +148,7 @@ func TestPutOmittedStatusPreservesArchivedAndAppliesProgress(t *testing.T) {
}
func TestGzipCompressesTextNotFonts(t *testing.T) {
srv, _ := newWebTestServer(t, webConfig())
srv, _ := newWebTestServer(t, testConfig())
cases := []struct {
path string
+356
View File
@@ -0,0 +1,356 @@
package main
import (
"encoding/json"
"io"
"net/http"
"net/http/httptest"
"os"
"path/filepath"
"strings"
"testing"
"bookmarkmanager/backend/internal/store"
"bookmarkmanager/backend/internal/token"
)
// registerReader creates an extra Reader the way a first login does and
// returns its id. The credential is derived the same way the owner's is, so it
// authenticates through the real router.
func registerReader(t *testing.T, s *store.Store, discordID string) int64 {
t.Helper()
id, err := s.EnsureReader(discordID, token.Hash(readerCredential(discordID)))
if err != nil {
t.Fatalf("register reader %q: %v", discordID, err)
}
return id
}
// credRequest builds a request authenticated as the Reader whose credential
// is passed.
func credRequest(method, target, cred string) *http.Request {
req := httptest.NewRequest(method, target, nil)
req.Header.Set("Authorization", "Bearer "+cred)
return req
}
// readerCredential is the epoch-0 derived credential of an arbitrary Reader.
func readerCredential(discordID string) string {
return token.Token([]byte(testTokenKey), discordID, 0)
}
// withBody attaches a request body, for PUTs that carry a JSON payload.
func withBody(req *http.Request, body string) *http.Request {
req.Body = io.NopCloser(strings.NewReader(body))
req.ContentLength = int64(len(body))
return req
}
// A refused credential is refused however plausible it looks: only a hash the
// readers table holds authenticates anything.
func TestUnknownCredentialRejected(t *testing.T) {
srv := newRouter(newTestStore(t), testConfig(), nil)
rr := httptest.NewRecorder()
srv.ServeHTTP(rr, credRequest(http.MethodGet, "/bookmarks", readerCredential("never-registered")))
if rr.Code != http.StatusUnauthorized {
t.Fatalf("unregistered Reader's credential: status = %d, want 401", rr.Code)
}
rr = httptest.NewRecorder()
srv.ServeHTTP(rr, credRequest(http.MethodGet, "/bookmarks", ownerCredential()))
if rr.Code != http.StatusOK {
t.Fatalf("owner's derived credential: status = %d, want 200", rr.Code)
}
}
// A Reader's credential authenticates exactly that Reader: rows written under
// one credential are invisible to the other, on the same key.
func TestPerReaderIsolation(t *testing.T) {
s := newTestStore(t)
registerReader(t, s, "other-reader")
srv := newRouter(s, testConfig(), nil)
ownerKey := "asura:solo"
putBookmark(t, srv, ownerKey, store.Bookmark{
Key: ownerKey, Site: "asura", SeriesID: "solo",
Title: "Solo Leveling", UpdatedAt: 1,
})
// The other Reader's list is empty even though the owner holds the key.
rr := httptest.NewRecorder()
srv.ServeHTTP(rr, credRequest(http.MethodGet, "/bookmarks", readerCredential("other-reader")))
if rr.Code != http.StatusOK {
t.Fatalf("other reader list: status = %d, want 200", rr.Code)
}
var theirs []store.Bookmark
if err := json.Unmarshal(rr.Body.Bytes(), &theirs); err != nil {
t.Fatalf("decode: %v", err)
}
if len(theirs) != 0 {
t.Fatalf("other reader sees %d bookmarks, want 0 (owner's rows leaked)", len(theirs))
}
// The other Reader writes the same key; both rows coexist, each visible
// only to its owner. The series title is shared (ADR-0003) — the
// reader-owned fields are progress and updated_at.
req := credRequest(http.MethodPut, "/bookmarks/"+ownerKey, readerCredential("other-reader"))
req.Header.Set("Content-Type", "application/json")
body := `{"key":"asura:solo","site":"asura","series_id":"solo","title":"Theirs","last_chapter_num":3}`
rr = httptest.NewRecorder()
srv.ServeHTTP(rr, withBody(req, body))
if rr.Code != http.StatusOK {
t.Fatalf("other reader put: status = %d, want 200", rr.Code)
}
rr = httptest.NewRecorder()
srv.ServeHTTP(rr, credRequest(http.MethodGet, "/bookmarks", readerCredential("other-reader")))
var theirs2 []store.Bookmark
if err := json.Unmarshal(rr.Body.Bytes(), &theirs2); err != nil {
t.Fatalf("decode: %v", err)
}
if len(theirs2) != 1 || theirs2[0].LastChapterNum != 3 {
t.Fatalf("other reader list = %+v, want their own row with their progress", theirs2)
}
rr = httptest.NewRecorder()
srv.ServeHTTP(rr, credRequest(http.MethodGet, "/bookmarks", ownerCredential()))
var owners []store.Bookmark
if err := json.Unmarshal(rr.Body.Bytes(), &owners); err != nil {
t.Fatalf("decode: %v", err)
}
if len(owners) != 1 || owners[0].Title != "Solo Leveling" || owners[0].LastChapterNum != 0 {
t.Fatalf("owner list = %+v, want their own row at their own progress", owners)
}
// The mirror: the owner's write does not move the other Reader's progress
// either. Without it, isolation is only asserted in one direction.
req = credRequest(http.MethodPut, "/bookmarks/"+ownerKey, ownerCredential())
req.Header.Set("Content-Type", "application/json")
rr = httptest.NewRecorder()
srv.ServeHTTP(rr, withBody(req, `{"key":"asura:solo","site":"asura","series_id":"solo","title":"Solo Leveling","last_chapter_num":9}`))
if rr.Code != http.StatusOK {
t.Fatalf("owner put: status = %d, want 200", rr.Code)
}
rr = httptest.NewRecorder()
srv.ServeHTTP(rr, credRequest(http.MethodGet, "/bookmarks", readerCredential("other-reader")))
var theirs3 []store.Bookmark
if err := json.Unmarshal(rr.Body.Bytes(), &theirs3); err != nil {
t.Fatalf("decode: %v", err)
}
if len(theirs3) != 1 || theirs3[0].LastChapterNum != 3 {
t.Fatalf("other reader list = %+v, want progress 3 after the owner's write", theirs3)
}
// DELETE is scoped to its caller too, asserted in both directions: each
// Reader's delete on the shared key takes only their own row.
list := func(cred string) []store.Bookmark {
t.Helper()
rr := httptest.NewRecorder()
srv.ServeHTTP(rr, credRequest(http.MethodGet, "/bookmarks", cred))
var got []store.Bookmark
if err := json.Unmarshal(rr.Body.Bytes(), &got); err != nil {
t.Fatalf("decode: %v", err)
}
return got
}
del := func(cred string) {
t.Helper()
rr := httptest.NewRecorder()
srv.ServeHTTP(rr, credRequest(http.MethodDelete, "/bookmarks/"+ownerKey, cred))
if rr.Code != http.StatusNoContent {
t.Fatalf("delete: status = %d, want 204", rr.Code)
}
}
del(readerCredential("other-reader"))
if got := list(ownerCredential()); len(got) != 1 {
t.Fatalf("owner's row was deletable by the other Reader: %+v", got)
}
if got := list(readerCredential("other-reader")); len(got) != 0 {
t.Fatalf("other Reader's own delete left %+v behind", got)
}
// The mirror: the other Reader takes the key again, the owner deletes
// theirs, and the other's row is untouched.
req = credRequest(http.MethodPut, "/bookmarks/"+ownerKey, readerCredential("other-reader"))
req.Header.Set("Content-Type", "application/json")
rr = httptest.NewRecorder()
srv.ServeHTTP(rr, withBody(req, body))
if rr.Code != http.StatusOK {
t.Fatalf("other reader re-put: status = %d, want 200", rr.Code)
}
del(ownerCredential())
if got := list(readerCredential("other-reader")); len(got) != 1 {
t.Fatalf("other Reader's row was deletable by the owner: %+v", got)
}
if got := list(ownerCredential()); len(got) != 0 {
t.Fatalf("owner's own delete left %+v behind", got)
}
}
// The install endpoints are session-gated and render the script directly
// with the Reader's credential inside: the credential never appears in the
// address bar, the page markup, or any Location header.
func TestInstallServesScriptWithCredential(t *testing.T) {
cfg := testConfig()
dir := t.TempDir()
path := filepath.Join(dir, "manga-bookmark.user.js")
novelPath := filepath.Join(dir, "novel-bookmark.user.js")
for _, p := range []string{path, novelPath} {
if err := os.WriteFile(p, []byte("const API_TOKEN = \"__API_TOKEN__\";\n"), 0o644); err != nil {
t.Fatalf("write script: %v", err)
}
}
cfg.UserscriptPath = path
cfg.NovelUserscriptPath = novelPath
srv, st := newWebTestServer(t, cfg)
for _, script := range []string{"manga-bookmark.user.js", "novel-bookmark.user.js"} {
rr := httptest.NewRecorder()
srv.ServeHTTP(rr, httptest.NewRequest(http.MethodGet, "/install/"+script, nil))
if rr.Code != http.StatusUnauthorized {
t.Fatalf("%s without session: status = %d, want 401", script, rr.Code)
}
req := httptest.NewRequest(http.MethodGet, "/install/"+script, nil)
req.AddCookie(sessionCookie(t, st))
rr = httptest.NewRecorder()
srv.ServeHTTP(rr, req)
if rr.Code != http.StatusOK {
t.Fatalf("%s with session: status = %d, want 200", script, rr.Code)
}
body := rr.Body.String()
if strings.Contains(body, "__API_TOKEN__") {
t.Fatalf("%s served with an unsubstituted placeholder", script)
}
// The credential rides inside the served script — nowhere visible in
// the UI — and is the session holder's own.
if !strings.Contains(body, `API_TOKEN = "`+ownerCredential()+`"`) {
t.Fatalf("%s does not carry the owner's credential:\n%s", script, body)
}
if loc := rr.Header().Get("Location"); loc != "" {
t.Fatalf("%s answered with a redirect, credential in Location %q", script, loc)
}
// The plain link must stay inline: Violentmonkey's updater polls the
// /u/ path and an attachment disposition there would break updates.
if cd := rr.Header().Get("Content-Disposition"); cd != "" {
t.Fatalf("%s served as %q, want inline", script, cd)
}
// ?download=1 is the mobile path: Violentmonkey on Chromium ignores a
// .user.js navigation, so the Reader saves the file and adds it by hand.
req = httptest.NewRequest(http.MethodGet, "/install/"+script+"?download=1", nil)
req.AddCookie(sessionCookie(t, st))
rr = httptest.NewRecorder()
srv.ServeHTTP(rr, req)
if rr.Code != http.StatusOK {
t.Fatalf("%s?download=1: status = %d, want 200", script, rr.Code)
}
if got, want := rr.Header().Get("Content-Disposition"), `attachment; filename="`+script+`"`; got != want {
t.Fatalf("%s?download=1: Content-Disposition = %q, want %q", script, got, want)
}
if !strings.Contains(rr.Body.String(), `API_TOKEN = "`+ownerCredential()+`"`) {
t.Fatalf("%s?download=1 does not carry the owner's credential", script)
}
}
}
// Rotation through the web UI invalidates the old credential immediately,
// mints one that authenticates the API and the script path, and warns that
// every device must reinstall.
func TestRotateCredentialViaWebUI(t *testing.T) {
s, _ := newTestStoreURL(t)
path := filepath.Join(t.TempDir(), "manga-bookmark.user.js")
if err := os.WriteFile(path, []byte("const API_TOKEN = \"__API_TOKEN__\";\n"), 0o644); err != nil {
t.Fatalf("write script: %v", err)
}
cfg := testConfig()
cfg.UserscriptPath = path
srv := newRouter(s, cfg, nil)
oldCred := ownerCredential()
rr := httptest.NewRecorder()
srv.ServeHTTP(rr, credRequest(http.MethodGet, "/bookmarks", oldCred))
if rr.Code != http.StatusOK {
t.Fatalf("old credential before rotation: status = %d, want 200", rr.Code)
}
req := httptest.NewRequest(http.MethodPost, "/rotate-token", nil)
req.AddCookie(sessionCookie(t, s))
rr = httptest.NewRecorder()
srv.ServeHTTP(rr, req)
if rr.Code != http.StatusOK {
t.Fatalf("rotate: status = %d, want 200", rr.Code)
}
if !strings.Contains(rr.Body.String(), "Credential rotated") {
t.Fatalf("rotation response does not warn about reinstall:\n%s", rr.Body.String())
}
// The old credential is dead on the API and on the script path.
rr = httptest.NewRecorder()
srv.ServeHTTP(rr, credRequest(http.MethodGet, "/bookmarks", oldCred))
if rr.Code != http.StatusUnauthorized {
t.Fatalf("old credential after rotation: status = %d, want 401", rr.Code)
}
rr = httptest.NewRecorder()
srv.ServeHTTP(rr, httptest.NewRequest(http.MethodGet, "/u/"+oldCred+"/manga-bookmark.user.js", nil))
if rr.Code != http.StatusNotFound {
t.Fatalf("old credential script path after rotation: status = %d, want 404", rr.Code)
}
// The new credential authenticates the API and the script path, and is
// substituted into the served script.
newCred := token.Token([]byte(testTokenKey), testDiscordID, 1)
rr = httptest.NewRecorder()
srv.ServeHTTP(rr, credRequest(http.MethodGet, "/bookmarks", newCred))
if rr.Code != http.StatusOK {
t.Fatalf("new credential after rotation: status = %d, want 200", rr.Code)
}
rr = httptest.NewRecorder()
srv.ServeHTTP(rr, httptest.NewRequest(http.MethodGet, "/u/"+newCred+"/manga-bookmark.user.js", nil))
if rr.Code != http.StatusOK {
t.Fatalf("new credential script path: status = %d, want 200", rr.Code)
}
if got := rr.Body.String(); !strings.Contains(got, `API_TOKEN = "`+newCred+`"`) {
t.Fatalf("served script does not carry the rotated credential:\n%s", got)
}
// The install link now renders the script with the new credential.
req = httptest.NewRequest(http.MethodGet, "/install/manga-bookmark.user.js", nil)
req.AddCookie(sessionCookie(t, s))
rr = httptest.NewRecorder()
srv.ServeHTTP(rr, req)
if rr.Code != http.StatusOK {
t.Fatalf("install after rotation: status = %d, want 200", rr.Code)
}
if got := rr.Body.String(); !strings.Contains(got, `API_TOKEN = "`+newCred+`"`) {
t.Fatalf("install after rotation does not carry the new credential:\n%s", got)
}
}
// The app page offers the install links; the credential never appears in its
// markup.
func TestIndexShowsSetupPanelWithoutCredential(t *testing.T) {
srv, st := newWebTestServer(t, testConfig())
req := httptest.NewRequest(http.MethodGet, "/", nil)
req.AddCookie(sessionCookie(t, st))
rr := httptest.NewRecorder()
srv.ServeHTTP(rr, req)
body := rr.Body.String()
for _, want := range []string{
`href="/install/manga-bookmark.user.js"`,
`href="/install/novel-bookmark.user.js"`,
`href="/install/manga-bookmark.user.js?download=1"`,
`href="/install/novel-bookmark.user.js?download=1"`,
"Rotate credential",
} {
if !strings.Contains(body, want) {
t.Errorf("app page lacks %q", want)
}
}
if strings.Contains(body, ownerCredential()) {
t.Fatal("app page leaks the credential")
}
}
+957 -144
View File
File diff suppressed because it is too large Load Diff
+28
View File
@@ -0,0 +1,28 @@
# Copy to chrome/.env on the home machine. Never commit the real .env.
#
# This file configures the browser unit only. It is separate from the API
# stack's ../.env on purpose: the two run on different machines.
# The address the CDP port is published on — required, no default.
#
# Use this machine's **tailnet IP**, e.g. 100.x.y.z (`tailscale ip -4`). Not
# 0.0.0.0, not the LAN address: CDP has no authentication of its own, so
# anything that can reach this port has full control of the browser and a
# foothold on this host. Tailscale device identity plus an ACL is the access
# control; the bind address is what enforces it.
#
# For a throwaway local test, 127.0.0.1 is fine — but then only this machine
# can reach it, so the API must run here too.
# Left commented so `cp .env.example .env && docker compose up` fails with the
# variable's own message telling you what to set, rather than Docker rejecting
# "100.x.y.z" as an invalid IP.
# BROWSER_BIND_ADDR=100.x.y.z
# Clock zone the browser reports. Any real zone works and it need not match
# the egress IP's country — but it must not be UTC, which is itself the bot
# signal that stops the challenge clearing. The measurement is in entrypoint.sh.
#
# Unset falls back to the host's /etc/timezone, which is a real zone whenever
# the host clock is set to local time. Set this when the host runs UTC — a UTC
# server is exactly the case that fails.
# BROWSER_TZ=Asia/Jakarta
+45
View File
@@ -0,0 +1,45 @@
# syntax=docker/dockerfile:1
# Real Google Chrome for the latest-chapter poller and the kagane cover proxy.
#
# Not chromedp/headless-shell, which this replaces. headless-shell is a stripped
# Chrome build and Cloudflare's managed challenge on kagane.to never clears for
# it: measured 2026-08-08, 60s of a held-open tab still served the interstitial,
# while stock Chrome from the same IP cleared in ~4s. The tells are structural
# rather than a header — navigator.webdriver true, an empty plugin list, and
# Chromium- rather than Chrome-branded client hints. Overriding webdriver alone
# was tried and did not move it, so the browser build itself is the fix.
#
# zenika/alpine-chrome was also tried: its Chrome is 124 (2024), old enough that
# Cloudflare refuses it outright and old enough to break chromedp's CDP structs.
FROM debian:trixie-slim
# Chrome is deliberately unpinned, against the usual rule. A pinned build goes
# stale, and a stale browser is exactly what Cloudflare turns away — the 124 in
# alpine-chrome is the worked example. Rebuild is the upgrade path.
RUN apt-get update \
&& apt-get install -y --no-install-recommends ca-certificates wget gnupg util-linux \
&& wget -qO- https://dl.google.com/linux/linux_signing_key.pub \
| gpg --dearmor -o /usr/share/keyrings/google-chrome.gpg \
&& echo "deb [arch=amd64 signed-by=/usr/share/keyrings/google-chrome.gpg] https://dl.google.com/linux/chrome/deb/ stable main" \
> /etc/apt/sources.list.d/google-chrome.list \
&& apt-get update \
&& apt-get install -y --no-install-recommends google-chrome-stable socat \
&& rm -rf /var/lib/apt/lists/*
# Keep the profile path present so Docker initializes the named volume with
# the unprivileged user's ownership.
RUN useradd --create-home --shell /usr/sbin/nologin chrome \
&& mkdir -p /home/chrome/profile /home/chrome/state \
&& chown -R chrome:chrome /home/chrome
# Unprivileged: Chrome refuses to run as root, and the CDP endpoint is a shell
# on whatever user owns it.
USER chrome
WORKDIR /home/chrome
COPY entrypoint.sh /entrypoint.sh
EXPOSE 9222
ENTRYPOINT ["/entrypoint.sh"]
+62
View File
@@ -0,0 +1,62 @@
# The browser, as its own deployable unit.
#
# This does NOT run beside the API. It runs on the home machine, reached from
# the VPS over the tailnet, and is updated without touching the API stack:
#
# cd chrome && docker compose up -d --build
#
# Set BROWSER_BIND_ADDR in chrome/.env to this machine's tailnet IP. See
# ../DEPLOY.md §7 for the full first-time procedure and ../docs/adr/
# 0006-browser-on-the-home-machine.md for why the browser lives here at all.
name: bookmark-browser
services:
browser:
build: .
image: bookmarkmanager-chrome:latest
container_name: bookmark-browser
restart: unless-stopped
environment:
# Any real zone works, but a UTC clock is itself the bot signal and the
# challenge then never clears — measurement in entrypoint.sh. Unset falls
# back to the host's /etc/timezone below, which is a real zone whenever
# the host clock is local; set BROWSER_TZ when the host runs UTC.
TZ: ${BROWSER_TZ:-}
volumes:
# The zone *name*, which is what Chrome's ICU needs — see entrypoint.sh.
# Absent on a non-Debian host, which the entrypoint handles by falling back to UTC.
- /etc/timezone:/etc/timezone:ro
# Cloudflare clearance must survive Chrome reaping and image recreation.
- chrome-profile:/home/chrome/profile
# Bound to the tailnet address only, never 0.0.0.0. CDP authenticates
# nothing: whatever reaches this port drives the browser and, through it,
# this host. On the VPS the safety was Docker network membership; here the
# machine has a real LAN, so the bind address *is* the access control,
# backed by Tailscale device identity. No default — an unset variable must
# fail the deploy rather than silently publish CDP to the LAN.
ports:
- "${BROWSER_BIND_ADDR:?set BROWSER_BIND_ADDR to this machine's tailnet IP}:9222:9222"
# Reaps zombie renderer processes, which otherwise accumulate for the
# container's lifetime.
init: true
# Chrome allocates shared memory per tab and dies on Docker's 64MB default.
# 128MB against a measured 19MB peak: the old 1GB reservation was sized by
# superstition, and this box has 1.8GB total.
shm_size: '128mb'
# The browser is the newcomer on a machine where a Gitea runner already
# holds ~1.2GiB of 1.8GiB. Load-bearing, not decorative: untuned Chrome
# peaked at 645MiB cgroup, which is more than is free here.
#
# memswap_limit is memory+swap combined, so this allows 512MiB of swap —
# Chrome reclaims its own cold pages onto this box's 5.9GiB of SATA swap
# instead of taking resident memory from the runner.
mem_limit: 512m
memswap_limit: 1g
# If the box does run out, the kernel takes the browser and never CI.
oom_score_adj: 800
# A challenge solve yields to a running build. Cold start degrades to ~3s
# at half a CPU, immaterial against a 45-second challenge budget.
cpu_shares: 512
volumes:
chrome-profile:
+225
View File
@@ -0,0 +1,225 @@
#!/bin/sh
set -eu
# A UTC clock is itself the bot signal — Cloudflare treats it as the datacenter
# default — and kagane's challenge then never clears. Measured 2026-08-08 with
# an identical container on one Indonesian egress IP: UTC never cleared in 60s
# (twice), while Asia/Jakarta and America/New_York both cleared in 4s. Any real
# zone will do; the zone does not have to match the IP's country, it just must
# not be UTC. It does have to be right the way Chrome reads it.
#
# TZ must carry the zone *name*. Chrome resolves the zone through ICU, which
# takes the name from /etc/localtime's symlink target and ignores the file's
# contents; bind-mounting the host's /etc/localtime therefore lands on the
# image's own symlink to Etc/UTC and leaves glibc reporting +07 while Chrome
# still reports UTC. /etc/timezone, mounted by docker-compose.yml, is the name.
[ -n "${TZ:-}" ] || TZ=$(cat /etc/timezone 2>/dev/null || echo UTC)
export TZ
state=/home/chrome/state
profile=/home/chrome/profile
lock_file=$state/lock
pid_file=$state/chrome.pid
connections_dir=$state/connections
last_use_file=$state/last-use
idle_seconds=300
mkdir -p "$state" "$profile" "$connections_dir"
exec 9>>"$lock_file"
# Chrome's own UA advertises "HeadlessChrome" under --headless=new, and that
# one token is the difference between kagane.to's challenge clearing in ~4s and
# never clearing at all. Read the installed major version so client hints and
# the UA stay aligned after an image rebuild.
major=$(google-chrome-stable --version | sed -E 's/[^0-9]*([0-9]+)\..*/\1/')
ua="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/${major}.0.0.0 Safari/537.36"
lock() {
flock 9
}
unlock() {
flock -u 9
}
browser_alive() {
[ -s "$pid_file" ] || return 1
pid=$(cat "$pid_file")
[ -n "$pid" ] && kill -0 "$pid" 2>/dev/null
}
has_connections() {
for marker in "$connections_dir"/*; do
[ -e "$marker" ] || continue
pid=${marker##*/}
if kill -0 "$pid" 2>/dev/null; then
return 0
fi
# A SIGKILLed helper cannot run its cleanup trap. Reconcile its marker
# here so one dead client cannot pin Chrome forever.
rm -f "$marker"
done
return 1
}
start_browser() {
# No --enable-automation: it sets navigator.webdriver, the first thing a
# bot check reads. setsid gives Chrome a process group so the reaper can
# terminate its renderer children with the browser.
# --no-sandbox avoids granting SYS_ADMIN solely for Docker's unavailable
# user namespaces; containment is the unprivileged user and private network.
setsid google-chrome-stable \
--headless=new \
--no-sandbox \
--remote-debugging-port=9223 \
--user-agent="$ua" \
--user-data-dir="$profile" \
--no-first-run \
--no-default-browser-check \
--disable-gpu \
about:blank >/dev/null &
printf '%s\n' "$!" >"$pid_file"
}
stop_browser() {
pid=$(cat "$pid_file")
kill -TERM -- "-$pid" 2>/dev/null || kill -TERM "$pid" 2>/dev/null || true
i=0
while kill -0 "$pid" 2>/dev/null && [ "$i" -lt 100 ]; do
i=$((i + 1))
sleep 0.1
done
if kill -0 "$pid" 2>/dev/null; then
kill -KILL -- "-$pid" 2>/dev/null || kill -KILL "$pid" 2>/dev/null || true
fi
rm -f "$pid_file"
}
wait_for_browser() {
i=0
while [ "$i" -lt 300 ]; do
if wget -qO /dev/null http://127.0.0.1:9223/json/version; then
return 0
fi
browser_alive || return 1
i=$((i + 1))
sleep 0.1
done
return 1
}
finish_connection() {
lock
rm -f "$connection_marker"
date +%s >"$last_use_file"
unlock
}
connection_signal() {
trap - INT TERM HUP
finish_connection
exit 143
}
connection() {
connection_marker=$connections_dir/$$
lock
: >"$connection_marker"
if ! browser_alive; then
rm -f "$pid_file"
start_browser
fi
date +%s >"$last_use_file"
unlock
trap connection_signal INT TERM HUP
if wait_for_browser; then
if socat STDIO TCP:127.0.0.1:9223; then
result=0
else
result=$?
fi
else
# The client only ever sees a bare connection reset here, so this is
# the sole record that the browser, not the network, was the problem.
echo "browser did not come up; dropping connection" >&2
result=1
fi
finish_connection
return "$result"
}
reaper() {
while :; do
sleep 10
lock
if ! has_connections && browser_alive; then
now=$(date +%s)
last=$(cat "$last_use_file" 2>/dev/null || printf '%s' "$now")
if [ $((now - last)) -ge "$idle_seconds" ]; then
stop_browser
fi
fi
unlock
done
}
if [ "${1:-}" = connection ]; then
connection
exit $?
fi
# The files are process state, not the Chrome profile. The profile is a named
# volume in Compose, so clearance survives both a reap and a container rebuild.
for marker in "$connections_dir"/*; do
[ -e "$marker" ] || continue
rm -f "$marker"
done
rm -f "$pid_file" "$last_use_file"
# Chrome's singleton lock names the hostname and pid that took it, and a
# container rebuild changes both — so a Chrome killed uncleanly (OOM, docker
# kill) leaves a lock the next container reads as "another computer holds this
# profile" and refuses to start behind, permanently, with the only symptom a
# bare connection reset at 9222. Clearing it here is safe precisely because
# container_name pins this volume to one container: nothing can be holding the
# profile at the moment this line runs. The lock is process state; the
# clearance cookies it sits beside are not, and are left alone.
rm -f "$profile"/Singleton*
# Chrome binds DevTools to loopback and silently ignores
# --remote-debugging-address. socat remains the network front-end, but each
# accepted connection now starts a browser on demand and is tracked by a
# per-helper marker. A connection held by Go's transport delays reap by its
# idle timeout; the 300-second threshold starts once the last connection closes.
reaper &
reaper_pid=$!
socat TCP-LISTEN:9222,reuseaddr,fork EXEC:'/entrypoint.sh connection',nofork &
front_pid=$!
stop_browser_gracefully() {
lock
if browser_alive; then
# Chrome is a separate session, so stop its process group explicitly;
# this gives its cookie batch time to flush before the container exits.
stop_browser
fi
unlock
}
shutdown() {
trap - INT TERM HUP
stop_browser_gracefully
kill "$front_pid" "$reaper_pid" 2>/dev/null || true
exit 143
}
trap shutdown INT TERM HUP
if wait "$front_pid"; then
status=0
else
status=$?
fi
stop_browser_gracefully
kill "$reaper_pid" 2>/dev/null || true
exit "$status"
+10 -15
View File
@@ -17,19 +17,12 @@ services:
bookmark-api:
# Traffic arrives over the Traefik network, not a published port.
ports: !reset []
environment:
# Must be an IP, not the DNS name — see the base file's comment on this
# same key: Chrome's DevTools HTTP handler 500s any Host header that
# isn't an IP or "localhost".
BROWSER_WS_URL: ${BROWSER_WS_URL:-ws://172.28.0.10:9222}
depends_on:
- headless-shell
# `networks:` here replaces the base file's list entirely, so both must be
# named: `proxy` for Traefik routing, `browser` (defined in the base file)
# to keep reaching headless-shell without putting it on `proxy` too.
# Compose *merges* this list with the base file's, so the service ends up on
# `default`, `db` and `proxy` — only the addition is named here. Do not
# "tidy" the base file down to `db` on the strength of `proxy` being present:
# `db` is `internal: true`, and egress comes from `default`.
networks:
- proxy
- browser
labels:
- "traefik.enable=true"
- "traefik.docker.network=${PROXY_NETWORK:-proxy}"
@@ -47,10 +40,12 @@ services:
- "traefik.http.routers.bmweb.tls.certresolver=${TRAEFIK_CERTRESOLVER:-le}"
- "traefik.http.routers.bmweb.service=bmapi"
# headless-shell is untouched here: it keeps its `browser` network membership
# from the base file and must never join `proxy` — that network is shared
# with whatever else sits behind Traefik on this host, and an exposed
# CDP endpoint on it would be remote code execution for any of them.
# No browser service here. It runs on the home machine as its own unit
# (chrome/docker-compose.yml) and is reached over the tailnet — see
# docs/adr/0006-browser-on-the-home-machine.md. It must never be given a
# service on this host: `proxy` is shared with whatever else sits behind
# Traefik, and an unauthenticated CDP endpoint on it is remote code
# execution for any of them.
networks:
proxy:
+94 -57
View File
@@ -5,93 +5,130 @@
# If your proxy runs in Docker on its own network, use the prod override which
# attaches to that network instead of publishing a port:
# docker compose -f docker-compose.yml -f docker-compose.prod.yml up -d
#
# The browser is not here. It is its own unit on the home machine —
# chrome/docker-compose.yml — reached over the tailnet via BROWSER_WS_URL.
services:
bookmark-api:
build: ./backend
build:
context: ./backend
args:
COVER_DIR: ${COVER_DIR:?set COVER_DIR in .env}
image: bookmarkmanager-backend:latest
container_name: bookmark-api
restart: unless-stopped
environment:
# API_TOKEN is required — compose refuses to start without it.
API_TOKEN: ${API_TOKEN:?set API_TOKEN in .env}
ALLOWED_ORIGINS: ${ALLOWED_ORIGINS:-https://asuracomic.net,https://asurascans.com,https://demonicscans.org,https://comix.to,https://kagane.to,https://novelfull.com,https://lightnovelworld.net}
DB_PATH: /data/bookmarks.db
# TOKEN_KEY derives every Reader's userscript credential (issue #24) —
# compose refuses to start without it.
TOKEN_KEY: ${TOKEN_KEY:?set TOKEN_KEY in .env}
# Owner's Discord user ID — required. Seeds the owner Reader (the
# administrator); every other Reader registers on their first login.
OWNER_DISCORD_ID: ${OWNER_DISCORD_ID:?set OWNER_DISCORD_ID in .env}
ALLOWED_ORIGINS: ${ALLOWED_ORIGINS:-https://asurascans.com,https://demonicscans.org,https://comix.to,https://kagane.to,https://novelfull.com,https://lightnovelworld.net}
# The bookmarks database. Host is the compose service name; the password
# comes from .env so it is never committed.
DATABASE_URL: ${DATABASE_URL:-postgres://bookmarks:${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD in .env}@postgres:5432/bookmarks?sslmode=disable}
# Required path inside the API. The build seeds ownership at this path
# and the named volume below mounts there.
COVER_DIR: ${COVER_DIR:?set COVER_DIR in .env}
# Public origin of this deployment, no trailing slash. Required: the
# Cover URLs on the wire are absolute, since the userscript renders them
# on a Site's origin rather than ours (ADR-0007).
PUBLIC_BASE_URL: ${PUBLIC_BASE_URL:?set PUBLIC_BASE_URL in .env}
PORT: "8080"
# Gates the browser UI. Unset means the web routes are not served at all.
WEB_PASSWORD: ${WEB_PASSWORD:-}
# Log timestamps only. Go's `log` stamps lines in local time, and this
# service has no other use for a zone: bookmark timestamps are unix ms
# and the two real time columns are timestamptz, both absolute instants.
# Purely so these lines read on the same clock as the browser's. Named
# API_TZ rather than TZ so an operator's exported shell TZ cannot leak
# in; distroless already carries tzdata, so the name just resolves.
TZ: ${API_TZ:-Asia/Jakarta}
# Discord OAuth for the browser UI (ADR-0002). The first four are
# required; DISCORD_REQUIRED_ROLE is optional and empty by default.
# Guild membership is the whole gate: any member becomes a Reader.
DISCORD_CLIENT_ID: ${DISCORD_CLIENT_ID:?set DISCORD_CLIENT_ID in .env}
DISCORD_CLIENT_SECRET: ${DISCORD_CLIENT_SECRET:?set DISCORD_CLIENT_SECRET in .env}
DISCORD_GUILD_ID: ${DISCORD_GUILD_ID:?set DISCORD_GUILD_ID in .env}
DISCORD_REQUIRED_ROLE: ${DISCORD_REQUIRED_ROLE:-}
DISCORD_API_BASE: ${DISCORD_API_BASE:-https://discord.com/api/v10}
DISCORD_REDIRECT_URI: ${DISCORD_REDIRECT_URI:?set DISCORD_REDIRECT_URI in .env}
# Path inside the container; matches the bindmount above.
USERSCRIPT_PATH: ${USERSCRIPT_PATH:-/userscript/manga-bookmark.user.js}
# Second script from the same bindmount; the novel library is a separate
# Violentmonkey install.
NOVEL_USERSCRIPT_PATH: ${NOVEL_USERSCRIPT_PATH:-/userscript/novel-bookmark.user.js}
# Latest-chapter poller. LATEST_CHAPTER_POLL_ENABLED=0 in .env is the kill
# switch; it only takes effect because these are listed here.
# Latest-chapter poller. LATEST_CHAPTER_POLL_ENABLED=0 in .env is the
# kill switch; it only takes effect because it is listed here. Pace is
# per Site in the registry (one Poll Lane per Site, issue #100) — the
# cooldown/interval/stagger/batch knobs are gone with the shared pace.
LATEST_CHAPTER_POLL_ENABLED: ${LATEST_CHAPTER_POLL_ENABLED:-1}
LATEST_CHAPTER_POLL_COOLDOWN: ${LATEST_CHAPTER_POLL_COOLDOWN:-1h}
LATEST_CHAPTER_POLL_INTERVAL: ${LATEST_CHAPTER_POLL_INTERVAL:-10m}
LATEST_CHAPTER_POLL_BATCH: ${LATEST_CHAPTER_POLL_BATCH:-14}
LATEST_CHAPTER_POLL_STAGGER: ${LATEST_CHAPTER_POLL_STAGGER:-20s}
# CDP endpoint for sites behind a JavaScript challenge (kagane). Unset
# disables browser polling for those sites; the userscript still covers them.
# Must be an IP, not the "headless-shell" DNS name: Chrome's DevTools HTTP
# handler rejects the discovery request (GET /json/version) with a 500
# unless the Host header is an IP address or "localhost" — confirmed
# 2026-08-03 against chromedp/headless-shell:stable, independent of
# chromedp's own dial logic. The sidecar's static address below exists so
# this URL survives container recreation.
BROWSER_WS_URL: ${BROWSER_WS_URL:-ws://172.28.0.10:9222}
# CDP endpoint for sites behind a JavaScript challenge (kagane,
# novelfull). The browser is not part of this stack — it runs on the home
# machine as its own unit (chrome/docker-compose.yml) and is reached over
# the tailnet. Unset disables browser polling for those sites and serves
# 404 from the cover proxy for covers not already stored; the userscript
# still covers them. Set it in .env to ws://<home machine tailnet IP>:9222.
#
# Must be an IP, not a MagicDNS hostname: Chrome's DevTools HTTP handler
# rejects the discovery request (GET /json/version) with a 500 unless the
# Host header is an IP address or "localhost" — confirmed 2026-08-03,
# independent of chromedp's own dial logic. The same trap that used to
# force a pinned Docker IP now forbids the tailnet name.
BROWSER_WS_URL: ${BROWSER_WS_URL:-}
depends_on:
- headless-shell
# The migration runner is the first thing the binary does, so a Postgres
# that is still initialising means a crash-loop until it is not.
postgres:
condition: service_healthy
volumes:
- bookmarks-data:/data
# The userscript is served from here, read fresh on every request. Editing
# the file in this checkout takes effect on the next Violentmonkey poll —
# no rebuild, no restart. `git pull` restores the committed version, which
# is why a redeploy always ships the repo's script.
- ./userscript:/userscript:ro
# Content-addressed cover bytes survive API restarts and redeploys.
- cover-data:${COVER_DIR:?set COVER_DIR in .env}
# Bound to loopback only: the proxy (or curl during smoke test) reaches it,
# the public internet does not.
ports:
- "127.0.0.1:8080:8080"
# `default` is not decoration: `db` is `internal: true`, and a container on
# nothing but an internal network gets neither a published port nor egress
# — which would silently kill every poller fetch.
networks:
- browser
- default
- db
headless-shell:
image: chromedp/headless-shell:stable
postgres:
image: postgres:17-alpine
restart: unless-stopped
# Chrome allocates shared memory per tab and dies on Docker's 64MB default.
shm_size: '1gb'
# Reaps zombie renderer processes, which otherwise accumulate for the
# container's lifetime.
init: true
# Deliberately no `ports:` — an exposed CDP endpoint is remote code
# execution. Only bookmark-api, via the `browser` network below, may reach it.
# Don't pass --remote-debugging-address/--remote-debugging-port here: the
# image's own entrypoint (/headless-shell/run.sh) already starts Chrome on
# 127.0.0.1:9223 and fronts it with a socat proxy listening on 0.0.0.0:9222.
# Redeclaring the port flag here overrides Chrome's, so it binds 9222
# directly (IPv6 loopback only) instead of 9223 — collides with socat's own
# bind on 9222 and leaves nothing listening on 9223, so every external
# connection to headless-shell:9222 fails with EOF. Only pass flags the
# entrypoint doesn't already set.
command:
- --disable-gpu
- --no-sandbox
environment:
POSTGRES_DB: bookmarks
POSTGRES_USER: bookmarks
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD in .env}
healthcheck:
test: ["CMD-SHELL", "pg_isready -U bookmarks -d bookmarks"]
interval: 5s
timeout: 3s
retries: 10
volumes:
- postgres-data:/var/lib/postgresql/data
# Deliberately no `ports:` — only bookmark-api, over the `db` network,
# reaches it. Use `docker compose exec postgres psql` for a shell.
networks:
browser:
# Pinned so BROWSER_WS_URL can name an IP (required, see above) that
# survives `docker compose up` recreating this container.
ipv4_address: 172.28.0.10
- db
volumes:
bookmarks-data:
postgres-data:
cover-data:
# The pre-Postgres SQLite volume (bookmarks-data) is deliberately no longer
# declared here: undeclared means `docker compose down -v` cannot take it
# with the rest, so the old database survives the cutover until someone
# removes it by hand.
networks:
# Not `internal: true`: headless Chrome still needs outbound access to reach
# kagane.to. Isolation here comes from membership (only bookmark-api and
# headless-shell join it), not from cutting egress.
browser:
ipam:
config:
- subnet: 172.28.0.0/24
# Postgres needs no egress and nothing outside bookmark-api needs to reach
# it, so this one really can be cut off from the outside world.
db:
internal: true
+40
View File
@@ -0,0 +1,40 @@
# Postgres replaces SQLite as the primary datastore
Status: accepted
The project is moving from a single-reader tracker to a service published to a community,
so we replaced `modernc.org/sqlite` with Postgres (`jackc/pgx/v5`, still pure Go, so
`CGO_ENABLED=0` and the distroless image are unaffected). The deciding reason is future
supportability — managed hosting, a datastore that survives the app outgrowing one box —
**not** concurrency, which was measured and found to be a non-issue.
## Considered options
**Stay on SQLite.** Benchmarked against the real store at 1,500 rows (≈50 readers × 30
series): ~9,700 upserts/sec single-writer, plateauing at ~780/sec under 8–50 concurrent
writers, with `List()` holding 90–103 calls/sec under continuous write load. Projected
real load at 50 readers is ~0.02 writes/sec — roughly four and a half orders of magnitude
of headroom. `SetMaxOpenConns(1)` serialises writes but was shown not to starve reads;
an apparent read collapse traced to row count and per-row scanning, not lock contention.
SQLite would have worked. It was rejected for where the project is going, not for what
it does today.
**Postgres.** Chosen. Migrating is cheapest now — 29 rows in one table — and gets
materially harder once there are live readers and a multi-tenant schema.
## Consequences
- Every statement in `internal/store` is rewritten: `?` → `$N`, `IS NOT` →
`IS DISTINCT FROM` (this one is load-bearing; it implements the `updated_at`
ordering rule), `pragma_table_info` → `information_schema.columns`,
`INTEGER`/`REAL` → `bigint`/`double precision`, `favorite` int-as-bool → `boolean`.
- ~75 tests currently get a free isolated database from `t.TempDir()`. They now need a
live server, which makes Docker a hard prerequisite for `go test ./...`. This is the
permanent cost of the decision and the main reason it was close.
- Backups get worse, not better: `VACUUM INTO` produced one self-contained file;
restoring now means `pg_dump`/`pg_restore`, a role, and a password.
- A second stateful container joins the VPS alongside the existing headless-shell.
- **Postgres does not address the real scaling limit.** At batch 14 per 10-minute tick
the poller checks at most 84 series/hour; 50 readers × 30 series is 1,500 bookmarks,
an 18-hour sweep against a configured 1-hour cooldown. That ceiling is an outbound
fetch budget and is fixed by deduplicating polls per Series, not by the datastore.
@@ -0,0 +1,44 @@
# Identity comes from Discord OAuth; we store no passwords and send no email
Status: accepted
The service is being published to a community that already lives on Discord, and we have
no transactional email infrastructure. Rather than build email verification and password
reset to get accounts, Readers sign in with Discord OAuth2 (authorization code grant,
`identify` + `guilds.members.read`), and guild membership replaces both the invite gate
and the email-verification step. No password is ever stored and no mail is ever sent.
## Considered options
**Email + password with invite codes, no verification.** Viable and dependency-free:
an invite code proves community membership, which is what email verification was
standing in for anyway. Rejected because it still requires password hashing, a manual
admin-driven reset path, and a credential store — all of which Discord removes.
**Email + password with a transactional provider** (Resend, Brevo). Rejected as
premature: it builds verification and self-serve reset before anyone has asked for them,
and adds deliverability as an operational concern.
**Discord OAuth.** Chosen. It is less code than either alternative — no hashing, no
reset flow, no invite table — and the authorization question ("is this person in my
community?") is answered by the same call that answers the authentication question.
## Consequences
- **Availability is now coupled to Discord.** If Discord's OAuth endpoint is down,
nobody can start a new session. Existing sessions are unaffected, which bounds the
blast radius.
- **Identity is a Discord snowflake.** Migrating off Discord later means re-identifying
every Reader, because we hold no other credential for them. This is the lock-in the
decision buys, and it is the reason this ADR exists.
- **`guilds.members.read` is checked at login, not continuously.** Someone who leaves
the guild keeps their session until it expires. Acceptable; revocation is a session
delete, not an architectural change.
- **The userscripts cannot use OAuth.** They run in an isolated world on third-party
pages with no redirect surface, so they keep a bearer token — now issued per Reader by
the backend rather than a single shared `API_TOKEN` literal. OAuth gates the web UI;
the web UI is where a Reader obtains their personal userscript.
- `WEB_PASSWORD` disappears, and with it the session HMAC key derivation
(`sha256(API_TOKEN | WEB_PASSWORD | …)`), which needs a replacement secret.
- Seeding the first Reader during migration requires knowing the owner's Discord user
ID up front — a stable snowflake, copied from the Discord client.
@@ -0,0 +1,51 @@
# Series is a shared entity, and only the Poll may update it
Status: accepted
Facts about a Series that are true regardless of who is reading — title, cover, canonical
URL, Latest Chapter — moved off the Bookmark onto a shared `series` row keyed
`(site, series_id)`. A Bookmark now holds only what differs between Readers: Progress,
Favourite, Lifecycle bucket. Fifty Readers tracking one Series produce fifty Bookmarks
and one Series, so the Series is polled once rather than fifty times.
## Why
The poller checks at most 84 series/hour (batch 14 per 10-minute tick). With ~50 Readers
holding ~30 Series each, polling per Bookmark means a 1,500-item sweep — roughly 18 hours
against a configured 1-hour cooldown, quietly breaking the New Chapter signal that is the
product's reason to exist. Deduplicating to distinct Series cuts the sweep several-fold,
and because the Series row now knows how many Readers hold it, the poll queue is ordered
`reader_count DESC, latest_checked_at ASC` — popular Series stay fresh and the long tail
absorbs the shortfall. That ordering is only expressible because the split happened.
Raising throughput instead was rejected: sweeping 400 Series hourly needs the stagger
cut from 20s to ~9s, doubling request rate against sites already fronted by Cloudflare
from the single VPS IP.
Corrected 2026-08-12: the original wording said those sites "bot-score" the VPS IP.
They do not — the 1-99 bot score is Enterprise Bot Management only, and no per-IP
request rate is documented as an input to challenge issuance
(`docs/research/cloudflare-bot-scoring-and-poll-cadence.md`). The decision stands on
its first argument, sweep depth versus the 1-hour cooldown; the rate-limit fear was
never evidenced.
## Only the Poll writes Series fields
A client may supply `title`, `cover` and `series_url` only when creating a Series nobody
has bookmarked yet. After that, client-supplied values are ignored; only the backend's
own fetch updates them.
This is a security boundary, not tidiness. Those values are scraped from third-party
pages, which `AGENTS.md` requires be treated as attacker-controlled. Before the split, a
hostile or compromised site could corrupt exactly one Reader's row. After it, the same
write lands on a row every Reader sees — one Reader's browser becomes a write path into
everyone else's UI, and a cover URL can point anywhere. The backend's own fetch is the
higher-trust source: its network, its parser, no third-party JavaScript in the path.
## Consequences
- `Store.Upsert` decomposes one incoming flat body across two tables and enforces the
ownership rule at that seam.
- The `updated_at` ordering rule stays on the Bookmark, where Progress lives. Unchanged.
- Per-Reader title overrides are deliberately not supported; they would reintroduce the
duplication this removes.
@@ -0,0 +1,32 @@
# The wire format stays flat and deliberately does not mirror the schema
Status: accepted
Storage splits a tracked series across two tables (ADR-0003), but `GET /bookmarks` and
`PUT /bookmarks/{key}` keep emitting and accepting one **flat** JSON object with `title`,
`cover`, `last_chapter` and `latest_chapter` as siblings — exactly the shape they had
when there was one table. The server joins on the way out and decomposes on the way in.
## Why a future reader will find this surprising
The obvious move after splitting a table is to nest the JSON to match. Don't "fix" this.
**A nested payload would have broken every installed userscript instantly.** Scripts read
`b.title` directly; moving it to `b.series.title` yields `undefined` — no error, just
blank rows and a New Chapter signal that silently reports nothing forever. Because
Violentmonkey updates roughly once a day per device, the migration relies on old scripts
continuing to work during a 14-day grace window. A nested format and that grace window
are mutually exclusive.
**It is also the better contract independently of compatibility.** A client rendering one
row needs the title and the reading position together; nesting exports the re-stitching
to every browser to mirror a decision about disk layout it should not know about. Keeping
them separate lets storage change again later without a client release — which is the
whole reason this ADR is worth the paragraph.
## Consequence
The flat shape is a contract, not an implementation detail. Changing the storage schema
must not change it. It follows the rule already in force for `updated_at`: the server
owns the truth and returns the row **as stored**, and clients adopt the response rather
than their own payload.
+51
View File
@@ -0,0 +1,51 @@
# ADR-0005: On-demand browser sidecar
Date: 2026-08-09
Status: accepted
Superseded in part by ADR-0006: the lifecycle below is unchanged, but the
service no longer lives in the API stack and the name `headless-shell` is gone.
## Decision
Keep the `headless-shell` service and its CDP port alive, but start Google Chrome
only when the first CDP connection arrives. The entrypoint supervises a `socat`
front-end, serializes browser start/reap state with `flock`, and tracks each
connection with a marker named for its helper PID. A reaper stops Chrome after
300 seconds with no live markers. Marker reconciliation covers a helper killed
before its cleanup trap runs.
Chrome runs in its own process group so reap sends the termination signal to
Chrome and its renderer children. The explicit `/home/chrome/profile` user-data
directory remains: Chrome remaps remote debugging to loopback on modern builds,
and Chrome ignores remote-debugging flags on a default profile. `socat` therefore
continues to front Chrome's loopback CDP port.
The profile is a named Compose volume. Clearance cookies survive both a reap and
`docker compose up --build`; the browser still starts with a fresh debugger UUID,
so chromedp must keep endpoint discovery enabled and must not use
`chromedp.NoModifyURL`.
The socat front-end and explicit profile are retained because Chromium remaps a
non-loopback debugging address to loopback since M113, while Chrome ignores the
remote-debugging flags on a default profile since Chrome 136. Flag tuning is
deliberately not adopted: its roughly 30% idle-footprint saving is irrelevant
to a browser that exists for seconds per wake and risks an untested fingerprint.
## Constraints
The 300-second floor is deliberate. Chromium batches cookie persistence on a
roughly 31-second timer, and Go's default HTTP transport can keep the discovery
connection parked for about 90 seconds after use. Reaping only with zero live
connections holds Chrome through both windows and through the poller's staggered
batch plus cover prefetch.
The anti-bot properties remain unchanged: a plausible non-UTC timezone, a
Chrome-version-derived User-Agent without `HeadlessChrome`, and no automation
flag. A remote browser restart can surface as `context.Canceled`, the same error
as a caller deadline, so the backend wraps cancellation observed with a closed
CDP connection as `browser interrupted`; the focused test asserts that
classification without killing a real browser.
The same process-group stop runs during supervisor shutdown, not only during
idle reap, so Chrome can flush its cookie batch before a container rebuild or
graceful stop.
@@ -0,0 +1,81 @@
# ADR-0006: The browser runs on the home machine, over the tailnet
Date: 2026-08-09
Status: accepted
## Decision
The headless browser is no longer part of the API stack. It is its own compose
unit (`chrome/docker-compose.yml`), deployed on the home machine, and the API on
the VPS reaches it over the existing tailnet through `BROWSER_WS_URL`. No
fallback sidecar remains on the VPS.
The backend needs no code change for this. The CDP endpoint was already a
configuration seam and the fetcher only ever holds the endpoint URL, so
relocation — and reversal — is one environment variable.
## Why
The sidecar held 471 MiB working set (645 MiB peak) on a 1974 MiB VPS with no
swap, which also hosts Traefik, Gitea and its Postgres. That is 24% of the host
and 86% of this project's memory, for a service that at the time answered zero
requests: the poller's due query joins bookmarks, production held four kagane
series and no bookmarks on any of them, and with no kagane bookmark the web UI
never rendered a kagane cover either.
The home machine has 5.9 GiB of swap and a residential egress, which avoids the
cloud-hosting-IP signature Cloudflare's Bot Fight Mode documentedly challenges. Both
machines were already on the tailnet.
Corrected 2026-08-12: the original wording said Cloudflare "scores" a residential
egress better than a datacenter IP. There is no score on a free-plan zone; what is
documented is signature matching, and hosting-provider IP space is one of the
signatures (`docs/research/cloudflare-bot-scoring-and-poll-cadence.md`). Memory was
the load-bearing reason regardless.
This move is only safe because covers are persisted (ADR-0005's sibling work,
issue #43/#45) and the browser is on-demand (ADR-0005). Without stored covers a
sleeping home machine would blank the library; without on-demand start the CI
runner that already holds ~1.2 GiB of that box's 1.8 GiB would be squeezed
around the clock.
## Constraints
**`BROWSER_WS_URL` must be the tailnet IP, never a MagicDNS hostname.** Chrome's
DevTools HTTP handler answers `/json/version` with a 500 for any `Host` header
that is not an IP or `localhost`. This is the same trap that previously forced a
pinned Docker IP; the pinned subnet is gone, the constraint is not.
**The CDP port binds to the tailnet address only, never `0.0.0.0`.** CDP
authenticates nothing: whatever reaches the port drives the browser and, through
it, the host. On the VPS the safety came from Docker network membership; the
home machine has a real LAN, so a `0.0.0.0` bind is a hole punched into it. The
bind address is the enforcement and Tailscale device identity plus a per-device
ACL is the policy. `BROWSER_BIND_ADDR` deliberately has no default, so an unset
value fails the deploy instead of publishing CDP to the LAN.
No bearer-token proxy is added in front of CDP. It would only defend against a
device already inside the tailnet, and it would be one more thing between the
poller and a browser that is already hard enough to keep clearing challenges.
**Resource limits are load-bearing, not decorative.** The browser is the
newcomer on that box, not the incumbent. A hard 512 MiB cap with 1 GiB
memory+swap makes Chrome reclaim its own cold pages onto the machine's SATA swap
instead of taking resident memory from the runner; untuned Chrome peaked at
645 MiB cgroup, which is more than is free there. `oom_score_adj` biases the
kernel to kill the browser first and never CI. Reduced CPU weight makes a
challenge solve yield to a running build — cold start degrades to about 3 s at
half a CPU, immaterial against a 45-second challenge budget. The shared-memory
reservation drops from 1 GiB to 128 MiB against a measured 19 MiB peak.
## Consequences
An unreachable browser degrades exactly as an unset `BROWSER_WS_URL` already
does: plain-TLS libraries are unaffected, kagane and novelfull log and skip, the
series waits out its cooldown, and stored covers keep serving. A power outage at
home costs chapter freshness on two sites, never the appearance of the library.
The two units are deployed and updated independently. `REDEPLOY.md` §8 covers
the browser; everything before it covers the API stack. A local `docker compose
up` now brings up two services, not three, and polls kagane only if
`BROWSER_WS_URL` is pointed somewhere.
+105
View File
@@ -0,0 +1,105 @@
# ADR-0007: The backend hosts every Site's Cover bytes
Date: 2026-08-09
Status: accepted
## Decision
A Cover is the image a Reader's browser can display for a Series. Where the Site
keeps the picture is the backend's problem, not the client's: the backend fetches
the bytes, stores them, and serves them from its own origin. No client ever
renders a third-party URL, and no client ever supplies one.
Concretely:
- **Acquisition is server-side.** The Cover is extracted from the same series-page
fetch that already yields Latest Chapter. It runs once at Series creation rather
than waiting for the poll queue, so a newly bookmarked Series has both facts in
seconds instead of up to a queue's depth. The poll fills a blank Cover and never
overwrites a non-blank one.
- **Bytes live on a filesystem volume**, content-addressed by the SHA-256 of the
source URL, sharded `${COVER_DIR}/ab/cd/<sha256>`. The database holds the path
and content type, not the bytes.
- **One public route** serves them. No session, no credential.
- **The wire carries an absolute URL** built from a configured public base, and
carries `""` until the bytes exist.
## Why a future reader will find this surprising
Four of the six Sites let anyone hot-link their covers — `static.comix.to` even
answers `access-control-allow-origin: *`. Hosting copies looks like work we were
not obliged to do.
We were obliged. kagane serves covers with `cross-origin-resource-policy:
same-origin` behind a JavaScript challenge (measured 2026-08-08), so no `<img>`
outside kagane.to can load one under any combination of referrer policy and
`crossorigin` attribute. The first fix for that was a kagane-only proxy applied in
the web templates — and it produced issue #47, because the JSON API kept emitting
the raw kagane URL and the userscript rendered it into a broken-image glyph. A
per-Site exception that only one of two clients knows about is not a fix; it is a
bug with a delay on it. Uniformity is the property being bought: every client
renders every Cover the same way, and a Site changing its CORP header or its CDN
cannot break a client again.
## Considered options
**Per-Site exceptions, proxying only what must be proxied.** Cheapest, and what we
had. Rejected: it is what produced #47, and it requires every current and future
client to know which Sites are special.
**A host allowlist for the outbound fetch**, mirroring `fetchableSeriesURL`.
Rejected in favour of destination-class control — see below.
**Cover bytes in Postgres `bytea`**, extending the existing `covers` table.
Rejected: covers are immutable blobs served straight to browsers, which is what a
filesystem is for. The cost is real and accepted — durability is now two things to
back up instead of one, against ADR-0001's grain.
**Per-Reader Cover overrides.** Rejected, consistent with ADR-0003's rejection of
per-Reader title overrides. A Cover is a fact about the Series.
## Two deliberate relaxations
**Destination control is deny-class, not an allowlist.** The outbound fetch
requires `https`, resolves DNS first and refuses loopback, private, link-local and
CGNAT addresses, re-checks on every redirect hop, and caps body size and content
type. It does *not* pin a host set, which is what `fetchableSeriesURL` does for
`series_url`. Cover hosts are CDNs that move: `demonicscans.org` serves its covers
from `readermc.org`, a host with no visible relationship to the Site. An allowlist
would silently stop producing Covers the day a Site switched CDN, and the failure
would look like this bug. The resolved-IP check is the load-bearing part; without
it, an attacker-controlled page need only publish a DNS name pointing at
`127.0.0.1`.
**The cover route is public, where the kagane proxy was session-gated.** An `<img>`
in the userscript panel cannot send a bearer token, and it cannot be given one: the
panel's shadow root is `mode: "open"`, so the host page's own JavaScript can read
any `src` we set. A credential in an image URL is a credential handed to a
third-party site. The route serves public artwork from public Sites and its path
reveals nothing about which Reader holds what. The residual cost is that we can be
hot-linked by others.
## Consequences
- The poll's cover prefetch, today guarded by `sr.Site != "kagane"`, applies to
every Site in both Libraries. Nothing about Covers is conditioned on `kind`.
- Only kagane still needs the CDP browser for its bytes. The other five Sites fetch
over plain TLS — including novelfull, whose HTML answers `cf-mitigated: challenge`
while its image paths answer 200 with `access-control-allow-origin: *`
(measured 2026-08-09).
- Client-side cover scraping is deleted from both userscripts. It could not help: a
scraped URL has no render path left, and it is absent exactly when a Series is
created — neither comix nor lightnovelworld exposes a cover on a chapter page,
which is where a Reader bookmarks mid-read.
- `PUT /bookmarks/{key}` still accepts a `cover` field and ignores it. This extends
ADR-0003's "ignored after creation" to "ignored always", and keeps the flat wire
contract ADR-0004 requires so installed scripts keep working. The field is
therefore permanently inert rather than pending removal, and says so at the
decode site.
- A Cover that fails to load falls back to the placeholder in both clients. The
broken-image glyph reported in #47 is not a state we render.
- The existing kagane `covers` rows are dropped rather than migrated; that path
re-fetches on demand already.
- Two Sites deserve a note for whoever writes the extractor: asura's `.webp` cover
URL answers `Content-Type: image/jpeg`, so trust the header; demonic's `og:image`
carries a raw unencoded space and must be percent-encoded before fetching.
@@ -0,0 +1,102 @@
# ADR-0008: A Series identity is discovered from the Site's links, never derived from an address
Date: 2026-08-11
Status: accepted
## Decision
On lightnovelworld, a Series is identified by the slug in its `/novel/<slug>/`
address, and the userscript obtains that address by reading the chapter page's
`a[aria-label="All Chapter"]` anchor. It no longer constructs the address by
string manipulation of the chapter path. When neither that anchor nor the
microdata breadcrumb is present, the page resolves to `type: "other"` and no
Bookmark is offered.
A Chapter Slug — the slug a chapter address is built from — is not an identity
and is not stored. The backend finds chapters by matching
`lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/` against the series page
body **truncated at the first `wpd-threads`**.
## Why
Measured on lightnovelworld between 2026-08-10 and 2026-08-11.
The chapter slug and the series slug are two independent facts. In a 41-novel
sample, 3 diverged (~7%). `/my-longevity-simulation-chapter-1/` returns 200
while `/novel/my-longevity-simulation/` returns 404 with no redirect, and that
novel's real address is `/novel/immortality-simulator/`. Divergence runs in both
directions: `the-sword-illuminates-the-great-wilderness` is served by chapter
slug `radiant-blade-of-the-wilderness`. Neither slug is computable from the
other, and the Site publishes no alternative-names field, so the mapping exists
only in the chapter page's own markup.
Storing the Chapter Slug beside the identity does not work, because a Series may
have more than one. `/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/`
serves chapters 1–99 under `…-one-skill-not` and 100–423 under
`…-one-skill-not-them-all`. Both resolve, and both chapter pages point back at
the same Series.
Deriving the identity from the chapter path also made one Series produce two
rows: bookmarking from the series page yielded `lightnovelworld:immortality-simulator`,
and from a chapter page `lightnovelworld:my-longevity-simulation`.
The pointer is reliable. Across 8 chapter pages — chapter 1, chapter 1200, the
latest chapter, both slugs of the split novel, two divergent novels, and a novel
with a number in its title — the `All Chapter` anchor and breadcrumb position 2
were both present and agreed every time, including on the old-slug pages. Three
narrower selectors were rejected on evidence: `a[href*="/novel/"]` matches the
header nav index first; matching the text "All Chapter" false-matches the novel
titled "…Not Them All Chapter 200"; and the JSON-LD breadcrumb's position 2 is
the chapter, not the Series.
The scan is truncated because a series page server-renders a wpdiscuz comment
thread below the chapter list, and comment bodies are HTML that can carry an
anchor. Verified on `/novel/the-sword-illuminates-the-great-wilderness/`:
comment `#wpd-comm-358_0` rendered in the initial HTML, corroborated by that
page's comment RSS feed. The scanner takes the maximum chapter number with no
upper bound, and a Series row is shared by every Reader (ADR-0003), so one
comment containing a link to a high-numbered chapter would pin that Series'
Latest Chapter for everyone. `wpd-threads` occurs exactly once per page and
follows every chapter anchor on all 4 series pages measured. `wpdcom` and
`wpdiscuz` are unusable: they occur 111 to 143 times per page, including in
`<head>` before the chapter list.
## Considered options
**Correct `series_url` only, leaving the Chapter Slug as the identity.** This
repairs the Poll with no migration, because the Poll reads the stored address
rather than the identity. Rejected: it keeps an identity that the Site does not
guarantee to be stable, and leaves the duplicate-row hazard in place.
**Scope the match to the chapter-list container.** Rejected on measurement. The
prior research named `ul.clstyle`; on all 4 pages sampled that is the hidden,
empty "Latest Reading" template, and the real list is a classless `<ul>` in
`div.eplister.eplisterfull`. The backend scans raw HTML with no parser, so
container extraction means a second regex against class names, which is more
fragile than the one-off truncation and protects nothing extra.
**Repair the stored rows with a SQL migration.** Rejected as impossible: the
database holds no source for the correct slug. The `series` table has no chapter
address, and no endpoint on the Site maps one slug to the other.
## Consequences
Existing Bookmarks on divergent novels stop matching their own chapter pages,
because `keyOf` changes. The userscript therefore migrates a row in place when
it sees the mismatch: it rewrites the row's key, identity and address in the
cache, the retry-queue entry, and the last-checked map, then syncs. This heals
only when the Reader next opens a chapter page of that novel.
A migrated row leaves its old `series` row behind. Nothing deletes it, but
`DueForLatestCheck` is an inner join against bookmarks, so a Series with no
Bookmarks is never polled again. The old row is permanently stored and
permanently inert.
A Bookmark whose Reader never opens a chapter page of that novel does not heal.
It continues to return 404 at every cooldown, as it does today.
If the truncation marker disappears, the scan is skipped and logged rather than
run against the whole page. An env-gated live test, `TestSmokeLnwCommentBoundary`,
asserts that `wpd-threads` still occurs exactly once and still follows the last
chapter anchor. It skips when its environment variable is unset, matching the
existing `TestSmokeKagane*` convention.
@@ -0,0 +1,93 @@
# ADR-0009: A Site answers fixed questions; an unusual Site owns its own fetch
Date: 2026-08-11
Status: accepted
## Decision
Per-Site knowledge lives in one registry in `backend/internal/latest/sites.go`,
keyed by the stored site string. An entry answers a fixed set of questions: the
hostname a `series_url` must carry, how to find the Latest Chapter in a body,
how to find the Cover address in a body, and — for a Site behind a JavaScript
challenge — how to read its payload from a cleared tab, how to tell that the
payload arrived, and whether the plain-TLS fetcher may take over when no
browser is configured (`Fallback`; kagane never falls back, novelfull does,
each on measured evidence).
The question set does not grow to accommodate one Site. When a Site needs
something the set cannot express, that Site gets an optional override and
performs its own fetch, leaving the other entries untouched. The override is a
per-Site escape hatch, not a stage every Site passes through, and it is added to
the registry type when the first Site needs it rather than in advance.
Two things stay outside the registry. Cover *bytes* are routed by address shape
in `fetchCoverBytes`, never by Site name, so the Poll and the Acquisition cannot
drift apart. And the host pin inside a browser entry's read is kept even though
`fetchableSeriesURL` has already pinned the same host: a headless browser is a
strong SSRF primitive and `series_url` arrives in a client-supplied PUT body, so
the second check is deliberate and must not be deduplicated.
## Why
Before the registry, the site string was compared in six places across three
files: the Latest Chapter switch (`sites.go:102`), the Cover switch
(`sites.go:263`), the browser-backed list and the fetcher choice
(`poller.go:60`, `poller.go:132`), the host pins (`poller.go:314-325`), and the
payload read (`browser.go:96-119`). Nothing tied them together, so adding a
seventh Site meant finding all six unaided, and a Site added to five of them
failed at the sixth in production rather than at compile time.
The Sites are not alike and the registry does not ask them to be. asura strips a
rotating build hash from its slug before scoping a regex; comix reads a JSON
blob embedded in server-rendered HTML; kagane's chapter list exists only in its
JSON API, which must be called from inside the page so the request carries the
clearance cookie; lightnovelworld must truncate the body at the comment thread
first. What they have in common is not behaviour, it is the questions they
answer. Arbitrary behaviour behind one entry is the point.
Making a browser Site contribute a read and a completion test, rather than
letting it drive the browser, was chosen because the tab lifecycle in
`BrowserFetcher.run` is load-bearing and shared. It holds one tab open across
re-reads, because a Cloudflare interstitial needs several seconds of live page
to solve itself and write clearance into the shared cookie jar; reading once and
closing the tab, which is what this did before 2026-08-08, never clears
anything. It also serialises the browser, binds the caller's deadline to the
tab, distinguishes a lost browser from a retryable read, and paces re-reads.
Spreading that across per-Site adapters would put one subtle, measured loop
behind six doors.
This costs the adapters little, because `chromedp.Run` takes an Action and
`chromedp.Tasks` is an Action. A Site that must click, wait on a selector, and
then evaluate expresses all of it as its read. Only a Site needing something
outside the per-tab loop — its own cadence, two tabs, a tab held between calls,
cookies set before navigation — falls outside, and that Site takes the override.
## Considered options
**Widen the shared interface whenever a Site needs something new.** Rejected:
one Site's requirement becomes a field on all seven entries, and the entries
that ignore it still have to be read and understood by anyone adding the eighth.
**Give every Site the whole fetch.** Rejected: it makes the browser lifecycle
above a per-Site concern, and pulls `chromedp` into adapters for five Sites that
never open a browser.
**A Go `interface` with a method set instead of a registry of records.**
Rejected: most Sites differ in one or two answers, and three share a single
Cover implementation, so a method set produces near-empty types. A missing
answer is a nil value caught at dispatch, which is where an unknown Site is
already handled.
## Consequences
Adding a Site is one registry entry. The existing dispatch functions —
`latestChapterFrom`, `coverFrom`, `fetchableSeriesURL` — become registry
lookups, so the table tests that drive them by site string are unchanged.
A future architecture review will see an override that only one Site uses and
read it as an inconsistency to collapse. It is not. Collapsing it means either
widening the question set for every Site or moving the shared tab lifecycle into
the adapters, and both were rejected here on the evidence above.
An unknown site string resolves to the zero entry and fails the existing
not-fetchable and no-fetcher paths, which log and skip. That is unchanged.
+91
View File
@@ -0,0 +1,91 @@
# ADR-0010: Poll Lanes — one independent Poll stream per Site
Date: 2026-08-16
Status: accepted
## Decision
Replace the single shared polling pace with one **Poll Lane** per Site: an
independent goroutine that polls only that Site's Series, paced by that Site's
registry entry. Pace moves out of config and into the Site registry
(`internal/latest/sites.go`): every entry carries a `Rest` (how long a Series
rests between Polls) and a `Gap` (how long the Lane waits between fetches).
Rest is enforced by the due query's WHERE clause (`latest_checked_at <= now -
Rest`), never by a timer — the same mechanism that enforced the old cooldown.
The Lane enforces its own gap by sleeping between fetches. `effectiveGap` is
the registry gap, or one hour divided by the Site's eligible Series count when
that is smaller, never below one second.
The five environment settings that used to size the shared pace —
`LATEST_CHAPTER_POLL_COOLDOWN`, `_BROWSER_COOLDOWN`, `_INTERVAL`, `_BATCH`,
`_STAGGER` — are deleted. Only the kill switch `LATEST_CHAPTER_POLL_ENABLED`
remains. No deployed `.env` may carry the deleted knobs.
## Why
The shared pace capped the whole backend at roughly 180 Polls an hour (one
20-second stagger across one queue). ~60 Series today, scaling to hundreds or
thousands, would stretch the hour beyond what the New Chapter signal can
tolerate. Worse, the queue mixed Sites with very different costs: kagane and
comix pay seconds of a serialized single-tab Chrome per Poll (a challenged
page, ADR-0005), and one hostile Site burning its challenge timeout made every
other Site's Series wait — "one hostile Site can eat most of an hour".
Lanes fix both at once:
- **Throughput scales per Site.** The six Lanes fetch concurrently; a Lane's
own gap paces it. The browser Lanes' combined ceiling stays about 360 Polls
an hour (one tab), and when they cannot keep up the wait past Rest grows and
is logged every pass — the "behind by X" measurement, so the decision to
give browser Sites more pages is made from data.
- **Hostility is contained.** A refusal (two challenge-held reads in one
pass) stops only that Site's Lane for `refuseBackoff` (15m); the rest of
that Lane's Series stay unstamped and due. A lost browser gates the other
browser Lanes' passes for the same window — the flag is shared Poller
state, so the loss is noticed once instead of once per Lane per pass, and
decays after 15m so the Lanes probe again. One Site can no longer tax the
others.
## Tradeoffs and rejections
- **Per-Site env knobs** (e.g. `KAGANE_POLL_GAP`) rejected: the registry is
the single place pace lives, testable and reviewable; config knobs would
recreate the shared-pace sprawl with six times the surface. All six entries
are deliberately uniform at first — rest an hour, gap ten seconds — so the
structure exists to differ without inventing numbers for Sites that have
not earned them.
- **Dynamic gap** (`rest / eligible`) is the one knob that stays automatic:
a Site with more Series than one per ten seconds would otherwise back up
behind its own gap, and the per-Series share of the hour is the natural
pace. The ten-second default is not arbitrary: one request per ten seconds
is the strictest rate rule a free-plan Site can even express (per-zone
rate limiting, as documented in
`docs/research/cloudflare-bot-scoring-and-poll-cadence.md`), so the
default pace is exactly what the most restrictive Site would demand of us.
The computed gap never goes below one second and logs loudly when the
floor engages.
- **Timer-based pacing** rejected: the old ticker made the poller's rate a
function of wall clock rather than of what was due. The due-query cutoff is
the only rate authority; the Lane sleep just prevents hammering.
- **Batch size** (the old `_BATCH` cap) is gone with the shared pace: a Lane
processes everything due, paced by its gap. There is no global queue left
to bound.
## Constraints preserved
- Stamp-before-fetch ("attempted" semantics): an untried Series stays due, so
a browser that appears after a restart finds its full queue waiting.
- Browser wake gate (ADR-0005): a browser Lane leaves Chrome asleep below
five due Series and 15 minutes of wait, per Lane.
- The browser is not in the API stack (ADR-0006): an unreachable browser
degrades a Lane exactly as an unset `BROWSER_WS_URL` — browser-only Sites
skipped, plain-TLS unaffected, stored covers still served.
- Cover heals moved to background goroutines (joined by the test suite via
`waitCovers`) so a slow cover CDN cannot consume a Lane's gap.
Supersedes the pace mechanics of ADR-0003's "raise throughput instead" note
(the stagger cut it rejected is what the per-Lane gap replaces) and the
6-hour browser cooldown introduced with the browser-backed Sites; the
1-hour browser rest was already cleared as safe by
`docs/research/cloudflare-bot-scoring-and-poll-cadence.md`.
+47
View File
@@ -0,0 +1,47 @@
# Domain Docs
How the engineering skills should consume this repo's domain documentation when exploring the
codebase. Layout: **single-context** — one `CONTEXT.md` plus `docs/adr/` at the repo root.
## Before exploring, read these
- **`CONTEXT.md`** at the repo root — the glossary / ubiquitous language.
- **`docs/adr/`** — read ADRs that touch the area you're about to work in.
If any of these files don't exist, **proceed silently**. Don't flag their absence; don't suggest
creating them upfront. The `/domain-modeling` skill (reached via `/grill-with-docs` and
`/improve-codebase-architecture`) creates them lazily when terms or decisions actually get resolved.
Neither exists yet in this repo. The existing `AGENTS.md` / `CLAUDE.md` and `docs/design-system.md`
carry the current architecture and design law — read those regardless.
## File structure
```
/
├── CONTEXT.md
├── docs/adr/
│ ├── 0001-....md
│ └── 0002-....md
├── backend/
└── userscript/
```
If this repo ever splits into genuinely separate contexts, add a root `CONTEXT-MAP.md` pointing at
one `CONTEXT.md` per context and update this file.
## Use the glossary's vocabulary
When your output names a domain concept (in an issue title, a refactor proposal, a hypothesis, a
test name), use the term as defined in `CONTEXT.md`. Don't drift to synonyms the glossary
explicitly avoids.
If the concept you need isn't in the glossary yet, that's a signal — either you're inventing
language the project doesn't use (reconsider) or there's a real gap (note it for
`/domain-modeling`).
## Flag ADR conflicts
If your output contradicts an existing ADR, surface it explicitly rather than silently overriding:
> _Contradicts ADR-0002 (…) — but worth reopening because…_
+60
View File
@@ -0,0 +1,60 @@
# Issue tracker: Gitea (`tea` CLI)
Issues and specs for this repo live as issues on the self-hosted Gitea instance
`gitea.violetcrown.my.id` (repo `sulthan/mangaBookmark`). **`gh` does not work here** — use
[`tea`](https://gitea.com/gitea/tea) for everything past plain git. Auth lives in `tea login`,
not a `GH_TOKEN` env var. `tea` infers the repo from the local clone's `origin`.
`tea` prints rendered boxes rather than plain text; pass `--output json` (or `-o json`) when a
skill needs to parse the result.
## Conventions
- **Create an issue**: `tea issue create --title "..." --description "..."` (`--labels`,
`--assignees` optional). Multi-line bodies: pass the body through a shell variable or heredoc.
- **Read an issue**: `tea issue <number> --comments` (add `-o json` for machine-readable output).
- **List issues**: `tea issue list --state open -o json --fields index,title,body,labels,state,author`;
filter with `--labels "..."`, `--state open|closed|all`, `--assignee`, `--keyword`.
- **Comment**: `tea comment <number> "..."` (alias of `tea comments add`).
- **Apply / remove labels**: `tea issue edit <number> --add-labels "..."` / `--remove-labels "..."`.
Labels must exist first — see `tea labels list` / `tea labels create --name "..." --color "#rrggbb"`.
- **Close**: `tea issue close <number>` (comment separately with `tea comment`; `close` takes no
`--comment` flag).
## Pull requests as a triage surface
**PRs as a request surface: no.** _(Set to `yes` if this repo treats external PRs as feature
requests; `/triage` reads this flag.)_
When set to `yes`, PRs run through the same labels and states as issues, using the `tea pr`
equivalents: `tea pr <number> --comments`, `tea pr list --state open -o json`,
`tea pr create --head <branch> --base main --title "..." --description "..."`, `tea comment`,
`tea pr close`. Gitea shares one index space across issues and PRs, so a bare `#42` may be either
— resolve with `tea pr 42` and fall back to `tea issue 42`.
## When a skill says "publish to the issue tracker"
Create a Gitea issue with `tea issue create`.
## When a skill says "fetch the relevant ticket"
Run `tea issue <number> --comments`.
## Wayfinding operations
Used by `/wayfinder`. The **map** is a single issue; **tickets** are child issues.
- **Map**: one issue labelled `wayfinder:map` holding the Notes / Decisions-so-far / Fog body.
`tea issue create --labels wayfinder:map --title "..." --description "..."`.
- **Child ticket**: an issue labelled `wayfinder:<type>` (`research`/`prototype`/`grilling`/`task`)
with `Part of #<map>` as the first body line, and a task-list entry in the map body. `tea` has no
sub-issue command, so the task list plus the `Part of` line is the canonical link.
- **Blocking**: a `Blocked by: #<n>, #<n>` line at the top of the child body. Gitea's native issue
dependencies exist in the API but `tea` does not expose them; the body line is the source of
truth. A ticket is unblocked when every listed blocker is closed.
- **Frontier query**: `tea issue list --state open -o json` scoped to the map's task list; drop any
ticket with an open blocker or an assignee; first in map order wins.
- **Claim**: `tea issue edit <n> --add-assignees <your-username>` — the session's first write.
(`tea` has no `@me` shorthand; use the Gitea username from `tea login list`.)
- **Resolve**: `tea comment <n> "<answer>"`, then `tea issue close <n>`, then append a context
pointer to the map's Decisions-so-far via `tea issue edit <map> --description "..."`.
+20
View File
@@ -0,0 +1,20 @@
# Triage Labels
The skills speak in terms of five canonical triage roles. This file maps those roles to the actual
label strings used in this repo's issue tracker (Gitea — see `docs/agents/issue-tracker.md`).
| Label in mattpocock/skills | Label in our tracker | Meaning |
| -------------------------- | -------------------- | ---------------------------------------- |
| `needs-triage` | `needs-triage` | Maintainer needs to evaluate this issue |
| `needs-info` | `needs-info` | Waiting on reporter for more information |
| `ready-for-agent` | `ready-for-agent` | Fully specified, ready for an AFK agent |
| `ready-for-human` | `ready-for-human` | Requires human implementation |
| `wontfix` | `wontfix` | Will not be actioned |
When a skill mentions a role (e.g. "apply the AFK-ready triage label"), use the corresponding label
string from this table.
Gitea will not auto-create labels on `tea issue edit --add-labels`; create a missing one first with
`tea labels create --name "<label>" --color "#rrggbb"`.
Edit the right-hand column to match whatever vocabulary you actually use.
+17 -4
View File
@@ -10,7 +10,7 @@ Implemented in:
| Surface | Files |
| --- | --- |
| Web UI (login, list, card, empty, errors) | `backend/internal/web/static/style.css`, `backend/internal/web/templates/{app,card,list,login,chrome,icons}.html`, `backend/internal/web/static/filter.js` |
| Web UI (login, list, card, empty, errors, admin) | `backend/internal/web/static/style.css`, `backend/internal/web/templates/{app,admin,lanes,readers,card,list,login,chrome,icons}.html`, `backend/internal/web/static/filter.js` |
| Userscript panel (Shadow DOM) | `userscript/manga-bookmark.user.js` — `TEMPLATE` and `CSS` at the bottom of the IIFE |
## 1. The one idea
@@ -73,6 +73,7 @@ Defined once in `backend/internal/web/static/style.css` `:root`, mirrored in the
| `--moss` | `#7fae86` | `#3d6c46` | finished accent |
| `--clay` | `#b5906f` | `#7c5533` | set-chapter accent |
| `--trash` | `#977671` | `#8c6558` | remove, at rest — icons need 3:1, not 4.5:1 |
| `--patina` | `#5fb3a6` | `#1f6f66` | admin page only — a Poll Lane needing attention, a Reader whose reports are blocked |
| `--play-hot-line` | `#3a1d18` | `#f0cfc6` | desktop cell border, play when `.is-new` |
| `--fav-line` | `#332b14` | `#e3d3a4` | desktop cell border, favourite when on |
| `--asura` | `#7d93a5` | `#4f6b80` | site tag |
@@ -81,9 +82,13 @@ Defined once in `backend/internal/web/static/style.css` `:root`, mirrored in the
| `--kagane` | `#9a8aa5` | `#6f5f7d` | site tag |
| `--hatch` / `--hatch-dim` | 135° 5px stripe | paper stripe | missing-cover slot |
`--slate`/`--moss`/`--clay`/`--brass` are held at the same weight deliberately:
one accent per action, so a press says which lane it belongs to, with none of
them competing with ember. Dark is the default (`color-scheme: dark light`);
`--slate`/`--moss`/`--clay`/`--brass`/`--patina` are held at the same weight
deliberately: one accent per meaning, so a press says which lane it belongs to,
with none of them competing with ember. `--patina` is the admin page's only
colour — a cool verdigris, the far side of the wheel from ember's crimson and
clear of the archive blue: system health is neither a new chapter nor
destruction, so it borrows neither `--ember` nor `--danger`.
Dark is the default (`color-scheme: dark light`);
light is a `@media (prefers-color-scheme: light)` override of the same names.
**Any new colour must be added in both branches** — light is not a filter over
dark, the hues are re-tuned.
@@ -140,6 +145,14 @@ Recurring specs (copy these rather than inventing sizes):
main#list article.card … | .empty
```
The owner's admin page (`admin.html`) is the same sheet with two sections in
place of the list — `.lanes` (Poll Lane rows) and `.readers` (the roster) —
and no library switch: it belongs to neither library, so its topbar carries a
plain `.ghost.back` link home. Both sections are eyebrow + hairline-separated
rows, the shape the roster already had as a fold-out. `.lanes` refreshes itself
every 30s via `hx-get="/ui/admin/lanes"` with `hx-swap="outerHTML"`; the roster
re-renders only in answer to an action.
**Brand mark**: an inline `<svg class="mark">` (`viewBox="0 0 200 172"`),
defined once in `chrome.html`'s `mark` template and reused by `app.html` and
`login.html` so it takes the page's `--ink`/`currentColor`/`--ember` rather
@@ -0,0 +1,243 @@
# Cloudflare bot scoring and poll cadence — what is actually documented
Research note for the browser-backed poller cadence decision (kagane.to, novelfull.com, comix.to). All pages were fetched live from **developers.cloudflare.com / blog.cloudflare.com on 2026-08-12**. Primary sources only: Cloudflare's own documentation, Cloudflare blog posts, and RFCs/standards where noted. Where Cloudflare does not publicly document something, this note says **`Not publicly documented`** instead of guessing. Repo-measured facts from the existing poller work are reused without re-derivation and marked as such.
The site configuration of the three challenged sites (which plan, which bot product, Challenge Passage setting, whether Precursor is enabled) is **not observable from outside** — Cloudflare does not expose a zone's security configuration to anonymous clients. Anything in this note that depends on those unknowns is flagged `[INFERENCE]`.
---
## Short answer
**No — polling once per hour per series, from one residential IP through one real Chrome holding a valid `cf_clearance`, carries no documented challenge risk beyond polling every six hours.** Challenge issuance on Free/Pro-grade protection (Bot Fight Mode, WAF rules) is signature- and fingerprint-driven (headless browsers, cloud-hosted IPs, browser signals); the only rate-aware detector — the per-request bot score — exists solely on Enterprise Bot Management, and free-plan Rate Limiting Rules count per-IP over 10-second windows, which 20–60 requests/hour cannot trip. Both cadences re-solve the challenge every visit anyway: `cf_clearance` expires after **30 minutes by default** (site-configurable), so a 1-hour gap always finds it expired. The documented lever that matters — already verified in this repo — is **fingerprint quality**: real Chrome + real timezone clears in ~4 s; headless variants never do. Residual, undocumented risk is site-specific: Challenge Passage, Precursor (behavior-bound re-challenge), and custom WAF rules are zone settings not observable from outside.
---
## Summary answer table
| Question | Answer | Section |
|---|---|---|
| What is a bot score? | Integer 1–99 = Cloudflare's certainty a request is automated; **Enterprise Bot Management only**; everyone else gets coarse "bot groupings" (Pro+ analytics) or nothing. | §1 |
| Is request frequency documented as a bot-score input? | Partially: ML inputs are "headers, session characteristics, and browser signals"; the `__cf_bm` cookie "measures a single user's request pattern". **No numeric rate threshold is documented.** Volume policing is Rate Limiting, a separate product. | §1, §5 |
| What can a free-plan site deploy? | Bot Fight Mode only: challenges *signatures* (headless browsers, cloud-hosting IPs) with a computational challenge; JavaScript Detections forced on; no scores, no analytics, cannot be skipped/customized. | §2 |
| Does the free tier score continuously? | **No.** No score exists on Free at all — granular scores need Enterprise Bot Management, groupings need Pro+. BFM just challenges signature matches. | §2 |
| What does `cf-mitigated: challenge` mean? | The response was a Cloudflare Challenge Page (any type); `challenge` is the only value; body is always `text/html`. | §3 |
| How is a successful solve remembered? | `cf_clearance` cookie, issued with `SameSite=None; Secure; Partitioned`; suppresses challenges while valid. | §3, §4 |
| `cf_clearance` lifetime? | **30 minutes by default**, configurable by the site via Challenge Passage (15–45 min recommended); +skew minutes; +1 h for XHR. | §4 |
| Is `cf_clearance` bound to IP / device? | Documented: "securely tied to the specific visitor and device it was issued to"; the *solve request* must come from the same IP that received the challenge (different IP → invalid solve → challenge loop). Replay from another machine/IP is therefore **not** valid. | §4 |
| What invalidates clearance early? | Precursor (if enabled): suspicious session → clearance reduced/invalidated, re-challenge even before expiry. Zone-level toggle; unobservable from outside. | §4 |
| Does polling more often raise challenge risk? | **No documented mechanism at 20–60 req/hour.** Free-plan rate limiting is 10 s/IP-only; DDoS thresholds are ~1,000 errors/sec. Scores (the only rate-aware thing) are Enterprise-only. | §5 |
| Is there a documented "legitimate poller" path? | Yes, but it requires **self-identification** (Web Bot Auth signature or published IP list + stable UA) via the verified-bots application — not anonymity. robots.txt is voluntary; nothing exempts anonymous scrapers. | §6 |
---
## 1. What a bot score is and what feeds it
### 1.1 The score itself
Cloudflare documents the bot score as "a score from _1_ to _99_ that indicates how likely that request came from a bot" — 1 = quite certain automated, 99 = quite certain human. Source: [Bot scores — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot-score/), read 2026-08-12.
Two access tiers, both gated:
- **Granular 1–99 scores are only available to Enterprise customers who purchased Bot Management.** "All other customers can only access this information through bot groupings in Bot Analytics" (categories: `Not computed` = 0, `Automated` = 1, `Likely automated` = 2–29, `Likely human` = 30–99, `Verified bot`). Bot groupings themselves require "a Pro plan or higher". Source: [Bot scores — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot-score/), read 2026-08-12.
- A score of 0 means "Bot Management did not evaluate the request" (redirected, handled by another feature) — "does not indicate the request is safe or human". Same source.
So: **on a Free-plan site there is no bot score at all, for anyone.** `[INFERENCE]` the three challenged sites are almost certainly not Enterprise Bot Management customers, but this is not externally verifiable.
### 1.2 The detection engines (Enterprise Bot Management)
Cloudflare documents four engines, all stated to apply to Enterprise Bot Management (the bot-score page: "The following detection engines only apply to Enterprise Bot Management"). Sources: [Bot scores — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot-score/) and [Bot detection engines — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot-detection-engines/), both read 2026-08-12.
| Engine | Documented behavior | Score it produces |
|---|---|---|
| **Heuristics** | "Processes all requests"; pattern matching against "a growing database of malicious fingerprints". | 1 for high-confidence deterministic detections; occasionally 29 "where Cloudflare has identified automated traffic and is still assessing traffic overlap" |
| **Machine learning** | Supervised model, trained on "billions of daily requests". Input variables: "headers, session characteristics, and browser signals". Output: "predicted probability that a client is human (such as the probability of successfully solving a Challenge)". | Most scores 2–99 |
| **Anomaly detection** | Unsupervised; learns a per-domain baseline, flags outlier requests; **deprecated, not onboarding new customers**. | 1 |
| **JavaScript detections** | "Identifies headless browsers and other automation tools" via "a lightweight, invisible JavaScript injection"; runs client-side; "blocks, challenges, or passes requests to other engines". Enabled by default (but optional) in Bot Management. | Pass/fail (`cf.bot_management.js_detection.passed`), not a score |
Crucially, the ML engine's documented inputs are *headers, session characteristics, and browser signals* — **no rate or per-IP volume parameter is listed.** The only place request patterns appear is the `__cf_bm` cookie note: "Cloudflare uses the `__cf_bm` cookie to smooth out the bot score and reduce false positives… The Bot Management cookie measures a single user's request pattern and applies it to the machine learning data to generate a reliable bot score for all of that user's requests." ([Bot scores — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot-score/), read 2026-08-12). So frequency is *a* signal inside Enterprise Bot Management via the session cookie — **but no numeric threshold, window, or per-IP rate is published anywhere.** `Not publicly documented`: any specific requests-per-hour / requests-per-IP value that raises or lowers a bot score.
### 1.3 Rate limiting is a separate product
Volume enforcement is not part of bot scoring at all. Rate Limiting Rules are a distinct WAF product with their own evaluation phase (`http_ratelimit`, running after custom rules and before SBFM). Sources: [Rate limiting rules — Cloudflare docs](https://developers.cloudflare.com/waf/rate-limiting-rules/) and [Security features interoperability — Cloudflare docs](https://developers.cloudflare.com/waf/feature-interoperability/), read 2026-08-12. See §5 for what the Free plan's version of that product can actually do.
---
## 2. The free-plan reality
### 2.1 What each plan gets
Cloudflare's plan table ([Plans — Cloudflare docs](https://developers.cloudflare.com/bots/plans/), read 2026-08-12):
| Plan | Bot product | Documented detections | Action | Control |
|---|---|---|---|---|
| **Free** | **Bot Fight Mode** (BFM) | "Simple bots from cloud hosting providers and headless browsers" | "Cloudflare issues a computationally expensive challenge" | Applied to all traffic across the domain; no exceptions possible |
| **Pro / Business / Enterprise (no BM)** | **Super Bot Fight Mode** (SBFM) | Configurable actions per bot category (Definitely automated / Likely automated / Verified bots) | Challenge or block | Runs on Ruleset Engine; **can** be skipped via custom rules |
| **Enterprise + Bot Management** | Bot Management | "Simple and sophisticated bots, headless browsers, and domain-specific anomalies" | Customer-chosen (block, challenges) | Per-path / per-IP rules; access to bot score, JA3/JA4, bot tags, detection IDs |
Sources: [Plans — Free](https://developers.cloudflare.com/bots/plans/free/), [Plans — Bot Management for Enterprise](https://developers.cloudflare.com/bots/plans/bm-subscription/), [Bot Fight Mode — Cloudflare docs](https://developers.cloudflare.com/bots/get-started/bot-fight-mode/), [Super Bot Fight Mode — Cloudflare docs](https://developers.cloudflare.com/bots/get-started/super-bot-fight-mode/) — all read 2026-08-12.
### 2.2 Bot Fight Mode specifics (the Free-plan product)
- Identifies "traffic matching patterns of known bots" and "issues computationally expensive challenges that force the requesting client to perform CPU-intensive calculations". ([Bot Fight Mode — Cloudflare docs](https://developers.cloudflare.com/bots/get-started/bot-fight-mode/), read 2026-08-12.)
- It "does not run on the Ruleset Engine — it operates in a separate evaluation pipeline where _Skip_, _Bypass_, and _Allow_ actions have no effect"; **you cannot bypass or skip BFM** with custom rules or Page Rules. ([Bot Fight Mode — Cloudflare docs](https://developers.cloudflare.com/bots/get-started/bot-fight-mode/) and [Security features interoperability — Cloudflare docs](https://developers.cloudflare.com/waf/feature-interoperability/), read 2026-08-12.)
- **JavaScript Detections is automatically enabled for BFM customers and cannot be disabled.** ([Bot Fight Mode — Cloudflare docs](https://developers.cloudflare.com/bots/get-started/bot-fight-mode/), read 2026-08-12.) This is the documented hook that explains the repo's `headless-shell` / `HeadlessChrome` failures: JSD "identifies headless browsers" ([JavaScript detections — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/javascript-detections/), read 2026-08-12).
- False positives on *legitimate automated* traffic are acknowledged as expected behavior: "false positives can occur where legitimate human or automated traffic is incorrectly challenged or blocked", and the only remedies are disabling BFM or upgrading to Bot Management. ([Handle False Positives — Cloudflare docs](https://developers.cloudflare.com/bots/troubleshooting/false-positives/), read 2026-08-12.)
### 2.3 Does the free tier "score" continuously?
**No.** The Free plan exposes no score, no bot analytics (groupings need Pro+), and no per-request decision data. BFM is a static on/off toggle that challenges signature matches; there is no continuous per-request score on Free. ([Plans — Free](https://developers.cloudflare.com/bots/plans/free/) and [Bot scores — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot-score/), read 2026-08-12.) JSD *runs* on every HTML request even on Free (forced by BFM), but the only documented way to act on its result — the `cf.bot_management.js_detection.passed` field — is gated behind an Enterprise Bot Management subscription ("Prerequisites: You must have an Enterprise Bot Management subscription"). ([JavaScript detections — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/javascript-detections/), read 2026-08-12.) On Free, JSD output feeds BFM's internal challenge decision; `Not publicly documented` exactly how.
`[INFERENCE]` The observed behavior on the three sites (real Chrome + real timezone passes in ~4 s; headless variants never do) is consistent with either BFM or a WAF custom rule using a challenge action — Cloudflare does not expose which product a site runs, and the failure signature (HeadlessChrome UA / headless-shell never passing) matches JSD's documented headless-browser detection either way.
---
## 3. Managed Challenge / JS challenge mechanics and `cf-mitigated`
### 3.1 What the observed response is
A `403` with `cf-mitigated: challenge`, `server: cloudflare`, and a "Just a moment…" body is a **Cloudflare Challenge Page**. Cloudflare documents: "the Challenge Page response (regardless of the Challenge Page type) will have the `cf-mitigated` header present and set to `challenge`… `challenge` is the only valid value. The header is set for all Challenge Page types", and "the content-type of a challenge will be `text/html`". ([Detect a Challenge Page response — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/detect-response/), read 2026-08-12.)
### 3.2 What a Challenge Page does
"An interstitial Challenge Page… acts as a gate between the visitor and your website… The Challenge Page intercepts the visitor… by holding the request and evaluating the browser environment for automated signals, and serving a challenge. The visitor cannot reach their destination without passing the challenge." ([Interstitial Challenge Pages — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/), read 2026-08-12.)
Three variants, in increasing severity ([Interstitial Challenge Pages — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/), read 2026-08-12):
- **Non-Interactive**: Cloudflare judges automation from browser signals gathered by injected JS; the page needs no human interaction, typically < 5 s of JS processing.
- **Managed Challenge**: "Cloudflare dynamically chooses the appropriate type of challenge… based on the characteristics of a request from the signals indicated by their browser. Most human visitors are automatically verified and the Challenge Page will display **Successful**. However, if Cloudflare detects non-human attributes… they may be required to interact." Cloudflare's stated recommendation for WAF rules.
- **Interactive**: requires explicit human interaction (CAPTCHA-style). Cloudflare's "End the CAPTCHA era" position is that Managed Challenges should make this rare.
Cloudflare's own framing matches the repo's ~4 s real-Chrome solve: a normal browser passes with no interaction (Managed Challenge auto-verify or Non-Interactive JS processing).
### 3.3 Which product issues which challenge
Documented mapping ([How Challenges work — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/concepts/how-challenges-work/), read 2026-08-12):
| Trigger | Challenge type |
|---|---|
| WAF custom rules, **rate limiting rules**, IP Access rules | Interstitial Challenge Page |
| Bot Management | JavaScript Detections (invisible, per-request) |
| **Bot Fight Mode / Super Bot Fight Mode** | Interstitial Challenge Page |
| Under Attack Mode | Managed Challenge |
### 3.4 How a successful solve is remembered
Solving issues the **`cf_clearance`** cookie: "Clearance Cookie stores the proof of challenge passed. It is used to no longer issue a challenge if present. It is required to reach an origin server." ([Cloudflare Cookies — Cloudflare docs](https://developers.cloudflare.com/fundamentals/reference/policies-compliances/cloudflare-cookies/), read 2026-08-12.) "When that visitor tries to access other parts of your website, Cloudflare evaluates the cookie before presenting another challenge. If the cookie is still valid, no challenges will be shown." ([Challenge Passage — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/challenge-passage/), read 2026-08-12.)
`cf_clearance` is set with `SameSite=None; Secure; Partitioned`; because of the `Partitioned` (CHIPS) attribute, "a clearance obtained in one top-level context is not reused in a different top-level context" — so a clearance from kagane.to does not carry to comix.to even on the same browser. ([SameSite cookie interaction — Cloudflare docs](https://developers.cloudflare.com/waf/troubleshooting/samesite-cookie-interaction/), read 2026-08-12.)
---
## 4. `cf_clearance` — lifetime, binding, invalidation (the key question)
### 4.1 Lifetime
- **Default: 30 minutes.** "By default, the `cf_clearance` cookie has a lifetime of 30 minutes. Cloudflare recommends a setting between 15 and 45 minutes." The site owner can change it via the **Challenge Passage** setting. ([Challenge Passage — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/challenge-passage/), read 2026-08-12; also [SameSite cookie interaction — Cloudflare docs](https://developers.cloudflare.com/waf/troubleshooting/samesite-cookie-interaction/), read 2026-08-12.)
- Validation grace: "a few extra minutes are included to account for clock skew. For XmlHTTP requests, an extra hour is added to the validation time." ([Challenge Passage — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/challenge-passage/), read 2026-08-12.)
- "The Challenge Passage does not apply to rate limiting rules." Same source.
- `Not publicly documented`: whether the three target sites have changed the Challenge Passage from the 30-minute default, and whether there is any maximum value Cloudflare enforces.
**Consequence for cadence:** with the default 30-minute TTL, *any* poll cadence ≥ 1 hour finds the cookie expired and re-solves the challenge on every visit. A 1-hour and a 6-hour cadence therefore differ only in *how many times per day* the browser re-solves (~4× for the same series), not in whether a re-solve happens. This repo already measured the re-solve cost: ~4 s with real Chrome + real timezone. `[INFERENCE]` a site could raise the Challenge Passage to hours/days, which would make a 1-hour cadence *cheaper* (cookie still valid, no re-solve) — but that setting is unobservable and unlikely to be long on free manga sites.
### 4.2 What it is bound to — can clearance be replayed from another IP?
Documented statements, both from Cloudflare's own docs:
1. **Device/visitor binding:** "The cookie is securely tied to the specific visitor and device it was issued to, preventing reuse across machines." ([Clearance — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/concepts/clearance/), read 2026-08-12.)
2. **IP binding of the solve:** under Challenge limitations, Cloudflare lists "Client software where the solve request of a Managed Challenge comes from a different IP than the original IP a Challenge request was issued to. For example, if you receive the Challenge from one IP and solve it using another IP, the solve is not valid and you may encounter a Challenge loop." ([How Challenges work — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/concepts/how-challenges-work/), read 2026-08-12.)
3. **Top-level-site binding** via CHIPS partitioning: clearance is not reused across embedding contexts. ([SameSite cookie interaction — Cloudflare docs](https://developers.cloudflare.com/waf/troubleshooting/samesite-cookie-interaction/), read 2026-08-12.)
So the repo's belief is **verified by primary sources**: a `cf_clearance` obtained on one machine cannot be replayed from a different IP — the solve is IP-bound and the cookie is device-bound. `Not publicly documented`: whether the cookie value is also cryptographically bound to the User-Agent or TLS/JA3 fingerprint. The only UA-adjacent documented statement is the reverse direction: challenge *solving* breaks when a browser extension modifies the User-Agent or Canvas/WebGL APIs ("Cloudflare Challenges cannot support… Browser extensions that modify the browser's User-Agent value or Web APIs such as Canvas and WebGL") — i.e., tampering with browser signals is documented to *fail* challenges ([How Challenges work — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/concepts/how-challenges-work/), read 2026-08-12).
### 4.3 Two-tier clearance, and the behavior-bound invalidation (Precursor)
`cf_clearance` now carries two kinds of clearance ([Clearance — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/concepts/clearance/), read 2026-08-12):
- **Challenge clearance** — granted by solving a challenge; level-gated (Interactive > Managed > Non-Interactive; higher clears bypass lower challenges); "remains valid for the duration configured by the customer (Challenge Passage), **unless Precursor determines the session is suspicious**".
- **Precursor clearance** — "continuously updated based on session behavior"; an ongoing client-side process that periodically reassesses behavior. "If Precursor determines that a session is suspicious: the visitor's effective Challenge clearance may be **reduced or invalidated**; the visitor may be **re-challenged, even if the cookie has not expired**."
Precursor is documented as "client-side, session-based verification that continuously evaluates visitor behavior to identify automation… to detect automation that appears legitimate in individual requests but exhibits non-human patterns across a session", writing session state back into `cf_clearance`. It is a **zone-level toggle** (Security → Settings → Precursor; modes: Minimize Friction default, Maximize Security recommended), and "Precursor supersedes JavaScript Detections (JSD)". ([Precursor — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/precursor/), read 2026-08-12.)
**Implication:** the only documented mechanism by which *behavior over time* (as opposed to a single request's fingerprint) can revoke a valid clearance is Precursor — and it is opt-in per zone. `[INFERENCE]` It is unlikely to be enabled on free manga sites, but this is not observable from outside. With Precursor off (the default posture `[INFERENCE]`), a valid `cf_clearance` is honored purely on TTL + device/IP binding.
---
## 5. Does polling more often raise challenge risk?
### 5.1 Is per-IP request rate an input to challenge issuance? (documented answer: no such lever on non-Enterprise protection)
- **Bot scores** (the only per-request automated-detection output) are Enterprise-Bot-Management-only (§1.1); the ML engine's documented inputs are headers/session/browser signals, with request *pattern* entering only via `__cf_bm` — and no numeric rate is published (§1.2). `Not publicly documented`: any requests-per-hour value that changes a bot score or challenge probability.
- **BFM/SBFM** match "patterns of known bots" — signatures, not volumes ([Bot Fight Mode — Cloudflare docs](https://developers.cloudflare.com/bots/get-started/bot-fight-mode/), read 2026-08-12).
- **Rate limiting** is the product that polices volume, and it is opt-in per zone with explicit per-plan constraints (§5.2). Cloudflare's only documented link between "one valid clearance + high volume" is a *recommendation to site owners*: "Cloudflare recommends that customers add a rate limiting rule based on the `cf_clearance` cookie value. This helps ensure that a single, valid cookie cannot be abused by one machine to send an excessive volume of requests." ([Clearance — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/concepts/clearance/), read 2026-08-12.) Note: counting by cookie value is only available on Enterprise (see table below); a Free/Pro site cannot even build that rule.
### 5.2 What Rate Limiting Rules would do to 20–60 requests/hour
Rate limiting rules are opt-in; nothing runs them unless the site creates a rule. The documented per-plan capabilities ([Rate limiting rules — Cloudflare docs](https://developers.cloudflare.com/waf/rate-limiting-rules/), read 2026-08-12):
| Capability | Free | Pro | Business |
|---|---|---|---|
| Number of rules | **1** | 2 | 5 |
| Counting characteristics | **IP only** | IP only | IP, IP w/ NAT |
| Counting period | **10 s only** | ≤ 1 min | ≤ 10 min |
| Mitigation timeout | **10 s** | ≤ 1 h | ≤ 1 day |
| Fields in expression | **Path, Verified Bot** | + Host, URI, Full URI, Query | + Method, Source IP, User Agent |
Even the most aggressive free-plan rule (1 request per 10 s = 360/hour) is 6–18× above our 20–60/hour volume; the counting window is 10 s, so a per-hour burst is invisible to it. At 1 request per minute worst-case, **20–60 requests/hour from one residential IP cannot trip any rate limiting rule the Free plan can express.** Also documented: rate limiting is approximate, not precise — "there may be a delay of up to a few seconds between detecting a request and updating rate counters… excess requests could still reach the origin", and counters are per-data-center. ([Rate limiting rules — Cloudflare docs](https://developers.cloudflare.com/waf/rate-limiting-rules/), read 2026-08-12.) `[INFERENCE]` a site could also challenge on rate via WAF custom rules, but that requires Pro+ (custom rules are not on Free) and its own configuration.
### 5.3 DDoS protection (always on, all plans)
HTTP DDoS Attack Protection "is always enabled" and can only be tuned, not disabled. The only *published numeric* thresholds are error-rate-based: origin-error floods mitigate at the default "High" sensitivity of **1,000 errors per second** (Pro+ also requires 5× normal origin traffic). Per-IP volumetric thresholds are adaptive and `Not publicly documented` in the managed ruleset docs. 20–60 requests/hour is ~9 orders of magnitude below the published figure. ([HTTP DDoS Attack Protection — Cloudflare docs](https://developers.cloudflare.com/ddos-protection/managed-rulesets/http/), read 2026-08-12.)
### 5.4 Execution order (which product fires first)
Documented phase order: `ddos_l7` → custom rules → `http_ratelimit` (rate limiting) → managed rules → `http_request_sbfm` (SBFM); BFM runs outside this pipeline and cannot be skipped; a terminating action (block/challenge) stops later phases. ([Security features interoperability — Cloudflare docs](https://developers.cloudflare.com/waf/feature-interoperability/), read 2026-08-12.) Practical reading: on the three sites, the challenge we see could come from any of these stages; none of them documents a volume input at our scale (§5.1–5.3).
---
## 6. The documented legitimate side
### 6.1 Verified bots — the only "treated well" path, and it requires self-identification
Cloudflare documents a Verified bot as one meeting two bars ([Verified bots — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot/verified-bots/), read 2026-08-12):
1. **Honest self-identification** — "through a cryptographic Web Bot Auth signature, a published IP list with a stable user-agent, or reverse DNS".
2. **Non-abusive behavior** — "it obeys `robots.txt` and crawl directives, **maintains reasonable request rates**, and has not been observed evading website owner preferences or attacking sites".
Relevant verified-bot *behavior classes* exist for exactly this kind of client: "**Feed Fetching** — RSS readers, podcast aggregators, and news feed bots" and "**Monitoring & Operations** — Uptime monitoring, webhooks, and health checks". Becoming verified requires an application via the dashboard and validation via Web Bot Auth or IP validation; breach of the policy (e.g. "An AI Crawler that does not respect the crawl-delay directive") removes the bot from the allowlist. ([Verified bots — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot/verified-bots/), read 2026-08-12.)
"Historically, Verified bots have been excluded in default bot configurations across all plans" (same source) — i.e., verified bots are *default-allowed* under SBFM/Bot Management. **But** this path is the opposite of what a scraper wants: it requires the poller to publicly identify itself (stable, published IPs or cryptographic signatures) and to have its identity vetted by Cloudflare — and the *site* still decides via verified-bot policy whether to allow the category. There is **no documented mechanism for an anonymous low-volume automated client to be treated well.** `[INFERENCE]` a manga-site scraper would never qualify (it would be classified as Data Collection / scraping behavior, which is not a default-allowed class).
### 6.2 robots.txt and crawl control
- `robots.txt` **compliance is voluntary** — "The file expresses your preferences, but it does not prevent crawlers from accessing your content at a technical level." Enforcement requires Cloudflare's AI Crawl Control. ([robots.txt setting — Cloudflare docs](https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/), read 2026-08-12.)
- The managed `robots.txt` feature (all plans) is aimed at AI crawlers; it prepends `Disallow` rules for AI bots and a Content Signals Policy. It does not create any allowance for generic scrapers. (Same source.)
- RFC-side: the robots exclusion standard is an unauthenticated convention; nothing in it grants access rights. ([RFC 9309 "Robots Exclusion Protocol"](https://www.rfc-editor.org/rfc/rfc9309.html) — read 2026-08-12.) The standard defines crawl-delay etc. as voluntary directives; Cloudflare's docs are the operative statement for CF-protected sites.
### 6.3 The documented takeaway for "slowing down vs. fingerprint quality"
Cloudflare's own documentation repeatedly points at **browser/device signals** as the decision input on non-Enterprise protection (JSD detecting headless browsers; Managed Challenge choosing based on "signals indicated by their browser"; challenges failing when UA/Canvas/WebGL are modified — §3.2, §4.2), and at **identity** (verified bots) as the only legitimacy signal for automation (§6.1). Request *rate* appears only as: (a) an unnamed component of Enterprise-ML "session characteristics", (b) a voluntary verified-bot behavioral bar, and (c) the separate, opt-in, Free-plan-impotent Rate Limiting product (§5). Nothing documented says "slow down and you'll be challenged less" for a free-plan site — **fingerprint quality is the lever that the documentation actually describes**, which matches this repo's measurements (real Chrome + real timezone passes; every headless variant fails regardless of rate).
---
## 7. Implications for cadence design
### Documented facts (with sources above)
1. **1-hour vs 6-hour cadence is not a documented risk lever.** Challenge issuance on Free/Pro-grade protection is signature-based; the rate-aware scoring only exists on Enterprise Bot Management; free-plan rate limiting cannot express a limit our volume could trip (§1, §2, §5).
2. **Every poll ≥ 1 hour re-solves the challenge anyway.** `cf_clearance` defaults to 30 minutes; Challenge Passage is site-configurable and unobservable. The re-solve cost is what this repo measured (~4 s, real Chrome + real timezone) (§4.1, repo measurements).
3. **The documented failure modes are fingerprint, not rate:** headless browsers (JSD), cloud-hosting IPs (BFM heuristics), modified UA/Canvas/WebGL (challenge solve failure) (§2.2, §3, §4.2).
4. **Clearance is not portable:** device-bound + solve-IP-bound + CHIPS-partitioned; replaying a cookie from another IP is documented invalid (§4.2).
5. **Anonymity has no documented "good citizen" path:** the only legitimate-automation route (verified bots) requires self-identification and site-side allowance (§6).
6. **The one behavior-bound revocation mechanism (Precursor) is opt-in per zone**, not a default documented behavior (§4.3).
### Inferences (not documented)
- `[INFERENCE]` The three sites run Free/Pro-grade protection (BFM, SBFM, or WAF challenge rules), not Enterprise Bot Management; therefore no continuous per-request bot score exists for our traffic.
- `[INFERENCE]` The sites have not changed Challenge Passage to hours/days (free manga sites default to the 30-minute default); if they had, hourly polling would get *cheaper* (valid cookie, no re-solve).
- `[INFERENCE]` Precursor is not enabled on these sites; if it were, hourly re-visits from an automated Chrome could accumulate session-behavior signals and trigger re-challenge even with a valid cookie — the only documented scenario in which polling *frequency* (via session behavior) could matter.
- `[INFERENCE]` 6-hour cooldowns buy nothing documented beyond raw request-count reduction (fewer challenge solves per day, less origin load); the risk profile at 1 request/hour/series is not documented to differ from 6 request/hour/series.
- `[INFERENCE]` If the owner wants belt-and-braces, the engineering levers that match the documentation are: keep the real-Chrome fingerprint (no UA spoofing, no headless-shell, real timezone — already done), keep a persistent user-data profile so `cf_clearance`/`__cf_bm` persist across visits, and treat any change of exit IP (e.g. home connection rebooting to a new IP) as a guaranteed re-solve, since clearance does not travel with the IP.
### Bottom line
Moving browser-backed sites from 6-hour to 1-hour per-series cooldown is **not contradicted by any documented Cloudflare mechanism** at 20–60 requests/hour from one residential IP through one real Chrome. The documented risk is carried by fingerprint quality (already solved in this repo) and by unobservable site configuration (Challenge Passage, Precursor, possible custom WAF rules). The residual, non-documented risk is that these sites sit behind Cloudflare's *proprietary* detection, and Cloudflare publishes neither its per-IP thresholds nor the ML feature set — so "no documented lever" is not "no lever".
@@ -0,0 +1,572 @@
# lightnovelworld.net — chapter slug vs. series slug
Research note for Gitea issue #77. All pages fetched live on **2026-08-11** with
plain `curl` and a desktop UA. **No Cloudflare challenge was encountered on any
request** — every fetch below returned real HTML on the first try, so the
Playwright fallback was never needed.
Every claim carries the URL it came from. Nothing here is inferred from the
existing code; where a claim is an interpretation rather than an observation it
is marked `[INFERENCE]`.
---
## 1. Summary answer table
| Question | Answer | Evidence |
|---|---|---|
| Does the chapter slug equal the series slug? | **No, not reliably.** 3 of 41 sampled novels diverge at their latest chapter (~7%); a 4th diverges historically. | §4 |
| Best chapter-page → series-URL pointer | `a[aria-label='All Chapter']` and the microdata breadcrumb `a[itemprop=item]` at `position 2`. Both carry the full absolute series URL. | §3 |
| Is JSON-LD `BreadcrumbList` usable? | **No.** The Rank Math JSON-LD breadcrumb on a chapter page contains only *Home* and *the chapter itself* — the series is absent. | §3.1 |
| Is `<link rel=canonical>` usable? | **No.** It points at the chapter page itself. | §3.2 |
| Are `og:` tags usable? | **No.** `og:url` is the chapter page; no `og:` tag names the series. | §3.3 |
| Does the series page list chapters in the initial HTML? | **Yes — all of them.** 1232 unique chapters (1…1232) present in one document. | §5 |
| Is there a paginated chapter endpoint? | **No.** `/novel/<slug>/chapters/` and `/novel/<slug>/chapters/page-2` both **404**. | §5 |
| Does the series page carry *other* novels' chapter anchors? | **No.** Across 30 series pages, every `-chapter-N/` URL belonged to the page's own novel. | §6 |
| Unscoped regex `lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/` on a series page | **SAFE** | §6 |
| Does `/novel/<chapter-slug>/` 404? | **Yes**, 404, no redirect. | §7 |
| Is there a chapter-slug → series-slug endpoint that avoids parsing a chapter page? | **No endpoint found.** But `/<chapter-slug>/` (no `-chapter-N`) **302s to chapter 1**, which still requires parsing a chapter page. | §7 |
| Is there an "Alternative names" field explaining divergence? | **No such field exists** on lightnovelworld series pages. | §4.3 |
---
## 2. The two reference pages
| URL | Status |
|---|---|
| `https://lightnovelworld.net/my-longevity-simulation-chapter-1/` | **200**, 0 redirects |
| `https://lightnovelworld.net/novel/immortality-simulator/` | **200**, 0 redirects |
| `https://lightnovelworld.net/novel/my-longevity-simulation/` | **404**, 0 redirects |
Confirmed: chapter slug `my-longevity-simulation`, series slug
`immortality-simulator`, and the naive `/novel/<chapter-slug>/` construction is
a hard 404.
---
## 3. Chapter page → series URL: every in-page pointer, in priority order
All snippets below are verbatim from
`https://lightnovelworld.net/my-longevity-simulation-chapter-1/`.
### Priority 1 — `a[aria-label='All Chapter']` ✅ WORKS
Single occurrence in the document, inside the chapter navigation bar:
```html
<div class="nvs nvsc"><a href='https://lightnovelworld.net/novel/immortality-simulator/' aria-label='All Chapter'><i class="fas fa-list-ul"></i> All Chapter</a></div>
```
Note the **single quotes** on both attributes — a regex written for `href="` will
miss it. This is the most narrowly-targeted pointer: exactly one element on the
page has `aria-label='All Chapter'`.
### Priority 2 — microdata `BreadcrumbList` / `a[itemprop=item]` ✅ WORKS
At line 363–379 of the served HTML. `position 2` is the series:
```html
<div itemscope="" itemtype="http://schema.org/BreadcrumbList">
<span itemprop="itemListElement" itemscope="" itemtype="http://schema.org/ListItem">
<a itemprop="item" href="https://lightnovelworld.net/"><span itemprop="name">Home</span></a>
<meta itemprop="position" content="1">
</span>
›
<span itemprop="itemListElement" itemscope="" itemtype="http://schema.org/ListItem">
<a itemprop="item" href="https://lightnovelworld.net/novel/immortality-simulator/"><span itemprop="name">Immortality Simulator</span></a>
<meta itemprop="position" content="2">
</span>
›
<span itemprop="itemListElement" itemscope="" itemtype="http://schema.org/ListItem">
<a itemprop="item" href="https://lightnovelworld.net/my-longevity-simulation-chapter-1/"><span itemprop="name">My Longevity Simulation Chapter 1</span></a>
<meta itemprop="position" content="3">
</span>
</div>
```
Bonus: `<span itemprop="name">Immortality Simulator</span>` gives the clean
series **title** as well, which the userscript currently derives by stripping
`Chapter <n>` off `h1.entry-title`.
Caveat: `itemprop="item"` is **not** unique to the breadcrumb — the site header
nav also uses `itemprop="url"` on anchors, and unrelated `itemprop="item"` usage
could appear. Scope any selector to the enclosing
`[itemtype="http://schema.org/BreadcrumbList"]`.
### Priority 3 — corroborating `/novel/` anchors ⚠️ AMBIGUOUS
The chapter page contains exactly **8** distinct `/novel/…` hrefs. Two are
navigational (`/novel/` index), one is the series, and **five are a
"recommended" strip of unrelated novels**:
```
href="/novel/"
href="https://lightnovelworld.net/novel/"
href="https://lightnovelworld.net/novel/dreadful-radio-game/"
href="https://lightnovelworld.net/novel/evil-god-average/"
href="https://lightnovelworld.net/novel/immortality-simulator/"
href="https://lightnovelworld.net/novel/omniscient-first-persons-viewpoint/"
href="https://lightnovelworld.net/novel/reverend-insanity/"
href="https://lightnovelworld.net/novel/surviving-in-a-romance-fantasy-novel/"
```
So **"grab the first `/novel/<slug>/` link on a chapter page" is unsafe** — the
recommendation strip contaminates it. Use Priority 1 or 2.
### ❌ NOT usable — JSON-LD `BreadcrumbList` (§3.1)
There is exactly one `application/ld+json` block on the page (Rank Math SEO).
Verbatim:
```html
<script type="application/ld+json" class="rank-math-schema">{"@context":"https://schema.org","@graph":[{"@type":"BreadcrumbList","@id":"https://lightnovelworld.net/my-longevity-simulation-chapter-1/#breadcrumb","itemListElement":[{"@type":"ListItem","position":"1","item":{"@id":"https://lightnovelworld.net","name":"Home"}},{"@type":"ListItem","position":"2","item":{"@id":"https://lightnovelworld.net/my-longevity-simulation-chapter-1/","name":"My Longevity Simulation Chapter 1"}}]}]}</script>
```
Only two items — Home and the chapter. **The series does not appear.** The
JSON-LD breadcrumb and the visible microdata breadcrumb disagree; the microdata
one is the richer of the two. Do not use the JSON-LD.
### ❌ NOT usable — `<link rel=canonical>` (§3.2)
```html
<link rel="canonical" href="https://lightnovelworld.net/my-longevity-simulation-chapter-1/" />
```
Self-referential.
### ❌ NOT usable — `og:` meta tags (§3.3)
```html
<meta property="og:type" content="article" />
<meta property="og:title" content="Read My Longevity Simulation Chapter 1 Online Free - LightNovelworld" />
<meta property="og:url" content="https://lightnovelworld.net/my-longevity-simulation-chapter-1/" />
<meta property="og:site_name" content="Light Novel World" />
```
`og:url` is the chapter. `og:title` carries the **chapter-slug-derived** title
("My Longevity Simulation"), *not* the series title ("Immortality Simulator") —
so og tags actively reinforce the wrong name.
### Also present — `rel=next` / `rel=prev` chapter navigation
```html
<a aria-label="next" href="https://lightnovelworld.net/my-longevity-simulation-chapter-2/" rel="next">Next <i class="fa fa-angle-right" aria-hidden="true"></i></a>
```
Chapter 1 has no `rel=prev`. Useful for walking chapters, useless for finding
the series.
---
## 4. The reverse direction, and how common divergence is
### 4.1 Sample method
Two independent samples, deduplicated:
- 14 novels taken from the latest-updates listing on `https://lightnovelworld.net/`
(distinct chapter slugs), each chapter page fetched and its
`aria-label='All Chapter'` href read for the series slug.
- 30 series URLs taken from `https://lightnovelworld.net/novel/`, each series
page fetched and its chapter anchors' slug prefix extracted. Two of the 30
(`/novel/feed/` and `/novel/list-mode/`) are not novels and are excluded,
leaving 28.
One novel (`investing-in-my-crippled-wife-every-return-makes-me-stronger`)
appears in both samples. **Distinct novels sampled: 41.**
### 4.2 Divergence results
| Series slug (`/novel/<…>/`) | Chapter slug (`/<…>-chapter-N/`) | Result |
|---|---|---|
| `immortality-simulator` | `my-longevity-simulation` | **DIVERGE** |
| `the-sword-illuminates-the-great-wilderness` | `radiant-blade-of-the-wilderness` | **DIVERGE** |
| `a-villains-will-to-survive` | `the-villain-wants-to-live` | **DIVERGE** |
| `all-jobs-and-classes-i-just-wanted-one-skill-not-them-all` | `all-jobs-and-classes-i-just-wanted-one-skill-not` (ch. 1–99) **and** `all-jobs-and-classes-i-just-wanted-one-skill-not-them-all` (ch. 100–423) | **SPLIT** (latest matches) |
| `86-eighty-six` | same | match |
| `a-dungeon-beneath-my-house-let-me-gain-800-exp` | same | match |
| `a-journey-of-black-and-red` | same | match |
| `a-knight-who-eternally-regresses` | same | match |
| `a-regressors-tale-of-cultivation` | same | match |
| `a-will-eternal` | same | match |
| `absolute-resonance` | same | match |
| `absolute-sword-sense` | same | match |
| `advent-of-the-three-calamities` | same | match |
| `after-changing-to-the-ruthless-way-the-brothers-cried-and-begged-for-forgiveness` | same | match |
| `against-the-gods` | same | match |
| `apocalypse-i-built-the-infinite-train` | same | match |
| `arcane-exfil` | same | match |
| `as-a-mafia-boss-i-refuse-to-be-an-extra` | same | match |
| `ascendance-of-a-bookworm` | same | match |
| `avatar-conquering-the-elements` | same | match |
| `battle-world-ascending-without-limits` | same | match |
| `became-the-patron-of-villains` | same | match |
| `ending-maker` | same | match |
| `got-dropped-into-a-ghost-story-still-gotta-work` | same | match |
| `greed-all-for-what` | same | match |
| `investing-in-my-crippled-wife-every-return-makes-me-stronger` | same | match |
| `lord-of-mysteries-2-circle-of-inevitability` | same | match |
| `lord-of-the-mysteries` | same | match |
| `my-vampire-system` | same | match |
| `cleaver-of-sin` | same | match |
| `greatest-legacy-of-the-magus-universe` | same | match |
| `magus-infinite` | same | match |
| `path-of-the-extra` | same | match |
| `regnum-aetern-dual-rebirth` | same | match |
| `shadow-slave` | same | match |
| `slime-evolution` | same | match |
| `sss-awakening-i-can-class-change-at-will` | same | match |
| `the-dao-of-reincarnation-ascending-through-a-hundred-lives` | same | match |
| `the-gamers-pov` | same | match |
| `the-insane-regressor-throne-of-pride` | same | match |
| `the-villains-pov` | same | match |
**Counts (41 distinct novels):**
- **37 match** (90.2%)
- **3 diverge at the latest chapter** (7.3%) — these are the ones that break the poller *today*
- **1 splits mid-series** but its newest chapters happen to match the series slug (2.4%)
`[INFERENCE]` The true site-wide divergence rate is probably in the same
ballpark, but 41 novels out of a catalogue of thousands is a small sample and
the listing page is alphabetically front-loaded. Treat ~7% as an order of
magnitude, not a precise figure.
### 4.3 Is there a derivable rule? **No.**
The divergence is **not directional**, so you cannot compute one slug from the
other:
- `immortality-simulator` — the *series* carries the polished English title
(`<h1 class="entry-title" itemprop="name">` is absent from the chapter page's
own naming; `og:title` says "My Longevity Simulation"), while the *chapters*
carry the literal translation.
- `the-sword-illuminates-the-great-wilderness` — the reverse: the *series* has
the literal title
(`<h1 class="entry-title" itemprop="name">The Sword Illuminates the Great Wilderness</h1>`,
`<title>Read The Sword Illuminates the Great Wilderness English Online Free - LightNovelworld</title>`)
while the *chapters* carry the polished one (`radiant-blade-of-the-wilderness`).
- `a-villains-will-to-survive` —
`<h1 class="entry-title" itemprop="name">A Villain’s Will to Survive</h1>`,
chapters at `the-villain-wants-to-live-chapter-N/`.
`[INFERENCE]` The consistent explanation is that a novel is retitled after
publication; WordPress updates the series post's slug but leaves the already-published
chapter posts' slugs alone. The direction of the retitle varies per novel, which
is why no rule exists. This is consistent with the split case in §4.4, but the
site exposes no field that states it.
**There is no "Alternative names" / "Associated names" field.** Scanning
`https://lightnovelworld.net/novel/a-villains-will-to-survive/` and
`https://lightnovelworld.net/novel/immortality-simulator/` for such a label
returns nothing; the series info panel exposes only **Author**, **Released**,
**Type**, **Status**. (The only `names?` match in the document is the `>Name*<`
label on the comment form.) So the old title is **not** recoverable from the
series page — the mapping only exists in the chapter anchors themselves.
### 4.4 The split case — a slug can change *mid-series*
`https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/`
carries **two** chapter slugs in one page:
| Chapter slug prefix | Anchors | Chapter range |
|---|---|---|
| `all-jobs-and-classes-i-just-wanted-one-skill-not` | 100 | 1 – 99 |
| `all-jobs-and-classes-i-just-wanted-one-skill-not-them-all` | 325 | 100 – 423 |
Both resolve, and **both point back at the same series**:
```
200 https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-chapter-1/
<a href='https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/' aria-label='All Chapter'>
200 https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-404/
<a href='https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/' aria-label='All Chapter'>
```
**Consequence:** a single stored chapter slug is not a sufficient key even for
one novel. A fix that captures "the chapter slug at bookmark time" and pins the
poller regex to it would still miss chapters published under a later slug. This
is the strongest argument for the unscoped regex over a stored-slug regex.
---
## 5. Series page → chapter list
From `https://lightnovelworld.net/novel/immortality-simulator/` (200, 655 037 bytes):
- **1234** anchors matching `href="https://lightnovelworld.net/<slug>-chapter-<n>/"`.
- **1232 unique** chapter URLs, covering chapter numbers **1 through 1232 with no gaps**.
- The 2 extras are the First/Last shortcut widget, which duplicates two chapters:
```html
<h2>Read Immortality Simulator</h2></div>
<div class="lastend">
<div class="inepcx">
<a href="https://lightnovelworld.net/my-longevity-simulation-chapter-1/">
<span>First Chapter</span>
<span class="epcur epcurfirst">Vol. 1 Ch. 1</span>
</a>
</div>
<div class="inepcx">
<a href="https://lightnovelworld.net/my-longevity-simulation-chapter-1232/">
```
**The entire chapter list is in the initial HTML.** No JS hydration, no
pagination, no separate endpoint. Confirmed by probing the shapes the issue
speculated about:
| URL | Status |
|---|---|
| `https://lightnovelworld.net/novel/immortality-simulator/chapters/` | **404** |
| `https://lightnovelworld.net/novel/immortality-simulator/chapters/page-2` | **404** |
The only `page-numbers` / pagination markup in the document belongs to
`class="wpdiscuz-comment-pagination"` (the comment widget), not the chapter list.
This holds for large novels too: `greed-all-for-what` served 2666 chapter
anchors and `my-vampire-system` 2547, all inline in one response.
**All chapter hrefs are absolute** (`https://lightnovelworld.net/…`). A grep for
relative `href="/<slug>-chapter-N/"` on the series page returned **0** matches,
so a pattern anchored on the literal host is safe.
---
## 6. Is the unscoped chapter regex SAFE on a series page?
### Verdict: **SAFE**
Proposed pattern: `lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/`
**Proof.** For each of the 30 series pages fetched from
`https://lightnovelworld.net/novel/`, every occurrence of `<slug>-chapter-<n>/`
anywhere in the document (href or not) was reduced to its slug prefix and
deduplicated:
- **27 of 28 real series pages had exactly 1 distinct chapter-slug prefix**, equal
to that novel's own chapter slug.
- **1 page had 2 prefixes** — `all-jobs-and-classes-…`, and §4.4 proves *both
belong to that same novel*. Not foreign.
- **0 pages carried any other novel's chapter URL.**
The recommendation and sidebar widgets on a series page link to **series** URLs
only, never chapter URLs. On
`https://lightnovelworld.net/novel/immortality-simulator/` the recommendation
strip renders as:
```html
<a href="https://lightnovelworld.net/novel/reverend-insanity/" itemprop="url" title="Reverend Insanity" class="tip" rel="292">
```
— `/novel/<slug>/`, which the pattern cannot match.
The one widget that *would* carry foreign chapter links, "Latest Reading", is an
**empty client-side template**. Its container is `display:none` and its `<ul>` is
empty in the served HTML; the row markup lives in an inert
`<span id="series-history-tpl" style='display:none'>` with mustache placeholders
and a `href="#/{{number}}"` — no real chapter URL:
```html
<div class="bixbox bxcl" id="series-history" style="display:none;">
<div class="releases"><h2>Latest Reading</h2></div>
<div class="series-history-pool">
<ul class="clstyle" id="series-history-ul"></ul>
</div>
</div>
<span id="series-history-tpl" style='display:none'>
<li data-id="{{id}}" data-num="{{chapter}}" data-vol="{{volume}}" data-title="{{title}}">
<div class="chbox"><div class="eph-num">
<a onclick="…return series_history.redirect({{id}});" href="#/{{number}}">
```
It is populated from the visitor's own local history, so a **server-side fetch
never sees content there** — the poller is immune. A browser-rendered fetch with
a fresh profile is likewise immune (no history to render).
### Caveats to record with the verdict
1. **`[INFERENCE]` This is an empirical guarantee, not a structural one.** Unlike
the asura/novelfull cases where scoping to the slug makes foreign contamination
*impossible*, here we only know that lightnovelworld's series template does not
currently emit foreign chapter anchors. If the theme ever adds a "latest site
updates" strip rendered server-side, the unscoped pattern breaks silently and
in the worst direction (a foreign chapter number *higher* than the real one
wins the maximum and the bookmark shows a phantom update).
**Correction, 2026-08-11.** This caveat understated the risk. A series page
server-renders a wpdiscuz comment thread below the chapter list, and comment
bodies are HTML that can carry an `<a href>`. Verified on
`/novel/the-sword-illuminates-the-great-wilderness/`: `#wpdcom` present,
comment `#wpd-comm-358_0` by "hasbi asy" rendered as `<p>where's everyone</p>`
inside `.wpd-comment-text`, corroborated by that page's comment RSS feed;
`firstLoadWithAjax: 0` confirms the thread is in the initial HTML. The sampled
pages carried 0 or 1 comments and no links, which is why §6 read clean. So the
unscoped pattern does not merely depend on the *theme* staying unchanged — it
reads a region any visitor can write to, and the scanner takes the maximum with
no upper bound. The scan must stop before the comment thread.
2. ~~Consider scoping the match to the chapter-list container rather than the whole
document, which would restore the structural guarantee at low cost. The list
items sit under `ul.clstyle` (`class="bixbox bxcl"`).~~
**Refuted, 2026-08-11**, measured on 4 series pages
(`the-sword-illuminates-the-great-wilderness`, `all-jobs-…-them-all`,
`a-will-eternal`, `immortality-simulator`). The single `ul.clstyle` per page is
the hidden, empty "Latest Reading" template (`#series-history-ul`, see §5), not
the chapter list — scoping to it would match nothing. The real chapter list is a
classless `<ul>` inside `div.eplister.eplisterfull`, in
`div.bixbox.bxcl.epcheck`.
The workable boundary is `wpd-threads`: it occurs exactly once per page, and on
all 4 pages every chapter anchor precedes it and no chapter-shaped href follows
it. `wpdcom` and `wpdiscuz` are unusable as markers — they occur 111 to 143
times per page, including in `<head>` roughly 15 KB *before* the chapter list,
so cutting there would discard the list itself. `id='comments'` also occurs once
but is single-quoted; the canonical `id="comments"` never appears.
3. `[0-9.]+` is fine: no decimal chapter numbers appeared anywhere in the sampled
pages, but the existing `strings.Trim(m[1], ".")` guard already handles them.
4. The current test fixture `lnwSeriesFixture` in
`backend/internal/latest/sites_test.go` includes
`<a href="https://lightnovelworld.net/overgeared-chapter-9999/">` and asserts it
is ignored. **That anchor is not representative of a real series page** — no
sampled page contained a foreign chapter anchor. Relaxing the regex will make
that assertion fail, and the correct response is to fix the fixture, not to
keep the scoping.
---
## 7. Redirects and reverse-lookup endpoints
| URL | Status | Redirects | Final |
|---|---|---|---|
| `https://lightnovelworld.net/novel/my-longevity-simulation/` | **404** | 0 | — |
| `https://lightnovelworld.net/novel/radiant-blade-of-the-wilderness/` | **404** | 0 | — |
| `https://lightnovelworld.net/immortality-simulator-chapter-1/` | **404** | 0 | — |
| `https://lightnovelworld.net/my-longevity-simulation/` | **200** | **1** | `https://lightnovelworld.net/my-longevity-simulation-chapter-1/` |
Two findings:
1. **`/novel/<chapter-slug>/` genuinely 404s.** Confirmed for both
`my-longevity-simulation` and `radiant-blade-of-the-wilderness`. There is no
alias, no 301, no fallback. Any stored series URL built by that construction is
permanently dead.
2. **The chapter slug alone (no `-chapter-N`) 302s to chapter 1** of that novel.
This is the closest thing to a reverse-lookup endpoint, but it lands on a
*chapter page*, so recovering the series URL still requires parsing that page's
`aria-label='All Chapter'` anchor or breadcrumb. **No endpoint maps chapter slug
→ series slug directly.**
3. The inverse construction also fails: `/<series-slug>-chapter-1/` 404s for a
divergent novel, so you cannot probe your way from a series slug to a chapter
slug either. The chapter slug must be read off the series page's anchors.
Also present but not a lookup path: the site is WordPress and exposes
`https://lightnovelworld.net/wp-json/wp/v2/posts/177695` (advertised via
`<link rel="alternate" title="JSON" type="application/json" …>` on the chapter
page). Whether the REST API exposes a chapter→series relation was **not tested**
— out of scope for this note, and it would still require fetching the chapter
page to learn the post ID.
---
## 8. Implications for issue #77
> The issue text itself could not be read: `gh` is not installed in this
> environment, so `issue://77` failed to resolve. The proposed regex quoted below
> is the one supplied in the task brief.
### 8.1 The bug
`backend/internal/latest/sites.go:128-137` derives the chapter-anchor pattern
from the **series slug**:
```go
m := lnwSlugRe.FindStringSubmatch(u.Path) // /novel/<series-slug>/
re = regexp.MustCompile(`lightnovelworld\.net/` + regexp.QuoteMeta(m[1]) + `-chapter-([0-9.]+)/`)
```
For `immortality-simulator` this compiles to a pattern matching
`…net/immortality-simulator-chapter-N/`, which appears **zero times** on the
series page — the anchors all say `my-longevity-simulation`. `latestChapterFrom`
returns `ok == false`, indistinguishable from a Cloudflare challenge or a site
redesign, and the bookmark silently stops tracking updates. Same for
`the-sword-illuminates-the-great-wilderness` and `a-villains-will-to-survive`.
### 8.2 The proposed fix is sound
Dropping the scoping to
`lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/` is sound in shape and fixes
all three currently-broken novels plus the split case in §4.4 — which no
stored-slug approach can fix, since that novel legitimately has two chapter
slugs. It must not be applied to the whole document, however: truncate the body
at `wpd-threads` first, per the corrected caveat 2 in §6. The `ul.clstyle`
suggestion originally recorded here is refuted.
Also worth removing: the `overgeared-chapter-9999` line in `lnwSeriesFixture`
and the `"lightnovelworld takes the max and ignores another series"` test case,
which encode a contamination scenario that §6 shows does not occur on this site.
Replacing them with a fixture that mirrors §4.4's two-slugs-one-novel reality
would be a better regression test.
### 8.3 The userscript has the same bug, and it is worse
`userscript/novel-bookmark.user.js` builds the series URL by string-concatenating
the chapter slug (lines 154 and 169):
```js
seriesUrl: "https://lightnovelworld.net/novel/" + m[1] + "/",
```
For a divergent novel this **writes a permanently-404 series URL into the
database at bookmark time**. Fixing only the backend regex leaves those rows
broken, and `fetchableSeriesURL` will happily accept the dead URL because it only
checks the hostname (`sites.go:321-324`).
The userscript should read the series URL off the page instead of constructing
it. On a chapter page, prefer in this order (§3):
```js
document.querySelector("a[aria-label='All Chapter']")?.href
?? document.querySelector("[itemtype$='BreadcrumbList'] a[itemprop=item][href*='/novel/']")?.href
```
The breadcrumb's `<span itemprop="name">` also supplies the correct series title
directly, replacing the current `h1.entry-title` + strip-`Chapter <n>` heuristic
— which today yields "My Longevity Simulation" where the series is actually
titled "Immortality Simulator".
A backfill for already-stored rows is likely needed: any `lightnovelworld` row
whose `series_url` 404s can be repaired by fetching
`https://lightnovelworld.net/<series_id>-chapter-1/` (or `/<series_id>/`, which
302s there per §7) and reading its `All Chapter` anchor.
### 8.4 Note on `series_id`
Rows are keyed `lightnovelworld:<series_id>` where `series_id` is currently the
**chapter slug** (from the userscript's `m[1]`). If the fix changes `series_id` to
the real series slug, existing divergent bookmarks change key and need migrating.
Whether that is worth it depends on whether `series_id` is used to rebuild URLs
anywhere — `sites.go` uses the stored `series_url`, not `series_id`, for
lightnovelworld, so storing the correct `series_url` may be sufficient without a
re-key.
---
## Reproduction
```sh
UA='Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0 Safari/537.36'
# §2 status codes
curl -sS -o /dev/null -w '%{http_code} %{url_effective}\n' -L -A "$UA" \
https://lightnovelworld.net/my-longevity-simulation-chapter-1/ \
https://lightnovelworld.net/novel/immortality-simulator/ \
https://lightnovelworld.net/novel/my-longevity-simulation/
# §3 the two working pointers
curl -sS -A "$UA" https://lightnovelworld.net/my-longevity-simulation-chapter-1/ \
| grep -E "aria-label='All Chapter'|itemprop=\"item\""
# §6 SAFE proof — distinct chapter-slug prefixes per series page (expect 1)
curl -sS -A "$UA" https://lightnovelworld.net/novel/immortality-simulator/ \
| grep -oE '[a-z0-9-]+-chapter-[0-9.]+/' | sed -E 's/-chapter-[0-9.]+\///' | sort | uniq -c
```
+221
View File
@@ -0,0 +1,221 @@
{
"0": "HTMX Library Internals",
"1": "Cover Fetch Test Helpers",
"2": "Manga Userscript Adapters",
"3": "Novel Userscript Adapters",
"4": "Series Acquisition Tests",
"5": "Bookmarks API Tests",
"6": "Storage Choice ADR",
"7": "Cover & Acquire Internals",
"8": "System Architecture Concepts",
"9": "Session Middleware",
"10": "Go Test Helpers",
"11": "Store Tests",
"12": "Bookmarks API Handler",
"13": "Web UI Handlers",
"14": "Go Error Handling",
"15": "CDP Browser Client",
"16": "Cloudflare bot scoring and poll cadence — what is actually documented",
"17": "Go Code Style Guide",
"18": "Agent Skills",
"19": "Store",
"20": "I/O Performance Patterns",
"21": "CPU Optimization",
"22": "Caching Patterns",
"23": "Browser Entrypoint",
"24": "Memory Allocation & GC",
"25": "Cover Fetcher Tests",
"26": "Open",
"27": "Find Skills Guide",
"28": "Allocation Patterns",
"29": "Observability & Alerting",
"30": "AGENTS.md",
"31": "Store",
"32": "Open",
"33": "Go Testing Guide",
"34": "Session Store",
"35": "Web UI Filter Logic",
"36": "Userscript Test Harness",
"37": "Product & Security Context",
"38": "novel-logic.test.js",
"39": "UI Critique 2026-07-26A",
"40": "UI Critique 2026-07-26B",
"42": "pgtest.go",
"43": "Issue Tracker & Triage",
"44": "Ticket Workflow",
"45": "Go Perf Alert Rules",
"46": "Userscript Display Logic",
"47": "Go Perf Skill Docs",
"48": "Login Page Art",
"49": "BookmarkManager Logo",
"50": "Skills CLI",
"51": "Skills Leaderboard",
"52": "Complex Condition Extraction",
"53": "Sentinel Errors",
"54": "errors.As Patterns",
"55": "errors.Is Patterns",
"56": "errors.Join Patterns",
"57": "Error Wrapping",
"58": "Single Error Handling",
"59": "SIMD Optimizations",
"60": "GOGC Tuning",
"61": "GOMEMLIMIT",
"62": "Bottleneck Decision Tree",
"63": "pprof Profiling",
"64": "Test Timeout Helper",
"65": "httptest Patterns",
"66": "testify Suite Pattern",
"67": "go:embed Fixtures",
"68": "clockwork Time Mocking",
"69": "testify Mocking",
"70": "t.ArtifactDir Helper",
"71": "Subtests Pitfall",
"72": "golang-benchmark Skill",
"73": "golang-concurrency Skill",
"74": "golang-ci Skill",
"75": "golang-database Skill",
"76": "golang-lint Skill",
"77": "testify Skill",
"78": "Build Tag Integration Tests",
"79": "Test Naming Convention",
"80": "UI Critique A Finding",
"81": "UI Critique B Finding",
"82": "P0 Overflow Bug",
"83": "P1 hx-indicator Gap",
"84": "golang-benchmark Skill (ext)",
"85": "golang-concurrency Skill (ext)",
"86": "golang-ci Skill (ext)",
"87": "golang-data-structures Skill (ext)",
"88": "golang-database Skill (ext)",
"89": "golang-design-patterns Skill (ext)",
"90": "golang-documentation Skill (ext)",
"91": "golang-gopls Skill (ext)",
"92": "golang-lint Skill (ext)",
"93": "golang-naming Skill (ext)",
"94": "golang-observability Skill (ext)",
"95": "golang-refactoring Skill (ext)",
"96": "golang-safety Skill (ext)",
"97": "golang-samber-oops Skill (ext)",
"98": "golang-samber-slog Skill (ext)",
"99": "golang-structs-interfaces Skill (ext)",
"100": "golang-troubleshooting Skill (ext)",
"101": "promql-cli Skill",
"102": "Backend Module",
"103": "bookmark-api Service",
"104": "AGENTS.md",
"105": "reviewer.md",
"106": "Redeploy runbook",
"107": "1. Backend",
"108": "Deployment",
"109": "Cinder — BookmarkManager design system",
"110": "Implement tickets",
"111": "SQLite → Postgres cutover runbook",
"112": "Testing the userscript",
"113": "ADR-0007: The backend hosts every Site's Cover bytes",
"114": "Issue tracker: Gitea (`tea` CLI)",
"115": "ADR-0006: The browser runs on the home machine, over the tailnet",
"116": "ADR-0008: A Series identity is discovered from the Site's links, never derived from an address",
"117": "Domain Docs",
"118": "ticket-implementer.md",
"119": "implementer.md",
"120": "Series is a shared entity, and only the Poll may update it",
"121": "Postgres replaces SQLite as the primary datastore",
"122": "Identity comes from Discord OAuth; we store no passwords and send no email",
"123": "The wire format stays flat and deliberately does not mirror the schema",
"124": "ADR-0005: On-demand browser sidecar",
"125": "sessions_test.go",
"126": "Bookmark Manager",
"127": "triage-labels.md",
"128": "Cross-Ticket Contract",
"129": "Implement Tickets Skill",
"130": "Orchestrator Role",
"131": "resolving-merge-conflicts Skill",
"132": "tdd Skill",
"133": "Ticket Wave Batching",
"134": "Four-Object Browser Stub",
"135": "Module Export Hook",
"136": "logic.test.js Test Harness",
"137": "manga-bookmark.user.js",
"138": "stripBuildHash",
"139": "Testing the Userscript Skill",
"140": "cr-spec Agent",
"141": "cr-standards Agent",
"142": "Escalate Rather Than Guess",
"143": "Status Contract",
"144": "Ticket Implementer Agent",
"145": "Worktree Isolation",
"146": "Escalate Rather Than Guess (opencode)",
"147": "Implementer Subagent (opencode)",
"148": "Subagent-Driven Development",
"149": "Code Quality Review",
"150": "Reviewer Subagent (opencode)",
"151": "Finding Severity Rubric",
"152": "Spec Compliance Review",
"156": "ResponseWriter",
"161": "Why Use samber/oops",
"162": "singleflight Cache Stampede Prevention",
"163": "Struct Field Alignment",
"164": "testing/synctest Deterministic Goroutine Testing",
"177": "Backend CLAUDE.md Guidance",
"178": "Graphify Knowledge Graph (graphify-out/)",
"179": "CLAUDE.md (Symlink to AGENTS.md)",
"192": "ADR-0001 (Drop modernc.org/sqlite)",
"193": "ADR-0003 (Split Shared Series Facts)",
"194": "SQLite-to-Postgres Cutover Runbook",
"195": "Import SQL Generation Rules",
"196": "Throwaway Import Generator",
"206": "Real scaling limit is the poller outbound fetch budget",
"207": "PostgreSQL (jackc/pgx/v5)",
"208": "SQLite (modernc.org/sqlite)",
"209": "Postgres chosen for future supportability, not concurrency",
"210": "Per-Reader bearer token for userscripts",
"211": "Discord OAuth2 (authorization code grant)",
"212": "ADR-0002: Discord OAuth, no passwords, no email",
"213": "Discord snowflake is the sole identity (lock-in)",
"214": "Bookmark (per-Reader state: Progress, Favourite, Lifecycle)",
"215": "Deduplicate polling per Series (reader_count DESC queue)",
"216": "ADR-0003: Series is shared, only the Poll updates it",
"217": "Only the Poll writes Series fields (security boundary)",
"218": "Series (shared entity keyed site+series_id)",
"219": "ADR-0004: Wire format stays flat, does not mirror schema",
"220": "Flat wire shape is a contract, not an implementation detail",
"221": "Installed userscripts must keep working (14-day grace window)",
"222": "CDP (Chrome DevTools Protocol) endpoint",
"223": "headless-shell service (socat-fronted CDP)",
"224": "Start Chrome on first CDP connection, reap after 300s idle",
"225": "BROWSER_WS_URL configuration seam",
"226": "ADR-0006: Browser runs on the home machine over the tailnet",
"227": "Browser moved home: VPS memory pressure, no requests served",
"228": "Tailnet (Tailscale network)",
"229": "Content-addressed filesystem storage (SHA-256 of source URL)",
"230": "Cover (Series image bytes)",
"231": "Deny-class destination control for outbound fetch",
"232": "ADR-0007: Backend hosts every Site's Cover bytes",
"233": "kagane CORP same-origin cover restriction",
"234": "Backend acquires, stores, serves every Cover (uniformity)",
"235": "a[aria-label='All Chapter'] anchor pointer",
"236": "Series identity is discovered from the Site's links",
"237": "ADR-0008: Series identity discovered, never derived",
"238": "Chapter slug vs series slug divergence (~7% measured)",
"239": "Scan truncated at first wpd-threads marker",
"240": "Surface ADR conflicts explicitly rather than silently overriding",
"241": "Domain docs: single-context layout guidance",
"242": "/domain-modeling skill (lazy CONTEXT.md creation)",
"243": "CONTEXT.md glossary (ubiquitous language)",
"244": "Gitea (tea CLI, gitea.violetcrown.my.id)",
"245": "wayfinder map/ticket mechanism",
"246": "Triage labels: canonical roles to tracker labels",
"247": "Canonical triage role labels (needs-triage ... wontfix)",
"252": "a[aria-label='All Chapter'] priority pointer",
"253": "Research: lightnovelworld chapter slug vs series slug",
"254": "Gitea issue #77 (chapter vs series slug)",
"255": "Slug divergence measurements (3/41 diverge, 1 split)",
"256": "Unscoped chapter regex is SAFE, truncated at wpd-threads",
"257": "BookmarkManager",
"258": "Bromite (Primary Device)",
"259": "Dark-First Design Constraint",
"260": "Discord Guild Membership",
"261": "Reader Isolation Invariant",
"273": "AGENTS.md",
"279": "Userscript CLAUDE guidance"
}

Some files were not shown because too many files have changed in this diff Show More