Compare commits

..

3 Commits

Author SHA1 Message Date
sulthan 5d330c2ff4 docs: make every AGENTS.md cite code, not docs or issues
A spec, ADR, plan file, or Gitea issue records what was true when it was
written and then goes stale silently, so an agent that follows the pointer
reads a decision that may already have been reversed. Code is the only
source true at read time.

Strip every non-code citation from the three AGENTS.md files (ADRs, spec
and plan files, docs/research, docs/agents/*, DEPLOY/REDEPLOY, and issue
numbers), restating inline any fact the linked doc actually carried: the
tea command set and triage label strings move into the root Forge section.
The Domain docs subsection goes entirely, as it pointed only at CONTEXT.md
and docs/adr/, neither of which exists.

Then rewrite the backend and userscript files around derivability, since
prose that restates mechanism rots the same way a doc link does. Structure
and mechanism now name a symbol and stop; rationale, rejected alternatives
and dated measurements stay written out, because code cannot carry them.
Record that split as a rule in the root file.

Verified by extracting all 118 backticked identifiers and checking each
against the Go, JS, SQL, HTML and CSS sources. That caught one claim that
was already lying: the old cover text said CoverFetcher was gone, but
NewCoverFetcher, TLSCoverFetcher and BrowserCoverFetcher are all live in
internal/latest, so the sentence now names only the dead /img/kagane route.

Also drop the "Guidance for OpenCode (and Claude Code)" openers, so the
files read the same under any harness.
2026-08-17 13:39:28 +07:00
sulthan 550b258c59 fix: give covers their own 10 MiB byte cap (#71) (#112)
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-17 12:06:14 +07:00
sulthan 3ac865cd08 chore: remove graphify (#111)
Removes the graphify integration. It was measured against this repo rather than assumed.

## Why

`graphify query` returns a keyword-seeded BFS neighbourhood, not a location. Asked where CORS origin reflection is implemented, it returned 73 nodes — mostly `api_test.go` helpers, plus a `Reflection and Type Assertions` section from `.agents/skills/golang-performance/references/cpu.md` matched on the word "reflection" — and never named `httpmw/middleware.go:135` or `main.go:121`. `grep` returned both in 39ms. Same shape asking how the poller skips kagane: 145 nodes, top hits `poller_test.go` helpers and two nodes named `T`.

`graphify explain "BrowserFetcher"` is sound (`browser.go L52`, 9 `EXTRACTED` edges), but that is what `lsp references` already answers, against live files instead of a snapshot.

Staleness was never the problem — `graph.json` rebuilt 5s after `f568fb5`, so the git hooks worked. Retrieval quality was.

## What it cost

- Two `PreToolUse` hooks injecting a "MANDATORY: run graphify query first" paragraph into context on **every** grep/find and every source-file read.
- 685k input tokens across 5 build runs (`cost.json`).
- 3.4MB of `graph.json` + `graph.html` tracked, across 11 commits of map-refresh churn.

`AGENTS.md` is the stronger orientation artifact for a repo this size: it carries the CDP constraints, the UTC-clock finding, the per-site adapter list, and the security invariants — none of which an AST graph derives. Graphify earns its keep on repos too large to grep coherently and without curated docs; not this one.

## Changes

- Delete the committed map (`graphify-out/`, -58k lines).
- Drop the `## graphify` rules block from `AGENTS.md` (`CLAUDE.md` is a symlink, so both).
- Drop the five `graphify-out/*` entries from `.gitignore`.
- Empty the two `PreToolUse` hooks in `.claude/settings.json`.
- Remove the stale `graphify query` instruction from `.claude/skills/implement-tickets/SKILL.md` — it pointed dispatched ticket-implementer agents at a binary that no longer exists.

Uninstalled outside the tree (not in this diff): the `graphifyy` CLI, `~/.claude/skills/graphify/`, the global `~/.claude/CLAUDE.md` block, the `Bash(graphify query *)` permission in the git-ignored `.claude/settings.local.json`, and the `post-commit` / `post-checkout` git hooks.

## Verification

`grep -ri graphify` over the worktree is clean; remaining hits are inside `.git/` (commit messages, two stale branch configs). No code touched — backend and userscript are untouched, so `go test ./...` is unaffected.

Reviewed-on: #111
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-17 11:43:15 +07:00
6 changed files with 815 additions and 400 deletions
+75 -40
View File
@@ -1,6 +1,6 @@
# AGENTS.md # AGENTS.md
Guidance for OpenCode (and Claude Code) working in this repo. Repo-wide guidance for coding agents.
## What this is ## What this is
@@ -13,15 +13,18 @@ One backend, one `bookmarks` table: a `kind` column (`manga`|`novel`) splits the
## Hard constraints (drive design — don't violate) ## Hard constraints (drive design — don't violate)
Nothing below is derivable from reading the code — it is why the code looks the
way it does, plus dated measurements against services we don't control.
Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free where plain web APIs suffice — keeps portability across engines: Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free where plain web APIs suffice — keeps portability across engines:
- **Avoid `GM_*` unless needed.** Prefer page `localStorage` over `GM_setValue`/`GM_getValue`, on-page UI over `GM_registerMenuCommand`, plain `fetch()` over `GM_xmlhttpRequest` for cross-origin. - **Avoid `GM_*` unless needed.** Prefer page `localStorage` over `GM_setValue`/`GM_getValue`, on-page UI over `GM_registerMenuCommand`, plain `fetch()` over `GM_xmlhttpRequest` for cross-origin.
- Cross-origin `fetch()` work **only** against CORS-enabled backend. Manga sites `https://`, so backend **must be HTTPS** (else mixed-content block). - Cross-origin `fetch()` work **only** against CORS-enabled backend. Manga sites `https://`, so backend **must be HTTPS** (else mixed-content block).
- Every site is its **own origin with its own `localStorage`** — a shared remote store is the only way to unify bookmarks. Cloud sync required, not optional. - Every site is its **own origin with its own `localStorage`** — a shared remote store is the only way to unify bookmarks. Cloud sync required, not optional.
- Userscript run in **isolated world**, so embedded API token safe from site's JS. - Userscript run in **isolated world**, so embedded API token safe from site's JS.
- Cloudflare's block on manga sites is **per-zone configuration plus request fingerprint, not IP reputation — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — a Site can turn its protection on overnight, which is exactly what comix.to did on 2026-08-12. An earlier version of this line blamed "Cloudflare's bot scoring"; that was wrong. The 1-99 bot score is Enterprise Bot Management only and does not exist for a free-plan zone, and no per-IP request rate is documented as an input to challenge issuance — `docs/research/cloudflare-bot-scoring-and-poll-cadence.md`. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test. - Cloudflare's block on manga sites is **per-zone configuration plus request fingerprint, not IP reputation — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — a Site can turn its protection on overnight, which is exactly what comix.to did on 2026-08-12. An earlier version of this line blamed "Cloudflare's bot scoring"; that was wrong. The 1-99 bot score is Enterprise Bot Management only and does not exist for a free-plan zone, and no per-IP request rate is documented as an input to challenge issuance. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test.
- **kagane.to, comix.to and novelfull.com are the exception to the above** — all three sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`). When that's unset, kagane and comix are skipped entirely (a plain fetch would only retrieve a challenge page) while novelfull pages are still attempted over plain TLS — its challenge is a live time-varying fact and its cover bytes never need the browser. comix turned hostile on 2026-08-12 (#98): its cover host `static.comix.to` is gated too, so its cover bytes go through the browser as well, and its page is read as an in-tab `fetch()` of the series URL rather than a rendered DOM — comix is an SPA, and rendering costs ~65 requests for the same server-rendered HTML one fetch returns. The three other sites poll fine over plain TLS. - **kagane.to, comix.to and novelfull.com are the exception to the above** — all three sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`). When that's unset, kagane and comix are skipped entirely (a plain fetch would only retrieve a challenge page) while novelfull pages are still attempted over plain TLS — its challenge is a live time-varying fact and its cover bytes never need the browser. comix turned hostile on 2026-08-12: its cover host `static.comix.to` is gated too, so its cover bytes go through the browser as well, and its page is read as an in-tab `fetch()` of the series URL rather than a rendered DOM — comix is an SPA, and rendering costs ~65 requests for the same server-rendered HTML one fetch returns. The three other sites poll fine over plain TLS.
- **The CDP browser must look like a real browser, and stock headless images don't.** Measured 2026-08-08 against kagane.to, all from the same IP: `chromedp/headless-shell:stable` never cleared the challenge in 90s (`navigator.webdriver` true, empty plugin list, Chromium-branded client hints — suppressing `webdriver` alone changed nothing); `zenika/alpine-chrome` ships Chrome 124, refused outright; real Chrome with the default `--headless=new` UA never cleared, because the UA says `HeadlessChrome`; real Chrome with a stock UA **and** a non-UTC clock zone cleared in ~4s. Hence `chrome/` — a Debian image with `google-chrome-stable`, a version-derived UA, and `TZ`/`BROWSER_TZ`. Chrome reads the zone *name* through ICU from `/etc/localtime`'s symlink target, ignoring the file's contents, so mounting the host's `/etc/localtime` does **not** work; `/etc/timezone` is mounted instead. - **The CDP browser must look like a real browser, and stock headless images don't.** Measured 2026-08-08 against kagane.to, all from the same IP: `chromedp/headless-shell:stable` never cleared the challenge in 90s (`navigator.webdriver` true, empty plugin list, Chromium-branded client hints — suppressing `webdriver` alone changed nothing); `zenika/alpine-chrome` ships Chrome 124, refused outright; real Chrome with the default `--headless=new` UA never cleared, because the UA says `HeadlessChrome`; real Chrome with a stock UA **and** a non-UTC clock zone cleared in ~4s. Hence `chrome/` — a Debian image with `google-chrome-stable`, a version-derived UA, and `TZ`/`BROWSER_TZ`. Chrome reads the zone *name* through ICU from `/etc/localtime`'s symlink target, ignoring the file's contents, so mounting the host's `/etc/localtime` does **not** work; `/etc/timezone` is mounted instead.
- **The browser is not in the API stack and must not be put back.** It's its own compose unit (`chrome/docker-compose.yml`) on a second machine, reached over the tailnet — it held 471 MiB on a 1974 MiB swapless VPS, and a residential egress avoids the cloud-hosting-IP signature Bot Fight Mode documentedly challenges (ADR-0006; not a better "score" — free-plan zones have no score). Consequences that constrain code: `BROWSER_WS_URL` must be a tailnet **IP** (a MagicDNS name 500s at `/json/version`, same trap as the old Docker service name); the CDP port binds to the tailnet address only, since CDP authenticates nothing and that host has a real LAN; and the browser is on-demand (ADR-0005), so an unreachable or asleep one must degrade exactly as an unset `BROWSER_WS_URL` — plain-TLS libraries unaffected, kagane/comix logged and skipped, stored covers still served. Never add `chromedp.NoModifyURL`: discovery per fetch is what makes a restarted Chrome invisible. - **The browser is not in the API stack and must not be put back.** It's its own compose unit (`chrome/docker-compose.yml`) on a second machine, reached over the tailnet — it held 471 MiB on a 1974 MiB swapless VPS, and a residential egress avoids the cloud-hosting-IP signature Bot Fight Mode documentedly challenges (not a better "score" — free-plan zones have no score). Consequences that constrain code: `BROWSER_WS_URL` must be a tailnet **IP** (a MagicDNS name 500s at `/json/version`, same trap as the old Docker service name); the CDP port binds to the tailnet address only, since CDP authenticates nothing and that host has a real LAN; and the browser is on-demand, so an unreachable or asleep one must degrade exactly as an unset `BROWSER_WS_URL` — plain-TLS libraries unaffected, kagane/comix logged and skipped, stored covers still served. Never add `chromedp.NoModifyURL`: discovery per fetch is what makes a restarted Chrome invisible.
- **UTC is the tell, not a country mismatch.** A UTC clock is the datacenter default, and the challenge refuses it; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while `Asia/Jakarta` **and** `America/New_York` both cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (`Asia/Bangkok`) rather than the measured egress. A second earlier claim, that Cloudflare "scores" a UTC clock, was also wrong: the measurement is real but the mechanism is not documented anywhere — Cloudflare publishes no timezone signal, and free-plan zones carry no score at all. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one. - **UTC is the tell, not a country mismatch.** A UTC clock is the datacenter default, and the challenge refuses it; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while `Asia/Jakarta` **and** `America/New_York` both cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (`Asia/Bangkok`) rather than the measured egress. A second earlier claim, that Cloudflare "scores" a UTC clock, was also wrong: the measurement is real but the mechanism is not documented anywhere — Cloudflare publishes no timezone signal, and free-plan zones carry no score at all. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one.
- **A challenged page needs the tab kept open.** The interstitial takes seconds to solve and only then writes clearance into the browser's shared cookie jar. Navigate-read-close never clears anything; `BrowserFetcher.run` holds one tab and re-reads until the payload arrives. - **A challenged page needs the tab kept open.** The interstitial takes seconds to solve and only then writes clearance into the browser's shared cookie jar. Navigate-read-close never clears anything; `BrowserFetcher.run` holds one tab and re-reads until the payload arrives.
@@ -36,7 +39,7 @@ Two Violentmonkey userscripts (isolated world, per-site adapters, localStorage c
on-demand Chrome, separate machine (chrome/) on-demand Chrome, separate machine (chrome/)
``` ```
Two deployable units on two machines: the API stack (`docker-compose.yml` + `docker-compose.prod.yml`, on the VPS) and the browser (`chrome/docker-compose.yml`, on the home machine). They share nothing but `BROWSER_WS_URL` and update independently. Backend-specific architecture (packages, endpoints, poller, config env vars) lives in `backend/AGENTS.md`. Userscript-specific structure (adapters, retry queue, UI, live URL shapes) lives in `userscript/AGENTS.md`. Deploy order `DEPLOY.md` (§7 for the browser), redeploy `REDEPLOY.md` (§8 for the browser). Two deployable units on two machines: the API stack (`docker-compose.yml` + `docker-compose.prod.yml`, on the VPS) and the browser (`chrome/docker-compose.yml`, on the home machine). They share nothing but `BROWSER_WS_URL` and update independently. Backend-specific detail lives in `backend/AGENTS.md`, userscript-specific detail in `userscript/AGENTS.md`.
## Commands ## Commands
@@ -64,28 +67,29 @@ Smoke test: `curl` endpoints with `Authorization: Bearer <token>`; confirm `OPTI
## Forge: Gitea, not GitHub ## Forge: Gitea, not GitHub
`origin` is self-hosted Gitea instance (`gitea.violetcrown.my.id`), so **`gh` don't work here — use `tea` (Gitea CLI) for anything past plain git.** Common ones: `origin` is self-hosted Gitea instance (`gitea.violetcrown.my.id`, repo `sulthan/mangaBookmark`), so **`gh` don't work here — use `tea` (Gitea CLI) for anything past plain git.** `tea` infers the repo from `origin`; auth lives in `tea login`, not a `GH_TOKEN` env var. It prints rendered boxes rather than plain text, so pass `-o json` when parsing; a PR URL lands on the last line.
- Open PR: `tea pr create --head <branch> --base main --title "..." --description "..."` - PRs: `tea pr create --head <branch> --base main --title "..." --description "..."`, `tea pr list`, `tea pr <n>`, `tea pr checkout <n>`.
- List / view / check out: `tea pr list`, `tea pr <n>`, `tea pr checkout <n>` - Issues: `tea issue create --title "..." --description "..."` (`--labels`, `--assignees` optional), `tea issue <n> --comments`, `tea issue list --state open|closed|all -o json`, `tea issue close <n>`.
- Issues: `tea issue create`, `tea issue list` - Comments: `tea comment <n> "..."` — `tea issue close` takes no `--comment` flag.
- Auth lives in `tea login`, not `GH_TOKEN` env var. - Labels: `tea issue edit <n> --add-labels "..."` / `--remove-labels "..."`. Gitea will **not** auto-create a label, so `tea labels create --name "..." --color "#rrggbb"` first.
- Triage vocabulary is `needs-triage`, `needs-info`, `ready-for-agent`, `ready-for-human`, `wontfix`.
`tea` print output as rendered boxes rather than plain text; PR URL lands on last line. - Gitea shares one index space across issues and PRs, so a bare `#42` may be either — try `tea pr 42`, fall back to `tea issue 42`.
## Design system ## Design system
Web UI + userscript panel follow **Cinder**, rules in `docs/design-system.md` Web UI + userscript panel follow **Cinder**. Tokens are the `:root` block in
— source of truth Claude Design project `BookmarkManager Web UI` `backend/internal/web/static/style.css`; that file, `backend/internal/web/templates/*`,
(`969ac210-fe02-4c01-ae1b-9a271dcc779a`). Read it before touching and the userscript `TEMPLATE`/`CSS` are the only places it is expressed.
`backend/internal/web/static/style.css`, `backend/internal/web/templates/*`, or userscript Source of truth for the visual language is the Claude Design project
`TEMPLATE`/`CSS`. Core law: **ember means new chapter only** — no other `BookmarkManager Web UI` (`969ac210-fe02-4c01-ae1b-9a271dcc779a`).
state (busy, error, destruction) may use `--ember`; destruction gets
`--danger`. No cards/corners/shadows, one `--measure: 760px` column, tokens Core law: **ember means new chapter only** — no other state (busy, error,
only (never hardcode hex outside `:root`), both colour branches touched destruction) may use `--ember`; destruction gets `--danger`. No
together. Any move that pulls series out of list (archive/finish/remove) cards/corners/shadows, one `--measure: 760px` column, tokens only (never
must be confirm-gated via its own `.confirm-row`; only restore fires hardcode hex outside `:root`), both colour branches touched together. Any move
instantly. that pulls a series out of the list (archive/finish/remove) must be
confirm-gated via its own `.confirm-row`; only restore fires instantly.
## Security invariants ## Security invariants
@@ -127,13 +131,19 @@ Review gate: auth, CORS, session, crypto, and the fetch gate are security-critic
## Comments ## Comments
Comment only if code alone can't carry info. Cost per read — must earn spot. Comment only if code alone can't carry info. Cost per read — must earn spot.
Wrong comment worse than none: it misleads readers and measurably degrades
LLM performance on the file. Missing comment costs little. Bias to fewer.
Write for: Docstring on public/exported surface — exception, near-always worth it.
- Why not what. Tradeoffs, non-obvious decisions. Contract only: what it takes, returns, throws, mutates; units; pre/post
- Load-bearing detail looking incidental — say so if "simplify" breaks it. conditions. Not a restatement of the body. Skip on private/obvious.
Inline — write for:
- Why not what. Tradeoffs, non-obvious decisions, rejected alternatives.
- heavy detail looking incidental — say so if "simplify" breaks it.
- Non-local consequence, invisible from function alone. - Non-local consequence, invisible from function alone.
- Wire format / encoding / interface contract — save callers re-deriving. - Wire format / encoding / ordering / invariant — save callers re-deriving.
- Gotcha/workaround, with ref if exists. - Gotcha/workaround, with ref (issue, RFC, vendor bug) if exists.
- Domain/business rule not derivable from code. - Domain/business rule not derivable from code.
Skip: Skip:
@@ -142,24 +152,49 @@ Skip:
- Banners, dividers, `// helpers`. - Banners, dividers, `// helpers`.
- Change narration (`// fix bug`, `// as requested`, `// new impl`) — git's job. - Change narration (`// fix bug`, `// as requested`, `// new impl`) — git's job.
- Commented-out code — delete. - Commented-out code — delete.
- TODO without concrete action. - TODO without concrete action + owner.
- Narrating the plan you just reasoned through. Plan in prose or in your head;
ship the code, not the transcript.
- Anything restating a name that could be fixed by renaming instead.
Style: one dense comment over function beats one per line inside. Tight, no worked example unless bug subtle. Wrong comment worse than none — update/delete on change. Default fewer — sparse+high-signal beats comprehensive. Staleness filter: if the comment describes something likely to change
independently of this line, it will rot and start lying. Either anchor it to
something stable, assert it in a test, or leave it out.
Test: "competent reader get this from code in few sec?" Yes → skip. Needs detour through another file/spec/git-blame → write it. Style: one dense comment over a function beats one per line inside. Tight; no
worked example unless the bug is subtle. On edit, update or delete stale
comments in the code you touch — silence beats a lie.
## Agent skills Test: "competent reader get this from code in a few sec?" Yes → skip.
Needs detour through another file/spec/git-blame/external doc → write it.
`AGENTS.md` is the single source of truth for agent guidance; every `CLAUDE.md` in this repo is a symlink to the `AGENTS.md` beside it. Edit `AGENTS.md`. ## Writing an AGENTS.md
### Issue tracker `AGENTS.md` is the single source of truth for agent guidance; every `CLAUDE.md`
in this repo is a symlink to the `AGENTS.md` beside it. Edit `AGENTS.md`.
Issues live as Gitea issues on `gitea.violetcrown.my.id` (`sulthan/mangaBookmark`), driven by the `tea` CLI — not `gh`. See `docs/agents/issue-tracker.md`. **Cite code, never docs, issues, or plans.** A spec, ADR, plan file, or Gitea
issue records what was true when it was written and then goes stale silently;
an agent that follows the pointer reads a decision that may already have been
reversed. Code is the only source true at read time — cite a package, file,
symbol, env var, or route. The sole non-code exception is a sibling
`AGENTS.md`. If a doc holds a fact an agent needs, restate the fact here rather
than linking to it.
### Triage labels **State a fact in prose only if the code cannot answer it.** Split by
derivability:
Default five-role vocabulary, label strings unchanged (`needs-triage`, `needs-info`, `ready-for-agent`, `ready-for-human`, `wontfix`). See `docs/agents/triage-labels.md`. - *Structure* — packages, routes, env vars, columns, struct fields. Rots fast,
cheap to re-read. **Name the symbol, write nothing else.**
- *Mechanism* — what a function does, how a flow proceeds. **Name the symbol
plus at most one line of orientation.**
- *Rationale* — why it is this way, what a "simplify" would break, what was
tried and rejected. Not in the code and cannot be re-derived. **Write it out.**
- *Measurement* — an observation against something we don't control. **Write it
out with the date**; a dated fact is honest, an undated one pretends to be
permanent.
### Domain docs Restating mechanism in prose is how these files rot: the code changes, the
paragraph doesn't, and the next agent trusts the paragraph. A pointer degrades
Single-context: one root `CONTEXT.md` plus `docs/adr/`, both created lazily. See `docs/agents/domain.md`. more honestly — and every symbol you name must actually exist, since a dead
pointer is a bug, not a stale sentence.
+264 -270
View File
@@ -1,274 +1,268 @@
Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGENTS.md` for the project-wide architecture diagram, hard constraints, and design system. Scope: `backend/`.
- **Backend** (`backend/`): stdlib `net/http` (handful routes, no framework) + Postgres over `jackc/pgx/v5` (pure Go, `CGO_ENABLED=0` -> static binary -> distroless/scratch image). Reverse proxy terminates TLS; Go service listens plain `:8080`. Each entry names the code that holds the truth — read that for *what it does*.
Single binary, split into packages under `backend/internal/`: `store` The prose here is only what code cannot tell you: rationale, rejected
(Bookmark type, Postgres persistence, migration runner), `latest` (background alternatives, dated measurements, and invariants a plausible refactor would
poller, site parsers, TLS fetcher), `session` (cookie signing, login silently break.
rate limiter), `httpmw` (Auth/Gzip/CORS middleware), `api` (JSON
bookmark handlers), `userscript` (userscript-serving handler), `web` ### Layout
(browser UI handler + `templates/` + `static/`, `go:embed`-ed).
`backend/main.go` is the composition root — the only place that wires `backend/main.go` → `newRouter` is the composition root, the only place
packages together into `newRouter`. Root-level `*_test.go` hold packages are wired. Packages under `backend/internal/`: `store`, `latest`,
integration tests that exercise the full router; unit tests for a `session`, `httpmw`, `api`, `userscript`, `web`, `token`, `pgtest`. Root-level
package live beside it under `internal/`. `*_test.go` exercise the full router; unit tests live beside their package.
- **Schema is migration-owned.** `internal/store/migrations/*.sql` is
`go:embed`-ed and applied on every start by `store.migrate`: one numbered Not visible from any single file: stdlib `net/http` with no framework,
file per change, one transaction each, versions recorded in Postgres over `jackc/pgx/v5`, `CGO_ENABLED=0` static binary into a distroless
`schema_migrations`. Files are **append-only** — editing an applied one image, TLS terminated by the reverse proxy so the service listens plain `:8080`.
changes nothing on a database that already ran it. No column probing, no
data-fixup migrations: both were SQLite-era machinery and are gone. ### Schema — `internal/store/migrations/*.sql`, run by `store.migrate`
- **Tests need Docker.** `internal/pgtest` starts one `postgres:17-alpine`
container per test binary (`TestMain` -> `pgtest.Main`) and hands each test - Migration files are **append-only**. Editing an applied one changes nothing
its own database (`pgtest.URL(t)`). A package whose tests touch the store on a database that already recorded its version in `schema_migrations`, so
must have that `TestMain`. the fix silently applies to new deployments only.
- **Reader-owned store, four tables.** `readers` is keyed by Discord user ID - No column probing, no data-fixup migrations. Both were SQLite-era machinery
and carries the SHA-256 of the Reader's userscript credential plus a and were removed deliberately — don't reintroduce either.
`token_epoch` (issue #24). Credentials are derived, never stored: `token.Token(TOKEN_KEY, discord_id, epoch)` (HMAC, `internal/token`), and only its SHA-256 sits in `readers.token_sha256`, so install URLs can be rebuilt after any restart while a database leak yields nothing but hashes. The seed creates the **owner** row at startup; its epoch-0 hash is refreshed on every start **only while the row has never been rotated**, so a restart can never resurrect a rotated-away credential. Every other row is created by that Reader's own first login (`Store.EnsureReader`, idempotent on `discord_id`, and it never rewrites an existing row's hash). Rotation is `Store.RotateToken` (epoch bump + hash rewrite in one transaction), driven by the web UI.
`series` keyed `(site, series_id)` ### Tests need Docker — `internal/pgtest`
(`asura`|`demonic`|`comix`|`kagane`|`novelfull`|`lightnovelworld`) owns the
shared facts — title, cover, canonical URL, `kind` (`manga`|`novel`), `pgtest.Main` from `TestMain` starts one `postgres:17-alpine` per test binary;
Latest Chapter, `latest_checked_at`, and the Sighting pair `pgtest.URL` hands each test its own database. A package whose tests touch the
`latest_sighted_at`/`latest_raised_by` (issue #103) — and `bookmarks` holds only what store must have that `TestMain` or it has no database at all.
differs between readers: progress, favourite, lifecycle bucket,
`updated_at`. A bookmark is keyed `(reader_id, site, series_id)` — no ### Reader-owned store — `internal/store`, `internal/token`
surrogate id; the wire `key` is derived as `site:series_id` on read — and
every store read/write is scoped to the reader it names. Auth resolves the Four tables; shape is in the migrations, behaviour in `Store`'s methods.
acting Reader from the presented credential (`httpmw.Auth`) and nothing
else — there is no unauthenticated-by-Reader route and no global token; the - **Credentials are derived, never stored.** `token.Token(TOKEN_KEY, discord_id, epoch)`
reader id travels in the request context. Sync **last-write-wins**; the wire format is an HMAC; only its SHA-256 reaches `readers.token_sha256`. So install URLs
stays flat (ADR-0004). `Store.Upsert` decomposes one flat body across two can be rebuilt after any restart, and a database leak yields nothing usable.
tables and enforces the ownership rule: client `title`/`series_url`/`cover` - **The owner's epoch-0 hash is refreshed at startup only while the row has
are written only when the series row is new (ADR-0003). never been rotated.** Drop that condition and a restart resurrects a
- **Endpoints:** `GET /bookmarks`, `PUT /bookmarks/{key}` (upsert; see `updated_at` rule below), `DELETE /bookmarks/{key}`, `GET /healthz` (no auth). rotated-away credential.
- **Web UI:** same binary serve the browser UI on a second - `Store.EnsureReader` never rewrites an existing row's hash — a returning
hostname — `GET /` (list, or login page when no session), Reader's login must not invalidate their installed scripts.
`GET /auth/discord` + `GET /auth/discord/callback` (Discord OAuth, - **Every read and write is scoped to the acting Reader**, resolved from the
ADR-0002), `POST /logout`, `GET /static/*`, htmx fragment endpoints presented credential by `httpmw.Auth` and carried in the request context.
under `/ui/*`. Templates + assets `go:embed`-ed under There is no unauthenticated-by-Reader route and no global token.
`backend/internal/web/`, so `backend/Dockerfile` must copy the whole - **`series` holds what readers share, `bookmarks` only what differs.** A
`internal/` tree, not just `*.go`. Sessions are rows in the `sessions` bookmark key is `(reader_id, site, series_id)` with no surrogate id; the wire
table: the cookie carries only an opaque id, looked up (and expiry- `key` is derived as `site:series_id` on read.
checked) on every request, and deleting the row revokes the session. - `Store.Upsert` splits one flat body across both tables and enforces the
Guild membership *is* registration (issue #27): `discordCallback` gates on ownership rule: client `title`/`series_url`/`cover` are written **only when
membership (and `DISCORD_REQUIRED_ROLE` when set) and then calls the series row is new**, so one reader cannot retitle a shared series.
`Store.EnsureReader`, so a refusal creates nothing and a returning Reader - Sync is last-write-wins and the wire format stays flat — clients depend on
reuses their row. The owner is the only Reader with administrative reach: both; neither is an implementation detail to tidy up.
`POST /readers/{id}/revoke` (404 for anyone else) drops that Reader's
sessions, and the `readers` panel renders only on the owner's page. ### Web UI — `internal/web`
A Reader with no bookmarks at all sees `listView.Fresh`, whose empty state
offers both install links instead of describing a filter. Routes, templates and assets are all in that package; `AdminPatterns()` and
UI mutations read-modify-write `adminRoutes()` enumerate the privileged ones.
through `Store.Get` + `Store.Upsert` so `updated_at` rule stays one
place. See `docs/superpowers/specs/2026-07-25-web-ui-design.md`. - **`backend/Dockerfile` must copy the whole `internal/` tree**, not just
**Design-tool caveat:** templates link `/static/style.css` root-absolutely `*.go`: templates and static assets are `go:embed`-ed from
(correct — served from `/`), but impeccable detector resolves `internal/web/`.
stylesheet href with `path.resolve(fileDir, href)`, drops directory - **Guild membership *is* registration.** `discordCallback` gates on membership
on leading `/` and silently skip file. Relative href don't help (plus `DISCORD_REQUIRED_ROLE` when set) and only then calls
either: template's directory isn't its served path. So `Store.EnsureReader`, so a refusal creates nothing.
`detect.mjs backend/internal/web/templates` reports **false clean** — - Sessions are rows, not signatures: the cookie carries an opaque id and
always pass `backend/internal/web/static` too. One finding there, expiry is checked on lookup, which is what makes deleting the row an instant
`overused-font` on "Instrument Serif", deliberate identity choice, not debt. revocation.
- **Every action that moves series out of list is confirm-gated.** - UI mutations go through `Store.Get` + `Store.Upsert` so the `updated_at` rule
Archive, finish, remove each open own `.confirm-row` disclosure below stays in exactly one place.
(`toggleConfirmRow(key, kind)` in `filter.js`, `kind` ∈ - `listView.Fresh` exists because a Reader with no bookmarks at all needs
`archive|finish|remove`); restore fire instantly since it's the reversal. install links, not an empty-filter message.
Remove's row wear ember wash, two reversible ones wear `.calm` grey. - **Design-tool caveat:** `detect.mjs backend/internal/web/templates` reports a
`--ember` stay reserved for new-chapter signal: busy bar and inline **false clean**. Templates link `/static/style.css` root-absolutely (correct —
error use `--mute`. it is served from `/`), but the detector resolves hrefs with
- **Latest-chapter poller:** one goroutine per Site (a Poll Lane, issue #100), `path.resolve(fileDir, href)`, which drops the directory on a leading `/` and
each re-checking that Site's bookmarked series' newest published chapter from skips the file silently; a relative href doesn't help either, since a
backend's own network access, so `latest_chapter` stays fresh when the user template's directory isn't its served path. Always pass
isn't browsing. Second, parallel signal — the userscript keeps its own `backend/internal/web/static` too. The one finding there, `overused-font` on
`maybeCaptureLatestOnSeriesPage`/`backgroundRefreshLatest` schedule, and its "Instrument Serif", is a deliberate identity choice, not debt.
`reportLatestChapter` PUTs every read, unchanged numbers included, because an
unchanged read is exactly the Sighting worth deferring a Poll on (#103). ### Confirm gating — `internal/web/static/filter.js`, `toggleConfirmRow(key, kind)`
Two independent clocks: per-series rest (`series.latest_checked_at`,
enforced by `Store.DueForLatestCheck`'s WHERE clause — `now - Rest`) and Every action that pulls a series out of the list (`archive|finish|remove`) opens
per-Lane gap (the Lane sleeping between fetches, `effectiveGap`). Both live its own `.confirm-row`; restore fires instantly because it is the reversal.
in the Site registry (`internal/latest/sites.go`), not config: the five env Remove wears the ember wash, the two reversible ones wear `.calm` grey.
knobs that used to size a shared pace are gone. **`--ember` is reserved for the new-chapter signal** — the busy bar and inline
The poller walks **Series, not Bookmarks** — a series referenced by several errors must use `--mute`, or the one colour that means "something to read"
bookmarks is fetched once per cycle, and the due queue orders stops meaning it.
`reader_count DESC, latest_checked_at ASC` (ADR-0003). Series row stamped
*before* fetch so broken series wait out the rest instead of retrying ### Latest-chapter poller — `internal/latest`, Site registry in `sites.go`
every tick; found chapter written straight to the series row via
`Store.SetLatestChapter`, so a bookmark's `updated_at` — and the list One goroutine per Site (a Poll Lane) re-checks that Site's bookmarked series
order — is never touched. from the backend's own network position, so `latest_chapter` stays fresh while
**Sightings** (issue #103, ADR-0011) let a Reader's own page read defer a nobody is browsing. The userscript's `reportLatestChapter` is a second,
Poll: `Store.RecordSighting` — called by the PUT handler *before* the Upsert, parallel signal — it PUTs every read, unchanged numbers included, because an
because the raise test needs the row as it stands — stamps unchanged read is exactly the Sighting worth deferring a Poll on.
`series.latest_sighted_at` and, when the report raises the stored number,
names its Reader in `series.latest_raised_by`. The due query's HAVING clause - **Pace lives in the Site registry, not config.** Two clocks: per-series rest
is where deferral lives: a Series is skipped only while it has exactly one (`series.latest_checked_at`, enforced in `Store.DueForLatestCheck`'s WHERE)
Bookmark, was sighted within one Rest, and is under the ceiling and per-Lane gap (`effectiveGap`). The five env knobs that used to size one
(`sightingCeilingRests`, six of that Site's rests) since its last Poll. So a shared pace are gone; don't add them back.
shared Series is never deferred, and no Series goes six hours unpolled - **The poller walks Series, not Bookmarks** — a series several readers hold is
whatever arrives. `checkOne` judges the named Reader off the comparison it fetched once per cycle, and the due queue orders `reader_count DESC,
already makes: a lower number is a contradiction (logged with the Reader and latest_checked_at ASC` so the widely-read ones win contention.
both numbers), the same number an agreement, a higher number the Site - **The series row is stamped *before* the fetch**, so a permanently broken
publishing and neither — that last one clears the attribution instead, since series waits out its rest instead of being retried every tick.
the value the Poll then stores is its own and a later retraction is not the - `Store.SetLatestChapter` is a single-column UPDATE, deliberately not a
Reader's fault. Three contradictions read-modify-write of the bookmark: it therefore cannot revert read progress
(`store.SightingDisagreementLimit`) stop that Reader deferring — their or move `updated_at`. The old stale-re-read race died with the Get+Upsert
reports still write the Latest Chapter — and twenty consecutive agreements flow — don't restore one here.
(`store.SightingAgreementsToClear`) forgive them, as does the owner's
clear-marks control. Deferral is recomputed from live facts every round, so **Sightings** (`Store.RecordSighting`, the due query's HAVING clause,
nothing needs invalidating when a Series gains a second Bookmark; the one `latest.checkOne`) let a Reader's own page read defer a Poll.
input read earlier is the Reader's marks, checked when the Sighting is
recorded, so crossing the threshold or being cleared takes effect from that - Recorded by the PUT handler **before** the Upsert, because the raise test
Reader's next Sighting and the standing already bought lasts out its rest. needs the row as it stands.
Refusals and browser loss are Lane-local: two `errChallengeHeld` in one pass - A Series is deferred only while it has exactly one Bookmark, was sighted
stop that Site for `refuseBackoff` (15m) while other Lanes continue; an within one Rest, and is under `sightingCeilingRests` since its last Poll — so
`errBrowserInterrupted` (remote Chrome restart) sets a shared Poller flag a shared Series is never deferred and nothing goes six hours unpolled
that makes the other browser Lanes skip their passes for the same 15m, so a whatever arrives.
restarting Chrome doesn't stamp one Series per Lane per pass — after the - A *higher* report clears the attribution rather than crediting it: the value
window the flag decays and they probe again. Browser Lanes wake Chrome only the Poll then stores is its own, so a later retraction isn't the Reader's
when 5+ Series are due or one has waited 15m (ADR-0005 on-demand browser), fault.
and cover work (both healing a stored source URL and filling a blank from - `store.SightingDisagreementLimit` contradictions stop a Reader deferring —
the series page) runs in the background so a slow CDN can't consume a their reports still write the Latest Chapter — and
Lane's gap. `store.SightingAgreementsToClear` agreements forgive them, as does the
A refusal is only ever the challenge *page*: `isInterstitial` matches the owner's clear-marks control.
orchestration path `/cdn-cgi/challenge-platform/h/`, never the bare prefix. - Deferral is recomputed from live facts each round, so nothing needs
Cloudflare injects `/cdn-cgi/challenge-platform/scripts/jsd/main.js` into invalidating when a Series gains a second Bookmark. The one input read
ordinary 200 pages once a zone turns JS detections on, which demonic did on earlier is the Reader's marks, so crossing or clearing a threshold takes
2026-08-16 — the prefix match then read every real demonic page as a refusal effect from their next Sighting and the standing already bought lasts out its
and parked that Lane in 15m backoff while plain TLS was returning the full rest.
series page.
Fetches use `bogdanfinn/tls-client` with Chrome profile as defence in depth **Refusals and browser loss are Lane-local.** Two `errChallengeHeld` in a pass
against fingerprint-based blocking; any failure log and skip. kagane, comix stop that Site for `refuseBackoff` while other Lanes continue. An
and novelfull sit behind Cloudflare JavaScript challenges the TLS client `errBrowserInterrupted` (remote Chrome restarted) sets a shared Poller flag so
can't clear, so they are fetched over CDP via `BROWSER_WS_URL`; kagane and the *other* browser Lanes skip their passes for the same window — otherwise a
comix are simply not polled when that's unset, while novelfull falls back to restarting Chrome stamps one Series per Lane per pass, burning rests on
a plain-TLS attempt — its challenge is a live time-varying fact, and its failures. The flag decays and they probe again.
cover bytes never need the browser. comix's browser read is an in-tab
`fetch()` of the Series URL, not a DOM render: it is an SPA, so rendering - **`isInterstitial` matches the orchestration path
costs ~65 requests for the same server-rendered HTML one fetch returns `/cdn-cgi/challenge-platform/h/`, never the bare prefix.** Cloudflare injects
(measured 2026-08-12, issue #98). See `/cdn-cgi/challenge-platform/scripts/jsd/main.js` into ordinary 200 pages
`docs/superpowers/specs/2026-07-26-server-latest-chapter-polling-design.md`. once a zone turns JS detections on, which demonic did on 2026-08-16: the
The poller's series write is a single-column UPDATE prefix match read every real demonic page as a refusal and parked the Lane in
(`Store.SetLatestChapter`), not a read-modify-write of the whole bookmark: backoff while plain TLS was returning full series pages.
it cannot revert read progress or move `updated_at`, so the old - Fetches use `bogdanfinn/tls-client` with a Chrome profile as defence in depth
stale-re-read race is gone with the Get+Upsert flow. against fingerprint blocking; any failure logs and skips.
- **Covers are acquired at creation, then served from our own origin - kagane, comix and novelfull sit behind Cloudflare JS challenges the TLS
(ADR-0007):** the first Bookmark of a Series fires `Store.OnSeriesCreated`, client can't clear, so they go over CDP (`BROWSER_WS_URL`). kagane and comix
which `latest.Acquirer` turns into one series-page fetch yielding both the are simply not polled when it's unset — a plain fetch would only retrieve a
Latest Chapter and the cover URL; the bytes then go through challenge page — while novelfull still attempts plain TLS, because its
`latest.CoverBytesFetcher` into `Store.SetSeriesCover`. It runs in a challenge is a live time-varying fact and its cover bytes never need a browser.
goroutine — the Reader's PUT must neither block on a Site nor fail with one - **comix's browser read is an in-tab `fetch()` of the Series URL, not a DOM
— and every failure is logged and dropped, leaving the Bookmark intact. The render.** It is an SPA: rendering cost ~65 requests for the same
wire's `cover` is the absolute `PUBLIC_BASE_URL + /covers/{sha256}` once server-rendered HTML one fetch returns (measured 2026-08-12).
bytes exist and `""` before, never an address that 404s. `GET /covers/{addr}` - Browser Lanes wake Chrome only when 5+ Series are due or one has waited 15m,
is public and uncredentialed: the userscript renders it on a Site's origin, and cover work runs in the background so a slow CDN can't eat a Lane's gap.
where no cookie or token of ours travels. A client-sent `cover` is decoded
and discarded, permanently (ADR-0004 compatibility). ### Covers — `Store.OnSeriesCreated`, `latest.Acquirer`, `latest.CoverBytesFetcher`, `Store.SetSeriesCover`
Browser-backed Sites join the same pipeline (issue #62, extended to comix by
#98): kagane and comix pages *and* cover bytes go through the browser sidecar Acquired once when the first Bookmark of a Series is created, then served from
(nothing falls back to a plain fetch, which would only retrieve a challenge our own origin by the public `GET /covers/{addr}`.
page), while novelfull needs the browser only for its HTML — the cover URL
comes out of the browser-fetched page and the bytes go over plain TLS. With - Acquisition runs in a goroutine: the Reader's PUT must neither block on a
no browser configured, kagane and comix Covers are simply absent; novelfull Site nor fail with one. Every failure is logged and dropped, leaving the
still gets one — at creation and on the poll — when its page body happens to Bookmark intact.
answer a plain request (the challenge is a live time-varying fact). comix - The wire `cover` is the absolute `PUBLIC_BASE_URL + /covers/{sha256}` once
cover bytes must arrive by direct navigation, not an in-page fetch: its bytes exist and `""` before — **never an address that 404s**. Absolute
Series page sets `cross-origin-embedder-policy: require-corp`, which fails a because the userscript renders it on a Site's origin.
page-context fetch of `static.comix.to`. The old kagane-only - `GET /covers/{addr}` is public and uncredentialed by design: no cookie or
serving path (`/img/kagane/{id}`, template rewrite, `CoverFetcher`) is gone token of ours may travel to a Site's origin.
(issue #63): the one public route serves every Site. - A client-sent `cover` is decoded and discarded, permanently — wire
- **`updated_at` drives list order, so moves only on real reading progress:** server apply its timestamp when row new or `last_chapter_num` changes, else keep stored value — favouriting series or recording newly published chapter must not reorder list. `PUT` therefore returns row **as stored**, clients must adopt that response rather than own payload. See `plans/2026-07-25-bookmark-list-favorites-design.md` §4. compatibility, not an oversight.
- **Lifecycle buckets:** `status` on each bookmark is `reading` | `archived` | - **One route serves all six Sites.** No proxy, no per-Site rewrite, no second
`finished`, orthogonal to `favorite`. Archived and finished appear only in place that decides a renderable address: the wire `cover` is it. Templates
own tab — not in All, Updated, Favourites, or recent strip. Poller keeps render `.Cover` and nothing else. The old kagane-only serving path
checking archived series and skip finished ones. `finished` settable (`/img/kagane/{id}` plus a template rewrite) is gone; don't reintroduce a
only from web UI; `PUT /bookmarks/{key}` reject it with 400. per-Site route because one Site's CDN misbehaves.
**Empty incoming status means "keep stored one"** — resolved on the - The only Site names left in cover code are in `browserOnlyCoverURL`
`VALUES` side of `Store.Upsert`, not conflict clause, since (`internal/latest`): kagane answers a plain fetch with a challenge *and*
`excluded.*` is post-evaluation row and default applied there would
wipe bucket on every PUT from client that predates column. See
`docs/superpowers/specs/2026-07-27-status-buckets-design.md`.
- **Config via env:** `TOKEN_KEY` (derives every Reader's userscript credential;
required), `OWNER_DISCORD_ID` (seeds the owner Reader — the administrator and
the owner of every pre-registration bookmark; required),
`ALLOWED_ORIGINS` (comma list),
`DATABASE_URL` (Postgres connection URL, required — no default),
`COVER_DIR` (required filesystem volume for content-addressed Cover bytes),
`PUBLIC_BASE_URL` (required origin this deployment answers on, trailing
slash trimmed; every Cover URL on the wire is built from it, absolute
because the userscript renders on a Site's origin — ADR-0007),
`PORT` (default `8080`), `DISCORD_CLIENT_ID`/`_CLIENT_SECRET`/`_GUILD_ID`/
`_REDIRECT_URI` (required; Discord OAuth for the browser UI),
`DISCORD_REQUIRED_ROLE` (optional role gate, empty by default),
`DISCORD_API_BASE` (default `https://discord.com/api/v10`),
`LATEST_CHAPTER_POLL_ENABLED` (background latest-chapter poller kill
switch, default on). Pace is per Site in the registry (issue #100): every
Site rests an hour and gaps ten seconds, a Site with more eligible Series
than 360 tightens its own gap toward the 1s floor, and browser Lanes wake
Chrome only on demand (ADR-0005). The `_COOLDOWN`/`_BROWSER_COOLDOWN`/
`_INTERVAL`/`_BATCH`/`_STAGGER` knobs that used to size a shared pace are
gone. The 1h rest for browser Sites is safe on documented grounds: a
challenged page costs seconds of a serialized single-tab browser, free-plan
zones have no bot score and no published per-IP rate input, and
`cf_clearance` expires in 30 minutes so every cadence at or above 1h
re-solves anyway —
`docs/research/cloudflare-bot-scoring-and-poll-cadence.md`.
`USERSCRIPT_PATH` and `NOVEL_USERSCRIPT_PATH` (files served at
`/u/{token}/manga-bookmark.user.js` and `/u/{token}/novel-bookmark.user.js`,
defaults `/userscript/manga-bookmark.user.js` and
`/userscript/novel-bookmark.user.js`, both supplied by bindmount; the
`__API_TOKEN__` placeholder inside them is substituted with the requesting
Reader's credential at serve time).
`BROWSER_WS_URL` (CDP endpoint of the browser, which runs on a **separate
machine** and is reached over the tailnet — ADR-0006, `chrome/docker-compose.yml`.
Used by the poller for kagane, comix and novelfull page fetches and by the
cover pipeline for kagane's and comix's image bytes (the browser is the only
route that clears the challenge those two serve their covers behind); unset —
the default — disables browser polling and leaves kagane and comix Covers
blank until stored bytes
exist. Must be a tailnet IP, never a hostname: Chrome's DevTools handler 500s
`/json/version` for any Host that isn't an IP or `localhost`).
- **No per-Site cover path (issue #63):** every Cover — all six Sites — is
served by the one public `GET /covers/{addr}` route from content-addressed
bytes. There is no proxy, no per-Site rewrite, no second place that decides
a Cover's renderable address: the wire `cover` is it. The only place a Site
name still appears in cover code is the extraction module (`latest`), where
kagane's and comix's image URLs are claimed by `browserOnlyCoverURL` — kagane
answers a plain fetch with a challenge and
`cross-origin-resource-policy: same-origin`, and `static.comix.to` answers `cross-origin-resource-policy: same-origin`, and `static.comix.to` answers
one with the same Cloudflare challenge its pages serve; with the same challenge its pages serve. Every other Site's CDN answers plain
every other Site's CDN answers plain TLS. Templates render `.Cover` — the TLS.
wire value — never anything else. - **comix cover bytes must arrive by direct navigation, not an in-page fetch:**
- **Web UI also owns:** session-gated `GET /install/{manga,novel}-bookmark.user.js` its Series page sets `cross-origin-embedder-policy: require-corp`, which
(renders the bindmounted script with the acting Reader's derived credential fails a page-context fetch of `static.comix.to`.
substituted in — the credential never appears in page markup, the address - With no browser configured, kagane and comix Covers are simply absent;
bar, or a redirect; `?download=1` adds `Content-Disposition: attachment` for novelfull still gets one whenever its page answers a plain request.
mobile Violentmonkey, which ignores a `.user.js` navigation) and
`POST /rotate-token` (atomic epoch bump + hash ### `updated_at` drives list order — `Store.Upsert`
rewrite; invalidates every installed copy, so the panel warns to reinstall
on all devices). The server applies its own timestamp only when the row is new or
- **Owner-only admin page (`internal/web/admin.go`, issue #102):** `GET /admin` `last_chapter_num` changes, else it keeps the stored value. **Favouriting a
carries the Reader roster (sessions, Sighting counters, `POST series, or a newly published chapter arriving, must not reorder the list** —
/readers/{id}/revoke` and `POST /readers/{id}/clear-marks`) and Poll Lane only real reading progress moves a row. Consequently `PUT` returns the row **as
status (`GET /ui/admin/lanes`, self-refreshing every 30s). Every route that stored** and clients must adopt that response rather than their own payload.
reaches past the acting Reader is listed in `adminRoutes()` and wrapped in
`requireOwner` at registration — add a route there, not a check inside a ### Lifecycle buckets — `status` on each bookmark
handler; `web.AdminPatterns()` is what the gate test walks. A non-owner gets
404, never 403. Lane figures come from the running poller through the `reading` | `archived` | `finished`, orthogonal to `favorite`. Archived and
`web.LaneReporter` seam (`latest.Poller.LaneStatus`), never from a table: a finished appear only in their own tab, never in All, Updated, Favourites or the
nil reporter or a Lane that has not finished a pass renders "no data yet" recent strip. The poller keeps checking archived series and skips finished ones.
rather than zeroes. `main.newRouter` takes the reporter as an interface and
converts a nil `*Poller` to a nil interface — a typed nil would make the page - `finished` is settable only from the web UI; `PUT /bookmarks/{key}` rejects
claim a poller exists. it with 400.
The one owner comparison left outside `requireOwner` is in `index` - **An empty incoming status means "keep the stored one"**, and it is resolved
on the `VALUES` side of `Store.Upsert`, not in the conflict clause:
`excluded.*` is the post-evaluation row, so a default applied there would
wipe the bucket on every PUT from a client predating the column.
### Config — `Config` / `loadConfig` / `loadLatestPoll` in `backend/main.go`
That function is the complete list of env vars, their defaults, and which are
required. What it can't tell you:
- `PUBLIC_BASE_URL` must be an absolute origin because every Cover URL on the
wire is built from it and the userscript renders on a Site's origin.
- `BROWSER_WS_URL` **must be a tailnet IP, never a hostname** — Chrome's
DevTools handler 500s `/json/version` for any Host that isn't an IP or
`localhost`. Unset (the default) disables browser polling.
- `USERSCRIPT_PATH` / `NOVEL_USERSCRIPT_PATH` are bindmounted files; the
`__API_TOKEN__` placeholder inside them is substituted with the requesting
Reader's credential at serve time.
- Pace is per Site in the registry, not env. The
`_COOLDOWN`/`_BROWSER_COOLDOWN`/`_INTERVAL`/`_BATCH`/`_STAGGER` knobs are
gone on purpose.
- The 1h rest for browser Sites is safe on documented grounds: a challenged
page costs seconds of a serialized single-tab browser, free-plan zones carry
no bot score and no published per-IP rate input, and `cf_clearance` expires
in 30 minutes, so every cadence at or above 1h re-solves anyway.
### Userscript install & rotation — `internal/userscript`, `internal/token`
Session-gated `GET /install/{manga,novel}-bookmark.user.js` renders the
bindmounted script with the acting Reader's derived credential substituted in,
so the credential never appears in page markup, the address bar, or a redirect.
`?download=1` adds `Content-Disposition: attachment` for mobile Violentmonkey,
which ignores a `.user.js` navigation. `POST /rotate-token` is an atomic epoch
bump plus hash rewrite and invalidates every installed copy — the panel must
keep warning to reinstall on all devices.
### Owner-only admin — `internal/web/admin.go`
- **Every route reaching past the acting Reader is listed in `adminRoutes()`
and wrapped in `requireOwner` at registration** — add it there, not as a
check inside a handler; `web.AdminPatterns()` is what the gate test walks. A
non-owner gets 404, never 403.
- The one owner comparison left outside the gate is in `index`
(`view.Owner = readerID == h.store.OwnerID()`): it gates a link, not an (`view.Owner = readerID == h.store.OwnerID()`): it gates a link, not an
endpoint, so it is a rendering decision a registration-time wrapper cannot endpoint, so it is a rendering decision a registration-time wrapper cannot
express — do not "unify" it into the gate. express. Do not "unify" it into the gate.
A Lane pass that returns before computing its figures (refusal backoff, - Lane figures come through the `web.LaneReporter` seam
sidecar down) carries the previous pass's due count and gap forward rather (`latest.Poller.LaneStatus`), never a table. `main.newRouter` takes the
than recording zeroes; a Lane that has never reached a pace renders no gap at reporter as an interface and converts a nil `*Poller` to a nil interface — a
all. `Checked` next to `Due` is what separates a stopped Lane from a quiet typed nil would make the page claim a poller exists.
one, so neither figure may be dropped from the row. - A pass that returns before computing figures (refusal backoff, sidecar down)
Due-without-Checked is *not* by itself a stall: a browser Lane under both carries the previous pass's numbers forward rather than recording zeroes.
wake thresholds sets `LaneState.Asleep` at the on-demand gate and renders - **`Checked` next to `Due` is what separates a stopped Lane from a quiet one**,
"browser asleep" instead of "not checking", and never counts toward so neither may be dropped from the row.
`Attention`. That is the commonest healthy state for kagane, comix and - Due-without-Checked is **not** by itself a stall: a browser Lane under both
novelfull — one due Series, nothing checked — so spending the stall mark on wake thresholds sets `LaneState.Asleep` and renders "browser asleep", and
it would train the owner to ignore the mark that matters. never counts toward `Attention`. That is the commonest healthy state for
kagane, comix and novelfull, so spending the stall mark on it would train the
owner to ignore the mark that matters.
+14 -5
View File
@@ -47,6 +47,15 @@ func fetchCoverBytes(ctx context.Context, cover string, browser BrowserCoverFetc
// inject it to exercise hostile DNS results without touching the live network. // inject it to exercise hostile DNS results without touching the live network.
type CoverResolver func(context.Context, string) ([]netip.Addr, error) type CoverResolver func(context.Context, string) ([]netip.Addr, error)
// maxCoverBytes caps one cover, separately from the series-page maxBodyBytes:
// a cover is a bounded binary asset, not a text page, and 4 MiB rejected 12%
// of asurascans covers measured 2026-08-17 (p90 4.52 MB, max 8.57 MB — two of
// the three over-cap files were JPEGs, not the animated GIF of issue #71).
// 10 MiB is ~18% headroom over that worst case and matches the GitHub and
// Discord image limits; see docs/research/gif-maximum-byte-size.md. GIF itself
// has no maximum size, so this number is policy, not format.
const maxCoverBytes = 10 << 20
// TLSCoverFetcher retrieves image bytes with the standard HTTPS client. Unlike // TLSCoverFetcher retrieves image bytes with the standard HTTPS client. Unlike
// TLSFetcher, it does not need a browser fingerprint: cover hosts are public // TLSFetcher, it does not need a browser fingerprint: cover hosts are public
// CDNs and the response is accepted only after the destination gate passes. // CDNs and the response is accepted only after the destination gate passes.
@@ -152,15 +161,15 @@ func (f *TLSCoverFetcher) Fetch(ctx context.Context, sourceURL string) ([]byte,
if !ok { if !ok {
return nil, "", fmt.Errorf("fetch cover: unsupported content type %q", raw) return nil, "", fmt.Errorf("fetch cover: unsupported content type %q", raw)
} }
if resp.ContentLength > maxBodyBytes { if resp.ContentLength > maxCoverBytes {
return nil, "", fmt.Errorf("fetch cover: response exceeds %d bytes", maxBodyBytes) return nil, "", fmt.Errorf("fetch cover: response exceeds %d bytes", maxCoverBytes)
} }
body, err := io.ReadAll(io.LimitReader(resp.Body, maxBodyBytes+1)) body, err := io.ReadAll(io.LimitReader(resp.Body, maxCoverBytes+1))
if err != nil { if err != nil {
return nil, "", fmt.Errorf("read cover: %w", err) return nil, "", fmt.Errorf("read cover: %w", err)
} }
if len(body) > maxBodyBytes { if len(body) > maxCoverBytes {
return nil, "", fmt.Errorf("fetch cover: response exceeds %d bytes", maxBodyBytes) return nil, "", fmt.Errorf("fetch cover: response exceeds %d bytes", maxCoverBytes)
} }
return body, contentType, nil return body, contentType, nil
} }
+24 -1
View File
@@ -171,7 +171,7 @@ func TestCoverFetcherRejectsOversizedBody(t *testing.T) {
var calls int var calls int
client := &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) { client := &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) {
calls++ calls++
response := coverResponse(http.StatusOK, "image/webp", "", bytes.Repeat([]byte("x"), maxBodyBytes+1)) response := coverResponse(http.StatusOK, "image/webp", "", bytes.Repeat([]byte("x"), maxCoverBytes+1))
response.ContentLength = -1 response.ContentLength = -1
return response, nil return response, nil
})} })}
@@ -187,6 +187,29 @@ func TestCoverFetcherRejectsOversizedBody(t *testing.T) {
} }
} }
// Covers between the series-page cap and the cover cap must be accepted: the
// 4 MiB page cap rejected 12% of asurascans covers (issue #71).
func TestCoverFetcherAcceptsCoverOverPageCap(t *testing.T) {
body := bytes.Repeat([]byte("x"), maxBodyBytes+1)
client := &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) {
return coverResponse(http.StatusOK, "image/gif", "", body), nil
})}
fetcher := newCoverFetcher(client, func(context.Context, string) ([]netip.Addr, error) {
return []netip.Addr{netip.MustParseAddr("198.51.100.10")}, nil
})
got, contentType, err := fetcher.Fetch(context.Background(), "https://cdn.example/big.gif")
if err != nil {
t.Fatalf("Fetch rejected a %d-byte cover: %v", len(body), err)
}
if len(got) != len(body) {
t.Fatalf("body = %d bytes, want %d", len(got), len(body))
}
if contentType != "image/gif" {
t.Fatalf("content type = %q, want image/gif", contentType)
}
}
func TestCoverFetcherRejectsNonImage(t *testing.T) { func TestCoverFetcherRejectsNonImage(t *testing.T) {
var calls int var calls int
client := &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) { client := &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) {
+344
View File
@@ -0,0 +1,344 @@
# GIF — maximum byte size of a file
Research note for Gitea issue #71 (backend `maxBodyBytes` = 4 MiB rejects the
8,571,192-byte animated cover GIF at
`https://cdn.asurascans.com/asura-images/covers/a-dragonslayers-peerless-regression.gif`).
All facts fetched live on **2026-08-17**: the GIF89a spec at
`https://www.w3.org/Graphics/GIF/spec-gif89a.txt`, Go stdlib `image/gif`
sources at `/usr/local/go/src/image/gif/reader.go` (Go 1.26.5), Chromium
`blink/renderer/platform/image-decoders/` sources via
`chromium.googlesource.com`, Firefox `image/decoders/nsGIFDecoder2.cpp` via
`hg.mozilla.org`, and cover bytes probed with plain `curl` (desktop Chrome UA;
`HEAD`/ranged `GET`). **No Cloudflare challenge was encountered on any CDN
probe** — every request returned real headers, consistent with the AGENTS.md
note of 2026-07-26 that plain `curl` works against both scan sites from the
dev machine and the VPS.
Every claim carries the URL it came from, or a reproducible command.
Interpretation rather than observation is marked `[INFERENCE]`.
---
## 1. Summary answer table
| Question | Answer | Evidence |
|---|---|---|
| Does the GIF89a spec define a maximum file size? | **No.** There is no file-size field anywhere in the format; the only numeric ceilings are per-field (16-bit screen/image dimensions, 255-byte sub-blocks, 12-bit LZW codes). | §2 |
| Maximum logical screen | 65535 × 65535 pixels (unsigned 16-bit width/height). | §2.1 |
| Number of frames / image descriptors | Unbounded — "An unlimited number of images may be present per Data Stream." | §2.2 |
| Formal max byte size of any single GIF | None. Single-frame worst case ≈ **6.44 GB** (12-bit LZW, max canvas); animated GIFs are **unbounded** because frames are unbounded. | §3 |
| Does the backend's decoder (Go `image/gif`) bound size? | **No.** It reads 16-bit dimensions and allocates `width×height` bytes per frame; a 65535² frame forces a ~4 GiB allocation. No total-size or dimension guard. | §4.1 |
| Do browsers bound on-disk GIF size? | Chromium and Firefox: no on-wire size cap in their GIF readers; Chromium caps *decoded* memory at min(4 B × pixels, platform budget). | §4.3, §4.4 |
| Real cover sizes (asurascans, n=25) | min 190,410 B · median 1,275,082 B · p90 4,524,788 B · max 8,571,192 B · **3/25 > 4 MiB** (two JPEGs and the animated GIF) | §5 |
| Real cover sizes (demonicscans/readermc, n=78) | min 13,298 B · median 63,061 B · max 801,200 B · 0/78 > 4 MiB | §5 |
| Comparable service caps | GitHub: 10 MB for images/GIFs. Discord API: default 10 MiB per file. Wikimedia: 100 MiB upload / 5 GiB host. | §6 |
| Recommended cover cap for #71 | **10 MiB** (separate from the 4 MiB series-page cap). Covers 100% of the 103 observed covers; matches GitHub/Discord calibration; ≤ 20 MiB worst-case transient per concurrent fetch+serve on a 1974 MiB swapless VPS. | §7 |
---
## 2. What the GIF89a specification actually bounds
Source: `https://www.w3.org/Graphics/GIF/spec-gif89a.txt` (fetched 2026-08-17).
### 2.1 Fixed-width fields — the only hard ceilings
The format is a stream of fixed-width blocks; the numeric fields that *do*
have a ceiling are all 16-bit unsigned, little-endian ("multi-byte numeric
fields are ordered Least Significant Byte first", §4 of the spec):
- **Logical Screen Width / Height** — "Unsigned" 2-byte fields (§18, Logical
Screen Descriptor) → maximum **65535 × 65535** pixels.
- **Image Left / Top Position, Image Width / Height** — "Unsigned" 2-byte
fields (§20, Image Descriptor). Each image "must fit within the boundaries
of the Logical Screen" (§20a), so an image cannot exceed the 65535² canvas
even though its own fields would allow it.
- **Data sub-blocks** — "A data sub-block may contain from 0 to 255 data
bytes" (§15); each sub-block is preceded by a 1-byte size field and the
stream is terminated by a 0x00 Block Terminator (§16). This bounds a
*chunk*, not the stream.
- **Global/Local Color Tables** — optional, "3 x 2^(Size of Global Color
Table+1)" bytes with a 3-bit size field → at most 3 × 2⁸ = **768 bytes**
each (§19, §21).
- **LZW codes** — "The output codes are of variable length, starting at
<code size>+1 bits per code, **up to 12 bits per code**. This defines a
maximum code value of 4095 (0xFFF)" (Appendix F, COMPRESSION, rule 4).
- **Trailer** — a single byte, fixed value 0x3B, "indicating the end of the
GIF Data Stream" (§27).
### 2.2 What is unbounded
- **Number of images (frames).** §20a, verbatim: "This block is REQUIRED for
an image. Exactly one Image Descriptor must be present per image in the
Data Stream. **An unlimited number of images may be present per Data
Stream.**"
- **The Data Stream itself.** The grammar in Appendix B is
`<GIF Data Stream> ::= Header <Logical Screen> <Data>* Trailer`, and the
spec states "the entity Data … may be repeated any number of times,
including 0 times." There is **no field anywhere that carries a file size,
byte count, frame count, or total-length value**. §13 (Block Sizes) only
defines sizes *within* blocks.
### 2.3 Verdict
**The GIF89a specification defines no maximum file size.** The only hard
bounds are per-field: 65535×65535 pixels per screen/image, 255 bytes per
sub-block, 12 bits per LZW code, and one trailer byte. A compliant decoder
must process whatever stream the blocks describe. Any byte ceiling a
particular GIF actually hits is therefore *implicit* — 16-bit dimensions,
LZW code width, decoder memory, or an external policy — never something the
format itself enforces. `[INFERENCE]` This is why real-world GIFs cap out at
"a few GB at most" and every service that wants a bound has to impose one
itself (see §6; Wikimedia explicitly documents that a 4 GiB host limit was a
storage-representation artifact of 32-bit integers, `phab:T191805`, not a
format limit).
---
## 3. Theoretical worst case
### 3.1 Single frame, maximal canvas, 8-bit pixels
| Quantity | Value | Derivation |
|---|---|---|
| Max pixels | 4,294,836,225 | 65535 × 65535 |
| Raw 8-bit palette-index raster | 4,294,836,225 B ≈ **4.29 GB / 4.00 GiB** | 1 byte per pixel (Table Based Image Data, §22; Go's `image.Paletted` uses exactly 1 byte/pixel) |
| LZW worst case | ≈ **6.44 GB / 6.00 GiB** | codes ≤ 12 bits each (Appendix F), at most ~1 code per pixel for incompressible data → ≤ 12 bits/px = 1.5 B/px → 4,294,836,225 × 1.5 B |
| Sub-block overhead | ≈ +25.3 MB | every ≤255-byte chunk carries a 1-byte size field (§15): ⌈6,442,254,338 / 255⌉ ≈ 25,263,743 size bytes, + 1 block terminator |
| Fixed overhead | ≈ +1.6 KB | header 6 B (§17) + logical screen descriptor 7 B (§18) + global color table ≤ 768 B (§19) + image descriptor 10 B (§20) + local color table ≤ 768 B (§21) + LZW minimum code size 1 B (§22) |
So a **single maximal-frame GIF cannot exceed ≈ 6.47 GB on the wire**
(12-bit LZW bound), and LZW being lossless means the real byte count depends
entirely on image content — the same canvas can be a few KB (flat color) or
~6 GB (noise).
Two caveats, both marked `[INFERENCE]`:
- The "1.5 B/px" figure assumes ~one emitted code per pixel. An encoder is
permitted to emit a Clear code at any point (Appendix F: "The Clear code
can appear at any point in the image data stream"), so a
pathological-but-compliant encoder emitting clear+pixel per pixel reaches
~24 bits/px ≈ 12.9 GB for the max canvas. Real encoders do not do this;
12-bit/px is the practical bound.
- The spec's deferred-clear note (cover sheet) explicitly allows an encoder
to keep using a full table at 12-bit codes without clearing, so the 12-bit
cap holds for the whole stream, it cannot "grow" past 12 bits.
### 3.2 Animated GIFs: unbounded
Every frame is one Image Descriptor, each bounded by the 65535² canvas, but
the *count* of frames is unbounded (§2.2). Total bytes = sum over frames —
therefore **there is no finite maximum byte size for an animated GIF** in
the format. The only thing that stops a real one is decoder memory, a
service cap, or disk space. `[INFERENCE]` This is the category the issue #71
cover falls into: it is an animated GIF (NETSCAPE2.0 loop extension found at
offset 0x310 of the file, verified 2026-08-17 by a ranged GET), and its
8,571,192 bytes are ~2.04× the current 4 MiB backend cap.
---
## 4. Decoder-side real limits
### 4.1 Go `image/gif` (the backend's decoder path, stdlib)
Source: `/usr/local/go/src/image/gif/reader.go`, Go 1.26.5.
- Dimensions are read as little-endian uint16 — `left/top/width/height :=
int(d.tmp[N]) + int(d.tmp[N+1])<<8` (reader.go:490-493) — so the format
ceiling 65535 applies, and nothing smaller is enforced.
- The only geometric check is that each frame fits inside the logical
screen: `if left+width > d.width || top+height > d.height` →
`errors.New("gif: frame bounds larger than image bounds")` (reader.go:512-513).
- **There is no file-size, byte-count, frame-count, or pixel-count guard.**
Each frame allocates `image.NewPaletted(...)` (reader.go:515) — a
`[]byte` of width×height — so decoding one legal 65535² frame attempts a
**~4.29 GB allocation**. `DecodeAll` (reader.go:603-605) additionally
retains every frame's `Pix` slice for the lifetime of the returned `*GIF`.
- `[INFERENCE]` On the 1974 MiB swapless VPS (root AGENTS.md), decoding such
a file would OOM rather than error cleanly; nothing in stdlib protects
the process. This matters for §7: the backend stores cover bytes without
decoding them (see §5.3), so the fetch path never triggers this — but any
future "validate/re-encode server-side" scheme would.
- Grep for `MaxInt|limit|too large|bounds` in reader.go: the only hits are
the frame-bounds check above and the `tmp [1024]byte` scratch buffer
(reader.go:109); no size caps exist.
### 4.2 giflib / libgif
**Not verified from source.** On 2026-08-17 the giflib sources were not
reachable from this network: `github.com/giflib/giflib` returns 404 (repo
gone/moved), `gitlab.com/giflib/giflib/-/raw/...` answers a Cloudflare
"Just a moment…" challenge, and the SourceForge project download path
404s. No limit claim about giflib is made here. `[INFERENCE]` giflib is
widely known to be allocation-driven with no dimension cap, but that is not
checked against source and is not needed for issue #71 (the backend uses Go
stdlib, not giflib).
### 4.3 Chromium (browser behaviour, first-party source)
- `third_party/blink/renderer/platform/image-decoders/gif/gif_image_reader.cc`
(via `chromium.googlesource.com/chromium/src/+/main/...`, fetched
2026-08-17): **no GIF byte-size or dimension cap found** — grep for
`max|limit|too large|dimension|65535|overflow` matches only license text.
- The base `ImageDecoder` caps *decoded memory*, not transfer size:
`CalculateMaxDecodedBytes` computes `min(4 * num_pixels, platform_max_decoded_bytes)`
(8 bytes/pixel for high-bit-depth), and the header comment says "Ignoring
this limit can cause excessive memory use or even crashes on low-memory
devices"
(`image_decoder.cc:94-117`, `image_decoder.h:545-549`). The GIF reader
itself is untouched by this — it is a decoded-buffer budget.
- Practical consequence `[INFERENCE]`: a browser will happily download and
store a multi-GB GIF from its own cache perspective; Chromium only limits
what it *decodes* into pixels.
### 4.4 Firefox
`image/decoders/nsGIFDecoder2.cpp` (via `hg.mozilla.org/mozilla-central/
raw-file/tip/...`, fetched 2026-08-17): **no dimension or size limit**; the
only guards are on LZW code width (`MAX_BITS` = 12, "maximum codeword size
of 12 bits") and the decode stack. Nothing bounds the on-disk byte size.
### 4.5 Summary
No mainstream decoder enforces a byte-size ceiling; they stop at the 16-bit
dimension ceiling (Go, by construction) or at decoded-memory budgets
(Chromium) or nowhere (Firefox). A GIF's byte size is policed only by
*storage* policies — which is what §6 calibrates and §7 sets.
---
## 5. Practical distribution — what real manga covers weigh
Probed **2026-08-17** with `curl -sI` (HEAD) and ranged GETs, desktop Chrome
UA. No Cloudflare block on any request. Sample = covers *as the backend
would fetch them* (the `og:image`/page-listed cover URL), not thumbnails we
chose by hand.
### 5.1 Exact commands
```sh
# asurascans.com — harvest cover URLs from the homepage, then HEAD each
curl -s -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/126.0" https://asurascans.com/ -o home.html
grep -oE 'https://cdn\.asurascans\.com/asura-images/covers/[^"&\\< ]+\.(webp|gif|jpg|jpeg|png)' home.html \
| sort -u | grep -v '\-400\.' | head -25 > sample.txt # one full-res cover per series, no -400 thumbs
while read -r u; do curl -s -A "…Chrome/126.0" -I "$u" | tr -d '\r' \
| grep -iE '^content-length:'; done < sample.txt
# demonicscans.org — covers live on readermc.org (ADR-0007), URLs contain spaces/UTF-8
curl -s -A "…Chrome/126.0" https://demonicscans.org/ -o demonic.html
grep -oE 'src="https://readermc\.org/images/thumbnails/[^"]+"' demonic.html | tr -d 'src="' > demonic.txt
# …plus og:image from 5 manga pages (Catastrophic-Necromancer, Magic-Emperor, …)
# each URL percent-encoded per path segment (urllib.parse.quote, safe=':/') before HEAD
```
### 5.2 asurascans — 25 full-res covers (mixed formats)
Homepage fetched 200 (664,700 B). All 25 returned `200` with a real
`Content-Length`. Distribution:
| Statistic | Bytes |
|---|---|
| n | 25 |
| min | 190,410 |
| median | 1,275,082 |
| p90 | 4,524,788 |
| max | 8,571,192 |
| mean | 1,943,651 |
| **> 4 MiB (4,194,304)** | **3 (12%)** — `a-dragonslayers-peerless-regression.gif` 8,571,192 (the issue #71 cover, animated: NETSCAPE2.0 at 0x310, 550×733, 256 colors); `bad-born-blood.3008f6.webp` 4,524,788 `image/jpeg`; `ending-maker.cfbf53.webp` 4,619,303 `image/jpeg` |
Notes: the CDN serves `Content-Type` by stored bytes, not by URL extension
(the `.webp` URLs return `image/png`, `image/jpeg`, or `image/webp` — the
sample spans all four of `png/jpeg/webp/gif`). Two of the three over-cap
files are **not GIFs**, so the current 4 MiB cap already silently drops 12%
of asura covers of any format. p90 itself (4.52 MB) exceeds the cap.
### 5.3 demonicscans — 78 covers on readermc.org
78 unique cover URLs (73 from the homepage's `/images/thumbnails/` plus 5
`og:image` values from manga pages — demonicscans publishes the thumbnail
file as the full cover, so that is exactly what the backend would fetch).
**78/78 returned 200 with a real Content-Length** (spaces and UTF-8 in the
filenames were percent-encoded per path segment; the homepage's raw HTML
carries `’`-style mojibake for curly quotes, which was repaired by
latin-1→utf-8 re-encoding before probing).
| Statistic | Bytes |
|---|---|
| n | 78 |
| min | 13,298 |
| median | 63,061 |
| p90 | 206,994 |
| max | 801,200 |
| mean | 110,038 |
| > 4 MiB | 0 |
### 5.4 Reading
`[INFERENCE]` asurascans covers are the heavy tail (median 1.3 MB, top
decile > 4 MiB, occasional ~5–9 MB), demonicscans covers are tiny (all
< 0.8 MB). A cover cap must be chosen against the *asura* distribution —
the 8.57 MB animated GIF is not a freak one-off outlier; the 90th
percentile already crosses 4 MiB and two JPEGs sit between 4.5–4.7 MB.
---
## 6. Comparable documented byte caps (first-party docs only)
| Service | Cap | Source (fetched 2026-08-17) |
|---|---|---|
| GitHub (issues/PR comments) | **10 MB for images and gifs**; 25 MB other files; 10/100 MB video | `https://docs.github.com/en/get-started/writing-on-github/working-with-advanced-formatting/attaching-files` — "The maximum file size is: 10MB for images and gifs … 25MB for all other files" |
| Discord (API uploads) | default **10 MiB per file**, higher with Nitro / boost tier | `https://discord.com/developers/docs/reference#uploading-files` — "The file upload size limit applies to each file in a request. The default limit is `10 MiB` for all users" (help-center article `support.discord.com/hc/en-us/articles/115002935588` exists but answered 403 from this network on the probe date, so its figures were not verified here) |
| Wikimedia Commons | **100 MiB** upload limit; hosting up to **5 GiB**; GIF thumbnails limited to **100 megapixels**; prior 4 GiB host cap was a 32-bit storage artifact (phab:T191805) | `https://commons.wikimedia.org/wiki/Commons:Maximum_file_size` |
| MDN | nothing — MDN documents no byte-size limit for images; browsers impose none (see §4.3–4.4) | `[INFERENCE]` from absence in the platform docs read in §4 |
Calibration takeaway: two major platforms independently land on **~10 MB**
as the ceiling for an uploadable image/GIF (GitHub exactly 10 MB, Discord
exactly 10 MiB), with Wikimedia the outlier at 100 MiB/5 GiB because it is a
media *archive*. A 10 MiB cover cap is therefore squarely inside industry
normal.
---
## 7. Recommendation for issue #71
**Raise the cover cap to 10 MiB (10,485,760 B) — as a separate constant, not
by moving the shared one.**
Why:
- **Fits the measured reality.** The largest observed cover is 8,571,192 B
(the issue's animated GIF) = 82% of 10 MiB; 10 MiB covers **100% of the
103 sampled covers** and the *entire* asura distribution, including its
heavy tail. 4 MiB rejects 12% of asura covers (two of them plain JPEGs).
- **Matches industry calibration** (§6): GitHub 10 MB images/GIFs, Discord
10 MiB default. A 10 MiB cap is a number every engineer recognizes, and
it leaves ~18% headroom over the current worst observed file.
- **Costs little on the target hardware.** The backend buffers cover bytes
whole during fetch (`backend/internal/latest/cover.go`: `ContentLength >
maxBodyBytes` rejection at :155, then `io.ReadAll(io.LimitReader(…,
maxBodyBytes+1))` at :158) and loads the full body per `GET /covers/…`
(`backend/internal/api/handlers.go`, `Cover` → `w.Write(body)`). Worst
case per concurrent fetch **+** serve is therefore 2 × cap = 20 MiB; even
ten of each concurrently is ~200 MiB of a 1974 MiB swapless VPS (~10%),
and the browser unit (471 MiB, root AGENTS.md) is no longer on that box.
The 4 MiB series-page cap is *not* the issue — measured pages run
100 KB–1.2 MB (`backend/internal/latest/fetch.go` comment) — so keep it.
- **The cap is a separate knob.** Today one `const maxBodyBytes = 4 << 20`
(`backend/internal/latest/fetch.go:17`) gates *both* series pages and
covers (`cover.go` references it). Raising it wholesale would loosen the
page-side memory guard for no benefit; a cover-specific constant (e.g.
`maxCoverBytes = 10 << 20`) keeps the two policies independent. The fetch
already double-checks `ContentLength` and the post-`LimitReader` length,
so a larger constant changes nothing else.
Alternatives and their costs:
| Option | Cost |
|---|---|
| Keep 4 MiB | 12% of asura covers (incl. non-GIF JPEGs) never stored — current bug, silent missing covers. |
| 16 MiB cap | 2× headroom over the observed max for future GIFs; +60% worst-case transient memory vs 10 MiB; diverges from the GitHub/Discord 10 MB calibration. |
| Server-side re-encode / downscale covers | Requires decoding → Go `image/gif` allocates width×height per frame with **no guard** (§4.1); a legal 65535² GIF forces a ~4.29 GB allocation on a 1974 MiB swapless box — OOM, not an error. Also mutates bytes, which the store treats as immutable/content-addressed (ADR-0007). Highest risk, no upside at this scale. |
| No cap | Unbounded transient memory and disk; rejected outright. |
Decision is the user's; on the evidence, **10 MiB for covers, 4 MiB for
pages** is the defensible middle.
+94 -84
View File
@@ -1,93 +1,103 @@
Guidance for OpenCode (and Claude Code) working under `userscript/`. See root `AGENTS.md` for the project-wide architecture diagram, hard constraints, and design system. Scope: `userscript/`.
### Userscript structure (single IIFE, `manga-bookmark.user.js`) Each entry names the code that holds the truth. The prose is only what the code
cannot tell you: rationale, invariants a refactor would break, and dated
observations about sites we don't control.
1. **Site adapters** — one per host, `detect(location, document)` return page `type` + IDs. Identify type/IDs from **URL regex** (most stable); pull `title` from **`og:title`** (or the page heading where a site ships no og: tags), not CSS classes. **No adapter reads a cover**: the backend acquires, stores and serves every Cover from its own origin (ADR-0007), the wire's `cover` is already an address on our origin, and `apiPut` strips any `cover` off an outgoing body. ### Structure — single IIFE, `manga-bookmark.user.js`
2. **API client** — `apiGet/apiPut/apiDelete` with bearer header; `localStorage` key `bmgr:manga:cache` for instant render + offline fallback.
3. **Progress logic** — auto-upsert `last_chapter` only when `chapterNum >= stored last_chapter_num` (re-reading old chapters must not regress progress; unparseable -> set current). Manual panel override forces any value.
4. **Retry queue** — every write go through `pushBookmark`/`pushDelete`, so
failed mutation park in `localStorage` (`bmgr:manga:queue`) and replayed on
next navigation, reconnect, or `refresh()`. Entries are markers
(`{key, op, sendStatus, attempts}`), never payloads — body read from
cache at send time, so one entry per key give ordering and coalescing for
free. `sendStatus` is **sticky**: while archive pending, later writes to
that key keep carrying bucket, which stop successful
in-between write from silently un-archiving series. `refresh()` drains
before it fetches and overlays anything still pending, so list never
flaps. 400 drops entry, 401 abort pass and keep queue, and
transient failures retry to cap of 10. Latest-chapter writes deliberately
stay out of queue. See
`docs/superpowers/specs/2026-07-27-offline-retry-queue-design.md`.
5. **UI** — rendered inside **Shadow DOM** root to isolate from site CSS
(critical on mobile). Three tabs (All / Favourites / Archived) and row of
link chips to web UI and both manga sites; `WEB_BASE` sits in CONFIG
block next to `API_BASE`. FAB is `7 × 44` edge tab whose *hit* area
widened to `28 × 72` by invisible `#hit` child; `#fab` must keep
`touch-action: none` and must **not** regain `overflow: hidden`. Since
`touch-action` resolved at gesture start, strip can't be both
browser-scrolled and script-dragged, so `makeDraggable` splits by intent: swipe
from `#hit` scrolls via `window.scrollBy`, hold of `ARM_MS` arms
reposition drag, visible sliver drags with no hold. See
`docs/superpowers/specs/2026-07-28-edge-tab-hitbox-design.md`.
6. **SPA navigation** — Asura is Astro, client-routed on comic/chapter pages: patch `history.pushState`/`replaceState` + listen `popstate`, re-run `detect()` on URL change so auto-update fire without reload. Demonic uses classic reloads (initial `document-idle` run suffice).
### Live URL shapes (verified 2026-07-26, may drift — re-check against live pages before trust) Six parts, in file order: site adapters, API client, progress logic, retry
queue, UI, SPA navigation.
- **asurascans.com**: series `/comics/<slug>` (slug carries trailing **Site adapters** — one per host, `detect(location, document)` returning page
site-wide build-hash suffix, e.g. `-059befe1`, that **rotates on every `type` + IDs.
redeploy**), chapter `/comics/<slug>/chapter/<n>`. `seriesId` must strip
hash (`/-[0-9a-f]{8}$/`, `stripBuildHash` in userscript, - Identify type and IDs from **URL regex**, which is the most stable surface a
`asuraBuildHash` in backend); URLs keep full slug — stale-hash site exposes; take `title` from **`og:title`** (or the page heading where a
URLs 302 to current ones. Astro-rendered; chapter links present in raw site ships no og: tags), never CSS classes.
server HTML. - **No adapter reads a cover.** The backend acquires, stores and serves every
- **demonicscans.org**: series `/manga/<slug>` (slug may URL-encode punctuation, e.g. `%2527` for `'`), chapter `/title/<slug>/chapter/<n>/<page>` (older `chaptered.php?manga=<id>&chapter=<n>` form still exists as redirect, what series-page chapter-list anchors link through). Cover from its own origin, the wire `cover` is already an address there, and
Encodings (incl. triple-encoded punctuation like `%25252D`) identical `apiPut` strips any `cover` off an outgoing body.
on /manga/ and /title/ pages, so decode-once seriesIds match — verified
2026-07-28. **Progress logic** — auto-upsert `last_chapter` only when
- **comix.to**: series `/title/<id>-<slug>`, chapter `chapterNum >= stored last_chapter_num`; unparseable sets the current value.
`/title/<id>-<slug>/<uploadId>-chapter-<n>`. Only the leading `<id>` is Re-reading an old chapter must not regress progress. A manual panel override
identity — the slug re-renders when a series is renamed (`comixSeriesId`). forces any value.
An SPA that **never rewrites `og:title`**: the server-rendered head keeps
whatever document loaded first, so on a cold load `og:title` is the homepage's **Retry queue** — every write goes through `pushBookmark`/`pushDelete`.
"Comix — Read Comics online for free" and after an in-page hop it is the
*previous* series' name. `document.title` is the one thing client routing does - Entries are markers (`{key, op, sendStatus, attempts}`), **never payloads**:
update, so titles come from there, with the chapter page's `" · Ch.<n>"` tail the body is read from cache at send time, so one entry per key gives ordering
stripped. It publishes no `og:image` either, which is one of the reasons cover and coalescing for free.
- `sendStatus` is **sticky** — while an archive is pending, later writes to that
key keep carrying the bucket. Without it a successful in-between write
silently un-archives the series.
- `refresh()` drains before it fetches and overlays anything still pending, so
the list never flaps.
- 400 drops the entry, 401 aborts the pass and keeps the queue, transient
failures retry to a cap. Latest-chapter writes deliberately stay out of the
queue.
**UI** — rendered inside a **Shadow DOM** root to isolate it from site CSS,
which is critical on mobile.
- The FAB's *hit* area is widened by an invisible `#hit` child. `#fab` must keep
`touch-action: none` and must **not** regain `overflow: hidden`.
- `touch-action` is resolved at gesture start, so the strip cannot be both
browser-scrolled and script-dragged. `makeDraggable` therefore splits by
intent: a swipe from `#hit` scrolls via `window.scrollBy`, a hold of `ARM_MS`
arms a reposition drag, and the visible sliver drags with no hold.
**SPA navigation** — Asura is Astro and client-routes on comic/chapter pages, so
`history.pushState`/`replaceState` are patched and `popstate` listened to, and
`detect()` re-runs on URL change. Demonic uses classic reloads, where the
initial `document-idle` run suffices.
### Live URL shapes
Encoded in the adapters; the notes below are the parts a reader of the regex
would get wrong. **Verified 2026-07-26 unless dated otherwise — sites drift, so
re-check against a live page before trusting any of it.**
- **asurascans.com** — the series slug carries a site-wide build-hash suffix
(e.g. `-059befe1`) that **rotates on every redeploy**, so `seriesId` must
strip it (`stripBuildHash` here, `asuraBuildHash` in the backend) while URLs
keep the full slug — stale-hash URLs 302 to current ones.
- **demonicscans.org** — slugs may URL-encode punctuation, and the older
`chaptered.php?manga=<id>&chapter=<n>` form still exists as a redirect, which
is what series-page chapter-list anchors link through. Encodings (including
triple-encoded punctuation like `%25252D`) are identical on `/manga/` and
`/title/` pages, so decode-once seriesIds match (verified 2026-07-28).
- **comix.to** — only the leading `<id>` is identity; the slug re-renders when a
series is renamed (`comixSeriesId`). It is an SPA that **never rewrites
`og:title`**: the server-rendered head keeps whatever document loaded first,
so on a cold load `og:title` is the homepage's name and after an in-page hop
it is the *previous* series'. `document.title` is the one thing client routing
updates, hence titles come from there with the chapter page's `" · Ch.<n>"`
tail stripped. It publishes no `og:image` either, one of the reasons cover
acquisition moved to the backend. acquisition moved to the backend.
- **kagane.to**: series `/series/<uuid>`, reader - **kagane.to** — reader URLs carry no chapter number, so the number comes out
`/series/<uuid>/reader/<bookUuid>`. Reader URLs carry no chapter number, so of `og:title`. Two shapes exist, `"<Series> - Chapter <n>[ - Episode <n>]"`
the number comes out of `og:title`. Two shapes exist: `"<Series> - Chapter and `"<Series> - Volume <v> Chapter <n>"`; both must yield a bare series
<n>[ - Episode <n>]"` and, for volume-numbered series, `"<Series> - Volume <v> title, or the volume tail lands in the bookmark's title. Its covers are
Chapter <n>"` with no episode name — both must yield a bare series title, or challenge- and CORP-protected, so nothing outside kagane.to can load one —
the volume tail lands in the bookmark's title. the panel renders the backend's cover address like every other Site.
Its covers are challenge- and CORP-protected, so nothing outside kagane.to can - **novelfull.com** (novel script) — no `og:*` tags at all, so the title comes
load one directly; the panel renders the backend's own cover address like every from `h3.title` (series) or `a.truyen-title` (chapter).
other Site. Behind a Cloudflare JS challenge, so the backend polls it - **lightnovelworld.net** (novel script) — chapter paths are flat at the site
through the headless browser. root and their slug is a **Chapter Slug, not an identity**: a Series may
- **novelfull.com** (novel script): series `/<slug>.html`, chapter publish under several. The Series address is read off the page's
`/<slug>/chapter-<n>[-<title-slug>].html`. No `og:*` tags at all — title from
`h3.title` (series) or `a.truyen-title` (chapter); the script reads no cover.
Behind a Cloudflare JS challenge no TLS fingerprint
clears, so the backend polls it through the headless browser.
- **lightnovelworld.net** (novel script): series `/novel/<slug>/`, chapter
`/<slug>-chapter-<n>/` — flat, at the site root. The chapter path's slug is a
Chapter Slug, not an identity: the Series address is read off the page's
`a[aria-label='All Chapter']` (fallback: the BreadcrumbList's second crumb), `a[aria-label='All Chapter']` (fallback: the BreadcrumbList's second crumb),
and a Series may publish under several Chapter Slugs. A chapter page with no and a chapter page with no pointer resolves to `other` so no Bookmark is
pointer resolves to `other`, so no Bookmark is offered. `h1.entry-title` is offered. The client runs **no latest-chapter scan** for this Site —
the clean title on a series page and `<Title> Chapter <n>` on a chapter page. `computeLatestChapter` yields null and `backgroundRefreshLatest` skips it
Its series page lists every chapter with an before any fetch — because the backend Poll's one-hour cooldown dominates the
absolute href, so the backend polls it with the plain TLS client. client's four-hour throttle, so a scan would add no freshness while having to
The client performs no latest-chapter scan for this Site: the Poll's truncate at the page's wpdiscuz thread, a public write surface.
one-hour cooldown dominates the client's four-hour throttle, so a scan
would add no freshness, and the page's wpdiscuz thread is a public write
surface a scan would have to truncate at. `computeLatestChapter` yields
null here and `backgroundRefreshLatest` skips the Site before any fetch.
### Second script: `novel-bookmark.user.js` ### Second script — `novel-bookmark.user.js`
A copy of the manga script with two adapters, `LIBRARY = "novel"` and A copy of the manga script with two adapters, `LIBRARY = "novel"` and
`STORE_PREFIX = "bmgr:novel:"`. No migration loop (this script has no previous `STORE_PREFIX = "bmgr:novel:"`. No migration loop, because this script has no
installation to carry keys over from). Installed alongside the manga script; previous installation to carry keys over from. Installed alongside the manga
both write to the same backend with the same `LIBRARY` column discriminating script; both write to the same backend, discriminated by `LIBRARY`.
them.