Fixes five reported symptoms across comix.to and kagane.to. Diagnosing them turned up two latent bugs underneath, both of which had to be fixed for the kagane cover work to function at all.
## Reported symptoms and their causes
| # | Symptom | Cause |
|---|---------|-------|
| 1 | comix bookmark titled `Comix - Read Comics online for free` | comix is an SPA that rewrites `document.title` on client routing but never touches the server-rendered `og:title`. The adapter read `og:title`, so a cold load stored the homepage's title. |
| 2 | next comix bookmark gets the *previous* series' title | Same cause. After an in-page hop, `og:title` still holds whatever page loaded first. |
| 3 | comix cover shows the placeholder | comix serves no `og:image` at all, so `coverFromPage()` had nothing to read. |
| 4 | kagane chapter never appears in the bookmark list | Reader URLs carry no chapter number, so it is parsed out of `og:title`. Volume-numbered series render `"<Series> - Volume <v> Chapter <n>"`, which the suffix regex did not match, so `chapterNum` came back null and nothing was recorded. |
| 5 | kagane title includes the chapter, e.g. `SP Baby - Volume 1 Chapter 1` | Same unmatched regex — the tail was never stripped. One fix covers 4 and 5. |
| 6 | kagane cover blocked in the web UI | kagane serves covers behind its Cloudflare challenge **and** with `cross-origin-resource-policy: same-origin`. No `<img>` on the UI's origin can load one even from a browser holding the clearance cookie. Hot-linking cannot be made to work. |
## What changed
**Userscript.** comix titles now come from `document.title` with the chapter page's `" - Ch.<n>"` tail stripped, and the cover is the `img` whose `alt` matches the cleaned title. comix fills `document.title` a beat *after* the URL changes — later than the nav watcher's 300 ms snapshot — so the watcher also re-detects when the `detect()` signature changes, not only when the URL does. The kagane suffix regex takes an optional `Volume <v> ` segment. All three page shapes were captured live on 2026-08-08 and pinned as regression tests.
**Cover proxy.** `Bookmark.CoverURL()` rewrites a stored kagane `og:image` to `/img/kagane/{id}`; templates render `.CoverURL` instead of `.Cover`. The endpoint is session-gated like every other UI route and fetches through the shared headless browser, which is same-origin with kagane and so satisfies both the challenge and the CORP header. Results are memoised in-process, so a cover costs one navigation per deployment lifetime. With `BROWSER_WS_URL` unset the endpoint answers 404 rather than reaching for a nil fetcher — the same degrade-to-userscript behaviour the poller already has.
The image id is matched against a UUID regex before it reaches the browser. That gate is load-bearing rather than tidiness: the cover is a stored client-supplied string, so an unvalidated one turns this endpoint into an SSRF primitive aimed at the deployment's own network. `ServeMux` path-cleans a traversal into a redirect before the handler runs, but the handler does not depend on that, and a test pins it.
## Two latent bugs found underneath
**`BrowserFetcher.run` never let a challenge solve.** It navigated, waited for `body`, read once, and closed the tab — roughly half a second end to end. The Cloudflare interstitial has a `body` too, so `WaitReady` was satisfied by the challenge page itself. This made the challenge *unclearable* rather than merely slow: an interstitial needs several seconds of a live page to solve itself and write clearance into the browser's shared cookie jar, so tearing the tab down first means every subsequent call is challenged exactly like the one before it. `run` now holds one tab and re-reads until the caller's predicate reports an answer, bounded by `challengeTimeout` and the caller's own deadline. Exhausting the budget maps back to the 403 the poller already expects, keeping a challenged site distinct from a broken transport.
**`chromedp/headless-shell` cannot clear kagane's challenge at all.** It is a stripped Chrome build and the tells are structural rather than a header: `navigator.webdriver` is true, the plugin list is empty, and the client hints are Chromium- rather than Chrome-branded. Overriding `webdriver` through CDP was tried on its own and changed nothing.
All measured 2026-08-08 from one IP against the same cover, so the comparisons are like for like:
| Browser | Result |
|---------|--------|
| `chromedp/headless-shell:stable` | never cleared (90 s) |
| `zenika/alpine-chrome` | never cleared — ships Chrome 124, old enough that Cloudflare refuses it and old enough to break chromedp's CDP structs |
| `google-chrome`, default UA | never cleared (60 s) — `--headless=new` advertises `HeadlessChrome` |
| `google-chrome`, stock UA, `TZ=UTC` | never cleared (90 s) |
| `google-chrome`, stock UA, any non-UTC `TZ` | **cleared in ~4 s** |
Both remaining tells are load-bearing, and each was tested in isolation. `chrome/` is a Debian image with `google-chrome-stable`, a UA whose version is read back out of the binary at startup (a hardcoded one would drift out of step with the `Sec-CH-UA` hints on the next Chrome update and become a fresh tell), and no `--enable-automation`.
### The timezone tell: UTC, not a country mismatch
The first pass concluded the zone had to match the egress IP's country. Re-measuring against the actual deployment case shows that was wrong, and the correction is in `1552dd1`.
The original inference read the host's `/etc/timezone` (`Asia/Bangkok`) and assumed a Thai egress. It isn't — this host egresses from an Indonesian IP. `Asia/Bangkok` cleared not because it matched a country but because it simply isn't UTC, and the two share +07, which hid the distinction. Same container, same Indonesian IP:
| `TZ` | Result |
|------|--------|
| `UTC` | never cleared (60 s, **twice**) |
| `Asia/Jakarta` | cleared in 4 s |
| `America/New_York` | cleared in 4 s |
`America/New_York` matches neither the country nor the offset nor the hemisphere and clears just as fast. A UTC clock is itself the bot signal — Cloudflare scores it as the datacenter default — and any real zone satisfies the check. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one, and a deployment that changes region need not keep it in sync.
One sharp edge remains: the usual `-v /etc/localtime:/etc/localtime:ro` does **not** work. Chrome resolves the zone through ICU, which takes the name from that path's symlink target and ignores the file's contents, so glibc reports the host zone while Chrome still reports UTC. `/etc/timezone` carries the name and is mounted instead.
Chrome also binds its DevTools port to loopback and silently ignores `--remote-debugging-address`, which is why headless-shell fronted it with socat. This image does the same, so it stays a drop-in: the compose service keeps the `headless-shell` name and its pinned address, and `BROWSER_WS_URL` is unchanged.
## Verification
```
go test ./... all packages ok
node --test 37 + 12 pass, 0 fail
SMOKE_BROWSER_WS_URL=... go test -run TestSmokeKagane ./internal/latest
TestSmokeKaganeImage PASS (5.29s) fetched 56710 bytes of image/webp
TestSmokeKaganeGet PASS (1.17s) status=200, real chapter-list JSON
```
The smoke test ran against the exact compose configuration — built image, empty `BROWSER_TZ`, `/etc/timezone` mounted, cold profile — hitting real kagane.to. It skips unless `SMOKE_BROWSER_WS_URL` names a sidecar, so `go test ./...` stays hermetic and Docker-only.
A red smoke run means the challenge is not clearing from that IP, which is a live, time-varying fact to re-check rather than necessarily a defect.
## Security invariants
- Auth unchanged. `/img/kagane/{id}` is session-gated by `requireSession`, the same guard as every other UI route.
- Outbound fetch gated: the id is UUID-validated before it reaches the browser, keeping the existing rule that a client-supplied string never selects a fetch target unchecked.
- No new secrets, no new logging of credentials, no change to CORS, sessions, or crypto.
- Templates still escape everything; `.CoverURL` returns a plain string and is not wrapped in `template.HTML`/`URL`.
- One new dependency-free image (`chrome/`) built from Debian plus Google's own apt repo; no new Go modules.
## Deploying
Needs `docker compose build headless-shell`.
**A UTC host must set `BROWSER_TZ`, or kagane silently stops working.** With it unset the sidecar falls back to the host's `/etc/timezone`; on a UTC server that yields UTC, which is the one value that never clears. Any real zone works — `BROWSER_TZ=Asia/Jakarta` for the current deployment. `.env.example` now documents this; it previously did not mention the knob at all.
Only the browser sidecar reads `BROWSER_TZ`. The backend keeps its UTC clock, and stored timestamps are unix ms, so nothing else shifts.
## Deliberately not done
Retry/backoff around the cover proxy, and a panel-side cover fix. The panel renders no covers, and covers cache in-process after the first fetch. Worth adding if kagane starts rate-limiting.
## Correction after review of the deployment case
`1552dd1` was added after the branch was first pushed: the deployment host runs UTC with an Indonesian egress IP, which prompted re-measuring the timezone claim and falsifying it. The earlier commits' reasoning is left intact rather than rebased away, so the diagnostic trail — including the wrong turn and what disproved it — stays readable.
Reviewed-on: #37
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
14 KiB
AGENTS.md
Guidance for OpenCode (and Claude Code) working in this repo.
What this is
Read-progress tracker for two libraries — manga and novels — behind one self-hosted Go backend. Two separate Violentmonkey userscripts inject on-page UI (floating button + slide-in panel) and sync progress, so bookmarks unify across sites and devices:
manga-bookmark.user.js— asurascans.com (current domain; asuracomic.net 301s here), demonicscans.org, comix.to, kagane.to.novel-bookmark.user.js— novelfull.com, lightnovelworld.net.
One backend, one bookmarks table: a kind column (manga|novel) splits the libraries and the web UI switches between them. Rows are keyed <site>:<series_id>.
Hard constraints (drive design — don't violate)
Userscript targets Violentmonkey, so GM_* APIs available, but stay GM-free where plain web APIs suffice — keeps portability across engines:
- Avoid
GM_*unless needed. Prefer pagelocalStorageoverGM_setValue/GM_getValue, on-page UI overGM_registerMenuCommand, plainfetch()overGM_xmlhttpRequestfor cross-origin. - Cross-origin
fetch()work only against CORS-enabled backend. Manga siteshttps://, so backend must be HTTPS (else mixed-content block). - Every site is its own origin with its own
localStorage— a shared remote store is the only way to unify bookmarks. Cloud sync required, not optional. - Userscript run in isolated world, so embedded API token safe from site's JS.
- Cloudflare's block on manga sites IP-reputation-based, not universal — and not reliably reproducible. Verified 2026-07-26: plain
curlfrom both CGNAT dev machine and deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — Cloudflare's bot scoring can flip previously-clean IP without notice. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be verified against live pages (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test. - kagane.to and novelfull.com are the exception to the above — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (
BROWSER_WS_URL) and skips them entirely when that's unset. The four other sites poll fine over plain TLS. - The CDP sidecar must look like a real browser, and stock headless images don't. Measured 2026-08-08 against kagane.to, all from the same IP:
chromedp/headless-shell:stablenever cleared the challenge in 90s (navigator.webdrivertrue, empty plugin list, Chromium-branded client hints — suppressingwebdriveralone changed nothing);zenika/alpine-chromeships Chrome 124, refused outright; real Chrome with the default--headless=newUA never cleared, because the UA saysHeadlessChrome; real Chrome with a stock UA and a non-UTC clock zone cleared in ~4s. Hencechrome/— a Debian image withgoogle-chrome-stable, a version-derived UA, andTZ/BROWSER_TZ. Chrome reads the zone name through ICU from/etc/localtime's symlink target, ignoring the file's contents, so mounting the host's/etc/localtimedoes not work;/etc/timezoneis mounted instead. - UTC is the tell, not a country mismatch. A UTC clock is the datacenter default, so Cloudflare scores it as one; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while
Asia/JakartaandAmerica/New_Yorkboth cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (Asia/Bangkok) rather than the measured egress.BROWSER_TZtherefore needs a plausible zone, not a geolocated one. - A challenged page needs the tab kept open. The interstitial takes seconds to solve and only then writes clearance into the browser's shared cookie jar. Navigate-read-close never clears anything;
BrowserFetcher.runholds one tab and re-reads until the payload arrives.
Architecture
Two Violentmonkey userscripts (isolated world, per-site adapters, localStorage cache)
-- fetch() HTTPS --> reverse proxy (TLS + CORS) --> Go net/http --> Postgres (volume)
Backend-specific architecture (packages, endpoints, poller, config env vars) lives in backend/AGENTS.md. Userscript-specific structure (adapters, retry queue, UI, live URL shapes) lives in userscript/AGENTS.md.
Commands
Backend (cd backend):
- Test all:
go test ./...— needs Docker. Each test package starts a throwawaypostgres:17-alpinecontainer (internal/pgtest). - Single test:
go test -run TestName ./... - Build static binary:
CGO_ENABLED=0 go build
Local stack: docker compose up (bookmark-api + postgres + headless-shell; postgres-data named volume, restart: unless-stopped). The headless-shell service keeps its name but now builds chrome/ — real Google Chrome, for the reason in the hard constraints above.
Live CDP proof (needs a sidecar and network, skipped otherwise):
SMOKE_BROWSER_WS_URL=ws://<host>:<port> go test -run TestSmokeKagane ./internal/latest
— fetches a real kagane cover and chapter list. A red run means the challenge is
not clearing from this IP, which is a live fact to re-check, not necessarily a defect.
Smoke test: curl endpoints with Authorization: Bearer <token>; confirm OPTIONS preflight return CORS headers and /healthz return 200.
Forge: Gitea, not GitHub
origin is self-hosted Gitea instance (gitea.violetcrown.my.id), so gh don't work here — use tea (Gitea CLI) for anything past plain git. Common ones:
- Open PR:
tea pr create --head <branch> --base main --title "..." --description "..." - List / view / check out:
tea pr list,tea pr <n>,tea pr checkout <n> - Issues:
tea issue create,tea issue list - Auth lives in
tea login, notGH_TOKENenv var.
tea print output as rendered boxes rather than plain text; PR URL lands on last line.
Design system
Web UI + userscript panel follow Cinder, rules in docs/design-system.md
— source of truth Claude Design project BookmarkManager Web UI
(969ac210-fe02-4c01-ae1b-9a271dcc779a). Read it before touching
backend/internal/web/static/style.css, backend/internal/web/templates/*, or userscript
TEMPLATE/CSS. Core law: ember means new chapter only — no other
state (busy, error, destruction) may use --ember; destruction gets
--danger. No cards/corners/shadows, one --measure: 760px column, tokens
only (never hardcode hex outside :root), both colour branches touched
together. Any move that pulls series out of list (archive/finish/remove)
must be confirm-gated via its own .confirm-row; only restore fires
instantly.
Security invariants
Existing guarantees — don't regress:
- Auth on
/bookmarks*: requireAuthorization: Bearer <credential>— the acting Reader's credential, matched by SHA-256 againstreaders.token_sha256— constant-time compare (via the hash, never the secret itself), 401 otherwise. - CORS: reflect
Originonly when inALLOWED_ORIGINS; allowGET,PUT,DELETE,OPTIONS+ headersAuthorization,Content-Type; answer preflightOPTIONSwith204.
Secure coding rules (code you write here)
Anchored to OWASP Top 10 / ASVS. Every rule below already has a working example in-tree — match it, don't start a second convention. AI-written backends fail on exactly these: broken access control, injection, weak session/error handling, invented dependencies.
Go backend:
- SQL always parameterized (
$N). Only compile-time constants (bookmarkColumns) may be concatenated into query text — never a request value, not even a validated one. html/templateonly for anything a browser parses, nevertext/template. Never wrap stored or fetched strings intemplate.HTML/JS/URL; that switches off the escaping every template depends on.- Any outbound fetch of a client-supplied URL passes
fetchableSeriesURL(site +https+ host check) first.series_urlarrives in a PUT body, so without the gate the poller will probe arbitrary hosts from the server's own network position. New fetch path reuses the gate rather than re-deriving one. - Cap every remote body with
io.LimitReader(maxBodyBytes). An unbounded read is an OOM handed to whatever is on the other end. - Compare secrets with
hmac.Equal/subtle.ConstantTimeCompare, never==. A credential is matched by the SHA-256 thereaderstable holds, which is already a fixed-width equality — a new secret comparison must not regress to==. - Errors: generic text to the client (
http.Error(w, "internal error", 500)), detail tolog.Printf. Never logTOKEN_KEY, a Reader's credential,DISCORD_CLIENT_SECRET, a session id, or a wholeAuthorizationheader. - Proxy headers are trusted only where they already are:
X-Forwarded-Protofor the Secure cookie flag, rightmostX-Forwarded-Forfor client IP (leftmost is attacker-supplied). Don't read either anywhere else. - Session cookies keep
HttpOnly,SameSite,Secure-when-HTTPS; expiry is enforced by thesessionstable lookup, not a signature. - Stdlib crypto only. No hand-rolled hashing, no MD5/SHA-1 anywhere security-bearing.
- Validate at the handler boundary before storing: body capped by
http.MaxBytesReader(64 KB), emptykeyand unknownstatus/kindrejected with400. A bad value that reaches the store becomes every later reader's problem.
Userscript:
- Site-derived and stored strings render via
el(..., {text})/textContent.{html}andinnerHTMLare for author-written literal markup only (TEMPLATE,CSS) — never a title, chapter label, or API response field. The page DOM belongs to a third-party site; treat it as attacker-controlled. - Isolated world protects the credential from the site's JS. It does not protect anything from an
innerHTMLsink you add yourself. - The userscripts carry
__API_TOKEN__placeholders, substituted at serve time with the requesting Reader's credential (internal/userscript). Never put a real credential in the repo, docs, commit messages, or issues. Rotation is a web-UI action (epoch bump,internal/token);TOKEN_KEYin backend env is what derives every credential — never log it. fetch()targetsAPI_BASEonly — no dynamic origin, no site-supplied URL.authHeaders()goes nowhere but the backend.localStorageis shared with the site's own JS: cache and queue live there, credentials never do.- Wrap every
localStorageread/write andJSON.parsein try/catch (quota, private mode, corrupt entry), as the existing helpers do.
Dependencies: stdlib first; a new module needs a stated reason. Confirm a package actually exists before adding it — a plausible name may be fiction (~20% of LLM-proposed packages don't resolve, which is how slopsquatting lands). Pin exact versions.
Review gate: auth, CORS, session, crypto, and the fetch gate are security-critical. Editing one is not a drive-by change — say which invariant you preserved and run go test ./... before calling it done.
Comments
Comment only if code alone can't carry info. Cost per read — must earn spot.
Write for:
- Why not what. Tradeoffs, non-obvious decisions.
- Load-bearing detail looking incidental — say so if "simplify" breaks it.
- Non-local consequence, invisible from function alone.
- Wire format / encoding / interface contract — save callers re-deriving.
- Gotcha/workaround, with ref if exists.
- Domain/business rule not derivable from code.
Skip:
- Restating code (no
// increment iabovei++). - Trivial getter/setter/pass-through.
- Banners, dividers,
// helpers. - Change narration (
// fix bug,// as requested,// new impl) — git's job. - Commented-out code — delete.
- TODO without concrete action.
Style: one dense comment over function beats one per line inside. Tight, no worked example unless bug subtle. Wrong comment worse than none — update/delete on change. Default fewer — sparse+high-signal beats comprehensive.
Test: "competent reader get this from code in few sec?" Yes → skip. Needs detour through another file/spec/git-blame → write it.
Relevant skills
multi-stage-dockerfile and docker-compose-orchestration for container work (referenced in plan).
golang-code-style, golang-error-handling, golang-performance, golang-testing for backend Go work.
Agent skills
AGENTS.md is the single source of truth for agent guidance; every CLAUDE.md in this repo is a symlink to the AGENTS.md beside it. Edit AGENTS.md.
Issue tracker
Issues live as Gitea issues on gitea.violetcrown.my.id (sulthan/mangaBookmark), driven by the tea CLI — not gh. See docs/agents/issue-tracker.md.
Triage labels
Default five-role vocabulary, label strings unchanged (needs-triage, needs-info, ready-for-agent, ready-for-human, wontfix). See docs/agents/triage-labels.md.
Domain docs
Single-context: one root CONTEXT.md plus docs/adr/, both created lazily. See docs/agents/domain.md.
graphify
Project has knowledge graph at graphify-out/ with god nodes, community structure, cross-file relationships.
Rules:
- For codebase questions, first run
graphify query "<question>"when graphify-out/graph.json exists. Usegraphify path "<A>" "<B>"for relationships andgraphify explain "<concept>"for focused concepts. Return scoped subgraph, usually much smaller than GRAPH_REPORT.md or raw grep output. - If graphify-out/wiki/index.md exists, use for broad navigation instead of raw source browsing.
- Read graphify-out/GRAPH_REPORT.md only for broad architecture review or when query/path/explain don't surface enough context.
- After modifying code, run
graphify update .to keep graph current (AST-only, no API cost).
Notes
- Keep comms terse — drop articles, fluff, pleasantries. Code/commits/security written normally.