# AGENTS.md Guidance for OpenCode (and Claude Code) working in this repo. ## What this is Read-progress tracker for two libraries — manga and novels — behind one self-hosted Go backend. Two separate Violentmonkey userscripts inject on-page UI (floating button + slide-in panel) and sync progress, so bookmarks unify across sites and devices: - `manga-bookmark.user.js` — **asurascans.com** (asuracomic.net is dropped: its deep links 301 to the asurascans.com root, discarding the path), **demonicscans.org**, **comix.to**, **kagane.to**. - `novel-bookmark.user.js` — **novelfull.com**, **lightnovelworld.net**. One backend, one `bookmarks` table: a `kind` column (`manga`|`novel`) splits the libraries and the web UI switches between them. Rows are keyed `:`. ## Hard constraints (drive design — don't violate) Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free where plain web APIs suffice — keeps portability across engines: - **Avoid `GM_*` unless needed.** Prefer page `localStorage` over `GM_setValue`/`GM_getValue`, on-page UI over `GM_registerMenuCommand`, plain `fetch()` over `GM_xmlhttpRequest` for cross-origin. - Cross-origin `fetch()` work **only** against CORS-enabled backend. Manga sites `https://`, so backend **must be HTTPS** (else mixed-content block). - Every site is its **own origin with its own `localStorage`** — a shared remote store is the only way to unify bookmarks. Cloud sync required, not optional. - Userscript run in **isolated world**, so embedded API token safe from site's JS. - Cloudflare's block on manga sites is **per-zone configuration plus request fingerprint, not IP reputation — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — a Site can turn its protection on overnight, which is exactly what comix.to did on 2026-08-12. An earlier version of this line blamed "Cloudflare's bot scoring"; that was wrong. The 1-99 bot score is Enterprise Bot Management only and does not exist for a free-plan zone, and no per-IP request rate is documented as an input to challenge issuance — `docs/research/cloudflare-bot-scoring-and-poll-cadence.md`. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test. - **kagane.to, comix.to and novelfull.com are the exception to the above** — all three sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`). When that's unset, kagane and comix are skipped entirely (a plain fetch would only retrieve a challenge page) while novelfull pages are still attempted over plain TLS — its challenge is a live time-varying fact and its cover bytes never need the browser. comix turned hostile on 2026-08-12 (#98): its cover host `static.comix.to` is gated too, so its cover bytes go through the browser as well, and its page is read as an in-tab `fetch()` of the series URL rather than a rendered DOM — comix is an SPA, and rendering costs ~65 requests for the same server-rendered HTML one fetch returns. The three other sites poll fine over plain TLS. - **The CDP browser must look like a real browser, and stock headless images don't.** Measured 2026-08-08 against kagane.to, all from the same IP: `chromedp/headless-shell:stable` never cleared the challenge in 90s (`navigator.webdriver` true, empty plugin list, Chromium-branded client hints — suppressing `webdriver` alone changed nothing); `zenika/alpine-chrome` ships Chrome 124, refused outright; real Chrome with the default `--headless=new` UA never cleared, because the UA says `HeadlessChrome`; real Chrome with a stock UA **and** a non-UTC clock zone cleared in ~4s. Hence `chrome/` — a Debian image with `google-chrome-stable`, a version-derived UA, and `TZ`/`BROWSER_TZ`. Chrome reads the zone *name* through ICU from `/etc/localtime`'s symlink target, ignoring the file's contents, so mounting the host's `/etc/localtime` does **not** work; `/etc/timezone` is mounted instead. - **The browser is not in the API stack and must not be put back.** It's its own compose unit (`chrome/docker-compose.yml`) on a second machine, reached over the tailnet — it held 471 MiB on a 1974 MiB swapless VPS, and a residential egress avoids the cloud-hosting-IP signature Bot Fight Mode documentedly challenges (ADR-0006; not a better "score" — free-plan zones have no score). Consequences that constrain code: `BROWSER_WS_URL` must be a tailnet **IP** (a MagicDNS name 500s at `/json/version`, same trap as the old Docker service name); the CDP port binds to the tailnet address only, since CDP authenticates nothing and that host has a real LAN; and the browser is on-demand (ADR-0005), so an unreachable or asleep one must degrade exactly as an unset `BROWSER_WS_URL` — plain-TLS libraries unaffected, kagane/comix logged and skipped, stored covers still served. Never add `chromedp.NoModifyURL`: discovery per fetch is what makes a restarted Chrome invisible. - **UTC is the tell, not a country mismatch.** A UTC clock is the datacenter default, and the challenge refuses it; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while `Asia/Jakarta` **and** `America/New_York` both cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (`Asia/Bangkok`) rather than the measured egress. A second earlier claim, that Cloudflare "scores" a UTC clock, was also wrong: the measurement is real but the mechanism is not documented anywhere — Cloudflare publishes no timezone signal, and free-plan zones carry no score at all. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one. - **A challenged page needs the tab kept open.** The interstitial takes seconds to solve and only then writes clearance into the browser's shared cookie jar. Navigate-read-close never clears anything; `BrowserFetcher.run` holds one tab and re-reads until the payload arrives. ## Architecture ``` Two Violentmonkey userscripts (isolated world, per-site adapters, localStorage cache) -- fetch() HTTPS --> reverse proxy (TLS + CORS) --> Go net/http --> Postgres (volume) | | CDP over tailnet v on-demand Chrome, separate machine (chrome/) ``` Two deployable units on two machines: the API stack (`docker-compose.yml` + `docker-compose.prod.yml`, on the VPS) and the browser (`chrome/docker-compose.yml`, on the home machine). They share nothing but `BROWSER_WS_URL` and update independently. Backend-specific architecture (packages, endpoints, poller, config env vars) lives in `backend/AGENTS.md`. Userscript-specific structure (adapters, retry queue, UI, live URL shapes) lives in `userscript/AGENTS.md`. Deploy order `DEPLOY.md` (§7 for the browser), redeploy `REDEPLOY.md` (§8 for the browser). ## Commands Backend (`cd backend`): - Test all: `go test ./...` — **needs Docker.** Each test package starts a throwaway `postgres:17-alpine` container (`internal/pgtest`). - Single test: `go test -run TestName ./...` - Build static binary: `CGO_ENABLED=0 go build` Local stack: `docker compose up` (bookmark-api + postgres only; `postgres-data` named volume, `restart: unless-stopped`). No browser — without `BROWSER_WS_URL` the poller logs and skips kagane and comix. To run one: `cd chrome && BROWSER_BIND_ADDR=172.17.0.1 docker compose up -d --build`, then `BROWSER_WS_URL=ws://172.17.0.1:9222` in the root `.env` (bridge gateway, so the API container can name it by IP). Live CDP proof (needs that browser and network, skipped otherwise): `SMOKE_BROWSER_WS_URL=ws://: go test -run 'TestSmokeKagane|TestSmokeComix' ./internal/latest` — fetches a real kagane and comix cover and chapter list. A red run means the challenge is not clearing from this IP, which is a live fact to re-check, not necessarily a defect. Note for this dev machine: comix.to is DNS-hijacked to an ISP block page here (`comix.to` CNAMEs to `aduankonten.id`, so Chrome fails `ERR_CERT_COMMON_NAME_INVALID`). Run the browser container with `--add-host comix.to: --add-host static.comix.to:` resolved over DoH to get a real reading. Smoke test: `curl` endpoints with `Authorization: Bearer `; confirm `OPTIONS` preflight return CORS headers and `/healthz` return 200. ## Forge: Gitea, not GitHub `origin` is self-hosted Gitea instance (`gitea.violetcrown.my.id`), so **`gh` don't work here — use `tea` (Gitea CLI) for anything past plain git.** Common ones: - Open PR: `tea pr create --head --base main --title "..." --description "..."` - List / view / check out: `tea pr list`, `tea pr `, `tea pr checkout ` - Issues: `tea issue create`, `tea issue list` - Auth lives in `tea login`, not `GH_TOKEN` env var. `tea` print output as rendered boxes rather than plain text; PR URL lands on last line. ## Design system Web UI + userscript panel follow **Cinder**, rules in `docs/design-system.md` — source of truth Claude Design project `BookmarkManager Web UI` (`969ac210-fe02-4c01-ae1b-9a271dcc779a`). Read it before touching `backend/internal/web/static/style.css`, `backend/internal/web/templates/*`, or userscript `TEMPLATE`/`CSS`. Core law: **ember means new chapter only** — no other state (busy, error, destruction) may use `--ember`; destruction gets `--danger`. No cards/corners/shadows, one `--measure: 760px` column, tokens only (never hardcode hex outside `:root`), both colour branches touched together. Any move that pulls series out of list (archive/finish/remove) must be confirm-gated via its own `.confirm-row`; only restore fires instantly. ## Security invariants Existing guarantees — don't regress: - Auth on `/bookmarks*`: require `Authorization: Bearer ` — the acting Reader's credential, matched by SHA-256 against `readers.token_sha256` — **constant-time compare** (via the hash, never the secret itself), 401 otherwise. - CORS: reflect `Origin` only when in `ALLOWED_ORIGINS`; allow `GET,PUT,DELETE,OPTIONS` + headers `Authorization,Content-Type`; answer preflight `OPTIONS` with `204`. ## Secure coding rules (code you write here) Anchored to OWASP Top 10 / ASVS. Every rule below already has a working example in-tree — match it, don't start a second convention. AI-written backends fail on exactly these: broken access control, injection, weak session/error handling, invented dependencies. Go backend: - SQL always parameterized (`$N`). Only compile-time constants (`bookmarkColumns`) may be concatenated into query text — never a request value, not even a validated one. - `html/template` only for anything a browser parses, never `text/template`. Never wrap stored or fetched strings in `template.HTML`/`JS`/`URL`; that switches off the escaping every template depends on. - Any outbound fetch of a client-supplied URL passes `fetchableSeriesURL` (site + `https` + host check) first. `series_url` arrives in a PUT body, so without the gate the poller will probe arbitrary hosts from the server's own network position. New fetch path reuses the gate rather than re-deriving one. - Cap every remote body with `io.LimitReader` (`maxBodyBytes`). An unbounded read is an OOM handed to whatever is on the other end. - Compare secrets with `hmac.Equal` / `subtle.ConstantTimeCompare`, never `==`. A credential is matched by the SHA-256 the `readers` table holds, which is already a fixed-width equality — a new secret comparison must not regress to `==`. - Errors: generic text to the client (`http.Error(w, "internal error", 500)`), detail to `log.Printf`. Never log `TOKEN_KEY`, a Reader's credential, `DISCORD_CLIENT_SECRET`, a session id, or a whole `Authorization` header. - Proxy headers are trusted only where they already are: `X-Forwarded-Proto` for the Secure cookie flag, **rightmost** `X-Forwarded-For` for client IP (leftmost is attacker-supplied). Don't read either anywhere else. - Session cookies keep `HttpOnly`, `SameSite`, `Secure`-when-HTTPS; expiry is enforced by the `sessions` table lookup, not a signature. - Stdlib crypto only. No hand-rolled hashing, no MD5/SHA-1 anywhere security-bearing. - Validate at the handler boundary before storing: body capped by `http.MaxBytesReader` (64 KB), empty `key` and unknown `status`/`kind` rejected with `400`. A bad value that reaches the store becomes every later reader's problem. Userscript: - Site-derived and stored strings render via `el(..., {text})` / `textContent`. `{html}` and `innerHTML` are for author-written literal markup only (`TEMPLATE`, `CSS`) — never a title, chapter label, or API response field. The page DOM belongs to a third-party site; treat it as attacker-controlled. - Isolated world protects the credential from the site's JS. It does not protect anything from an `innerHTML` sink you add yourself. - The userscripts carry `__API_TOKEN__` placeholders, substituted at serve time with the requesting Reader's credential (`internal/userscript`). Never put a real credential in the repo, docs, commit messages, or issues. Rotation is a web-UI action (epoch bump, `internal/token`); `TOKEN_KEY` in backend env is what derives every credential — never log it. - `fetch()` targets `API_BASE` only — no dynamic origin, no site-supplied URL. `authHeaders()` goes nowhere but the backend. - `localStorage` is shared with the site's own JS: cache and queue live there, credentials never do. - Wrap every `localStorage` read/write and `JSON.parse` in try/catch (quota, private mode, corrupt entry), as the existing helpers do. Dependencies: stdlib first; a new module needs a stated reason. Confirm a package actually exists before adding it — a plausible name may be fiction (~20% of LLM-proposed packages don't resolve, which is how slopsquatting lands). Pin exact versions. Review gate: auth, CORS, session, crypto, and the fetch gate are security-critical. Editing one is not a drive-by change — say which invariant you preserved and run `go test ./...` before calling it done. ## Comments Comment only if code alone can't carry info. Cost per read — must earn spot. Write for: - Why not what. Tradeoffs, non-obvious decisions. - Load-bearing detail looking incidental — say so if "simplify" breaks it. - Non-local consequence, invisible from function alone. - Wire format / encoding / interface contract — save callers re-deriving. - Gotcha/workaround, with ref if exists. - Domain/business rule not derivable from code. Skip: - Restating code (no `// increment i` above `i++`). - Trivial getter/setter/pass-through. - Banners, dividers, `// helpers`. - Change narration (`// fix bug`, `// as requested`, `// new impl`) — git's job. - Commented-out code — delete. - TODO without concrete action. Style: one dense comment over function beats one per line inside. Tight, no worked example unless bug subtle. Wrong comment worse than none — update/delete on change. Default fewer — sparse+high-signal beats comprehensive. Test: "competent reader get this from code in few sec?" Yes → skip. Needs detour through another file/spec/git-blame → write it. ## Agent skills `AGENTS.md` is the single source of truth for agent guidance; every `CLAUDE.md` in this repo is a symlink to the `AGENTS.md` beside it. Edit `AGENTS.md`. ### Issue tracker Issues live as Gitea issues on `gitea.violetcrown.my.id` (`sulthan/mangaBookmark`), driven by the `tea` CLI — not `gh`. See `docs/agents/issue-tracker.md`. ### Triage labels Default five-role vocabulary, label strings unchanged (`needs-triage`, `needs-info`, `ready-for-agent`, `ready-for-human`, `wontfix`). See `docs/agents/triage-labels.md`. ### Domain docs Single-context: one root `CONTEXT.md` plus `docs/adr/`, both created lazily. See `docs/agents/domain.md`. ## graphify Project has knowledge graph at graphify-out/ with god nodes, community structure, cross-file relationships. Rules: - For codebase questions and exploration, always first run `graphify query ""` when graphify-out/graph.json exists. Use `graphify path "" ""` for relationships and `graphify explain ""` for focused concepts. Return scoped subgraph, usually much smaller than GRAPH_REPORT.md or raw grep output. - If graphify-out/wiki/index.md exists, use for broad navigation instead of raw source browsing. - Read graphify-out/GRAPH_REPORT.md only for broad architecture review or when query/path/explain don't surface enough context. - After modifying code, run `graphify update .` to keep graph current (AST-only, no API cost).