741b23322b
Fixes five reported symptoms across comix.to and kagane.to. Diagnosing them turned up two latent bugs underneath, both of which had to be fixed for the kagane cover work to function at all.
## Reported symptoms and their causes
| # | Symptom | Cause |
|---|---------|-------|
| 1 | comix bookmark titled `Comix - Read Comics online for free` | comix is an SPA that rewrites `document.title` on client routing but never touches the server-rendered `og:title`. The adapter read `og:title`, so a cold load stored the homepage's title. |
| 2 | next comix bookmark gets the *previous* series' title | Same cause. After an in-page hop, `og:title` still holds whatever page loaded first. |
| 3 | comix cover shows the placeholder | comix serves no `og:image` at all, so `coverFromPage()` had nothing to read. |
| 4 | kagane chapter never appears in the bookmark list | Reader URLs carry no chapter number, so it is parsed out of `og:title`. Volume-numbered series render `"<Series> - Volume <v> Chapter <n>"`, which the suffix regex did not match, so `chapterNum` came back null and nothing was recorded. |
| 5 | kagane title includes the chapter, e.g. `SP Baby - Volume 1 Chapter 1` | Same unmatched regex — the tail was never stripped. One fix covers 4 and 5. |
| 6 | kagane cover blocked in the web UI | kagane serves covers behind its Cloudflare challenge **and** with `cross-origin-resource-policy: same-origin`. No `<img>` on the UI's origin can load one even from a browser holding the clearance cookie. Hot-linking cannot be made to work. |
## What changed
**Userscript.** comix titles now come from `document.title` with the chapter page's `" - Ch.<n>"` tail stripped, and the cover is the `img` whose `alt` matches the cleaned title. comix fills `document.title` a beat *after* the URL changes — later than the nav watcher's 300 ms snapshot — so the watcher also re-detects when the `detect()` signature changes, not only when the URL does. The kagane suffix regex takes an optional `Volume <v> ` segment. All three page shapes were captured live on 2026-08-08 and pinned as regression tests.
**Cover proxy.** `Bookmark.CoverURL()` rewrites a stored kagane `og:image` to `/img/kagane/{id}`; templates render `.CoverURL` instead of `.Cover`. The endpoint is session-gated like every other UI route and fetches through the shared headless browser, which is same-origin with kagane and so satisfies both the challenge and the CORP header. Results are memoised in-process, so a cover costs one navigation per deployment lifetime. With `BROWSER_WS_URL` unset the endpoint answers 404 rather than reaching for a nil fetcher — the same degrade-to-userscript behaviour the poller already has.
The image id is matched against a UUID regex before it reaches the browser. That gate is load-bearing rather than tidiness: the cover is a stored client-supplied string, so an unvalidated one turns this endpoint into an SSRF primitive aimed at the deployment's own network. `ServeMux` path-cleans a traversal into a redirect before the handler runs, but the handler does not depend on that, and a test pins it.
## Two latent bugs found underneath
**`BrowserFetcher.run` never let a challenge solve.** It navigated, waited for `body`, read once, and closed the tab — roughly half a second end to end. The Cloudflare interstitial has a `body` too, so `WaitReady` was satisfied by the challenge page itself. This made the challenge *unclearable* rather than merely slow: an interstitial needs several seconds of a live page to solve itself and write clearance into the browser's shared cookie jar, so tearing the tab down first means every subsequent call is challenged exactly like the one before it. `run` now holds one tab and re-reads until the caller's predicate reports an answer, bounded by `challengeTimeout` and the caller's own deadline. Exhausting the budget maps back to the 403 the poller already expects, keeping a challenged site distinct from a broken transport.
**`chromedp/headless-shell` cannot clear kagane's challenge at all.** It is a stripped Chrome build and the tells are structural rather than a header: `navigator.webdriver` is true, the plugin list is empty, and the client hints are Chromium- rather than Chrome-branded. Overriding `webdriver` through CDP was tried on its own and changed nothing.
All measured 2026-08-08 from one IP against the same cover, so the comparisons are like for like:
| Browser | Result |
|---------|--------|
| `chromedp/headless-shell:stable` | never cleared (90 s) |
| `zenika/alpine-chrome` | never cleared — ships Chrome 124, old enough that Cloudflare refuses it and old enough to break chromedp's CDP structs |
| `google-chrome`, default UA | never cleared (60 s) — `--headless=new` advertises `HeadlessChrome` |
| `google-chrome`, stock UA, `TZ=UTC` | never cleared (90 s) |
| `google-chrome`, stock UA, any non-UTC `TZ` | **cleared in ~4 s** |
Both remaining tells are load-bearing, and each was tested in isolation. `chrome/` is a Debian image with `google-chrome-stable`, a UA whose version is read back out of the binary at startup (a hardcoded one would drift out of step with the `Sec-CH-UA` hints on the next Chrome update and become a fresh tell), and no `--enable-automation`.
### The timezone tell: UTC, not a country mismatch
The first pass concluded the zone had to match the egress IP's country. Re-measuring against the actual deployment case shows that was wrong, and the correction is in `1552dd1`.
The original inference read the host's `/etc/timezone` (`Asia/Bangkok`) and assumed a Thai egress. It isn't — this host egresses from an Indonesian IP. `Asia/Bangkok` cleared not because it matched a country but because it simply isn't UTC, and the two share +07, which hid the distinction. Same container, same Indonesian IP:
| `TZ` | Result |
|------|--------|
| `UTC` | never cleared (60 s, **twice**) |
| `Asia/Jakarta` | cleared in 4 s |
| `America/New_York` | cleared in 4 s |
`America/New_York` matches neither the country nor the offset nor the hemisphere and clears just as fast. A UTC clock is itself the bot signal — Cloudflare scores it as the datacenter default — and any real zone satisfies the check. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one, and a deployment that changes region need not keep it in sync.
One sharp edge remains: the usual `-v /etc/localtime:/etc/localtime:ro` does **not** work. Chrome resolves the zone through ICU, which takes the name from that path's symlink target and ignores the file's contents, so glibc reports the host zone while Chrome still reports UTC. `/etc/timezone` carries the name and is mounted instead.
Chrome also binds its DevTools port to loopback and silently ignores `--remote-debugging-address`, which is why headless-shell fronted it with socat. This image does the same, so it stays a drop-in: the compose service keeps the `headless-shell` name and its pinned address, and `BROWSER_WS_URL` is unchanged.
## Verification
```
go test ./... all packages ok
node --test 37 + 12 pass, 0 fail
SMOKE_BROWSER_WS_URL=... go test -run TestSmokeKagane ./internal/latest
TestSmokeKaganeImage PASS (5.29s) fetched 56710 bytes of image/webp
TestSmokeKaganeGet PASS (1.17s) status=200, real chapter-list JSON
```
The smoke test ran against the exact compose configuration — built image, empty `BROWSER_TZ`, `/etc/timezone` mounted, cold profile — hitting real kagane.to. It skips unless `SMOKE_BROWSER_WS_URL` names a sidecar, so `go test ./...` stays hermetic and Docker-only.
A red smoke run means the challenge is not clearing from that IP, which is a live, time-varying fact to re-check rather than necessarily a defect.
## Security invariants
- Auth unchanged. `/img/kagane/{id}` is session-gated by `requireSession`, the same guard as every other UI route.
- Outbound fetch gated: the id is UUID-validated before it reaches the browser, keeping the existing rule that a client-supplied string never selects a fetch target unchecked.
- No new secrets, no new logging of credentials, no change to CORS, sessions, or crypto.
- Templates still escape everything; `.CoverURL` returns a plain string and is not wrapped in `template.HTML`/`URL`.
- One new dependency-free image (`chrome/`) built from Debian plus Google's own apt repo; no new Go modules.
## Deploying
Needs `docker compose build headless-shell`.
**A UTC host must set `BROWSER_TZ`, or kagane silently stops working.** With it unset the sidecar falls back to the host's `/etc/timezone`; on a UTC server that yields UTC, which is the one value that never clears. Any real zone works — `BROWSER_TZ=Asia/Jakarta` for the current deployment. `.env.example` now documents this; it previously did not mention the knob at all.
Only the browser sidecar reads `BROWSER_TZ`. The backend keeps its UTC clock, and stored timestamps are unix ms, so nothing else shifts.
## Deliberately not done
Retry/backoff around the cover proxy, and a panel-side cover fix. The panel renders no covers, and covers cache in-process after the first fetch. Worth adding if kagane starts rate-limiting.
## Correction after review of the deployment case
`1552dd1` was added after the branch was first pushed: the deployment host runs UTC with an Indonesian egress IP, which prompted re-measuring the timezone claim and falsifying it. The earlier commits' reasoning is left intact rather than rebased away, so the diagnostic trail — including the wrong turn and what disproved it — stays readable.
Reviewed-on: #37
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
6.1 KiB
6.1 KiB
Guidance for OpenCode (and Claude Code) working under userscript/. See root AGENTS.md for the project-wide architecture diagram, hard constraints, and design system.
Userscript structure (single IIFE, manga-bookmark.user.js)
- Site adapters — one per host,
detect(location, document)return pagetype+ IDs. Identify type/IDs from URL regex (most stable); pulltitle/coverfromog:title/og:imagemeta tags, not CSS classes. - API client —
apiGet/apiPut/apiDeletewith bearer header;localStoragekeybmgr:manga:cachefor instant render + offline fallback. - Progress logic — auto-upsert
last_chapteronly whenchapterNum >= stored last_chapter_num(re-reading old chapters must not regress progress; unparseable -> set current). Manual panel override forces any value. - Retry queue — every write go through
pushBookmark/pushDelete, so failed mutation park inlocalStorage(bmgr:manga:queue) and replayed on next navigation, reconnect, orrefresh(). Entries are markers ({key, op, sendStatus, attempts}), never payloads — body read from cache at send time, so one entry per key give ordering and coalescing for free.sendStatusis sticky: while archive pending, later writes to that key keep carrying bucket, which stop successful in-between write from silently un-archiving series.refresh()drains before it fetches and overlays anything still pending, so list never flaps. 400 drops entry, 401 abort pass and keep queue, and transient failures retry to cap of 10. Latest-chapter writes deliberately stay out of queue. Seedocs/superpowers/specs/2026-07-27-offline-retry-queue-design.md. - UI — rendered inside Shadow DOM root to isolate from site CSS
(critical on mobile). Three tabs (All / Favourites / Archived) and row of
link chips to web UI and both manga sites;
WEB_BASEsits in CONFIG block next toAPI_BASE. FAB is7 × 44edge tab whose hit area widened to28 × 72by invisible#hitchild;#fabmust keeptouch-action: noneand must not regainoverflow: hidden. Sincetouch-actionresolved at gesture start, strip can't be both browser-scrolled and script-dragged, somakeDraggablesplits by intent: swipe from#hitscrolls viawindow.scrollBy, hold ofARM_MSarms reposition drag, visible sliver drags with no hold. Seedocs/superpowers/specs/2026-07-28-edge-tab-hitbox-design.md. - SPA navigation — Asura is Astro, client-routed on comic/chapter pages: patch
history.pushState/replaceState+ listenpopstate, re-rundetect()on URL change so auto-update fire without reload. Demonic uses classic reloads (initialdocument-idlerun suffice).
Live URL shapes (verified 2026-07-26, may drift — re-check against live pages before trust)
- asurascans.com: series
/comics/<slug>(slug carries trailing site-wide build-hash suffix, e.g.-059befe1, that rotates on every redeploy), chapter/comics/<slug>/chapter/<n>.seriesIdmust strip hash (/-[0-9a-f]{8}$/,stripBuildHashin userscript,asuraBuildHashin backend); URLs keep full slug — stale-hash URLs 302 to current ones. Astro-rendered; chapter links present in raw server HTML. - demonicscans.org: series
/manga/<slug>(slug may URL-encode punctuation, e.g.%2527for'), chapter/title/<slug>/chapter/<n>/<page>(olderchaptered.php?manga=<id>&chapter=<n>form still exists as redirect, what series-page chapter-list anchors link through). Encodings (incl. triple-encoded punctuation like%25252D) identical on /manga/ and /title/ pages, so decode-once seriesIds match — verified 2026-07-28. - comix.to: series
/title/<id>-<slug>, chapter/title/<id>-<slug>/<uploadId>-chapter-<n>. Only the leading<id>is identity — the slug re-renders when a series is renamed (comixSeriesId). An SPA that never rewritesog:title: the server-rendered head keeps whatever document loaded first, so on a cold loadog:titleis the homepage's "Comix — Read Comics online for free" and after an in-page hop it is the previous series' name.document.titleis the one thing client routing does update, so titles come from there, with the chapter page's" · Ch.<n>"tail stripped. Covers likewise:og:imageis absent, so the cover is theimgwhosealtmatches the cleaned title — verified live 2026-08-08. - kagane.to: series
/series/<uuid>, reader/series/<uuid>/reader/<bookUuid>. Reader URLs carry no chapter number, so the number comes out ofog:title. Two shapes exist:"<Series> - Chapter <n>[ - Episode <n>]"and, for volume-numbered series,"<Series> - Volume <v> Chapter <n>"with no episode name — both must yield a bare series title, or the volume tail lands in the bookmark's title. Its covers are challenge- and CORP-protected, so the web UI proxies them; the userscript still stores the rawog:image. Behind a Cloudflare JS challenge, so the backend polls it through the headless browser. - novelfull.com (novel script): series
/<slug>.html, chapter/<slug>/chapter-<n>[-<title-slug>].html. Noog:*tags at all — title fromh3.title(series) ora.truyen-title(chapter), cover frommeta[name="image"]. Behind a Cloudflare JS challenge no TLS fingerprint clears, so the backend polls it through the headless browser. - lightnovelworld.net (novel script): series
/novel/<slug>/, chapter/<slug>-chapter-<n>/— flat, at the site root.h1.entry-titleis the clean title on a series page and<Title> Chapter <n>on a chapter page. Chapter pages carry noog:image. Its series page lists every chapter with an absolute href, so the backend polls it with the plain TLS client.
Second script: novel-bookmark.user.js
A copy of the manga script with two adapters, LIBRARY = "novel" and
STORE_PREFIX = "bmgr:novel:". No migration loop (this script has no previous
installation to carry keys over from). Installed alongside the manga script;
both write to the same backend with the same LIBRARY column discriminating
them.