Browser-backed Sites join the Cover pipeline (#62) (#72)

Fixes #62

Browser-backed Sites join the Cover pipeline: kagane and novelfull Series now get their Covers at creation, through the same acquisition path as every other Site, instead of waiting for a poll pass.

## What changed

`latest.Acquirer` (creation-time acquisition, fired by the first Bookmark of a Series) previously skipped kagane and novelfull entirely — their pages only yield a Cloudflare challenge to the TLS client, so the request was spent for nothing. It now routes them like the poller does, with the two Sites split exactly as the issue demands:

- **kagane** — page fetched through the browser sidecar, cover URL extracted from the API JSON, bytes fetched through the browser sidecar (the only path that clears the challenge) into the content-addressed store. With no `BROWSER_WS_URL` configured, acquisition is skipped entirely and nothing falls back to a plain fetch.
- **novelfull** — page fetched through the browser sidecar, cover URL extracted from the HTML, bytes fetched over plain TLS through the ordinary gated fetcher (its image paths answer 200 with `access-control-allow-origin: *`, measured 2026-08-09). With no browser configured, the page fetch falls back to the TLS client — novelfull's challenge is a live time-varying fact (AGENTS.md), so when the page body answers, the Cover still lands; when it is challenged, nothing happens.

The byte-routing rule (kagane → browser, every other Site → TLS) is now one shared function (`latest.fetchCoverBytes`) used by both the Poller and the Acquirer, so the two cannot drift apart.

## Acceptance criteria

- [x] kagane cover bytes are fetched through the browser sidecar and stored in the content-addressed store — `TestAcquireKaganeCoverThroughBrowser`
- [x] novelfull cover URLs are extracted from the browser-fetched HTML, and its bytes are fetched over plain TLS — `TestAcquireNovelfullCoverOverPlainTLS`
- [x] With no browser sidecar configured, kagane Covers are absent and nothing falls back to a plain fetch — `TestAcquireKaganeSkippedWithoutBrowser`
- [x] With no browser sidecar configured, novelfull Covers still work if its page body is available — `TestAcquireNovelfullCoverWithoutBrowser`
- [x] Manually verified on-device: a kagane Series shows its Cover in the panel, not a broken-image glyph — being run by a separate manual-verification agent against a mocked scenario (no prod data); not part of this PR
- [x] `go test ./...` is green, with live-network checks gated behind `SMOKE_BROWSER_WS_URL` like the existing kagane image smoke test — new `TestSmokeAcquireKaganeCover` proves the end-to-end acquire path against the real browser when the env var is set

## Verification

- `go test ./...` green across all packages
- New unit tests exercise every routing decision with fakes — no network in the default suite
- Smoke test gated behind `SMOKE_BROWSER_WS_URL`, skipped by default

## Post-review changes (a66491a)

- **One routing rule for pages too** — `fetcherFor` is now a shared function used by both the Poller and the Acquirer; novelfull falls back to the plain-TLS fetcher in *both* when no browser is configured, so pre-existing (client-scraped) novelfull rows get healed by the poll as well, not just Series created after this change (`TestNovelfullUsesTLSWhenNoBrowserFetcher`).
- **Byte-level no-fallback proof** — `TestAcquireKaganeBytesNeverFallBackToPlainTLS` pins that kagane cover bytes never route to the TLS fetcher even when the page came through a browser.
- **Acquirer wired independent of the TLS client** — if `NewTLSFetcher` fails, kagane/novelfull acquisition still works via the sidecar (`main.go`).
- AGENTS.md (root + backend) updated for the novelfull plain-TLS fallback.

Reviewed-on: #72
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
This commit was merged in pull request #72.
This commit is contained in:
2026-08-10 11:06:36 +07:00
committed by sulthan
parent b9220b3dfc
commit 78234f3c19
9 changed files with 406 additions and 53 deletions
+1 -1
View File
@@ -19,7 +19,7 @@ Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free
- Every site is its **own origin with its own `localStorage`** — a shared remote store is the only way to unify bookmarks. Cloud sync required, not optional.
- Userscript run in **isolated world**, so embedded API token safe from site's JS.
- Cloudflare's block on manga sites **IP-reputation-based, not universal — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — Cloudflare's bot scoring can flip previously-clean IP without notice. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test.
- **kagane.to and novelfull.com are the exception to the above** — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`) and skips them entirely when that's unset. The four other sites poll fine over plain TLS.
- **kagane.to and novelfull.com are the exception to the above** — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`). When that's unset, kagane is skipped entirely (a plain fetch would only retrieve a challenge page) while novelfull pages are still attempted over plain TLS — its challenge is a live time-varying fact and its cover bytes never need the browser. The four other sites poll fine over plain TLS.
- **The CDP browser must look like a real browser, and stock headless images don't.** Measured 2026-08-08 against kagane.to, all from the same IP: `chromedp/headless-shell:stable` never cleared the challenge in 90s (`navigator.webdriver` true, empty plugin list, Chromium-branded client hints — suppressing `webdriver` alone changed nothing); `zenika/alpine-chrome` ships Chrome 124, refused outright; real Chrome with the default `--headless=new` UA never cleared, because the UA says `HeadlessChrome`; real Chrome with a stock UA **and** a non-UTC clock zone cleared in ~4s. Hence `chrome/` — a Debian image with `google-chrome-stable`, a version-derived UA, and `TZ`/`BROWSER_TZ`. Chrome reads the zone *name* through ICU from `/etc/localtime`'s symlink target, ignoring the file's contents, so mounting the host's `/etc/localtime` does **not** work; `/etc/timezone` is mounted instead.
- **The browser is not in the API stack and must not be put back.** It's its own compose unit (`chrome/docker-compose.yml`) on a second machine, reached over the tailnet — it held 471 MiB on a 1974 MiB swapless VPS, and a residential egress scores better with Cloudflare anyway (ADR-0006). Consequences that constrain code: `BROWSER_WS_URL` must be a tailnet **IP** (a MagicDNS name 500s at `/json/version`, same trap as the old Docker service name); the CDP port binds to the tailnet address only, since CDP authenticates nothing and that host has a real LAN; and the browser is on-demand (ADR-0005), so an unreachable or asleep one must degrade exactly as an unset `BROWSER_WS_URL` — plain-TLS libraries unaffected, kagane/novelfull logged and skipped, stored covers still served. Never add `chromedp.NoModifyURL`: discovery per fetch is what makes a restarted Chrome invisible.
- **UTC is the tell, not a country mismatch.** A UTC clock is the datacenter default, so Cloudflare scores it as one; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while `Asia/Jakarta` **and** `America/New_York` both cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (`Asia/Bangkok`) rather than the measured egress. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one.