From 86160c164a60fa60e6018d8767c9cd6abe793b3c Mon Sep 17 00:00:00 2001 From: Sulthan Zaki Date: Sun, 16 Aug 2026 12:05:11 +0700 Subject: [PATCH] feat(latest): poll comix through the browser sidecar (#98) comix.to began answering plain-TLS fetches with a Cloudflare JavaScript challenge on 2026-08-12, so every poll got a 403 interstitial and its cover host static.comix.to is gated the same way. comix joins kagane and novelfull as a browser Site: one registry entry, no plain-TLS fallback, and cover bytes routed through the browser's image path behind a fully pinned URL pattern. The read is an in-tab fetch of the Series URL, not a DOM render: comix is an SPA, so rendering costs ~65 requests for the same server-rendered HTML one fetch returns (24.5 KB, ~480 ms). Parsers and stored Series identity are untouched. Verified live against the real browser unit: page 24793 bytes in one fetch, chapter 53, cover accepted by the pin and 26862 image bytes retrieved by direct navigation (comix's Series page sets cross-origin-embedder-policy: require-corp, so an in-page fetch of the cover host cannot work). --- AGENTS.md | 14 ++-- REDEPLOY.md | 4 +- backend/AGENTS.md | 49 +++++++------ backend/internal/latest/browser.go | 69 +++++++++++++----- backend/internal/latest/browser_test.go | 56 +++++++++++++++ backend/internal/latest/cover.go | 3 +- backend/internal/latest/poller.go | 10 +-- backend/internal/latest/poller_test.go | 80 +++++++++++++++++++++ backend/internal/latest/sites.go | 30 ++++++-- backend/internal/latest/sites_test.go | 4 ++ backend/internal/latest/smoke_comix_test.go | 72 +++++++++++++++++++ 11 files changed, 336 insertions(+), 55 deletions(-) create mode 100644 backend/internal/latest/smoke_comix_test.go diff --git a/AGENTS.md b/AGENTS.md index 8b1740b..7247db0 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -19,9 +19,9 @@ Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free - Every site is its **own origin with its own `localStorage`** — a shared remote store is the only way to unify bookmarks. Cloud sync required, not optional. - Userscript run in **isolated world**, so embedded API token safe from site's JS. - Cloudflare's block on manga sites is **per-zone configuration plus request fingerprint, not IP reputation — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — a Site can turn its protection on overnight, which is exactly what comix.to did on 2026-08-12. An earlier version of this line blamed "Cloudflare's bot scoring"; that was wrong. The 1-99 bot score is Enterprise Bot Management only and does not exist for a free-plan zone, and no per-IP request rate is documented as an input to challenge issuance — `docs/research/cloudflare-bot-scoring-and-poll-cadence.md`. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test. -- **kagane.to and novelfull.com are the exception to the above** — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`). When that's unset, kagane is skipped entirely (a plain fetch would only retrieve a challenge page) while novelfull pages are still attempted over plain TLS — its challenge is a live time-varying fact and its cover bytes never need the browser. The four other sites poll fine over plain TLS. +- **kagane.to, comix.to and novelfull.com are the exception to the above** — all three sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`). When that's unset, kagane and comix are skipped entirely (a plain fetch would only retrieve a challenge page) while novelfull pages are still attempted over plain TLS — its challenge is a live time-varying fact and its cover bytes never need the browser. comix turned hostile on 2026-08-12 (#98): its cover host `static.comix.to` is gated too, so its cover bytes go through the browser as well, and its page is read as an in-tab `fetch()` of the series URL rather than a rendered DOM — comix is an SPA, and rendering costs ~65 requests for the same server-rendered HTML one fetch returns. The three other sites poll fine over plain TLS. - **The CDP browser must look like a real browser, and stock headless images don't.** Measured 2026-08-08 against kagane.to, all from the same IP: `chromedp/headless-shell:stable` never cleared the challenge in 90s (`navigator.webdriver` true, empty plugin list, Chromium-branded client hints — suppressing `webdriver` alone changed nothing); `zenika/alpine-chrome` ships Chrome 124, refused outright; real Chrome with the default `--headless=new` UA never cleared, because the UA says `HeadlessChrome`; real Chrome with a stock UA **and** a non-UTC clock zone cleared in ~4s. Hence `chrome/` — a Debian image with `google-chrome-stable`, a version-derived UA, and `TZ`/`BROWSER_TZ`. Chrome reads the zone *name* through ICU from `/etc/localtime`'s symlink target, ignoring the file's contents, so mounting the host's `/etc/localtime` does **not** work; `/etc/timezone` is mounted instead. -- **The browser is not in the API stack and must not be put back.** It's its own compose unit (`chrome/docker-compose.yml`) on a second machine, reached over the tailnet — it held 471 MiB on a 1974 MiB swapless VPS, and a residential egress avoids the cloud-hosting-IP signature Bot Fight Mode documentedly challenges (ADR-0006; not a better "score" — free-plan zones have no score). Consequences that constrain code: `BROWSER_WS_URL` must be a tailnet **IP** (a MagicDNS name 500s at `/json/version`, same trap as the old Docker service name); the CDP port binds to the tailnet address only, since CDP authenticates nothing and that host has a real LAN; and the browser is on-demand (ADR-0005), so an unreachable or asleep one must degrade exactly as an unset `BROWSER_WS_URL` — plain-TLS libraries unaffected, kagane/novelfull logged and skipped, stored covers still served. Never add `chromedp.NoModifyURL`: discovery per fetch is what makes a restarted Chrome invisible. +- **The browser is not in the API stack and must not be put back.** It's its own compose unit (`chrome/docker-compose.yml`) on a second machine, reached over the tailnet — it held 471 MiB on a 1974 MiB swapless VPS, and a residential egress avoids the cloud-hosting-IP signature Bot Fight Mode documentedly challenges (ADR-0006; not a better "score" — free-plan zones have no score). Consequences that constrain code: `BROWSER_WS_URL` must be a tailnet **IP** (a MagicDNS name 500s at `/json/version`, same trap as the old Docker service name); the CDP port binds to the tailnet address only, since CDP authenticates nothing and that host has a real LAN; and the browser is on-demand (ADR-0005), so an unreachable or asleep one must degrade exactly as an unset `BROWSER_WS_URL` — plain-TLS libraries unaffected, kagane/comix logged and skipped, stored covers still served. Never add `chromedp.NoModifyURL`: discovery per fetch is what makes a restarted Chrome invisible. - **UTC is the tell, not a country mismatch.** A UTC clock is the datacenter default, and the challenge refuses it; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while `Asia/Jakarta` **and** `America/New_York` both cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (`Asia/Bangkok`) rather than the measured egress. A second earlier claim, that Cloudflare "scores" a UTC clock, was also wrong: the measurement is real but the mechanism is not documented anywhere — Cloudflare publishes no timezone signal, and free-plan zones carry no score at all. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one. - **A challenged page needs the tab kept open.** The interstitial takes seconds to solve and only then writes clearance into the browser's shared cookie jar. Navigate-read-close never clears anything; `BrowserFetcher.run` holds one tab and re-reads until the payload arrives. @@ -45,12 +45,16 @@ Backend (`cd backend`): - Single test: `go test -run TestName ./...` - Build static binary: `CGO_ENABLED=0 go build` -Local stack: `docker compose up` (bookmark-api + postgres only; `postgres-data` named volume, `restart: unless-stopped`). No browser — without `BROWSER_WS_URL` the poller logs and skips kagane and novelfull. To run one: `cd chrome && BROWSER_BIND_ADDR=172.17.0.1 docker compose up -d --build`, then `BROWSER_WS_URL=ws://172.17.0.1:9222` in the root `.env` (bridge gateway, so the API container can name it by IP). +Local stack: `docker compose up` (bookmark-api + postgres only; `postgres-data` named volume, `restart: unless-stopped`). No browser — without `BROWSER_WS_URL` the poller logs and skips kagane and comix. To run one: `cd chrome && BROWSER_BIND_ADDR=172.17.0.1 docker compose up -d --build`, then `BROWSER_WS_URL=ws://172.17.0.1:9222` in the root `.env` (bridge gateway, so the API container can name it by IP). Live CDP proof (needs that browser and network, skipped otherwise): -`SMOKE_BROWSER_WS_URL=ws://: go test -run TestSmokeKagane ./internal/latest` -— fetches a real kagane cover and chapter list. A red run means the challenge is +`SMOKE_BROWSER_WS_URL=ws://: go test -run 'TestSmokeKagane|TestSmokeComix' ./internal/latest` +— fetches a real kagane and comix cover and chapter list. A red run means the challenge is not clearing from this IP, which is a live fact to re-check, not necessarily a defect. +Note for this dev machine: comix.to is DNS-hijacked to an ISP block page here +(`comix.to` CNAMEs to `aduankonten.id`, so Chrome fails `ERR_CERT_COMMON_NAME_INVALID`). +Run the browser container with `--add-host comix.to: --add-host static.comix.to:` +resolved over DoH to get a real reading. Smoke test: `curl` endpoints with `Authorization: Bearer `; confirm `OPTIONS` preflight return CORS headers and `/healthz` return 200. diff --git a/REDEPLOY.md b/REDEPLOY.md index 307be8f..1d862a2 100644 --- a/REDEPLOY.md +++ b/REDEPLOY.md @@ -436,8 +436,8 @@ free -m # the Gitea runner should still have its headroom Nothing here needs doing during an API redeploy. The API stack does not `depends_on` the browser, and an unreachable one degrades exactly as an unset -`BROWSER_WS_URL`: plain-TLS libraries unaffected, kagane and novelfull logged -and skipped, stored covers still served. +`BROWSER_WS_URL`: plain-TLS libraries unaffected, kagane and comix logged +and skipped, novelfull attempted over plain TLS, stored covers still served. --- diff --git a/backend/AGENTS.md b/backend/AGENTS.md index 5dfbd16..cfc73a8 100644 --- a/backend/AGENTS.md +++ b/backend/AGENTS.md @@ -89,12 +89,15 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN `Store.SetLatestChapter`, so a bookmark's `updated_at` — and the list order — is never touched. Fetches use `bogdanfinn/tls-client` with Chrome profile as defence in depth - against fingerprint-based blocking; any failure log and skip. kagane and - novelfull sit behind Cloudflare JavaScript challenges the TLS client can't - clear, so they are fetched over CDP via `BROWSER_WS_URL`; kagane is simply - not polled when that's unset, while novelfull falls back to a plain-TLS - attempt — its challenge is a live time-varying fact, and its cover bytes - never need the browser. See + against fingerprint-based blocking; any failure log and skip. kagane, comix + and novelfull sit behind Cloudflare JavaScript challenges the TLS client + can't clear, so they are fetched over CDP via `BROWSER_WS_URL`; kagane and + comix are simply not polled when that's unset, while novelfull falls back to + a plain-TLS attempt — its challenge is a live time-varying fact, and its + cover bytes never need the browser. comix's browser read is an in-tab + `fetch()` of the Series URL, not a DOM render: it is an SPA, so rendering + costs ~65 requests for the same server-rendered HTML one fetch returns + (measured 2026-08-12, issue #98). See `docs/superpowers/specs/2026-07-26-server-latest-chapter-polling-design.md`. The poller's series write is a single-column UPDATE (`Store.SetLatestChapter`), not a read-modify-write of the whole bookmark: @@ -112,14 +115,17 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN is public and uncredentialed: the userscript renders it on a Site's origin, where no cookie or token of ours travels. A client-sent `cover` is decoded and discarded, permanently (ADR-0004 compatibility). - Browser-backed Sites join the same pipeline (issue #62): kagane pages *and* - cover bytes go through the browser sidecar (nothing falls back to a plain - fetch, which would only retrieve a challenge page), while novelfull needs - the browser only for its HTML — the cover URL comes out of the - browser-fetched page and the bytes go over plain TLS. With no browser - configured, kagane Covers are simply absent; novelfull still gets one — at - creation and on the poll — when its page body happens to answer a plain - request (the challenge is a live time-varying fact). The old kagane-only + Browser-backed Sites join the same pipeline (issue #62, extended to comix by + #98): kagane and comix pages *and* cover bytes go through the browser sidecar + (nothing falls back to a plain fetch, which would only retrieve a challenge + page), while novelfull needs the browser only for its HTML — the cover URL + comes out of the browser-fetched page and the bytes go over plain TLS. With + no browser configured, kagane and comix Covers are simply absent; novelfull + still gets one — at creation and on the poll — when its page body happens to + answer a plain request (the challenge is a live time-varying fact). comix + cover bytes must arrive by direct navigation, not an in-page fetch: its + Series page sets `cross-origin-embedder-policy: require-corp`, which fails a + page-context fetch of `static.comix.to`. The old kagane-only serving path (`/img/kagane/{id}`, template rewrite, `CoverFetcher`) is gone (issue #63): the one public route serves every Site. - **`updated_at` drives list order, so moves only on real reading progress:** server apply its timestamp when row new or `last_chapter_num` changes, else keep stored value — favouriting series or recording newly published chapter must not reorder list. `PUT` therefore returns row **as stored**, clients must adopt that response rather than own payload. See `plans/2026-07-25-bookmark-list-favorites-design.md` §4. @@ -164,10 +170,11 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN Reader's credential at serve time). `BROWSER_WS_URL` (CDP endpoint of the browser, which runs on a **separate machine** and is reached over the tailnet — ADR-0006, `chrome/docker-compose.yml`. - Used by the poller for kagane and novelfull page fetches and by the cover - pipeline for kagane's image bytes (the browser is the only route that clears - the challenge kagane serves its covers behind); unset — the default — - disables browser polling and leaves kagane Covers blank until stored bytes + Used by the poller for kagane, comix and novelfull page fetches and by the + cover pipeline for kagane's and comix's image bytes (the browser is the only + route that clears the challenge those two serve their covers behind); unset — + the default — disables browser polling and leaves kagane and comix Covers + blank until stored bytes exist. Must be a tailnet IP, never a hostname: Chrome's DevTools handler 500s `/json/version` for any Host that isn't an IP or `localhost`). - **No per-Site cover path (issue #63):** every Cover — all six Sites — is @@ -175,8 +182,10 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN bytes. There is no proxy, no per-Site rewrite, no second place that decides a Cover's renderable address: the wire `cover` is it. The only place a Site name still appears in cover code is the extraction module (`latest`), where - kagane's image URLs are claimed by `browserOnlyCoverURL` — they answer a - plain fetch with a challenge and `cross-origin-resource-policy: same-origin`; + kagane's and comix's image URLs are claimed by `browserOnlyCoverURL` — kagane + answers a plain fetch with a challenge and + `cross-origin-resource-policy: same-origin`, and `static.comix.to` answers + one with the same Cloudflare challenge its pages serve; every other Site's CDN answers plain TLS. Templates render `.Cover` — the wire value — never anything else. - **Web UI also owns:** session-gated `GET /install/{manga,novel}-bookmark.user.js` diff --git a/backend/internal/latest/browser.go b/backend/internal/latest/browser.go index 641b33a..31ce28e 100644 --- a/backend/internal/latest/browser.go +++ b/backend/internal/latest/browser.go @@ -24,12 +24,16 @@ const challengeTimeout = 45 * time.Second var kaganeSeriesRe = regexp.MustCompile(`^/series/([0-9a-f-]{36})/?$`) +// comixSeriesPathRe matches the one path shape comixRead will open: a Series +// page, "/title/-". Verified live 2026-08-12. +var comixSeriesPathRe = regexp.MustCompile(`^/title/[^/?#]+/?$`) + // BrowserFetcher retrieves pages through a remote headless Chrome over the // DevTools Protocol. // -// It exists for one reason: kagane.to and novelfull.com sit behind a -// Cloudflare JavaScript challenge. Verified 2026-08-03 (kagane) and 2026-08-05 -// (novelfull) from the deployment host, plain HTTP and bogdanfinn/tls-client +// It exists for one reason: kagane.to, novelfull.com and comix.to sit behind a +// Cloudflare JavaScript challenge. Verified 2026-08-03 (kagane), 2026-08-05 +// (novelfull) and 2026-08-12 (comix), plain HTTP and bogdanfinn/tls-client // with a Chrome_133 profile both get 403 with cf-mitigated: challenge on every // path, including the API, robots.txt and images. Clearing it requires // executing the challenge script, which only a real browser does. @@ -141,29 +145,47 @@ func novelfullRead(seriesURL string, out *string) (chromedp.Action, bool) { return chromedp.OuterHTML("html", out, chromedp.ByQuery), true } +// comixRead fetches the Series page from inside the cleared tab. comix is an +// SPA: rendering the page costs ~65 requests, while one same-origin fetch of +// the same address returns the server-rendered HTML — 24.5 KB, ~480 ms, +// carrying both parser anchors (measured 2026-08-12, issue #98). So this is +// kaganeRead's shape, not novelfullRead's, even though the payload is HTML. +// Refusing any other address is the per-Site half of the SSRF gate. +func comixRead(seriesURL string, out *string) (chromedp.Action, bool) { + pageURL, ok := comixSeriesPageURL(seriesURL) + if !ok { + return nil, false + } + return chromedp.Evaluate( + `fetch(`+jsString(pageURL)+`).then(r => r.ok ? r.text() : "")`, + out, awaitPromise), true +} + // Image retrieves one cover's bytes through the browser sidecar, and its // content type. // -// It exists because kagane serves covers behind the same challenge as its -// pages *and* with `cross-origin-resource-policy: same-origin`, so the bytes -// are only reachable from inside a browser that already holds the clearance -// cookie (verified 2026-08-08). Acquisition through the sidecar is the only -// route. +// It exists because kagane and comix serve covers behind the same challenge as +// their pages — kagane additionally with +// `cross-origin-resource-policy: same-origin` — so the bytes are only +// reachable from inside a browser that already holds the clearance cookie +// (verified 2026-08-08 for kagane, 2026-08-12 for comix). Acquisition through +// the sidecar is the only route. // -// The image URL is navigated to rather than fetched from some other kagane -// page: the challenge only runs on a top-level navigation, and once it clears +// The image URL is navigated to rather than fetched from another page of the +// Site: the challenge only runs on a top-level navigation, and once it clears // the document *is* the image, so a same-origin fetch of location.href reads -// it straight back out of the cache. +// it straight back out of the cache. For comix the navigation is also the only +// route that works at all — its Series page sets +// `cross-origin-embedder-policy: require-corp`, which fails a page-context +// fetch of the cover host. // // The challenge is not solved by the first read: WaitReady("body") is satisfied // by the interstitial too. run holds the tab open until the in-page fetch // succeeds, which is what gives the challenge script the seconds it needs. func (f *BrowserFetcher) Image(ctx context.Context, imageURL string) ([]byte, string, error) { - m := kaganeImageURLRe.FindStringSubmatch(imageURL) - if m == nil { + if !browserOnlyCoverURL(imageURL) { return nil, "", fmt.Errorf("not a browser-fetchable cover url: %q", imageURL) } - imageID := m[1] var dataURL string err := f.run(ctx, imageURL, chromedp.Evaluate(`fetch(location.href).then(r => r.ok @@ -175,16 +197,16 @@ func (f *BrowserFetcher) Image(ctx context.Context, imageURL string) ([]byte, st : "")`, &dataURL, awaitPromise), func() bool { return dataURL != "" }) if err != nil { - return nil, "", fmt.Errorf("browser image %s: %w", imageID, err) + return nil, "", fmt.Errorf("browser image %s: %w", imageURL, err) } // "data:image/webp;base64,". head, payload, ok := strings.Cut(dataURL, ";base64,") if !ok { - return nil, "", fmt.Errorf("browser image %s: not a data url", imageID) + return nil, "", fmt.Errorf("browser image %s: not a data url", imageURL) } raw, err := base64.StdEncoding.DecodeString(payload) if err != nil { - return nil, "", fmt.Errorf("browser image %s: %w", imageID, err) + return nil, "", fmt.Errorf("browser image %s: %w", imageURL, err) } return raw, strings.TrimPrefix(head, "data:"), nil } @@ -323,6 +345,19 @@ func novelfullSeriesURL(seriesURL string) bool { strings.HasSuffix(u.Path, ".html") } +// comixSeriesPageURL returns the address comixRead fetches inside the tab: the +// Series page itself, rebuilt from the pinned host and path so nothing else +// travels. Host-pinned here for the same reason kagane's is — series_url is +// client-supplied and a headless browser is a strong SSRF primitive. +func comixSeriesPageURL(seriesURL string) (string, bool) { + u, err := url.Parse(seriesURL) + if err != nil || u.Scheme != "https" || u.Hostname() != "comix.to" || + !comixSeriesPathRe.MatchString(u.Path) { + return "", false + } + return "https://comix.to" + u.Path, true +} + // awaitPromise makes Evaluate resolve the promise rather than returning a // serialised Promise object. func awaitPromise(p *runtime.EvaluateParams) *runtime.EvaluateParams { diff --git a/backend/internal/latest/browser_test.go b/backend/internal/latest/browser_test.go index 41dcf6b..34c20dc 100644 --- a/backend/internal/latest/browser_test.go +++ b/backend/internal/latest/browser_test.go @@ -61,6 +61,62 @@ func TestNovelfullSeriesURL(t *testing.T) { }) } } + +func TestComixSeriesPageURL(t *testing.T) { + const series = "https://comix.to/title/n8we-dungeons-and-crayons" + cases := []struct { + name string + url string + want string + }{ + {"series page", series, series}, + {"trailing slash kept", series + "/", series + "/"}, + // Query and fragment are dropped: only the pinned path travels. + {"query dropped", series + "?tab=chapters", series}, + {"foreign host", "https://evil.example/title/x", ""}, + {"lookalike host", "https://comix.to.evil.example/title/x", ""}, + {"not https", "http://comix.to/title/x", ""}, + {"not a series path", "https://comix.to/search", ""}, + {"chapter page", series + "/11139891-chapter-80", ""}, + {"garbage", "://nope", ""}, + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + got, ok := comixSeriesPageURL(tc.url) + if ok != (tc.want != "") || got != tc.want { + t.Fatalf("comixSeriesPageURL(%q) = %q, %v; want %q", tc.url, got, ok, tc.want) + } + }) + } +} + +// The browser is an SSRF primitive and a cover address can originate in a +// client-supplied PUT body, so this gate decides what it may navigate to. +func TestBrowserOnlyCoverURL(t *testing.T) { + cases := []struct { + url string + want bool + }{ + {"https://static.comix.to/039d/i/1/34/6a6742bf15736@280.jpg", true}, + {"https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed", true}, + // Every other Site's CDN answers plain TLS. + {"https://gg.asuracomic.net/covers/x.webp", false}, + {"http://static.comix.to/039d/x.jpg", false}, + {"https://static.comix.to.evil.example/039d/x.jpg", false}, + {"https://evil.example/static.comix.to/x.jpg", false}, + {"https://static.comix.to/039d/x.jpg?next=http://169.254.169.254/", false}, + {"https://static.comix.to/039d/x.svg", false}, + {"https://static.comix.to/../etc/passwd.jpg", false}, + {"https://static.comix.to/", false}, + } + for _, tc := range cases { + t.Run(tc.url, func(t *testing.T) { + if got := browserOnlyCoverURL(tc.url); got != tc.want { + t.Fatalf("browserOnlyCoverURL(%q) = %v, want %v", tc.url, got, tc.want) + } + }) + } +} func TestClassifyBrowserInterruption(t *testing.T) { if err := classifyBrowserError(context.Background(), true, context.Canceled); !errors.Is(err, errBrowserInterrupted) { t.Fatalf("classifyBrowserError(context.Canceled) = %v, want browser interruption", err) diff --git a/backend/internal/latest/cover.go b/backend/internal/latest/cover.go index fcee992..932e000 100644 --- a/backend/internal/latest/cover.go +++ b/backend/internal/latest/cover.go @@ -25,7 +25,8 @@ type CoverBytesFetcher interface { // fetchCoverBytes routes a cover's byte retrieval by URL shape, not by Site // name: the browser fetcher's module claims the addresses only it can fetch // (kagane's image route answers a plain fetch with a challenge and -// `cross-origin-resource-policy: same-origin`), and everything else goes over +// `cross-origin-resource-policy: same-origin`, static.comix.to answers one with +// the same challenge its pages serve), and everything else goes over // plain TLS. Missing fetchers degrade to an error the caller logs, never a // fallback onto a path that cannot succeed. One routing rule for the poll and // the acquirer, so the two cannot drift apart. diff --git a/backend/internal/latest/poller.go b/backend/internal/latest/poller.go index 8f31abc..057a2bd 100644 --- a/backend/internal/latest/poller.go +++ b/backend/internal/latest/poller.go @@ -17,8 +17,8 @@ type Fetcher interface { } // BrowserCoverFetcher retrieves one cover's bytes through the browser-backed -// path — the only route that clears the challenge kagane's image URLs answer -// a plain fetch with. Satisfied by BrowserFetcher. +// path — the only route that clears the challenge kagane's and comix's image +// URLs answer a plain fetch with. Satisfied by BrowserFetcher. type BrowserCoverFetcher interface { Image(ctx context.Context, imageURL string) (body []byte, contentType string, err error) } @@ -119,9 +119,9 @@ func (p *Poller) storeCover(ctx context.Context, sr store.Series, sourceURL stri // fetcherFor returns the fetcher a site's page needs, or nil when the site // cannot be fetched at all right now. A Site whose registry entry carries a -// Browser read — kagane and novelfull, both behind a Cloudflare JavaScript -// challenge no TLS fingerprint clears — prefers the browser; when it is -// absent, the entry's Fallback decides whether plain TLS may take over. One +// Browser read — kagane, comix and novelfull, all behind a Cloudflare +// JavaScript challenge no TLS fingerprint clears — prefers the browser; when it +// is absent, the entry's Fallback decides whether plain TLS may take over. One // routing rule for the poll and the acquirer, so the two cannot drift apart. func fetcherFor(site string, browser, tls Fetcher) Fetcher { s, known := sites[site] diff --git a/backend/internal/latest/poller_test.go b/backend/internal/latest/poller_test.go index 6bbb6e5..b1e0376 100644 --- a/backend/internal/latest/poller_test.go +++ b/backend/internal/latest/poller_test.go @@ -762,6 +762,86 @@ func TestKaganeUsesBrowserFetcher(t *testing.T) { } } +// comix joined kagane behind the challenge on 2026-08-12 (#98): its page goes +// to the browser, its Cover bytes go through the browser's image route because +// static.comix.to is gated the same way, and the TLS fetcher is never asked +// for either. +func TestComixUsesBrowserFetcher(t *testing.T) { + s, dbURL := newTestStore(t) + const ( + key = "comix:n8we-dungeons-and-crayons" + seriesID = "n8we-dungeons-and-crayons" + seriesURL = "https://comix.to/title/n8we-dungeons-and-crayons" + coverURL = "https://static.comix.to/039d/i/1/34/6a6742bf15736@280.jpg" + ) + if _, err := s.Upsert(s.OwnerID(), store.Bookmark{ + Key: key, Site: "comix", SeriesID: seriesID, SeriesURL: seriesURL, + UpdatedAt: 1000, + }); err != nil { + t.Fatalf("seed: %v", err) + } + seedCoverSource(t, dbURL, "comix", seriesID, coverURL) + + tlsF := &fakeFetcher{body: "", status: 200} + browserF := &fakeFetcher{body: comixSeriesFixture, status: 200} + covers := &fakeCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"} + tlsCovers := &fakeBytesCoverFetcher{body: []byte("tls-bytes"), contentType: "image/jpeg"} + p := &Poller{ + Store: s, Fetch: tlsF, BrowserFetch: browserF, + CoverFetch: covers, CoverBytesFetch: tlsCovers, + Now: func() time.Time { return time.UnixMilli(5_000_000) }, + Cooldown: time.Hour, BrowserCooldown: time.Hour, + Interval: time.Hour, Batch: 10, + } + p.runOnce(context.Background()) + + if len(tlsF.calls) != 0 { + t.Errorf("TLS fetcher was called for comix: %v", tlsF.calls) + } + if len(browserF.calls) != 1 { + t.Fatalf("browser fetcher calls = %v, want 1", browserF.calls) + } + if got := tlsCovers.callCount(); got != 0 { + t.Errorf("TLS cover fetches = %d, want 0: static.comix.to answers a challenge", got) + } + if got := covers.callCount(); got != 1 { + t.Fatalf("browser cover fetches = %d, want 1", got) + } + got, found, err := s.Get(s.OwnerID(), key) + if err != nil || !found { + t.Fatalf("Get: %v found=%v", err, found) + } + if got.LatestChapterNum == nil || *got.LatestChapterNum != 80 { + t.Errorf("LatestChapterNum = %v, want 80", got.LatestChapterNum) + } +} + +// Without a browser, comix is skipped outright rather than handed to plain +// TLS: a plain fetch retrieves only a challenge page (measured 2026-08-12). +func TestComixSkippedWhenNoBrowserFetcher(t *testing.T) { + s, _ := newTestStore(t) + if _, err := s.Upsert(s.OwnerID(), store.Bookmark{ + Key: "comix:n8we-dungeons-and-crayons", Site: "comix", + SeriesID: "n8we-dungeons-and-crayons", + SeriesURL: "https://comix.to/title/n8we-dungeons-and-crayons", + UpdatedAt: 1000, + }); err != nil { + t.Fatalf("seed: %v", err) + } + f := &fakeFetcher{body: comixSeriesFixture, status: 200} + p := &Poller{ + Store: s, Fetch: f, + Now: func() time.Time { return time.UnixMilli(5_000_000) }, + Cooldown: time.Hour, BrowserCooldown: time.Hour, + Interval: time.Hour, Batch: 10, + } + p.runOnce(context.Background()) + + if len(f.calls) != 0 { + t.Errorf("TLS fetcher was called for comix: %v", f.calls) + } +} + func TestRunOncePrefetchesKaganeCover(t *testing.T) { s, dbURL := newTestStore(t) const ( diff --git a/backend/internal/latest/sites.go b/backend/internal/latest/sites.go index 2c6333b..6c4d3ac 100644 --- a/backend/internal/latest/sites.go +++ b/backend/internal/latest/sites.go @@ -248,14 +248,25 @@ var comixInitialDataRe = regexp.MustCompile(`(?is)]*\bid\s*=\s*["']i // supplied, and a headless browser is a strong SSRF primitive. var kaganeImageURLRe = regexp.MustCompile(`^https://kagane\.to/api/v2/image/([0-9a-f-]{36})/compressed$`) +// comixImageURLRe matches comix's cover host and path shape. Pinned in full +// (scheme, host, path characters, image extension) for the same reason +// kaganeImageURLRe is: the address reaches a headless browser, and it can +// originate in a client-supplied PUT body. No dot is allowed inside the path, +// so no traversal or second extension can hide in it. Shape from a live page, +// 2026-08-10: /039d/i/1/34/6a6742bf15736@280.jpg. +var comixImageURLRe = regexp.MustCompile(`^https://static\.comix\.to/[A-Za-z0-9@/_-]+\.(?:jpg|jpeg|png|webp)$`) + // browserOnlyCoverURL reports whether the browser sidecar is the only fetcher // for cover bytes at imageURL. kagane's image route answers a plain fetch with -// a challenge and `cross-origin-resource-policy: same-origin`, so a TLS fetch -// would only ever retrieve a challenge page and must not be attempted -// (ADR-0007). This is the byte-fetch router's per-Site knowledge; it lives in -// the extraction module, which owns kagane's URL shapes. +// a challenge and `cross-origin-resource-policy: same-origin`, and +// static.comix.to answers one with the same Cloudflare challenge its pages +// serve (measured 2026-08-12, issue #98), so a TLS fetch would only ever +// retrieve a challenge page and must not be attempted (ADR-0007). This is the +// byte-fetch router's per-Site knowledge; it lives in the extraction module, +// which owns those URL shapes. func browserOnlyCoverURL(imageURL string) bool { - return kaganeImageURLRe.MatchString(imageURL) + return kaganeImageURLRe.MatchString(imageURL) || + comixImageURLRe.MatchString(imageURL) } // kagane's browser-fetched series response publishes cover image IDs under @@ -389,6 +400,15 @@ var sites = map[string]site{ Host: "comix.to", LatestChapter: comixLatestChapter, Cover: comixCoverEntry, + Browser: &browserRead{ + Read: comixRead, + // The interstitial is served in place of the page, so "arrived" + // has to exclude it explicitly, as novelfull's does. + Done: func(body string) bool { return body != "" && !isInterstitial(body) }, + // Never falls back: a plain fetch of a comix page or cover + // retrieves only a challenge page (measured 2026-08-12). + Fallback: false, + }, }, "kagane": { Host: "kagane.to", diff --git a/backend/internal/latest/sites_test.go b/backend/internal/latest/sites_test.go index d665b25..c852020 100644 --- a/backend/internal/latest/sites_test.go +++ b/backend/internal/latest/sites_test.go @@ -41,6 +41,10 @@ const challengeFixture = `Just a moment...</ti // https://comix.to/title/n8we-dungeons-and-crayons fetched 2026-08-03. comix is // an SPA: the page ships a JSON state blob rather than a list of chapter // anchors, and latestChapterUrl is where the newest chapter actually lives. +// +// Still the right fixture after comix moved behind the challenge (#98): the +// browser read is an in-tab fetch of the Series URL, so the body a poll parses +// is this same server-rendered HTML, not a rendered DOM. const comixSeriesFixture = ` {"firstChapterUrl":"/title/n8we-dungeons-and-crayons/5038739-chapter-1","latestChapterUrl":"/title/n8we-dungeons-and-crayons/11139891-chapter-80"}, {""manga","recommended","n8we",1]":{"items":[{"latestChapterUrl":"/title/qqwrm-full-time-awakening/99999999-chapter-999"}]} diff --git a/backend/internal/latest/smoke_comix_test.go b/backend/internal/latest/smoke_comix_test.go new file mode 100644 index 0000000..b3946b4 --- /dev/null +++ b/backend/internal/latest/smoke_comix_test.go @@ -0,0 +1,72 @@ +package latest + +import ( + "context" + "os" + "testing" + "time" + + "bookmarkmanager/backend/internal/store" +) + +// TestSmokeComix answers "is comix's challenge clearing from this browser right +// now" — a live, time-varying fact, so a red run is something to re-check +// before it is a defect. Needs the real browser unit with outbound network: +// +// cd chrome && BROWSER_BIND_ADDR=127.0.0.1 docker compose up -d --build +// SMOKE_BROWSER_WS_URL=ws://127.0.0.1:9222 go test -run TestSmokeComix ./internal/latest +// +// It walks the whole read: the in-tab page fetch, both parses, and the Cover +// bytes by direct navigation to static.comix.to. The Cover address comes out of +// the page rather than being pinned in the test, because a stored one rots. +func TestSmokeComix(t *testing.T) { + ws := os.Getenv("SMOKE_BROWSER_WS_URL") + if ws == "" { + t.Skip("SMOKE_BROWSER_WS_URL unset") + } + const seriesURL = "https://comix.to/title/m12d-classmate" + + f, err := NewBrowserFetcher(ws) + if err != nil { + t.Fatalf("NewBrowserFetcher: %v", err) + } + defer f.Close() + + ctx, cancel := context.WithTimeout(context.Background(), 120*time.Second) + defer cancel() + + body, status, err := f.Get(ctx, seriesURL) + if err != nil { + t.Fatalf("Get: %v", err) + } + t.Logf("status=%d bytes=%d", status, len(body)) + if status != 200 { + t.Fatalf("status = %d, want 200 — the sidecar is not clearing the challenge", status) + } + chapter, ok := latestChapterFrom("comix", seriesURL, body) + if !ok { + t.Fatalf("no latest chapter in %d bytes — page shape changed", len(body)) + } + t.Logf("latest chapter: %v %q", chapter.Num, chapter.Label) + + cover, ok := coverFrom("comix", seriesURL, body) + if !ok { + t.Fatalf("no cover address in %d bytes — page shape changed", len(body)) + } + t.Logf("cover: %s", cover) + if !browserOnlyCoverURL(cover) { + t.Fatalf("cover %q is not claimed by the browser gate: the pin and the live URL shape disagree", cover) + } + + bytes, contentType, err := f.Image(ctx, cover) + if err != nil { + t.Fatalf("Image: %v", err) + } + if len(bytes) < 1000 { + t.Fatalf("cover is %d bytes, want a real image", len(bytes)) + } + t.Logf("fetched %d bytes of %s", len(bytes), contentType) + if _, ok := store.CoverContentType(contentType); !ok { + t.Fatalf("content type %q is not storable", contentType) + } +}