feat(latest): poll comix through the browser sidecar (#98)

comix.to began answering plain-TLS fetches with a Cloudflare JavaScript
challenge on 2026-08-12, so every poll got a 403 interstitial and its cover
host static.comix.to is gated the same way. comix joins kagane and novelfull
as a browser Site: one registry entry, no plain-TLS fallback, and cover bytes
routed through the browser's image path behind a fully pinned URL pattern.

The read is an in-tab fetch of the Series URL, not a DOM render: comix is an
SPA, so rendering costs ~65 requests for the same server-rendered HTML one
fetch returns (24.5 KB, ~480 ms). Parsers and stored Series identity are
untouched.

Verified live against the real browser unit: page 24793 bytes in one fetch,
chapter 53, cover accepted by the pin and 26862 image bytes retrieved by
direct navigation (comix's Series page sets cross-origin-embedder-policy:
require-corp, so an in-page fetch of the cover host cannot work).
This commit is contained in:
2026-08-16 12:05:11 +07:00
parent 3f53c79cf4
commit 86160c164a
11 changed files with 336 additions and 55 deletions
+9 -5
View File
@@ -19,9 +19,9 @@ Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free
- Every site is its **own origin with its own `localStorage`** — a shared remote store is the only way to unify bookmarks. Cloud sync required, not optional.
- Userscript run in **isolated world**, so embedded API token safe from site's JS.
- Cloudflare's block on manga sites is **per-zone configuration plus request fingerprint, not IP reputation — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — a Site can turn its protection on overnight, which is exactly what comix.to did on 2026-08-12. An earlier version of this line blamed "Cloudflare's bot scoring"; that was wrong. The 1-99 bot score is Enterprise Bot Management only and does not exist for a free-plan zone, and no per-IP request rate is documented as an input to challenge issuance — `docs/research/cloudflare-bot-scoring-and-poll-cadence.md`. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test.
- **kagane.to and novelfull.com are the exception to the above** — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`). When that's unset, kagane is skipped entirely (a plain fetch would only retrieve a challenge page) while novelfull pages are still attempted over plain TLS — its challenge is a live time-varying fact and its cover bytes never need the browser. The four other sites poll fine over plain TLS.
- **kagane.to, comix.to and novelfull.com are the exception to the above** — all three sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`). When that's unset, kagane and comix are skipped entirely (a plain fetch would only retrieve a challenge page) while novelfull pages are still attempted over plain TLS — its challenge is a live time-varying fact and its cover bytes never need the browser. comix turned hostile on 2026-08-12 (#98): its cover host `static.comix.to` is gated too, so its cover bytes go through the browser as well, and its page is read as an in-tab `fetch()` of the series URL rather than a rendered DOM — comix is an SPA, and rendering costs ~65 requests for the same server-rendered HTML one fetch returns. The three other sites poll fine over plain TLS.
- **The CDP browser must look like a real browser, and stock headless images don't.** Measured 2026-08-08 against kagane.to, all from the same IP: `chromedp/headless-shell:stable` never cleared the challenge in 90s (`navigator.webdriver` true, empty plugin list, Chromium-branded client hints — suppressing `webdriver` alone changed nothing); `zenika/alpine-chrome` ships Chrome 124, refused outright; real Chrome with the default `--headless=new` UA never cleared, because the UA says `HeadlessChrome`; real Chrome with a stock UA **and** a non-UTC clock zone cleared in ~4s. Hence `chrome/` — a Debian image with `google-chrome-stable`, a version-derived UA, and `TZ`/`BROWSER_TZ`. Chrome reads the zone *name* through ICU from `/etc/localtime`'s symlink target, ignoring the file's contents, so mounting the host's `/etc/localtime` does **not** work; `/etc/timezone` is mounted instead.
- **The browser is not in the API stack and must not be put back.** It's its own compose unit (`chrome/docker-compose.yml`) on a second machine, reached over the tailnet — it held 471 MiB on a 1974 MiB swapless VPS, and a residential egress avoids the cloud-hosting-IP signature Bot Fight Mode documentedly challenges (ADR-0006; not a better "score" — free-plan zones have no score). Consequences that constrain code: `BROWSER_WS_URL` must be a tailnet **IP** (a MagicDNS name 500s at `/json/version`, same trap as the old Docker service name); the CDP port binds to the tailnet address only, since CDP authenticates nothing and that host has a real LAN; and the browser is on-demand (ADR-0005), so an unreachable or asleep one must degrade exactly as an unset `BROWSER_WS_URL` — plain-TLS libraries unaffected, kagane/novelfull logged and skipped, stored covers still served. Never add `chromedp.NoModifyURL`: discovery per fetch is what makes a restarted Chrome invisible.
- **The browser is not in the API stack and must not be put back.** It's its own compose unit (`chrome/docker-compose.yml`) on a second machine, reached over the tailnet — it held 471 MiB on a 1974 MiB swapless VPS, and a residential egress avoids the cloud-hosting-IP signature Bot Fight Mode documentedly challenges (ADR-0006; not a better "score" — free-plan zones have no score). Consequences that constrain code: `BROWSER_WS_URL` must be a tailnet **IP** (a MagicDNS name 500s at `/json/version`, same trap as the old Docker service name); the CDP port binds to the tailnet address only, since CDP authenticates nothing and that host has a real LAN; and the browser is on-demand (ADR-0005), so an unreachable or asleep one must degrade exactly as an unset `BROWSER_WS_URL` — plain-TLS libraries unaffected, kagane/comix logged and skipped, stored covers still served. Never add `chromedp.NoModifyURL`: discovery per fetch is what makes a restarted Chrome invisible.
- **UTC is the tell, not a country mismatch.** A UTC clock is the datacenter default, and the challenge refuses it; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while `Asia/Jakarta` **and** `America/New_York` both cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (`Asia/Bangkok`) rather than the measured egress. A second earlier claim, that Cloudflare "scores" a UTC clock, was also wrong: the measurement is real but the mechanism is not documented anywhere — Cloudflare publishes no timezone signal, and free-plan zones carry no score at all. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one.
- **A challenged page needs the tab kept open.** The interstitial takes seconds to solve and only then writes clearance into the browser's shared cookie jar. Navigate-read-close never clears anything; `BrowserFetcher.run` holds one tab and re-reads until the payload arrives.
@@ -45,12 +45,16 @@ Backend (`cd backend`):
- Single test: `go test -run TestName ./...`
- Build static binary: `CGO_ENABLED=0 go build`
Local stack: `docker compose up` (bookmark-api + postgres only; `postgres-data` named volume, `restart: unless-stopped`). No browser — without `BROWSER_WS_URL` the poller logs and skips kagane and novelfull. To run one: `cd chrome && BROWSER_BIND_ADDR=172.17.0.1 docker compose up -d --build`, then `BROWSER_WS_URL=ws://172.17.0.1:9222` in the root `.env` (bridge gateway, so the API container can name it by IP).
Local stack: `docker compose up` (bookmark-api + postgres only; `postgres-data` named volume, `restart: unless-stopped`). No browser — without `BROWSER_WS_URL` the poller logs and skips kagane and comix. To run one: `cd chrome && BROWSER_BIND_ADDR=172.17.0.1 docker compose up -d --build`, then `BROWSER_WS_URL=ws://172.17.0.1:9222` in the root `.env` (bridge gateway, so the API container can name it by IP).
Live CDP proof (needs that browser and network, skipped otherwise):
`SMOKE_BROWSER_WS_URL=ws://<ip>:<port> go test -run TestSmokeKagane ./internal/latest`
— fetches a real kagane cover and chapter list. A red run means the challenge is
`SMOKE_BROWSER_WS_URL=ws://<ip>:<port> go test -run 'TestSmokeKagane|TestSmokeComix' ./internal/latest`
— fetches a real kagane and comix cover and chapter list. A red run means the challenge is
not clearing from this IP, which is a live fact to re-check, not necessarily a defect.
Note for this dev machine: comix.to is DNS-hijacked to an ISP block page here
(`comix.to` CNAMEs to `aduankonten.id`, so Chrome fails `ERR_CERT_COMMON_NAME_INVALID`).
Run the browser container with `--add-host comix.to:<cloudflare-ip> --add-host static.comix.to:<same>`
resolved over DoH to get a real reading.
Smoke test: `curl` endpoints with `Authorization: Bearer <token>`; confirm `OPTIONS` preflight return CORS headers and `/healthz` return 200.
+2 -2
View File
@@ -436,8 +436,8 @@ free -m # the Gitea runner should still have its headroom
Nothing here needs doing during an API redeploy. The API stack does not
`depends_on` the browser, and an unreachable one degrades exactly as an unset
`BROWSER_WS_URL`: plain-TLS libraries unaffected, kagane and novelfull logged
and skipped, stored covers still served.
`BROWSER_WS_URL`: plain-TLS libraries unaffected, kagane and comix logged
and skipped, novelfull attempted over plain TLS, stored covers still served.
---
+29 -20
View File
@@ -89,12 +89,15 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
`Store.SetLatestChapter`, so a bookmark's `updated_at` — and the list
order — is never touched.
Fetches use `bogdanfinn/tls-client` with Chrome profile as defence in depth
against fingerprint-based blocking; any failure log and skip. kagane and
novelfull sit behind Cloudflare JavaScript challenges the TLS client can't
clear, so they are fetched over CDP via `BROWSER_WS_URL`; kagane is simply
not polled when that's unset, while novelfull falls back to a plain-TLS
attempt — its challenge is a live time-varying fact, and its cover bytes
never need the browser. See
against fingerprint-based blocking; any failure log and skip. kagane, comix
and novelfull sit behind Cloudflare JavaScript challenges the TLS client
can't clear, so they are fetched over CDP via `BROWSER_WS_URL`; kagane and
comix are simply not polled when that's unset, while novelfull falls back to
a plain-TLS attempt — its challenge is a live time-varying fact, and its
cover bytes never need the browser. comix's browser read is an in-tab
`fetch()` of the Series URL, not a DOM render: it is an SPA, so rendering
costs ~65 requests for the same server-rendered HTML one fetch returns
(measured 2026-08-12, issue #98). See
`docs/superpowers/specs/2026-07-26-server-latest-chapter-polling-design.md`.
The poller's series write is a single-column UPDATE
(`Store.SetLatestChapter`), not a read-modify-write of the whole bookmark:
@@ -112,14 +115,17 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
is public and uncredentialed: the userscript renders it on a Site's origin,
where no cookie or token of ours travels. A client-sent `cover` is decoded
and discarded, permanently (ADR-0004 compatibility).
Browser-backed Sites join the same pipeline (issue #62): kagane pages *and*
cover bytes go through the browser sidecar (nothing falls back to a plain
fetch, which would only retrieve a challenge page), while novelfull needs
the browser only for its HTML — the cover URL comes out of the
browser-fetched page and the bytes go over plain TLS. With no browser
configured, kagane Covers are simply absent; novelfull still gets one — at
creation and on the poll — when its page body happens to answer a plain
request (the challenge is a live time-varying fact). The old kagane-only
Browser-backed Sites join the same pipeline (issue #62, extended to comix by
#98): kagane and comix pages *and* cover bytes go through the browser sidecar
(nothing falls back to a plain fetch, which would only retrieve a challenge
page), while novelfull needs the browser only for its HTML — the cover URL
comes out of the browser-fetched page and the bytes go over plain TLS. With
no browser configured, kagane and comix Covers are simply absent; novelfull
still gets one — at creation and on the poll — when its page body happens to
answer a plain request (the challenge is a live time-varying fact). comix
cover bytes must arrive by direct navigation, not an in-page fetch: its
Series page sets `cross-origin-embedder-policy: require-corp`, which fails a
page-context fetch of `static.comix.to`. The old kagane-only
serving path (`/img/kagane/{id}`, template rewrite, `CoverFetcher`) is gone
(issue #63): the one public route serves every Site.
- **`updated_at` drives list order, so moves only on real reading progress:** server apply its timestamp when row new or `last_chapter_num` changes, else keep stored value — favouriting series or recording newly published chapter must not reorder list. `PUT` therefore returns row **as stored**, clients must adopt that response rather than own payload. See `plans/2026-07-25-bookmark-list-favorites-design.md` §4.
@@ -164,10 +170,11 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
Reader's credential at serve time).
`BROWSER_WS_URL` (CDP endpoint of the browser, which runs on a **separate
machine** and is reached over the tailnet — ADR-0006, `chrome/docker-compose.yml`.
Used by the poller for kagane and novelfull page fetches and by the cover
pipeline for kagane's image bytes (the browser is the only route that clears
the challenge kagane serves its covers behind); unset — the default —
disables browser polling and leaves kagane Covers blank until stored bytes
Used by the poller for kagane, comix and novelfull page fetches and by the
cover pipeline for kagane's and comix's image bytes (the browser is the only
route that clears the challenge those two serve their covers behind); unset —
the default — disables browser polling and leaves kagane and comix Covers
blank until stored bytes
exist. Must be a tailnet IP, never a hostname: Chrome's DevTools handler 500s
`/json/version` for any Host that isn't an IP or `localhost`).
- **No per-Site cover path (issue #63):** every Cover — all six Sites — is
@@ -175,8 +182,10 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
bytes. There is no proxy, no per-Site rewrite, no second place that decides
a Cover's renderable address: the wire `cover` is it. The only place a Site
name still appears in cover code is the extraction module (`latest`), where
kagane's image URLs are claimed by `browserOnlyCoverURL` — they answer a
plain fetch with a challenge and `cross-origin-resource-policy: same-origin`;
kagane's and comix's image URLs are claimed by `browserOnlyCoverURL` — kagane
answers a plain fetch with a challenge and
`cross-origin-resource-policy: same-origin`, and `static.comix.to` answers
one with the same Cloudflare challenge its pages serve;
every other Site's CDN answers plain TLS. Templates render `.Cover` — the
wire value — never anything else.
- **Web UI also owns:** session-gated `GET /install/{manga,novel}-bookmark.user.js`
+52 -17
View File
@@ -24,12 +24,16 @@ const challengeTimeout = 45 * time.Second
var kaganeSeriesRe = regexp.MustCompile(`^/series/([0-9a-f-]{36})/?$`)
// comixSeriesPathRe matches the one path shape comixRead will open: a Series
// page, "/title/<id>-<slug>". Verified live 2026-08-12.
var comixSeriesPathRe = regexp.MustCompile(`^/title/[^/?#]+/?$`)
// BrowserFetcher retrieves pages through a remote headless Chrome over the
// DevTools Protocol.
//
// It exists for one reason: kagane.to and novelfull.com sit behind a
// Cloudflare JavaScript challenge. Verified 2026-08-03 (kagane) and 2026-08-05
// (novelfull) from the deployment host, plain HTTP and bogdanfinn/tls-client
// It exists for one reason: kagane.to, novelfull.com and comix.to sit behind a
// Cloudflare JavaScript challenge. Verified 2026-08-03 (kagane), 2026-08-05
// (novelfull) and 2026-08-12 (comix), plain HTTP and bogdanfinn/tls-client
// with a Chrome_133 profile both get 403 with cf-mitigated: challenge on every
// path, including the API, robots.txt and images. Clearing it requires
// executing the challenge script, which only a real browser does.
@@ -141,29 +145,47 @@ func novelfullRead(seriesURL string, out *string) (chromedp.Action, bool) {
return chromedp.OuterHTML("html", out, chromedp.ByQuery), true
}
// comixRead fetches the Series page from inside the cleared tab. comix is an
// SPA: rendering the page costs ~65 requests, while one same-origin fetch of
// the same address returns the server-rendered HTML — 24.5 KB, ~480 ms,
// carrying both parser anchors (measured 2026-08-12, issue #98). So this is
// kaganeRead's shape, not novelfullRead's, even though the payload is HTML.
// Refusing any other address is the per-Site half of the SSRF gate.
func comixRead(seriesURL string, out *string) (chromedp.Action, bool) {
pageURL, ok := comixSeriesPageURL(seriesURL)
if !ok {
return nil, false
}
return chromedp.Evaluate(
`fetch(`+jsString(pageURL)+`).then(r => r.ok ? r.text() : "")`,
out, awaitPromise), true
}
// Image retrieves one cover's bytes through the browser sidecar, and its
// content type.
//
// It exists because kagane serves covers behind the same challenge as its
// pages *and* with `cross-origin-resource-policy: same-origin`, so the bytes
// are only reachable from inside a browser that already holds the clearance
// cookie (verified 2026-08-08). Acquisition through the sidecar is the only
// route.
// It exists because kagane and comix serve covers behind the same challenge as
// their pages — kagane additionally with
// `cross-origin-resource-policy: same-origin` — so the bytes are only
// reachable from inside a browser that already holds the clearance cookie
// (verified 2026-08-08 for kagane, 2026-08-12 for comix). Acquisition through
// the sidecar is the only route.
//
// The image URL is navigated to rather than fetched from some other kagane
// page: the challenge only runs on a top-level navigation, and once it clears
// The image URL is navigated to rather than fetched from another page of the
// Site: the challenge only runs on a top-level navigation, and once it clears
// the document *is* the image, so a same-origin fetch of location.href reads
// it straight back out of the cache.
// it straight back out of the cache. For comix the navigation is also the only
// route that works at all — its Series page sets
// `cross-origin-embedder-policy: require-corp`, which fails a page-context
// fetch of the cover host.
//
// The challenge is not solved by the first read: WaitReady("body") is satisfied
// by the interstitial too. run holds the tab open until the in-page fetch
// succeeds, which is what gives the challenge script the seconds it needs.
func (f *BrowserFetcher) Image(ctx context.Context, imageURL string) ([]byte, string, error) {
m := kaganeImageURLRe.FindStringSubmatch(imageURL)
if m == nil {
if !browserOnlyCoverURL(imageURL) {
return nil, "", fmt.Errorf("not a browser-fetchable cover url: %q", imageURL)
}
imageID := m[1]
var dataURL string
err := f.run(ctx, imageURL,
chromedp.Evaluate(`fetch(location.href).then(r => r.ok
@@ -175,16 +197,16 @@ func (f *BrowserFetcher) Image(ctx context.Context, imageURL string) ([]byte, st
: "")`, &dataURL, awaitPromise),
func() bool { return dataURL != "" })
if err != nil {
return nil, "", fmt.Errorf("browser image %s: %w", imageID, err)
return nil, "", fmt.Errorf("browser image %s: %w", imageURL, err)
}
// "data:image/webp;base64,<payload>".
head, payload, ok := strings.Cut(dataURL, ";base64,")
if !ok {
return nil, "", fmt.Errorf("browser image %s: not a data url", imageID)
return nil, "", fmt.Errorf("browser image %s: not a data url", imageURL)
}
raw, err := base64.StdEncoding.DecodeString(payload)
if err != nil {
return nil, "", fmt.Errorf("browser image %s: %w", imageID, err)
return nil, "", fmt.Errorf("browser image %s: %w", imageURL, err)
}
return raw, strings.TrimPrefix(head, "data:"), nil
}
@@ -323,6 +345,19 @@ func novelfullSeriesURL(seriesURL string) bool {
strings.HasSuffix(u.Path, ".html")
}
// comixSeriesPageURL returns the address comixRead fetches inside the tab: the
// Series page itself, rebuilt from the pinned host and path so nothing else
// travels. Host-pinned here for the same reason kagane's is — series_url is
// client-supplied and a headless browser is a strong SSRF primitive.
func comixSeriesPageURL(seriesURL string) (string, bool) {
u, err := url.Parse(seriesURL)
if err != nil || u.Scheme != "https" || u.Hostname() != "comix.to" ||
!comixSeriesPathRe.MatchString(u.Path) {
return "", false
}
return "https://comix.to" + u.Path, true
}
// awaitPromise makes Evaluate resolve the promise rather than returning a
// serialised Promise object.
func awaitPromise(p *runtime.EvaluateParams) *runtime.EvaluateParams {
+56
View File
@@ -61,6 +61,62 @@ func TestNovelfullSeriesURL(t *testing.T) {
})
}
}
func TestComixSeriesPageURL(t *testing.T) {
const series = "https://comix.to/title/n8we-dungeons-and-crayons"
cases := []struct {
name string
url string
want string
}{
{"series page", series, series},
{"trailing slash kept", series + "/", series + "/"},
// Query and fragment are dropped: only the pinned path travels.
{"query dropped", series + "?tab=chapters", series},
{"foreign host", "https://evil.example/title/x", ""},
{"lookalike host", "https://comix.to.evil.example/title/x", ""},
{"not https", "http://comix.to/title/x", ""},
{"not a series path", "https://comix.to/search", ""},
{"chapter page", series + "/11139891-chapter-80", ""},
{"garbage", "://nope", ""},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
got, ok := comixSeriesPageURL(tc.url)
if ok != (tc.want != "") || got != tc.want {
t.Fatalf("comixSeriesPageURL(%q) = %q, %v; want %q", tc.url, got, ok, tc.want)
}
})
}
}
// The browser is an SSRF primitive and a cover address can originate in a
// client-supplied PUT body, so this gate decides what it may navigate to.
func TestBrowserOnlyCoverURL(t *testing.T) {
cases := []struct {
url string
want bool
}{
{"https://static.comix.to/039d/i/1/34/6a6742bf15736@280.jpg", true},
{"https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed", true},
// Every other Site's CDN answers plain TLS.
{"https://gg.asuracomic.net/covers/x.webp", false},
{"http://static.comix.to/039d/x.jpg", false},
{"https://static.comix.to.evil.example/039d/x.jpg", false},
{"https://evil.example/static.comix.to/x.jpg", false},
{"https://static.comix.to/039d/x.jpg?next=http://169.254.169.254/", false},
{"https://static.comix.to/039d/x.svg", false},
{"https://static.comix.to/../etc/passwd.jpg", false},
{"https://static.comix.to/", false},
}
for _, tc := range cases {
t.Run(tc.url, func(t *testing.T) {
if got := browserOnlyCoverURL(tc.url); got != tc.want {
t.Fatalf("browserOnlyCoverURL(%q) = %v, want %v", tc.url, got, tc.want)
}
})
}
}
func TestClassifyBrowserInterruption(t *testing.T) {
if err := classifyBrowserError(context.Background(), true, context.Canceled); !errors.Is(err, errBrowserInterrupted) {
t.Fatalf("classifyBrowserError(context.Canceled) = %v, want browser interruption", err)
+2 -1
View File
@@ -25,7 +25,8 @@ type CoverBytesFetcher interface {
// fetchCoverBytes routes a cover's byte retrieval by URL shape, not by Site
// name: the browser fetcher's module claims the addresses only it can fetch
// (kagane's image route answers a plain fetch with a challenge and
// `cross-origin-resource-policy: same-origin`), and everything else goes over
// `cross-origin-resource-policy: same-origin`, static.comix.to answers one with
// the same challenge its pages serve), and everything else goes over
// plain TLS. Missing fetchers degrade to an error the caller logs, never a
// fallback onto a path that cannot succeed. One routing rule for the poll and
// the acquirer, so the two cannot drift apart.
+5 -5
View File
@@ -17,8 +17,8 @@ type Fetcher interface {
}
// BrowserCoverFetcher retrieves one cover's bytes through the browser-backed
// path — the only route that clears the challenge kagane's image URLs answer
// a plain fetch with. Satisfied by BrowserFetcher.
// path — the only route that clears the challenge kagane's and comix's image
// URLs answer a plain fetch with. Satisfied by BrowserFetcher.
type BrowserCoverFetcher interface {
Image(ctx context.Context, imageURL string) (body []byte, contentType string, err error)
}
@@ -119,9 +119,9 @@ func (p *Poller) storeCover(ctx context.Context, sr store.Series, sourceURL stri
// fetcherFor returns the fetcher a site's page needs, or nil when the site
// cannot be fetched at all right now. A Site whose registry entry carries a
// Browser read — kagane and novelfull, both behind a Cloudflare JavaScript
// challenge no TLS fingerprint clears — prefers the browser; when it is
// absent, the entry's Fallback decides whether plain TLS may take over. One
// Browser read — kagane, comix and novelfull, all behind a Cloudflare
// JavaScript challenge no TLS fingerprint clears — prefers the browser; when it
// is absent, the entry's Fallback decides whether plain TLS may take over. One
// routing rule for the poll and the acquirer, so the two cannot drift apart.
func fetcherFor(site string, browser, tls Fetcher) Fetcher {
s, known := sites[site]
+80
View File
@@ -762,6 +762,86 @@ func TestKaganeUsesBrowserFetcher(t *testing.T) {
}
}
// comix joined kagane behind the challenge on 2026-08-12 (#98): its page goes
// to the browser, its Cover bytes go through the browser's image route because
// static.comix.to is gated the same way, and the TLS fetcher is never asked
// for either.
func TestComixUsesBrowserFetcher(t *testing.T) {
s, dbURL := newTestStore(t)
const (
key = "comix:n8we-dungeons-and-crayons"
seriesID = "n8we-dungeons-and-crayons"
seriesURL = "https://comix.to/title/n8we-dungeons-and-crayons"
coverURL = "https://static.comix.to/039d/i/1/34/6a6742bf15736@280.jpg"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "comix", SeriesID: seriesID, SeriesURL: seriesURL,
UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
seedCoverSource(t, dbURL, "comix", seriesID, coverURL)
tlsF := &fakeFetcher{body: "", status: 200}
browserF := &fakeFetcher{body: comixSeriesFixture, status: 200}
covers := &fakeCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"}
tlsCovers := &fakeBytesCoverFetcher{body: []byte("tls-bytes"), contentType: "image/jpeg"}
p := &Poller{
Store: s, Fetch: tlsF, BrowserFetch: browserF,
CoverFetch: covers, CoverBytesFetch: tlsCovers,
Now: func() time.Time { return time.UnixMilli(5_000_000) },
Cooldown: time.Hour, BrowserCooldown: time.Hour,
Interval: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if len(tlsF.calls) != 0 {
t.Errorf("TLS fetcher was called for comix: %v", tlsF.calls)
}
if len(browserF.calls) != 1 {
t.Fatalf("browser fetcher calls = %v, want 1", browserF.calls)
}
if got := tlsCovers.callCount(); got != 0 {
t.Errorf("TLS cover fetches = %d, want 0: static.comix.to answers a challenge", got)
}
if got := covers.callCount(); got != 1 {
t.Fatalf("browser cover fetches = %d, want 1", got)
}
got, found, err := s.Get(s.OwnerID(), key)
if err != nil || !found {
t.Fatalf("Get: %v found=%v", err, found)
}
if got.LatestChapterNum == nil || *got.LatestChapterNum != 80 {
t.Errorf("LatestChapterNum = %v, want 80", got.LatestChapterNum)
}
}
// Without a browser, comix is skipped outright rather than handed to plain
// TLS: a plain fetch retrieves only a challenge page (measured 2026-08-12).
func TestComixSkippedWhenNoBrowserFetcher(t *testing.T) {
s, _ := newTestStore(t)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: "comix:n8we-dungeons-and-crayons", Site: "comix",
SeriesID: "n8we-dungeons-and-crayons",
SeriesURL: "https://comix.to/title/n8we-dungeons-and-crayons",
UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
f := &fakeFetcher{body: comixSeriesFixture, status: 200}
p := &Poller{
Store: s, Fetch: f,
Now: func() time.Time { return time.UnixMilli(5_000_000) },
Cooldown: time.Hour, BrowserCooldown: time.Hour,
Interval: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if len(f.calls) != 0 {
t.Errorf("TLS fetcher was called for comix: %v", f.calls)
}
}
func TestRunOncePrefetchesKaganeCover(t *testing.T) {
s, dbURL := newTestStore(t)
const (
+25 -5
View File
@@ -248,14 +248,25 @@ var comixInitialDataRe = regexp.MustCompile(`(?is)<script\b[^>]*\bid\s*=\s*["']i
// supplied, and a headless browser is a strong SSRF primitive.
var kaganeImageURLRe = regexp.MustCompile(`^https://kagane\.to/api/v2/image/([0-9a-f-]{36})/compressed$`)
// comixImageURLRe matches comix's cover host and path shape. Pinned in full
// (scheme, host, path characters, image extension) for the same reason
// kaganeImageURLRe is: the address reaches a headless browser, and it can
// originate in a client-supplied PUT body. No dot is allowed inside the path,
// so no traversal or second extension can hide in it. Shape from a live page,
// 2026-08-10: /039d/i/1/34/6a6742bf15736@280.jpg.
var comixImageURLRe = regexp.MustCompile(`^https://static\.comix\.to/[A-Za-z0-9@/_-]+\.(?:jpg|jpeg|png|webp)$`)
// browserOnlyCoverURL reports whether the browser sidecar is the only fetcher
// for cover bytes at imageURL. kagane's image route answers a plain fetch with
// a challenge and `cross-origin-resource-policy: same-origin`, so a TLS fetch
// would only ever retrieve a challenge page and must not be attempted
// (ADR-0007). This is the byte-fetch router's per-Site knowledge; it lives in
// the extraction module, which owns kagane's URL shapes.
// a challenge and `cross-origin-resource-policy: same-origin`, and
// static.comix.to answers one with the same Cloudflare challenge its pages
// serve (measured 2026-08-12, issue #98), so a TLS fetch would only ever
// retrieve a challenge page and must not be attempted (ADR-0007). This is the
// byte-fetch router's per-Site knowledge; it lives in the extraction module,
// which owns those URL shapes.
func browserOnlyCoverURL(imageURL string) bool {
return kaganeImageURLRe.MatchString(imageURL)
return kaganeImageURLRe.MatchString(imageURL) ||
comixImageURLRe.MatchString(imageURL)
}
// kagane's browser-fetched series response publishes cover image IDs under
@@ -389,6 +400,15 @@ var sites = map[string]site{
Host: "comix.to",
LatestChapter: comixLatestChapter,
Cover: comixCoverEntry,
Browser: &browserRead{
Read: comixRead,
// The interstitial is served in place of the page, so "arrived"
// has to exclude it explicitly, as novelfull's does.
Done: func(body string) bool { return body != "" && !isInterstitial(body) },
// Never falls back: a plain fetch of a comix page or cover
// retrieves only a challenge page (measured 2026-08-12).
Fallback: false,
},
},
"kagane": {
Host: "kagane.to",
+4
View File
@@ -41,6 +41,10 @@ const challengeFixture = `<!DOCTYPE html><html><head><title>Just a moment...</ti
// https://comix.to/title/n8we-dungeons-and-crayons fetched 2026-08-03. comix is
// an SPA: the page ships a JSON state blob rather than a list of chapter
// anchors, and latestChapterUrl is where the newest chapter actually lives.
//
// Still the right fixture after comix moved behind the challenge (#98): the
// browser read is an in-tab fetch of the Series URL, so the body a poll parses
// is this same server-rendered HTML, not a rendered DOM.
const comixSeriesFixture = `
{"firstChapterUrl":"/title/n8we-dungeons-and-crayons/5038739-chapter-1","latestChapterUrl":"/title/n8we-dungeons-and-crayons/11139891-chapter-80"},
{""manga","recommended","n8we",1]":{"items":[{"latestChapterUrl":"/title/qqwrm-full-time-awakening/99999999-chapter-999"}]}
@@ -0,0 +1,72 @@
package latest
import (
"context"
"os"
"testing"
"time"
"bookmarkmanager/backend/internal/store"
)
// TestSmokeComix answers "is comix's challenge clearing from this browser right
// now" — a live, time-varying fact, so a red run is something to re-check
// before it is a defect. Needs the real browser unit with outbound network:
//
// cd chrome && BROWSER_BIND_ADDR=127.0.0.1 docker compose up -d --build
// SMOKE_BROWSER_WS_URL=ws://127.0.0.1:9222 go test -run TestSmokeComix ./internal/latest
//
// It walks the whole read: the in-tab page fetch, both parses, and the Cover
// bytes by direct navigation to static.comix.to. The Cover address comes out of
// the page rather than being pinned in the test, because a stored one rots.
func TestSmokeComix(t *testing.T) {
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
if ws == "" {
t.Skip("SMOKE_BROWSER_WS_URL unset")
}
const seriesURL = "https://comix.to/title/m12d-classmate"
f, err := NewBrowserFetcher(ws)
if err != nil {
t.Fatalf("NewBrowserFetcher: %v", err)
}
defer f.Close()
ctx, cancel := context.WithTimeout(context.Background(), 120*time.Second)
defer cancel()
body, status, err := f.Get(ctx, seriesURL)
if err != nil {
t.Fatalf("Get: %v", err)
}
t.Logf("status=%d bytes=%d", status, len(body))
if status != 200 {
t.Fatalf("status = %d, want 200 — the sidecar is not clearing the challenge", status)
}
chapter, ok := latestChapterFrom("comix", seriesURL, body)
if !ok {
t.Fatalf("no latest chapter in %d bytes — page shape changed", len(body))
}
t.Logf("latest chapter: %v %q", chapter.Num, chapter.Label)
cover, ok := coverFrom("comix", seriesURL, body)
if !ok {
t.Fatalf("no cover address in %d bytes — page shape changed", len(body))
}
t.Logf("cover: %s", cover)
if !browserOnlyCoverURL(cover) {
t.Fatalf("cover %q is not claimed by the browser gate: the pin and the live URL shape disagree", cover)
}
bytes, contentType, err := f.Image(ctx, cover)
if err != nil {
t.Fatalf("Image: %v", err)
}
if len(bytes) < 1000 {
t.Fatalf("cover is %d bytes, want a real image", len(bytes))
}
t.Logf("fetched %d bytes of %s", len(bytes), contentType)
if _, ok := store.CoverContentType(contentType); !ok {
t.Fatalf("content type %q is not storable", contentType)
}
}