Poll comix.to through the browser sidecar (#98) (#105)

Closes #98.

comix.to began answering plain-TLS fetches with a Cloudflare JavaScript
challenge on 2026-08-12, so every poll got a 403 interstitial. Its cover host
`static.comix.to` is gated the same way. comix therefore joins kagane and
novelfull as a browser-backed Site.

## What changed

- **Registry** (`internal/latest/sites.go`): comix gains a `Browser` entry —
  `comixRead`, `Done: body != "" && !isInterstitial(body)`, `Fallback: false`.
  Skip-when-no-browser falls out of the existing routing; no site-string compare
  was added anywhere.
- **Read shape** (`internal/latest/browser.go`): an in-tab `fetch()` of the
  Series URL, not a DOM render. comix is an SPA — rendering it costs ~65
  requests for the same server-rendered HTML one fetch returns (24.5 KB,
  ~480 ms measured). `comixSeriesPageURL` pins scheme + host + `/title/<slug>`
  and rebuilds the address, so a client-supplied `series_url` cannot aim the
  browser anywhere else.
- **Cover bytes**: `comixImageURLRe` pins `https://static.comix.to/<path>.<ext>`;
  `BrowserFetcher.Image` now gates on `browserOnlyCoverURL` rather than a
  kagane-only regex, so both Sites' image URLs route through the one path.
  Bytes come from direct navigation, not a page-context fetch — comix's Series
  page sets `cross-origin-embedder-policy: require-corp`, which fails one.
- **Parsers and stored Series identity: untouched.** The in-tab body is the same
  server-rendered HTML the existing fixtures were cut from.

## Verification

- `go test ./...` green (needs Docker).
- New seam tests: comix routes to the browser when one is configured, and is
  not fetched at all when none is (`TestComixUsesBrowserFetcher`,
  `TestComixSkippedWhenNoBrowserFetcher`); URL-pin and cover-gate table tests.
- Live proof against the real browser unit, `TestSmokeComix` (env-gated):
  page 24793 bytes in one in-tab fetch, chapter 53, cover accepted by the pin,
  26862 bytes of `image/jpg` retrieved.
- Two-axis review run; findings were stale comments on `BrowserFetcher`, `Get`
  and the `Fallback` field, fixed in f000cc7.

Docs updated: root `AGENTS.md` (constraint + smoke command, including the note
that this dev machine's ISP DNS-hijacks `comix.to`), `backend/AGENTS.md`
(poller, cover pipeline, `BROWSER_WS_URL`), `REDEPLOY.md` §8 degrade note.

Reviewed-on: #105
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
This commit was merged in pull request #105.
This commit is contained in:
2026-08-16 12:13:59 +07:00
committed by sulthan
parent 3f53c79cf4
commit ddbd57070d
16 changed files with 1265 additions and 1006 deletions
+29 -20
View File
@@ -89,12 +89,15 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
`Store.SetLatestChapter`, so a bookmark's `updated_at` — and the list
order — is never touched.
Fetches use `bogdanfinn/tls-client` with Chrome profile as defence in depth
against fingerprint-based blocking; any failure log and skip. kagane and
novelfull sit behind Cloudflare JavaScript challenges the TLS client can't
clear, so they are fetched over CDP via `BROWSER_WS_URL`; kagane is simply
not polled when that's unset, while novelfull falls back to a plain-TLS
attempt — its challenge is a live time-varying fact, and its cover bytes
never need the browser. See
against fingerprint-based blocking; any failure log and skip. kagane, comix
and novelfull sit behind Cloudflare JavaScript challenges the TLS client
can't clear, so they are fetched over CDP via `BROWSER_WS_URL`; kagane and
comix are simply not polled when that's unset, while novelfull falls back to
a plain-TLS attempt — its challenge is a live time-varying fact, and its
cover bytes never need the browser. comix's browser read is an in-tab
`fetch()` of the Series URL, not a DOM render: it is an SPA, so rendering
costs ~65 requests for the same server-rendered HTML one fetch returns
(measured 2026-08-12, issue #98). See
`docs/superpowers/specs/2026-07-26-server-latest-chapter-polling-design.md`.
The poller's series write is a single-column UPDATE
(`Store.SetLatestChapter`), not a read-modify-write of the whole bookmark:
@@ -112,14 +115,17 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
is public and uncredentialed: the userscript renders it on a Site's origin,
where no cookie or token of ours travels. A client-sent `cover` is decoded
and discarded, permanently (ADR-0004 compatibility).
Browser-backed Sites join the same pipeline (issue #62): kagane pages *and*
cover bytes go through the browser sidecar (nothing falls back to a plain
fetch, which would only retrieve a challenge page), while novelfull needs
the browser only for its HTML — the cover URL comes out of the
browser-fetched page and the bytes go over plain TLS. With no browser
configured, kagane Covers are simply absent; novelfull still gets one — at
creation and on the poll — when its page body happens to answer a plain
request (the challenge is a live time-varying fact). The old kagane-only
Browser-backed Sites join the same pipeline (issue #62, extended to comix by
#98): kagane and comix pages *and* cover bytes go through the browser sidecar
(nothing falls back to a plain fetch, which would only retrieve a challenge
page), while novelfull needs the browser only for its HTML — the cover URL
comes out of the browser-fetched page and the bytes go over plain TLS. With
no browser configured, kagane and comix Covers are simply absent; novelfull
still gets one — at creation and on the poll — when its page body happens to
answer a plain request (the challenge is a live time-varying fact). comix
cover bytes must arrive by direct navigation, not an in-page fetch: its
Series page sets `cross-origin-embedder-policy: require-corp`, which fails a
page-context fetch of `static.comix.to`. The old kagane-only
serving path (`/img/kagane/{id}`, template rewrite, `CoverFetcher`) is gone
(issue #63): the one public route serves every Site.
- **`updated_at` drives list order, so moves only on real reading progress:** server apply its timestamp when row new or `last_chapter_num` changes, else keep stored value — favouriting series or recording newly published chapter must not reorder list. `PUT` therefore returns row **as stored**, clients must adopt that response rather than own payload. See `plans/2026-07-25-bookmark-list-favorites-design.md` §4.
@@ -164,10 +170,11 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
Reader's credential at serve time).
`BROWSER_WS_URL` (CDP endpoint of the browser, which runs on a **separate
machine** and is reached over the tailnet — ADR-0006, `chrome/docker-compose.yml`.
Used by the poller for kagane and novelfull page fetches and by the cover
pipeline for kagane's image bytes (the browser is the only route that clears
the challenge kagane serves its covers behind); unset — the default —
disables browser polling and leaves kagane Covers blank until stored bytes
Used by the poller for kagane, comix and novelfull page fetches and by the
cover pipeline for kagane's and comix's image bytes (the browser is the only
route that clears the challenge those two serve their covers behind); unset —
the default — disables browser polling and leaves kagane and comix Covers
blank until stored bytes
exist. Must be a tailnet IP, never a hostname: Chrome's DevTools handler 500s
`/json/version` for any Host that isn't an IP or `localhost`).
- **No per-Site cover path (issue #63):** every Cover — all six Sites — is
@@ -175,8 +182,10 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
bytes. There is no proxy, no per-Site rewrite, no second place that decides
a Cover's renderable address: the wire `cover` is it. The only place a Site
name still appears in cover code is the extraction module (`latest`), where
kagane's image URLs are claimed by `browserOnlyCoverURL` — they answer a
plain fetch with a challenge and `cross-origin-resource-policy: same-origin`;
kagane's and comix's image URLs are claimed by `browserOnlyCoverURL` — kagane
answers a plain fetch with a challenge and
`cross-origin-resource-policy: same-origin`, and `static.comix.to` answers
one with the same Cloudflare challenge its pages serve;
every other Site's CDN answers plain TLS. Templates render `.Cover` — the
wire value — never anything else.
- **Web UI also owns:** session-gated `GET /install/{manga,novel}-bookmark.user.js`
+60 -25
View File
@@ -24,12 +24,16 @@ const challengeTimeout = 45 * time.Second
var kaganeSeriesRe = regexp.MustCompile(`^/series/([0-9a-f-]{36})/?$`)
// comixSeriesPathRe matches the one path shape comixRead will open: a Series
// page, "/title/<id>-<slug>". Verified live 2026-08-12.
var comixSeriesPathRe = regexp.MustCompile(`^/title/[^/?#]+/?$`)
// BrowserFetcher retrieves pages through a remote headless Chrome over the
// DevTools Protocol.
//
// It exists for one reason: kagane.to and novelfull.com sit behind a
// Cloudflare JavaScript challenge. Verified 2026-08-03 (kagane) and 2026-08-05
// (novelfull) from the deployment host, plain HTTP and bogdanfinn/tls-client
// It exists for one reason: kagane.to, novelfull.com and comix.to sit behind a
// Cloudflare JavaScript challenge. Verified 2026-08-03 (kagane), 2026-08-05
// (novelfull) and 2026-08-12 (comix), plain HTTP and bogdanfinn/tls-client
// with a Chrome_133 profile both get 403 with cf-mitigated: challenge on every
// path, including the API, robots.txt and images. Clearing it requires
// executing the challenge script, which only a real browser does.
@@ -40,10 +44,11 @@ var kaganeSeriesRe = regexp.MustCompile(`^/series/([0-9a-f-]{36})/?$`)
// sync that break silently and separately. The browser's own cookie jar
// persists across polls, so the challenge is solved once every few hours.
//
// The two sites differ in how the chapter list is read: kagane serves it from
// a JSON API that must be called from inside the page (so the request carries
// the clearance cookie), while novelfull renders it into the HTML so the
// cleared DOM is the payload.
// The three sites differ in what a cleared tab is asked for: kagane fetches a
// JSON API from inside the page (the list exists nowhere else), comix fetches
// its own Series URL from inside the page (the served HTML carries the facts,
// and rendering the SPA costs ~65 requests instead of one), and novelfull
// renders its list into the HTML so the cleared DOM is the payload.
type BrowserFetcher struct {
allocCtx context.Context
cancel context.CancelFunc
@@ -87,10 +92,9 @@ func (f *BrowserFetcher) Close() {
}
// Get navigates to seriesURL, lets any challenge resolve, then reads the
// payload the Site's registry entry describes — kagane's chapter-list API from
// inside the page so the request carries the clearance cookie, novelfull's
// served HTML. The returned body is whatever the Site's chapter list lives in,
// which is what the entry's LatestChapter parse expects.
// payload the Site's registry entry describes (the shapes are listed on
// BrowserFetcher). The returned body is whatever the Site's chapter list lives
// in, which is what the entry's LatestChapter parse expects.
func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int, error) {
var body string
// Sorted order (browserBackedSites sorts) makes dispatch deterministic:
@@ -141,29 +145,47 @@ func novelfullRead(seriesURL string, out *string) (chromedp.Action, bool) {
return chromedp.OuterHTML("html", out, chromedp.ByQuery), true
}
// comixRead fetches the Series page from inside the cleared tab. comix is an
// SPA: rendering the page costs ~65 requests, while one same-origin fetch of
// the same address returns the server-rendered HTML — 24.5 KB, ~480 ms,
// carrying both parser anchors (measured 2026-08-12, issue #98). So this is
// kaganeRead's shape, not novelfullRead's, even though the payload is HTML.
// Refusing any other address is the per-Site half of the SSRF gate.
func comixRead(seriesURL string, out *string) (chromedp.Action, bool) {
pageURL, ok := comixSeriesPageURL(seriesURL)
if !ok {
return nil, false
}
return chromedp.Evaluate(
`fetch(`+jsString(pageURL)+`).then(r => r.ok ? r.text() : "")`,
out, awaitPromise), true
}
// Image retrieves one cover's bytes through the browser sidecar, and its
// content type.
//
// It exists because kagane serves covers behind the same challenge as its
// pages *and* with `cross-origin-resource-policy: same-origin`, so the bytes
// are only reachable from inside a browser that already holds the clearance
// cookie (verified 2026-08-08). Acquisition through the sidecar is the only
// route.
// It exists because kagane and comix serve covers behind the same challenge as
// their pages — kagane additionally with
// `cross-origin-resource-policy: same-origin` — so the bytes are only
// reachable from inside a browser that already holds the clearance cookie
// (verified 2026-08-08 for kagane, 2026-08-12 for comix). Acquisition through
// the sidecar is the only route.
//
// The image URL is navigated to rather than fetched from some other kagane
// page: the challenge only runs on a top-level navigation, and once it clears
// The image URL is navigated to rather than fetched from another page of the
// Site: the challenge only runs on a top-level navigation, and once it clears
// the document *is* the image, so a same-origin fetch of location.href reads
// it straight back out of the cache.
// it straight back out of the cache. For comix the navigation is also the only
// route that works at all — its Series page sets
// `cross-origin-embedder-policy: require-corp`, which fails a page-context
// fetch of the cover host.
//
// The challenge is not solved by the first read: WaitReady("body") is satisfied
// by the interstitial too. run holds the tab open until the in-page fetch
// succeeds, which is what gives the challenge script the seconds it needs.
func (f *BrowserFetcher) Image(ctx context.Context, imageURL string) ([]byte, string, error) {
m := kaganeImageURLRe.FindStringSubmatch(imageURL)
if m == nil {
if !browserOnlyCoverURL(imageURL) {
return nil, "", fmt.Errorf("not a browser-fetchable cover url: %q", imageURL)
}
imageID := m[1]
var dataURL string
err := f.run(ctx, imageURL,
chromedp.Evaluate(`fetch(location.href).then(r => r.ok
@@ -175,16 +197,16 @@ func (f *BrowserFetcher) Image(ctx context.Context, imageURL string) ([]byte, st
: "")`, &dataURL, awaitPromise),
func() bool { return dataURL != "" })
if err != nil {
return nil, "", fmt.Errorf("browser image %s: %w", imageID, err)
return nil, "", fmt.Errorf("browser image %s: %w", imageURL, err)
}
// "data:image/webp;base64,<payload>".
head, payload, ok := strings.Cut(dataURL, ";base64,")
if !ok {
return nil, "", fmt.Errorf("browser image %s: not a data url", imageID)
return nil, "", fmt.Errorf("browser image %s: not a data url", imageURL)
}
raw, err := base64.StdEncoding.DecodeString(payload)
if err != nil {
return nil, "", fmt.Errorf("browser image %s: %w", imageID, err)
return nil, "", fmt.Errorf("browser image %s: %w", imageURL, err)
}
return raw, strings.TrimPrefix(head, "data:"), nil
}
@@ -323,6 +345,19 @@ func novelfullSeriesURL(seriesURL string) bool {
strings.HasSuffix(u.Path, ".html")
}
// comixSeriesPageURL returns the address comixRead fetches inside the tab: the
// Series page itself, rebuilt from the pinned host and path so nothing else
// travels. Host-pinned here for the same reason kagane's is — series_url is
// client-supplied and a headless browser is a strong SSRF primitive.
func comixSeriesPageURL(seriesURL string) (string, bool) {
u, err := url.Parse(seriesURL)
if err != nil || u.Scheme != "https" || u.Hostname() != "comix.to" ||
!comixSeriesPathRe.MatchString(u.Path) {
return "", false
}
return "https://comix.to" + u.Path, true
}
// awaitPromise makes Evaluate resolve the promise rather than returning a
// serialised Promise object.
func awaitPromise(p *runtime.EvaluateParams) *runtime.EvaluateParams {
+56
View File
@@ -61,6 +61,62 @@ func TestNovelfullSeriesURL(t *testing.T) {
})
}
}
func TestComixSeriesPageURL(t *testing.T) {
const series = "https://comix.to/title/n8we-dungeons-and-crayons"
cases := []struct {
name string
url string
want string
}{
{"series page", series, series},
{"trailing slash kept", series + "/", series + "/"},
// Query and fragment are dropped: only the pinned path travels.
{"query dropped", series + "?tab=chapters", series},
{"foreign host", "https://evil.example/title/x", ""},
{"lookalike host", "https://comix.to.evil.example/title/x", ""},
{"not https", "http://comix.to/title/x", ""},
{"not a series path", "https://comix.to/search", ""},
{"chapter page", series + "/11139891-chapter-80", ""},
{"garbage", "://nope", ""},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
got, ok := comixSeriesPageURL(tc.url)
if ok != (tc.want != "") || got != tc.want {
t.Fatalf("comixSeriesPageURL(%q) = %q, %v; want %q", tc.url, got, ok, tc.want)
}
})
}
}
// The browser is an SSRF primitive and a cover address can originate in a
// client-supplied PUT body, so this gate decides what it may navigate to.
func TestBrowserOnlyCoverURL(t *testing.T) {
cases := []struct {
url string
want bool
}{
{"https://static.comix.to/039d/i/1/34/6a6742bf15736@280.jpg", true},
{"https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed", true},
// Every other Site's CDN answers plain TLS.
{"https://gg.asuracomic.net/covers/x.webp", false},
{"http://static.comix.to/039d/x.jpg", false},
{"https://static.comix.to.evil.example/039d/x.jpg", false},
{"https://evil.example/static.comix.to/x.jpg", false},
{"https://static.comix.to/039d/x.jpg?next=http://169.254.169.254/", false},
{"https://static.comix.to/039d/x.svg", false},
{"https://static.comix.to/../etc/passwd.jpg", false},
{"https://static.comix.to/", false},
}
for _, tc := range cases {
t.Run(tc.url, func(t *testing.T) {
if got := browserOnlyCoverURL(tc.url); got != tc.want {
t.Fatalf("browserOnlyCoverURL(%q) = %v, want %v", tc.url, got, tc.want)
}
})
}
}
func TestClassifyBrowserInterruption(t *testing.T) {
if err := classifyBrowserError(context.Background(), true, context.Canceled); !errors.Is(err, errBrowserInterrupted) {
t.Fatalf("classifyBrowserError(context.Canceled) = %v, want browser interruption", err)
+2 -1
View File
@@ -25,7 +25,8 @@ type CoverBytesFetcher interface {
// fetchCoverBytes routes a cover's byte retrieval by URL shape, not by Site
// name: the browser fetcher's module claims the addresses only it can fetch
// (kagane's image route answers a plain fetch with a challenge and
// `cross-origin-resource-policy: same-origin`), and everything else goes over
// `cross-origin-resource-policy: same-origin`, static.comix.to answers one with
// the same challenge its pages serve), and everything else goes over
// plain TLS. Missing fetchers degrade to an error the caller logs, never a
// fallback onto a path that cannot succeed. One routing rule for the poll and
// the acquirer, so the two cannot drift apart.
+5 -5
View File
@@ -17,8 +17,8 @@ type Fetcher interface {
}
// BrowserCoverFetcher retrieves one cover's bytes through the browser-backed
// path — the only route that clears the challenge kagane's image URLs answer
// a plain fetch with. Satisfied by BrowserFetcher.
// path — the only route that clears the challenge kagane's and comix's image
// URLs answer a plain fetch with. Satisfied by BrowserFetcher.
type BrowserCoverFetcher interface {
Image(ctx context.Context, imageURL string) (body []byte, contentType string, err error)
}
@@ -119,9 +119,9 @@ func (p *Poller) storeCover(ctx context.Context, sr store.Series, sourceURL stri
// fetcherFor returns the fetcher a site's page needs, or nil when the site
// cannot be fetched at all right now. A Site whose registry entry carries a
// Browser read — kagane and novelfull, both behind a Cloudflare JavaScript
// challenge no TLS fingerprint clears — prefers the browser; when it is
// absent, the entry's Fallback decides whether plain TLS may take over. One
// Browser read — kagane, comix and novelfull, all behind a Cloudflare
// JavaScript challenge no TLS fingerprint clears — prefers the browser; when it
// is absent, the entry's Fallback decides whether plain TLS may take over. One
// routing rule for the poll and the acquirer, so the two cannot drift apart.
func fetcherFor(site string, browser, tls Fetcher) Fetcher {
s, known := sites[site]
+80
View File
@@ -762,6 +762,86 @@ func TestKaganeUsesBrowserFetcher(t *testing.T) {
}
}
// comix joined kagane behind the challenge on 2026-08-12 (#98): its page goes
// to the browser, its Cover bytes go through the browser's image route because
// static.comix.to is gated the same way, and the TLS fetcher is never asked
// for either.
func TestComixUsesBrowserFetcher(t *testing.T) {
s, dbURL := newTestStore(t)
const (
key = "comix:n8we-dungeons-and-crayons"
seriesID = "n8we-dungeons-and-crayons"
seriesURL = "https://comix.to/title/n8we-dungeons-and-crayons"
coverURL = "https://static.comix.to/039d/i/1/34/6a6742bf15736@280.jpg"
)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key, Site: "comix", SeriesID: seriesID, SeriesURL: seriesURL,
UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
seedCoverSource(t, dbURL, "comix", seriesID, coverURL)
tlsF := &fakeFetcher{body: "", status: 200}
browserF := &fakeFetcher{body: comixSeriesFixture, status: 200}
covers := &fakeCoverFetcher{body: []byte("cover-bytes"), contentType: "image/jpeg"}
tlsCovers := &fakeBytesCoverFetcher{body: []byte("tls-bytes"), contentType: "image/jpeg"}
p := &Poller{
Store: s, Fetch: tlsF, BrowserFetch: browserF,
CoverFetch: covers, CoverBytesFetch: tlsCovers,
Now: func() time.Time { return time.UnixMilli(5_000_000) },
Cooldown: time.Hour, BrowserCooldown: time.Hour,
Interval: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if len(tlsF.calls) != 0 {
t.Errorf("TLS fetcher was called for comix: %v", tlsF.calls)
}
if len(browserF.calls) != 1 {
t.Fatalf("browser fetcher calls = %v, want 1", browserF.calls)
}
if got := tlsCovers.callCount(); got != 0 {
t.Errorf("TLS cover fetches = %d, want 0: static.comix.to answers a challenge", got)
}
if got := covers.callCount(); got != 1 {
t.Fatalf("browser cover fetches = %d, want 1", got)
}
got, found, err := s.Get(s.OwnerID(), key)
if err != nil || !found {
t.Fatalf("Get: %v found=%v", err, found)
}
if got.LatestChapterNum == nil || *got.LatestChapterNum != 80 {
t.Errorf("LatestChapterNum = %v, want 80", got.LatestChapterNum)
}
}
// Without a browser, comix is skipped outright rather than handed to plain
// TLS: a plain fetch retrieves only a challenge page (measured 2026-08-12).
func TestComixSkippedWhenNoBrowserFetcher(t *testing.T) {
s, _ := newTestStore(t)
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: "comix:n8we-dungeons-and-crayons", Site: "comix",
SeriesID: "n8we-dungeons-and-crayons",
SeriesURL: "https://comix.to/title/n8we-dungeons-and-crayons",
UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
f := &fakeFetcher{body: comixSeriesFixture, status: 200}
p := &Poller{
Store: s, Fetch: f,
Now: func() time.Time { return time.UnixMilli(5_000_000) },
Cooldown: time.Hour, BrowserCooldown: time.Hour,
Interval: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if len(f.calls) != 0 {
t.Errorf("TLS fetcher was called for comix: %v", f.calls)
}
}
func TestRunOncePrefetchesKaganeCover(t *testing.T) {
s, dbURL := newTestStore(t)
const (
+28 -8
View File
@@ -47,9 +47,9 @@ type browserRead struct {
// Done reports whether the payload arrived.
Done func(body string) bool
// Fallback allows the plain-TLS fetcher when no browser is configured.
// False skips the Site instead. kagane is false — a plain fetch would
// only ever retrieve a challenge page — and novelfull is true, because
// its challenge is a live time-varying fact (AGENTS.md).
// False skips the Site instead. kagane and comix are false — a plain fetch
// would only ever retrieve a challenge page — and novelfull is true,
// because its challenge is a live time-varying fact (AGENTS.md).
Fallback bool
}
@@ -248,14 +248,25 @@ var comixInitialDataRe = regexp.MustCompile(`(?is)<script\b[^>]*\bid\s*=\s*["']i
// supplied, and a headless browser is a strong SSRF primitive.
var kaganeImageURLRe = regexp.MustCompile(`^https://kagane\.to/api/v2/image/([0-9a-f-]{36})/compressed$`)
// comixImageURLRe matches comix's cover host and path shape. Pinned in full
// (scheme, host, path characters, image extension) for the same reason
// kaganeImageURLRe is: the address reaches a headless browser, and it can
// originate in a client-supplied PUT body. No dot is allowed inside the path,
// so no traversal or second extension can hide in it. Shape from a live page,
// 2026-08-10: /039d/i/1/34/6a6742bf15736@280.jpg.
var comixImageURLRe = regexp.MustCompile(`^https://static\.comix\.to/[A-Za-z0-9@/_-]+\.(?:jpg|jpeg|png|webp)$`)
// browserOnlyCoverURL reports whether the browser sidecar is the only fetcher
// for cover bytes at imageURL. kagane's image route answers a plain fetch with
// a challenge and `cross-origin-resource-policy: same-origin`, so a TLS fetch
// would only ever retrieve a challenge page and must not be attempted
// (ADR-0007). This is the byte-fetch router's per-Site knowledge; it lives in
// the extraction module, which owns kagane's URL shapes.
// a challenge and `cross-origin-resource-policy: same-origin`, and
// static.comix.to answers one with the same Cloudflare challenge its pages
// serve (measured 2026-08-12, issue #98), so a TLS fetch would only ever
// retrieve a challenge page and must not be attempted (ADR-0007). This is the
// byte-fetch router's per-Site knowledge; it lives in the extraction module,
// which owns those URL shapes.
func browserOnlyCoverURL(imageURL string) bool {
return kaganeImageURLRe.MatchString(imageURL)
return kaganeImageURLRe.MatchString(imageURL) ||
comixImageURLRe.MatchString(imageURL)
}
// kagane's browser-fetched series response publishes cover image IDs under
@@ -389,6 +400,15 @@ var sites = map[string]site{
Host: "comix.to",
LatestChapter: comixLatestChapter,
Cover: comixCoverEntry,
Browser: &browserRead{
Read: comixRead,
// The interstitial is served in place of the page, so "arrived"
// has to exclude it explicitly, as novelfull's does.
Done: func(body string) bool { return body != "" && !isInterstitial(body) },
// Never falls back: a plain fetch of a comix page or cover
// retrieves only a challenge page (measured 2026-08-12).
Fallback: false,
},
},
"kagane": {
Host: "kagane.to",
+7
View File
@@ -41,6 +41,13 @@ const challengeFixture = `<!DOCTYPE html><html><head><title>Just a moment...</ti
// https://comix.to/title/n8we-dungeons-and-crayons fetched 2026-08-03. comix is
// an SPA: the page ships a JSON state blob rather than a list of chapter
// anchors, and latestChapterUrl is where the newest chapter actually lives.
//
// Still the right fixture after comix moved behind the challenge (#98): the
// browser read is an in-tab fetch of the Series URL, so the body a poll parses
// is this same server-rendered HTML, not a rendered DOM. Confirmed against a
// live cleared tab 2026-08-16 (TestSmokeComix): the in-tab fetch returned
// 24793 bytes of server-rendered HTML that these same parses read a chapter
// and a cover out of.
const comixSeriesFixture = `
{"firstChapterUrl":"/title/n8we-dungeons-and-crayons/5038739-chapter-1","latestChapterUrl":"/title/n8we-dungeons-and-crayons/11139891-chapter-80"},
{""manga","recommended","n8we",1]":{"items":[{"latestChapterUrl":"/title/qqwrm-full-time-awakening/99999999-chapter-999"}]}
@@ -0,0 +1,72 @@
package latest
import (
"context"
"os"
"testing"
"time"
"bookmarkmanager/backend/internal/store"
)
// TestSmokeComix answers "is comix's challenge clearing from this browser right
// now" — a live, time-varying fact, so a red run is something to re-check
// before it is a defect. Needs the real browser unit with outbound network:
//
// cd chrome && BROWSER_BIND_ADDR=127.0.0.1 docker compose up -d --build
// SMOKE_BROWSER_WS_URL=ws://127.0.0.1:9222 go test -run TestSmokeComix ./internal/latest
//
// It walks the whole read: the in-tab page fetch, both parses, and the Cover
// bytes by direct navigation to static.comix.to. The Cover address comes out of
// the page rather than being pinned in the test, because a stored one rots.
func TestSmokeComix(t *testing.T) {
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
if ws == "" {
t.Skip("SMOKE_BROWSER_WS_URL unset")
}
const seriesURL = "https://comix.to/title/m12d-classmate"
f, err := NewBrowserFetcher(ws)
if err != nil {
t.Fatalf("NewBrowserFetcher: %v", err)
}
defer f.Close()
ctx, cancel := context.WithTimeout(context.Background(), 120*time.Second)
defer cancel()
body, status, err := f.Get(ctx, seriesURL)
if err != nil {
t.Fatalf("Get: %v", err)
}
t.Logf("status=%d bytes=%d", status, len(body))
if status != 200 {
t.Fatalf("status = %d, want 200 — the sidecar is not clearing the challenge", status)
}
chapter, ok := latestChapterFrom("comix", seriesURL, body)
if !ok {
t.Fatalf("no latest chapter in %d bytes — page shape changed", len(body))
}
t.Logf("latest chapter: %v %q", chapter.Num, chapter.Label)
cover, ok := coverFrom("comix", seriesURL, body)
if !ok {
t.Fatalf("no cover address in %d bytes — page shape changed", len(body))
}
t.Logf("cover: %s", cover)
if !browserOnlyCoverURL(cover) {
t.Fatalf("cover %q is not claimed by the browser gate: the pin and the live URL shape disagree", cover)
}
bytes, contentType, err := f.Image(ctx, cover)
if err != nil {
t.Fatalf("Image: %v", err)
}
if len(bytes) < 1000 {
t.Fatalf("cover is %d bytes, want a real image", len(bytes))
}
t.Logf("fetched %d bytes of %s", len(bytes), contentType)
if _, ok := store.CoverContentType(contentType); !ok {
t.Fatalf("content type %q is not storable", contentType)
}
}