Fix comix titles and covers, kagane volume chapters, and kagane cover rendering (#37)

Fixes five reported symptoms across comix.to and kagane.to. Diagnosing them turned up two latent bugs underneath, both of which had to be fixed for the kagane cover work to function at all.

## Reported symptoms and their causes

| # | Symptom | Cause |
|---|---------|-------|
| 1 | comix bookmark titled `Comix - Read Comics online for free` | comix is an SPA that rewrites `document.title` on client routing but never touches the server-rendered `og:title`. The adapter read `og:title`, so a cold load stored the homepage's title. |
| 2 | next comix bookmark gets the *previous* series' title | Same cause. After an in-page hop, `og:title` still holds whatever page loaded first. |
| 3 | comix cover shows the placeholder | comix serves no `og:image` at all, so `coverFromPage()` had nothing to read. |
| 4 | kagane chapter never appears in the bookmark list | Reader URLs carry no chapter number, so it is parsed out of `og:title`. Volume-numbered series render `"<Series> - Volume <v> Chapter <n>"`, which the suffix regex did not match, so `chapterNum` came back null and nothing was recorded. |
| 5 | kagane title includes the chapter, e.g. `SP Baby - Volume 1 Chapter 1` | Same unmatched regex — the tail was never stripped. One fix covers 4 and 5. |
| 6 | kagane cover blocked in the web UI | kagane serves covers behind its Cloudflare challenge **and** with `cross-origin-resource-policy: same-origin`. No `<img>` on the UI's origin can load one even from a browser holding the clearance cookie. Hot-linking cannot be made to work. |

## What changed

**Userscript.** comix titles now come from `document.title` with the chapter page's `" - Ch.<n>"` tail stripped, and the cover is the `img` whose `alt` matches the cleaned title. comix fills `document.title` a beat *after* the URL changes — later than the nav watcher's 300 ms snapshot — so the watcher also re-detects when the `detect()` signature changes, not only when the URL does. The kagane suffix regex takes an optional `Volume <v> ` segment. All three page shapes were captured live on 2026-08-08 and pinned as regression tests.

**Cover proxy.** `Bookmark.CoverURL()` rewrites a stored kagane `og:image` to `/img/kagane/{id}`; templates render `.CoverURL` instead of `.Cover`. The endpoint is session-gated like every other UI route and fetches through the shared headless browser, which is same-origin with kagane and so satisfies both the challenge and the CORP header. Results are memoised in-process, so a cover costs one navigation per deployment lifetime. With `BROWSER_WS_URL` unset the endpoint answers 404 rather than reaching for a nil fetcher — the same degrade-to-userscript behaviour the poller already has.

The image id is matched against a UUID regex before it reaches the browser. That gate is load-bearing rather than tidiness: the cover is a stored client-supplied string, so an unvalidated one turns this endpoint into an SSRF primitive aimed at the deployment's own network. `ServeMux` path-cleans a traversal into a redirect before the handler runs, but the handler does not depend on that, and a test pins it.

## Two latent bugs found underneath

**`BrowserFetcher.run` never let a challenge solve.** It navigated, waited for `body`, read once, and closed the tab — roughly half a second end to end. The Cloudflare interstitial has a `body` too, so `WaitReady` was satisfied by the challenge page itself. This made the challenge *unclearable* rather than merely slow: an interstitial needs several seconds of a live page to solve itself and write clearance into the browser's shared cookie jar, so tearing the tab down first means every subsequent call is challenged exactly like the one before it. `run` now holds one tab and re-reads until the caller's predicate reports an answer, bounded by `challengeTimeout` and the caller's own deadline. Exhausting the budget maps back to the 403 the poller already expects, keeping a challenged site distinct from a broken transport.

**`chromedp/headless-shell` cannot clear kagane's challenge at all.** It is a stripped Chrome build and the tells are structural rather than a header: `navigator.webdriver` is true, the plugin list is empty, and the client hints are Chromium- rather than Chrome-branded. Overriding `webdriver` through CDP was tried on its own and changed nothing.

All measured 2026-08-08 from one IP against the same cover, so the comparisons are like for like:

| Browser | Result |
|---------|--------|
| `chromedp/headless-shell:stable` | never cleared (90 s) |
| `zenika/alpine-chrome` | never cleared — ships Chrome 124, old enough that Cloudflare refuses it and old enough to break chromedp's CDP structs |
| `google-chrome`, default UA | never cleared (60 s) — `--headless=new` advertises `HeadlessChrome` |
| `google-chrome`, stock UA, `TZ=UTC` | never cleared (90 s) |
| `google-chrome`, stock UA, any non-UTC `TZ` | **cleared in ~4 s** |

Both remaining tells are load-bearing, and each was tested in isolation. `chrome/` is a Debian image with `google-chrome-stable`, a UA whose version is read back out of the binary at startup (a hardcoded one would drift out of step with the `Sec-CH-UA` hints on the next Chrome update and become a fresh tell), and no `--enable-automation`.

### The timezone tell: UTC, not a country mismatch

The first pass concluded the zone had to match the egress IP's country. Re-measuring against the actual deployment case shows that was wrong, and the correction is in `1552dd1`.

The original inference read the host's `/etc/timezone` (`Asia/Bangkok`) and assumed a Thai egress. It isn't — this host egresses from an Indonesian IP. `Asia/Bangkok` cleared not because it matched a country but because it simply isn't UTC, and the two share +07, which hid the distinction. Same container, same Indonesian IP:

| `TZ` | Result |
|------|--------|
| `UTC` | never cleared (60 s, **twice**) |
| `Asia/Jakarta` | cleared in 4 s |
| `America/New_York` | cleared in 4 s |

`America/New_York` matches neither the country nor the offset nor the hemisphere and clears just as fast. A UTC clock is itself the bot signal — Cloudflare scores it as the datacenter default — and any real zone satisfies the check. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one, and a deployment that changes region need not keep it in sync.

One sharp edge remains: the usual `-v /etc/localtime:/etc/localtime:ro` does **not** work. Chrome resolves the zone through ICU, which takes the name from that path's symlink target and ignores the file's contents, so glibc reports the host zone while Chrome still reports UTC. `/etc/timezone` carries the name and is mounted instead.

Chrome also binds its DevTools port to loopback and silently ignores `--remote-debugging-address`, which is why headless-shell fronted it with socat. This image does the same, so it stays a drop-in: the compose service keeps the `headless-shell` name and its pinned address, and `BROWSER_WS_URL` is unchanged.

## Verification

```
go test ./...        all packages ok
node --test          37 + 12 pass, 0 fail

SMOKE_BROWSER_WS_URL=... go test -run TestSmokeKagane ./internal/latest
  TestSmokeKaganeImage  PASS (5.29s)  fetched 56710 bytes of image/webp
  TestSmokeKaganeGet    PASS (1.17s)  status=200, real chapter-list JSON
```

The smoke test ran against the exact compose configuration — built image, empty `BROWSER_TZ`, `/etc/timezone` mounted, cold profile — hitting real kagane.to. It skips unless `SMOKE_BROWSER_WS_URL` names a sidecar, so `go test ./...` stays hermetic and Docker-only.

A red smoke run means the challenge is not clearing from that IP, which is a live, time-varying fact to re-check rather than necessarily a defect.

## Security invariants

- Auth unchanged. `/img/kagane/{id}` is session-gated by `requireSession`, the same guard as every other UI route.
- Outbound fetch gated: the id is UUID-validated before it reaches the browser, keeping the existing rule that a client-supplied string never selects a fetch target unchecked.
- No new secrets, no new logging of credentials, no change to CORS, sessions, or crypto.
- Templates still escape everything; `.CoverURL` returns a plain string and is not wrapped in `template.HTML`/`URL`.
- One new dependency-free image (`chrome/`) built from Debian plus Google's own apt repo; no new Go modules.

## Deploying

Needs `docker compose build headless-shell`.

**A UTC host must set `BROWSER_TZ`, or kagane silently stops working.** With it unset the sidecar falls back to the host's `/etc/timezone`; on a UTC server that yields UTC, which is the one value that never clears. Any real zone works — `BROWSER_TZ=Asia/Jakarta` for the current deployment. `.env.example` now documents this; it previously did not mention the knob at all.

Only the browser sidecar reads `BROWSER_TZ`. The backend keeps its UTC clock, and stored timestamps are unix ms, so nothing else shifts.

## Deliberately not done

Retry/backoff around the cover proxy, and a panel-side cover fix. The panel renders no covers, and covers cache in-process after the first fetch. Worth adding if kagane starts rate-limiting.

## Correction after review of the deployment case

`1552dd1` was added after the branch was first pushed: the deployment host runs UTC with an Indonesian egress IP, which prompted re-measuring the timezone claim and falsifying it. The earlier commits' reasoning is left intact rather than rebased away, so the diagnostic trail — including the wrong turn and what disproved it — stays readable.

Reviewed-on: #37
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
This commit was merged in pull request #37.
This commit is contained in:
2026-08-08 23:27:32 +07:00
committed by sulthan
parent 2ef769d421
commit 741b23322b
19 changed files with 886 additions and 94 deletions
+131 -30
View File
@@ -2,7 +2,9 @@ package latest
import (
"context"
"encoding/base64"
"encoding/json"
"errors"
"fmt"
"net/url"
"regexp"
@@ -22,6 +24,12 @@ const challengeTimeout = 45 * time.Second
var kaganeSeriesRe = regexp.MustCompile(`^/series/([0-9a-f-]{36})/?$`)
// kaganeImageIDRe pins the only path segment Image interpolates into an
// outbound URL. The id arrives from a stored cover URL, which a client
// supplied, so it is matched rather than trusted: a headless browser is a
// strong SSRF primitive.
var kaganeImageIDRe = regexp.MustCompile(`^[0-9a-f-]{36}$`)
// BrowserFetcher retrieves pages through a remote headless Chrome over the
// DevTools Protocol.
//
@@ -91,23 +99,6 @@ func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int
return "", 0, fmt.Errorf("not a fetchable browser series url: %q", seriesURL)
}
f.mu.Lock()
defer f.mu.Unlock()
ctx, cancel := context.WithTimeout(ctx, challengeTimeout)
defer cancel()
// A fresh tab per fetch, closed on return, so one wedged page cannot
// poison later polls.
tabCtx, cancelTab := chromedp.NewContext(f.allocCtx)
defer cancelTab()
// Bind the caller's deadline to the tab.
tabCtx, cancelDeadline := context.WithCancel(tabCtx)
defer cancelDeadline()
go func() {
<-ctx.Done()
cancelDeadline()
}()
var body string
// kagane's chapter list is only in its JSON API, which must be called from
// inside the page so the request carries the clearance cookie. novelfull
@@ -123,24 +114,134 @@ func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int
)
}
err := chromedp.Run(tabCtx,
chromedp.Navigate(seriesURL),
// The challenge reloads the page itself when it passes; waiting for the
// site's own root element is what tells us we are through it.
chromedp.WaitReady("body", chromedp.ByQuery),
read,
)
if err != nil {
// novelfull's payload is the DOM itself, and the interstitial has a DOM
// too, so "we have an answer" has to exclude it explicitly. kagane's
// in-page fetch just fails while challenged, which is already the signal.
done := func() bool { return body != "" && (isKagane || !isInterstitial(body)) }
if err := f.run(ctx, seriesURL, read, done); err != nil {
// Challenge never cleared, or the API refused. Indistinguishable from
// here and handled identically by the caller.
if errors.Is(err, errChallengeHeld) {
return "", 403, nil
}
return "", 0, fmt.Errorf("browser fetch %q: %w", seriesURL, err)
}
if body == "" {
// Challenge still up, or the API refused. Indistinguishable from here
// and handled identically by the caller.
return "", 403, nil
}
return body, 200, nil
}
// Image retrieves one kagane cover as raw bytes and its content type.
//
// It exists because kagane serves covers behind the same challenge as its
// pages *and* with `cross-origin-resource-policy: same-origin`, so an <img> on
// the web UI's origin cannot load one even from a browser that already holds
// the clearance cookie (verified 2026-08-08). Proxying is the only route.
//
// The image URL is navigated to rather than fetched from some other kagane
// page: the challenge only runs on a top-level navigation, and once it clears
// the document *is* the image, so a same-origin fetch of location.href reads
// it straight back out of the cache.
//
// The challenge is not solved by the first read: WaitReady("body") is satisfied
// by the interstitial too. run holds the tab open until the in-page fetch
// succeeds, which is what gives the challenge script the seconds it needs.
func (f *BrowserFetcher) Image(ctx context.Context, imageID string) ([]byte, string, error) {
if !kaganeImageIDRe.MatchString(imageID) {
return nil, "", fmt.Errorf("not a kagane image id: %q", imageID)
}
var dataURL string
err := f.run(ctx, "https://kagane.to/api/v2/image/"+imageID+"/compressed",
chromedp.Evaluate(`fetch(location.href).then(r => r.ok
? r.blob().then(b => new Promise(res => {
const fr = new FileReader();
fr.onload = () => res(fr.result);
fr.readAsDataURL(b);
}))
: "")`, &dataURL, awaitPromise),
func() bool { return dataURL != "" })
if err != nil {
return nil, "", fmt.Errorf("browser image %s: %w", imageID, err)
}
// "data:image/webp;base64,<payload>".
head, payload, ok := strings.Cut(dataURL, ";base64,")
if !ok {
return nil, "", fmt.Errorf("browser image %s: not a data url", imageID)
}
raw, err := base64.StdEncoding.DecodeString(payload)
if err != nil {
return nil, "", fmt.Errorf("browser image %s: %w", imageID, err)
}
return raw, strings.TrimPrefix(head, "data:"), nil
}
// errChallengeHeld reports that the budget ran out with the interstitial still
// up. Distinct from a transport failure: it means "this site said no", which
// the poller answers with a 403 and its ordinary cooldown.
var errChallengeHeld = errors.New("challenge held")
// challengePollInterval paces re-reads while a challenge solves itself.
const challengePollInterval = 2 * time.Second
// isInterstitial reports whether html is Cloudflare's challenge page rather
// than the site's own. Matched on the challenge runtime's script path, which is
// stable across the interstitial's wording and locale — the visible "Just a
// moment..." title is neither.
func isInterstitial(html string) bool {
return strings.Contains(html, "/cdn-cgi/challenge-platform/")
}
// run navigates to target and re-reads until done reports an answer, bounded by
// challengeTimeout and by the caller's own deadline, in a tab that is closed on
// return so one wedged page cannot poison later calls.
//
// Holding the tab open across re-reads is the whole point. A Cloudflare
// interstitial needs several seconds of a live page to solve itself and write
// clearance into the browser's shared cookie jar; reading once and closing the
// tab — which is what this did before 2026-08-08 — never gives it that window,
// so every fetch lands on the interstitial and the clearance that would have
// unblocked all the later ones is never obtained.
func (f *BrowserFetcher) run(ctx context.Context, target string, read chromedp.Action, done func() bool) error {
f.mu.Lock()
defer f.mu.Unlock()
ctx, cancel := context.WithTimeout(ctx, challengeTimeout)
defer cancel()
tabCtx, cancelTab := chromedp.NewContext(f.allocCtx)
defer cancelTab()
// Bind the caller's deadline to the tab.
tabCtx, cancelDeadline := context.WithCancel(tabCtx)
defer cancelDeadline()
go func() {
<-ctx.Done()
cancelDeadline()
}()
if err := chromedp.Run(tabCtx,
chromedp.Navigate(target),
chromedp.WaitReady("body", chromedp.ByQuery),
); err != nil {
return err
}
var lastErr error
for {
// The challenge reloads the page when it passes, which tears down the
// execution context mid-read. That is a retry, not a failure.
if err := chromedp.Run(tabCtx, read); err != nil {
lastErr = err
} else if done() {
return nil
}
select {
case <-ctx.Done():
if lastErr != nil {
return fmt.Errorf("%w (last read: %v)", errChallengeHeld, lastErr)
}
return errChallengeHeld
case <-time.After(challengePollInterval):
}
}
}
// kaganeAPIURL maps a stored series_url to the JSON endpoint carrying its
// chapter list. Returning false for anything else is a second line of defence
// behind fetchableSeriesURL: a headless browser is a strong SSRF primitive and
@@ -0,0 +1,91 @@
package latest
import (
"context"
"net/http"
"os"
"testing"
"time"
)
// TestSmokeKaganeImage is the live proof that the cover proxy's fetch actually
// clears Cloudflare and returns image bytes. It needs a real headless Chrome
// with outbound network, so it runs only when SMOKE_BROWSER_WS_URL is set:
//
// docker run --rm --shm-size=1gb -p 19222:9222 chromedp/headless-shell:stable
// SMOKE_BROWSER_WS_URL=ws://127.0.0.1:19222 go test -run TestSmokeKaganeImage ./internal/latest
func TestSmokeKaganeImage(t *testing.T) {
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
if ws == "" {
t.Skip("SMOKE_BROWSER_WS_URL unset")
}
const imageID = "019fe11a-84c3-7fc3-a84b-88787374b617" // SP Baby's cover
// The same URL through a plain client is what the web UI's <img> gets.
// Asserting on it keeps the test honest about why the browser is needed.
req, err := http.NewRequest(http.MethodGet,
"https://kagane.to/api/v2/image/"+imageID+"/compressed", nil)
if err != nil {
t.Fatal(err)
}
if res, err := (&http.Client{Timeout: 15 * time.Second}).Do(req); err == nil {
res.Body.Close()
if res.StatusCode == http.StatusOK {
t.Log("note: kagane answered a plain request 200 — the challenge is not up right now")
}
}
f, err := NewBrowserFetcher(ws)
if err != nil {
t.Fatalf("NewBrowserFetcher: %v", err)
}
defer f.Close()
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
defer cancel()
body, contentType, err := f.Image(ctx, imageID)
if err != nil {
t.Fatalf("Image: %v", err)
}
if len(body) < 1000 {
t.Fatalf("body is %d bytes, want a real image", len(body))
}
if contentType != "image/webp" {
t.Fatalf("content type = %q, want image/webp", contentType)
}
// WebP files start with "RIFF....WEBP".
if string(body[:4]) != "RIFF" || string(body[8:12]) != "WEBP" {
t.Fatalf("body is not a WebP: % x", body[:12])
}
t.Logf("fetched %d bytes of %s", len(body), contentType)
if _, _, err := f.Image(ctx, "not-a-uuid"); err == nil {
t.Fatal("Image accepted a non-uuid id")
}
}
// Control for the test above: the poller's own kagane path, same sidecar. If
// this fails too, the sidecar is not clearing the challenge at all and the
// image result says nothing about Image itself.
func TestSmokeKaganeGet(t *testing.T) {
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
if ws == "" {
t.Skip("SMOKE_BROWSER_WS_URL unset")
}
f, err := NewBrowserFetcher(ws)
if err != nil {
t.Fatalf("NewBrowserFetcher: %v", err)
}
defer f.Close()
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
defer cancel()
body, status, err := f.Get(ctx, "https://kagane.to/series/019fe11a-8670-7cf3-8343-0b02057d3787")
if err != nil {
t.Fatalf("Get: %v", err)
}
t.Logf("status=%d bytes=%d head=%.80q", status, len(body), body)
if status != 200 {
t.Fatalf("status = %d, want 200 — the sidecar is not clearing the challenge", status)
}
}