Fix comix titles and covers, kagane volume chapters, and kagane cover rendering (#37)

Fixes five reported symptoms across comix.to and kagane.to. Diagnosing them turned up two latent bugs underneath, both of which had to be fixed for the kagane cover work to function at all.

## Reported symptoms and their causes

| # | Symptom | Cause |
|---|---------|-------|
| 1 | comix bookmark titled `Comix - Read Comics online for free` | comix is an SPA that rewrites `document.title` on client routing but never touches the server-rendered `og:title`. The adapter read `og:title`, so a cold load stored the homepage's title. |
| 2 | next comix bookmark gets the *previous* series' title | Same cause. After an in-page hop, `og:title` still holds whatever page loaded first. |
| 3 | comix cover shows the placeholder | comix serves no `og:image` at all, so `coverFromPage()` had nothing to read. |
| 4 | kagane chapter never appears in the bookmark list | Reader URLs carry no chapter number, so it is parsed out of `og:title`. Volume-numbered series render `"<Series> - Volume <v> Chapter <n>"`, which the suffix regex did not match, so `chapterNum` came back null and nothing was recorded. |
| 5 | kagane title includes the chapter, e.g. `SP Baby - Volume 1 Chapter 1` | Same unmatched regex — the tail was never stripped. One fix covers 4 and 5. |
| 6 | kagane cover blocked in the web UI | kagane serves covers behind its Cloudflare challenge **and** with `cross-origin-resource-policy: same-origin`. No `<img>` on the UI's origin can load one even from a browser holding the clearance cookie. Hot-linking cannot be made to work. |

## What changed

**Userscript.** comix titles now come from `document.title` with the chapter page's `" - Ch.<n>"` tail stripped, and the cover is the `img` whose `alt` matches the cleaned title. comix fills `document.title` a beat *after* the URL changes — later than the nav watcher's 300 ms snapshot — so the watcher also re-detects when the `detect()` signature changes, not only when the URL does. The kagane suffix regex takes an optional `Volume <v> ` segment. All three page shapes were captured live on 2026-08-08 and pinned as regression tests.

**Cover proxy.** `Bookmark.CoverURL()` rewrites a stored kagane `og:image` to `/img/kagane/{id}`; templates render `.CoverURL` instead of `.Cover`. The endpoint is session-gated like every other UI route and fetches through the shared headless browser, which is same-origin with kagane and so satisfies both the challenge and the CORP header. Results are memoised in-process, so a cover costs one navigation per deployment lifetime. With `BROWSER_WS_URL` unset the endpoint answers 404 rather than reaching for a nil fetcher — the same degrade-to-userscript behaviour the poller already has.

The image id is matched against a UUID regex before it reaches the browser. That gate is load-bearing rather than tidiness: the cover is a stored client-supplied string, so an unvalidated one turns this endpoint into an SSRF primitive aimed at the deployment's own network. `ServeMux` path-cleans a traversal into a redirect before the handler runs, but the handler does not depend on that, and a test pins it.

## Two latent bugs found underneath

**`BrowserFetcher.run` never let a challenge solve.** It navigated, waited for `body`, read once, and closed the tab — roughly half a second end to end. The Cloudflare interstitial has a `body` too, so `WaitReady` was satisfied by the challenge page itself. This made the challenge *unclearable* rather than merely slow: an interstitial needs several seconds of a live page to solve itself and write clearance into the browser's shared cookie jar, so tearing the tab down first means every subsequent call is challenged exactly like the one before it. `run` now holds one tab and re-reads until the caller's predicate reports an answer, bounded by `challengeTimeout` and the caller's own deadline. Exhausting the budget maps back to the 403 the poller already expects, keeping a challenged site distinct from a broken transport.

**`chromedp/headless-shell` cannot clear kagane's challenge at all.** It is a stripped Chrome build and the tells are structural rather than a header: `navigator.webdriver` is true, the plugin list is empty, and the client hints are Chromium- rather than Chrome-branded. Overriding `webdriver` through CDP was tried on its own and changed nothing.

All measured 2026-08-08 from one IP against the same cover, so the comparisons are like for like:

| Browser | Result |
|---------|--------|
| `chromedp/headless-shell:stable` | never cleared (90 s) |
| `zenika/alpine-chrome` | never cleared — ships Chrome 124, old enough that Cloudflare refuses it and old enough to break chromedp's CDP structs |
| `google-chrome`, default UA | never cleared (60 s) — `--headless=new` advertises `HeadlessChrome` |
| `google-chrome`, stock UA, `TZ=UTC` | never cleared (90 s) |
| `google-chrome`, stock UA, any non-UTC `TZ` | **cleared in ~4 s** |

Both remaining tells are load-bearing, and each was tested in isolation. `chrome/` is a Debian image with `google-chrome-stable`, a UA whose version is read back out of the binary at startup (a hardcoded one would drift out of step with the `Sec-CH-UA` hints on the next Chrome update and become a fresh tell), and no `--enable-automation`.

### The timezone tell: UTC, not a country mismatch

The first pass concluded the zone had to match the egress IP's country. Re-measuring against the actual deployment case shows that was wrong, and the correction is in `1552dd1`.

The original inference read the host's `/etc/timezone` (`Asia/Bangkok`) and assumed a Thai egress. It isn't — this host egresses from an Indonesian IP. `Asia/Bangkok` cleared not because it matched a country but because it simply isn't UTC, and the two share +07, which hid the distinction. Same container, same Indonesian IP:

| `TZ` | Result |
|------|--------|
| `UTC` | never cleared (60 s, **twice**) |
| `Asia/Jakarta` | cleared in 4 s |
| `America/New_York` | cleared in 4 s |

`America/New_York` matches neither the country nor the offset nor the hemisphere and clears just as fast. A UTC clock is itself the bot signal — Cloudflare scores it as the datacenter default — and any real zone satisfies the check. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one, and a deployment that changes region need not keep it in sync.

One sharp edge remains: the usual `-v /etc/localtime:/etc/localtime:ro` does **not** work. Chrome resolves the zone through ICU, which takes the name from that path's symlink target and ignores the file's contents, so glibc reports the host zone while Chrome still reports UTC. `/etc/timezone` carries the name and is mounted instead.

Chrome also binds its DevTools port to loopback and silently ignores `--remote-debugging-address`, which is why headless-shell fronted it with socat. This image does the same, so it stays a drop-in: the compose service keeps the `headless-shell` name and its pinned address, and `BROWSER_WS_URL` is unchanged.

## Verification

```
go test ./...        all packages ok
node --test          37 + 12 pass, 0 fail

SMOKE_BROWSER_WS_URL=... go test -run TestSmokeKagane ./internal/latest
  TestSmokeKaganeImage  PASS (5.29s)  fetched 56710 bytes of image/webp
  TestSmokeKaganeGet    PASS (1.17s)  status=200, real chapter-list JSON
```

The smoke test ran against the exact compose configuration — built image, empty `BROWSER_TZ`, `/etc/timezone` mounted, cold profile — hitting real kagane.to. It skips unless `SMOKE_BROWSER_WS_URL` names a sidecar, so `go test ./...` stays hermetic and Docker-only.

A red smoke run means the challenge is not clearing from that IP, which is a live, time-varying fact to re-check rather than necessarily a defect.

## Security invariants

- Auth unchanged. `/img/kagane/{id}` is session-gated by `requireSession`, the same guard as every other UI route.
- Outbound fetch gated: the id is UUID-validated before it reaches the browser, keeping the existing rule that a client-supplied string never selects a fetch target unchecked.
- No new secrets, no new logging of credentials, no change to CORS, sessions, or crypto.
- Templates still escape everything; `.CoverURL` returns a plain string and is not wrapped in `template.HTML`/`URL`.
- One new dependency-free image (`chrome/`) built from Debian plus Google's own apt repo; no new Go modules.

## Deploying

Needs `docker compose build headless-shell`.

**A UTC host must set `BROWSER_TZ`, or kagane silently stops working.** With it unset the sidecar falls back to the host's `/etc/timezone`; on a UTC server that yields UTC, which is the one value that never clears. Any real zone works — `BROWSER_TZ=Asia/Jakarta` for the current deployment. `.env.example` now documents this; it previously did not mention the knob at all.

Only the browser sidecar reads `BROWSER_TZ`. The backend keeps its UTC clock, and stored timestamps are unix ms, so nothing else shifts.

## Deliberately not done

Retry/backoff around the cover proxy, and a panel-side cover fix. The panel renders no covers, and covers cache in-process after the first fetch. Worth adding if kagane starts rate-limiting.

## Correction after review of the deployment case

`1552dd1` was added after the branch was first pushed: the deployment host runs UTC with an Indonesian egress IP, which prompted re-measuring the timezone claim and falsifying it. The earlier commits' reasoning is left intact rather than rebased away, so the diagnostic trail — including the wrong turn and what disproved it — stays readable.

Reviewed-on: #37
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
This commit was merged in pull request #37.
This commit is contained in:
2026-08-08 23:27:32 +07:00
committed by sulthan
parent 2ef769d421
commit 741b23322b
19 changed files with 886 additions and 94 deletions
+20
View File
@@ -90,3 +90,23 @@ DISCORD_REDIRECT_URI=
# HTTP handler 500s any /json/version request whose Host header isn't an IP or
# "localhost", which silently breaks every kagane poll.
# BROWSER_WS_URL=ws://172.28.0.10:9222
# Clock zone the headless browser reports. A UTC clock is itself the bot
# signal — Cloudflare treats it as the datacenter default — and kagane's
# challenge then never clears. Measured 2026-08-08, identical container, one
# Indonesian egress IP: UTC never cleared in 60s (twice); Asia/Jakarta and
# America/New_York both cleared in 4s. So any real zone works; it does not
# have to match the IP's country, it just must not be UTC.
#
# Unset falls back to the host's /etc/timezone, which is a real zone whenever
# the host clock is set to local time. Set this when the host runs UTC — a UTC
# server is exactly the case that fails. Only the browser sidecar reads it —
# the backend's own zone is API_TZ below, and is cosmetic.
# BROWSER_TZ=Asia/Jakarta
# Zone the backend stamps its log lines in. Cosmetic only — it exists so the
# API's logs read on the same clock as the browser sidecar's. Nothing else in
# the service has a zone: bookmark timestamps are unix ms, and the two real
# time columns are timestamptz. Defaults to Asia/Jakarta; set to UTC for the
# conventional server default.
# API_TZ=Asia/Jakarta
+9 -1
View File
@@ -20,6 +20,9 @@ Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free
- Userscript run in **isolated world**, so embedded API token safe from site's JS.
- Cloudflare's block on manga sites **IP-reputation-based, not universal — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — Cloudflare's bot scoring can flip previously-clean IP without notice. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test.
- **kagane.to and novelfull.com are the exception to the above** — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`) and skips them entirely when that's unset. The four other sites poll fine over plain TLS.
- **The CDP sidecar must look like a real browser, and stock headless images don't.** Measured 2026-08-08 against kagane.to, all from the same IP: `chromedp/headless-shell:stable` never cleared the challenge in 90s (`navigator.webdriver` true, empty plugin list, Chromium-branded client hints — suppressing `webdriver` alone changed nothing); `zenika/alpine-chrome` ships Chrome 124, refused outright; real Chrome with the default `--headless=new` UA never cleared, because the UA says `HeadlessChrome`; real Chrome with a stock UA **and** a non-UTC clock zone cleared in ~4s. Hence `chrome/` — a Debian image with `google-chrome-stable`, a version-derived UA, and `TZ`/`BROWSER_TZ`. Chrome reads the zone *name* through ICU from `/etc/localtime`'s symlink target, ignoring the file's contents, so mounting the host's `/etc/localtime` does **not** work; `/etc/timezone` is mounted instead.
- **UTC is the tell, not a country mismatch.** A UTC clock is the datacenter default, so Cloudflare scores it as one; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while `Asia/Jakarta` **and** `America/New_York` both cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (`Asia/Bangkok`) rather than the measured egress. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one.
- **A challenged page needs the tab kept open.** The interstitial takes seconds to solve and only then writes clearance into the browser's shared cookie jar. Navigate-read-close never clears anything; `BrowserFetcher.run` holds one tab and re-reads until the payload arrives.
## Architecture
@@ -37,7 +40,12 @@ Backend (`cd backend`):
- Single test: `go test -run TestName ./...`
- Build static binary: `CGO_ENABLED=0 go build`
Local stack: `docker compose up` (bookmark-api + postgres + headless-shell; `postgres-data` named volume, `restart: unless-stopped`).
Local stack: `docker compose up` (bookmark-api + postgres + headless-shell; `postgres-data` named volume, `restart: unless-stopped`). The `headless-shell` service keeps its name but now builds `chrome/` — real Google Chrome, for the reason in the hard constraints above.
Live CDP proof (needs a sidecar and network, skipped otherwise):
`SMOKE_BROWSER_WS_URL=ws://<host>:<port> go test -run TestSmokeKagane ./internal/latest`
— fetches a real kagane cover and chapter list. A red run means the challenge is
not clearing from this IP, which is a live fact to re-check, not necessarily a defect.
Smoke test: `curl` endpoints with `Authorization: Bearer <token>`; confirm `OPTIONS` preflight return CORS headers and `/healthz` return 200.
+14 -3
View File
@@ -126,9 +126,20 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
`/userscript/novel-bookmark.user.js`, both supplied by bindmount; the
`__API_TOKEN__` placeholder inside them is substituted with the requesting
Reader's credential at serve time).
`BROWSER_WS_URL` (headless-shell CDP endpoint for kagane and novelfull;
unset disables browser polling and leaves those sites to the userscript
alone).
`BROWSER_WS_URL` (CDP endpoint of the `chrome/` sidecar, used by the poller
for kagane and novelfull *and* by the web UI's kagane cover proxy; unset
disables browser polling and serves 404 from the proxy, leaving those sites
to the userscript alone).
- **kagane covers are proxied, not hot-linked:** kagane serves cover images
behind the same challenge as its pages and with
`cross-origin-resource-policy: same-origin`, so no `<img>` on the web UI's
origin can load one — not even from a browser holding the clearance cookie
(verified 2026-08-08). `Bookmark.CoverURL` rewrites a stored kagane
`og:image` to `/img/kagane/{id}`, served by `internal/web/cover.go` through
`latest.BrowserFetcher.Image` and memoised in-process. The templates render
`.CoverURL`, never `.Cover`. The id is matched against a UUID regex before it
reaches the browser: the stored value is client-supplied, so an unchecked one
is an SSRF primitive pointed at the deployment's own network.
- **Web UI also owns:** session-gated `GET /install/{manga,novel}-bookmark.user.js`
(renders the bindmounted script with the acting Reader's derived credential
substituted in — the credential never appears in page markup, the address
+155
View File
@@ -0,0 +1,155 @@
package main
import (
"context"
"errors"
"net/http"
"net/http/httptest"
"sync/atomic"
"testing"
)
// fakeCovers stands in for the headless browser. It counts calls so the test
// can prove the cache spares the browser a second navigation.
type fakeCovers struct {
body []byte
contentType string
err error
calls atomic.Int32
lastID atomic.Value
}
func (f *fakeCovers) Image(_ context.Context, imageID string) ([]byte, string, error) {
f.calls.Add(1)
f.lastID.Store(imageID)
if f.err != nil {
return nil, "", f.err
}
return f.body, f.contentType, nil
}
const testCoverID = "019fe11a-84c3-7fc3-a84b-88787374b617"
func getCover(t *testing.T, srv http.Handler, path string, cookie *http.Cookie) *httptest.ResponseRecorder {
t.Helper()
req := httptest.NewRequest(http.MethodGet, path, nil)
if cookie != nil {
req.AddCookie(cookie)
}
rr := httptest.NewRecorder()
srv.ServeHTTP(rr, req)
return rr
}
// kagane serves its covers behind a Cloudflare challenge and with
// cross-origin-resource-policy: same-origin, so the UI can only show one by
// re-serving the bytes from its own origin.
func TestKaganeCoverProxiesAndCaches(t *testing.T) {
cf := &fakeCovers{body: []byte("\x00webp-bytes"), contentType: "image/webp"}
cfg := testConfig()
cfg.Covers = cf
srv, st := newWebTestServer(t, cfg)
cookie := sessionCookie(t, st)
for i := range 2 {
rr := getCover(t, srv, "/img/kagane/"+testCoverID, cookie)
if rr.Code != http.StatusOK {
t.Fatalf("request %d: status = %d, want 200", i, rr.Code)
}
if got := rr.Body.String(); got != string(cf.body) {
t.Fatalf("request %d: body = %q, want %q", i, got, cf.body)
}
if got := rr.Header().Get("Content-Type"); got != "image/webp" {
t.Fatalf("request %d: Content-Type = %q, want image/webp", i, got)
}
}
if got := cf.calls.Load(); got != 1 {
t.Fatalf("fetcher called %d times, want 1 — the second read must come from the cache", got)
}
if got := cf.lastID.Load(); got != testCoverID {
t.Fatalf("fetched image id = %v, want %s", got, testCoverID)
}
}
// The proxy reaches a headless browser, so it is not open to the internet.
func TestKaganeCoverRequiresSession(t *testing.T) {
cf := &fakeCovers{body: []byte("x"), contentType: "image/webp"}
cfg := testConfig()
cfg.Covers = cf
srv, _ := newWebTestServer(t, cfg)
rr := getCover(t, srv, "/img/kagane/"+testCoverID, nil)
if rr.Code != http.StatusUnauthorized {
t.Fatalf("status = %d, want 401", rr.Code)
}
if got := cf.calls.Load(); got != 0 {
t.Fatalf("fetcher called %d times for an unauthenticated request, want 0", got)
}
}
func TestKaganeCoverRejectsBadInput(t *testing.T) {
cases := []struct {
name string
id string
fetch *fakeCovers
}{
{
"an id that is not a uuid never reaches the browser",
"solo-leveling",
&fakeCovers{body: []byte("x"), contentType: "image/webp"},
},
{
"a uuid-shaped id with a trailing segment is rejected whole",
testCoverID + "x",
&fakeCovers{body: []byte("x"), contentType: "image/webp"},
},
{
"a challenged fetch is a missing cover",
testCoverID,
&fakeCovers{err: errors.New("challenge held")},
},
{
"a content type outside the image set is not echoed back",
testCoverID,
&fakeCovers{body: []byte("<script>"), contentType: "text/html"},
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
cfg := testConfig()
cfg.Covers = tc.fetch
srv, st := newWebTestServer(t, cfg)
rr := getCover(t, srv, "/img/kagane/"+tc.id, sessionCookie(t, st))
if rr.Code != http.StatusNotFound {
t.Fatalf("status = %d, want 404", rr.Code)
}
})
}
}
// ServeMux path-cleans a traversal into a redirect before the handler runs, so
// the guarantee to pin down is that no request shaped like one ever gets bytes.
func TestKaganeCoverTraversalServesNothing(t *testing.T) {
cf := &fakeCovers{body: []byte("secret"), contentType: "image/webp"}
cfg := testConfig()
cfg.Covers = cf
srv, st := newWebTestServer(t, cfg)
rr := getCover(t, srv, "/img/kagane/../../etc/passwd", sessionCookie(t, st))
if rr.Code == http.StatusOK {
t.Fatalf("status = 200, want anything but a served body")
}
if got := cf.calls.Load(); got != 0 {
t.Fatalf("fetcher called %d times for a traversal, want 0", got)
}
}
// Without BROWSER_WS_URL there is no fetcher, and the endpoint must answer
// rather than reach for a nil one.
func TestKaganeCoverWithoutFetcher(t *testing.T) {
srv, st := newWebTestServer(t, testConfig())
rr := getCover(t, srv, "/img/kagane/"+testCoverID, sessionCookie(t, st))
if rr.Code != http.StatusNotFound {
t.Fatalf("status = %d, want 404", rr.Code)
}
}
+131 -30
View File
@@ -2,7 +2,9 @@ package latest
import (
"context"
"encoding/base64"
"encoding/json"
"errors"
"fmt"
"net/url"
"regexp"
@@ -22,6 +24,12 @@ const challengeTimeout = 45 * time.Second
var kaganeSeriesRe = regexp.MustCompile(`^/series/([0-9a-f-]{36})/?$`)
// kaganeImageIDRe pins the only path segment Image interpolates into an
// outbound URL. The id arrives from a stored cover URL, which a client
// supplied, so it is matched rather than trusted: a headless browser is a
// strong SSRF primitive.
var kaganeImageIDRe = regexp.MustCompile(`^[0-9a-f-]{36}$`)
// BrowserFetcher retrieves pages through a remote headless Chrome over the
// DevTools Protocol.
//
@@ -91,23 +99,6 @@ func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int
return "", 0, fmt.Errorf("not a fetchable browser series url: %q", seriesURL)
}
f.mu.Lock()
defer f.mu.Unlock()
ctx, cancel := context.WithTimeout(ctx, challengeTimeout)
defer cancel()
// A fresh tab per fetch, closed on return, so one wedged page cannot
// poison later polls.
tabCtx, cancelTab := chromedp.NewContext(f.allocCtx)
defer cancelTab()
// Bind the caller's deadline to the tab.
tabCtx, cancelDeadline := context.WithCancel(tabCtx)
defer cancelDeadline()
go func() {
<-ctx.Done()
cancelDeadline()
}()
var body string
// kagane's chapter list is only in its JSON API, which must be called from
// inside the page so the request carries the clearance cookie. novelfull
@@ -123,24 +114,134 @@ func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int
)
}
err := chromedp.Run(tabCtx,
chromedp.Navigate(seriesURL),
// The challenge reloads the page itself when it passes; waiting for the
// site's own root element is what tells us we are through it.
chromedp.WaitReady("body", chromedp.ByQuery),
read,
)
if err != nil {
// novelfull's payload is the DOM itself, and the interstitial has a DOM
// too, so "we have an answer" has to exclude it explicitly. kagane's
// in-page fetch just fails while challenged, which is already the signal.
done := func() bool { return body != "" && (isKagane || !isInterstitial(body)) }
if err := f.run(ctx, seriesURL, read, done); err != nil {
// Challenge never cleared, or the API refused. Indistinguishable from
// here and handled identically by the caller.
if errors.Is(err, errChallengeHeld) {
return "", 403, nil
}
return "", 0, fmt.Errorf("browser fetch %q: %w", seriesURL, err)
}
if body == "" {
// Challenge still up, or the API refused. Indistinguishable from here
// and handled identically by the caller.
return "", 403, nil
}
return body, 200, nil
}
// Image retrieves one kagane cover as raw bytes and its content type.
//
// It exists because kagane serves covers behind the same challenge as its
// pages *and* with `cross-origin-resource-policy: same-origin`, so an <img> on
// the web UI's origin cannot load one even from a browser that already holds
// the clearance cookie (verified 2026-08-08). Proxying is the only route.
//
// The image URL is navigated to rather than fetched from some other kagane
// page: the challenge only runs on a top-level navigation, and once it clears
// the document *is* the image, so a same-origin fetch of location.href reads
// it straight back out of the cache.
//
// The challenge is not solved by the first read: WaitReady("body") is satisfied
// by the interstitial too. run holds the tab open until the in-page fetch
// succeeds, which is what gives the challenge script the seconds it needs.
func (f *BrowserFetcher) Image(ctx context.Context, imageID string) ([]byte, string, error) {
if !kaganeImageIDRe.MatchString(imageID) {
return nil, "", fmt.Errorf("not a kagane image id: %q", imageID)
}
var dataURL string
err := f.run(ctx, "https://kagane.to/api/v2/image/"+imageID+"/compressed",
chromedp.Evaluate(`fetch(location.href).then(r => r.ok
? r.blob().then(b => new Promise(res => {
const fr = new FileReader();
fr.onload = () => res(fr.result);
fr.readAsDataURL(b);
}))
: "")`, &dataURL, awaitPromise),
func() bool { return dataURL != "" })
if err != nil {
return nil, "", fmt.Errorf("browser image %s: %w", imageID, err)
}
// "data:image/webp;base64,<payload>".
head, payload, ok := strings.Cut(dataURL, ";base64,")
if !ok {
return nil, "", fmt.Errorf("browser image %s: not a data url", imageID)
}
raw, err := base64.StdEncoding.DecodeString(payload)
if err != nil {
return nil, "", fmt.Errorf("browser image %s: %w", imageID, err)
}
return raw, strings.TrimPrefix(head, "data:"), nil
}
// errChallengeHeld reports that the budget ran out with the interstitial still
// up. Distinct from a transport failure: it means "this site said no", which
// the poller answers with a 403 and its ordinary cooldown.
var errChallengeHeld = errors.New("challenge held")
// challengePollInterval paces re-reads while a challenge solves itself.
const challengePollInterval = 2 * time.Second
// isInterstitial reports whether html is Cloudflare's challenge page rather
// than the site's own. Matched on the challenge runtime's script path, which is
// stable across the interstitial's wording and locale — the visible "Just a
// moment..." title is neither.
func isInterstitial(html string) bool {
return strings.Contains(html, "/cdn-cgi/challenge-platform/")
}
// run navigates to target and re-reads until done reports an answer, bounded by
// challengeTimeout and by the caller's own deadline, in a tab that is closed on
// return so one wedged page cannot poison later calls.
//
// Holding the tab open across re-reads is the whole point. A Cloudflare
// interstitial needs several seconds of a live page to solve itself and write
// clearance into the browser's shared cookie jar; reading once and closing the
// tab — which is what this did before 2026-08-08 — never gives it that window,
// so every fetch lands on the interstitial and the clearance that would have
// unblocked all the later ones is never obtained.
func (f *BrowserFetcher) run(ctx context.Context, target string, read chromedp.Action, done func() bool) error {
f.mu.Lock()
defer f.mu.Unlock()
ctx, cancel := context.WithTimeout(ctx, challengeTimeout)
defer cancel()
tabCtx, cancelTab := chromedp.NewContext(f.allocCtx)
defer cancelTab()
// Bind the caller's deadline to the tab.
tabCtx, cancelDeadline := context.WithCancel(tabCtx)
defer cancelDeadline()
go func() {
<-ctx.Done()
cancelDeadline()
}()
if err := chromedp.Run(tabCtx,
chromedp.Navigate(target),
chromedp.WaitReady("body", chromedp.ByQuery),
); err != nil {
return err
}
var lastErr error
for {
// The challenge reloads the page when it passes, which tears down the
// execution context mid-read. That is a retry, not a failure.
if err := chromedp.Run(tabCtx, read); err != nil {
lastErr = err
} else if done() {
return nil
}
select {
case <-ctx.Done():
if lastErr != nil {
return fmt.Errorf("%w (last read: %v)", errChallengeHeld, lastErr)
}
return errChallengeHeld
case <-time.After(challengePollInterval):
}
}
}
// kaganeAPIURL maps a stored series_url to the JSON endpoint carrying its
// chapter list. Returning false for anything else is a second line of defence
// behind fetchableSeriesURL: a headless browser is a strong SSRF primitive and
@@ -0,0 +1,91 @@
package latest
import (
"context"
"net/http"
"os"
"testing"
"time"
)
// TestSmokeKaganeImage is the live proof that the cover proxy's fetch actually
// clears Cloudflare and returns image bytes. It needs a real headless Chrome
// with outbound network, so it runs only when SMOKE_BROWSER_WS_URL is set:
//
// docker run --rm --shm-size=1gb -p 19222:9222 chromedp/headless-shell:stable
// SMOKE_BROWSER_WS_URL=ws://127.0.0.1:19222 go test -run TestSmokeKaganeImage ./internal/latest
func TestSmokeKaganeImage(t *testing.T) {
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
if ws == "" {
t.Skip("SMOKE_BROWSER_WS_URL unset")
}
const imageID = "019fe11a-84c3-7fc3-a84b-88787374b617" // SP Baby's cover
// The same URL through a plain client is what the web UI's <img> gets.
// Asserting on it keeps the test honest about why the browser is needed.
req, err := http.NewRequest(http.MethodGet,
"https://kagane.to/api/v2/image/"+imageID+"/compressed", nil)
if err != nil {
t.Fatal(err)
}
if res, err := (&http.Client{Timeout: 15 * time.Second}).Do(req); err == nil {
res.Body.Close()
if res.StatusCode == http.StatusOK {
t.Log("note: kagane answered a plain request 200 — the challenge is not up right now")
}
}
f, err := NewBrowserFetcher(ws)
if err != nil {
t.Fatalf("NewBrowserFetcher: %v", err)
}
defer f.Close()
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
defer cancel()
body, contentType, err := f.Image(ctx, imageID)
if err != nil {
t.Fatalf("Image: %v", err)
}
if len(body) < 1000 {
t.Fatalf("body is %d bytes, want a real image", len(body))
}
if contentType != "image/webp" {
t.Fatalf("content type = %q, want image/webp", contentType)
}
// WebP files start with "RIFF....WEBP".
if string(body[:4]) != "RIFF" || string(body[8:12]) != "WEBP" {
t.Fatalf("body is not a WebP: % x", body[:12])
}
t.Logf("fetched %d bytes of %s", len(body), contentType)
if _, _, err := f.Image(ctx, "not-a-uuid"); err == nil {
t.Fatal("Image accepted a non-uuid id")
}
}
// Control for the test above: the poller's own kagane path, same sidecar. If
// this fails too, the sidecar is not clearing the challenge at all and the
// image result says nothing about Image itself.
func TestSmokeKaganeGet(t *testing.T) {
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
if ws == "" {
t.Skip("SMOKE_BROWSER_WS_URL unset")
}
f, err := NewBrowserFetcher(ws)
if err != nil {
t.Fatalf("NewBrowserFetcher: %v", err)
}
defer f.Close()
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
defer cancel()
body, status, err := f.Get(ctx, "https://kagane.to/series/019fe11a-8670-7cf3-8343-0b02057d3787")
if err != nil {
t.Fatalf("Get: %v", err)
}
t.Logf("status=%d bytes=%d head=%.80q", status, len(body), body)
if status != 200 {
t.Fatalf("status = %d, want 200 — the sidecar is not clearing the challenge", status)
}
}
+16
View File
@@ -137,6 +137,22 @@ func (b Bookmark) Initial() string {
return "?"
}
// kaganeCoverRe matches the cover URL kagane's og:image carries, which is what
// the userscript stores for that site.
var kaganeCoverRe = regexp.MustCompile(`^https://kagane\.to/api/v2/image/([0-9a-f-]{36})/compressed$`)
// CoverURL is the src the web UI puts in an <img>. For every site but kagane
// that is Cover as stored. kagane serves its images behind a Cloudflare
// challenge *and* with `cross-origin-resource-policy: same-origin`, so no page
// on another origin can load one however it asks (verified 2026-08-08); those
// go through the backend's own proxy instead.
func (b Bookmark) CoverURL() string {
if m := kaganeCoverRe.FindStringSubmatch(b.Cover); m != nil {
return "/img/kagane/" + m[1]
}
return b.Cover
}
// Library buckets. A bookmark is in exactly one. This cannot be derived from
// Site: asurascans serves manga and novels from the same /comics/ path, so the
// userscript that recorded the page is the only party that knows which.
+32
View File
@@ -543,6 +543,38 @@ func TestDisplayChapter(t *testing.T) {
}
}
func TestCoverURL(t *testing.T) {
cases := []struct {
name string
cover string
want string
}{
{
"kagane routes through the proxy",
"https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
"/img/kagane/019fe11a-84c3-7fc3-a84b-88787374b617",
},
{
"another site is served as stored",
"https://gg.asuracomic.net/storage/media/1/conversions/cover.webp",
"https://gg.asuracomic.net/storage/media/1/conversions/cover.webp",
},
{
"a lookalike host is not rewritten",
"https://evil.example/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
"https://evil.example/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
},
{"no cover stays empty", "", ""},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
if got := (Bookmark{Cover: tc.cover}).CoverURL(); got != tc.want {
t.Errorf("CoverURL() = %q, want %q", got, tc.want)
}
})
}
}
func TestUpsertKindDefaultsToManga(t *testing.T) {
store := newTestStore(t)
got, err := store.Upsert(store.OwnerID(), Bookmark{
+129
View File
@@ -0,0 +1,129 @@
package web
import (
"context"
"log"
"net/http"
"regexp"
"sync"
"time"
)
// CoverFetcher retrieves one kagane cover by image id. Satisfied by
// latest.BrowserFetcher, and nil when BROWSER_WS_URL is unset — which leaves
// kagane covers exactly as unavailable as they were before this endpoint
// existed, rather than hanging a request on a fetcher that cannot run.
type CoverFetcher interface {
Image(ctx context.Context, imageID string) (body []byte, contentType string, err error)
}
// coverIDRe matches the request path segment that becomes part of an outbound
// URL. The proxy is session-gated, but the id still reaches a headless browser,
// so it is validated at the boundary rather than passed through.
var coverIDRe = regexp.MustCompile(`^[0-9a-f-]{36}$`)
// coverTypes is the set of content types the proxy will echo back. A response
// header sourced from a third party is not repeated verbatim: anything outside
// this set is treated as "not a cover".
var coverTypes = map[string]bool{
"image/webp": true,
"image/jpeg": true,
"image/png": true,
"image/avif": true,
"image/gif": true,
}
// coverTimeout bounds one proxied cover. Shorter than the fetcher's own
// challenge budget on purpose: a browser page is waiting on this, and a cover
// that has not arrived by now is better left as a broken slot than as a request
// holding a connection open.
const coverTimeout = 20 * time.Second
// coverCacheMax caps the in-memory cover cache. Covers are immutable per image
// id and a library holds tens of series, so this is a ceiling that is never
// reached in practice; reaching it clears the map rather than evicting by age.
//
// ponytail: flush-on-full, not LRU. Swap it for an LRU if a library ever grows
// past this and the flush starts costing refetches.
const coverCacheMax = 500
type cachedCover struct {
body []byte
contentType string
}
type coverCache struct {
mu sync.Mutex
m map[string]cachedCover
}
func (c *coverCache) get(id string) (cachedCover, bool) {
c.mu.Lock()
defer c.mu.Unlock()
v, ok := c.m[id]
return v, ok
}
func (c *coverCache) put(id string, v cachedCover) {
c.mu.Lock()
defer c.mu.Unlock()
if c.m == nil || len(c.m) >= coverCacheMax {
c.m = make(map[string]cachedCover, coverCacheMax)
}
c.m[id] = v
}
// kaganeCover serves a kagane cover from the backend's own origin.
//
// kagane answers image requests with a Cloudflare challenge and
// `cross-origin-resource-policy: same-origin`, so the web UI cannot render one
// directly under any combination of referrer policy or crossorigin attribute
// (verified 2026-08-08). Fetching it through the headless browser that already
// clears the challenge, and re-serving it here, is what puts the bytes on an
// origin the page may load from.
//
// ponytail: covers are fetched on first view, one browser navigation at a time
// behind the fetcher's mutex, so a first load of a large kagane library
// trickles in over a few seconds. The cache makes it a one-off. Prefetching
// during the poll cycle is the upgrade if that ever grates.
func (h *Handler) kaganeCover(w http.ResponseWriter, r *http.Request) {
id := r.PathValue("id")
if !coverIDRe.MatchString(id) {
http.NotFound(w, r)
return
}
if h.covers == nil {
http.NotFound(w, r)
return
}
if v, ok := h.coverCache.get(id); ok {
writeCover(w, v)
return
}
ctx, cancel := context.WithTimeout(r.Context(), coverTimeout)
defer cancel()
body, contentType, err := h.covers.Image(ctx, id)
if err != nil {
log.Printf("kagane cover %s: %v", id, err)
http.NotFound(w, r)
return
}
if !coverTypes[contentType] {
log.Printf("kagane cover %s: unexpected content type %q", id, contentType)
http.NotFound(w, r)
return
}
v := cachedCover{body: body, contentType: contentType}
h.coverCache.put(id, v)
writeCover(w, v)
}
// writeCover sends the bytes with a long cache life: an image id names one
// immutable rendering, so a client that has it never needs to ask again.
func writeCover(w http.ResponseWriter, v cachedCover) {
w.Header().Set("Content-Type", v.contentType)
w.Header().Set("Cache-Control", "private, max-age=604800, immutable")
w.Write(v.body)
}
+1 -1
View File
@@ -6,7 +6,7 @@
<div class="row">
<a class="cover" href="{{.ContinueURL}}" target="_blank" rel="noopener noreferrer"
tabindex="-1" aria-hidden="true">
{{if .Cover}}<img src="{{.Cover}}" alt="" loading="lazy">
{{if .CoverURL}}<img src="{{.CoverURL}}" alt="" loading="lazy">
{{/* aria-hidden on the cover link is not enough — Chromium still exposes
the letter because the link is programmatically focusable — so the
monogram carries its own, same as the recent strip's. */}}
+1 -1
View File
@@ -14,7 +14,7 @@
<a class="recent-card {{if .HasNewChapter}}is-new{{end}}" href="{{.ContinueURL}}"
target="_blank" rel="noopener noreferrer">
<span class="recent-cover">
{{if .Cover}}<img src="{{.Cover}}" alt="" loading="lazy">
{{if .CoverURL}}<img src="{{.CoverURL}}" alt="" loading="lazy">
{{else}}<span class="monogram" aria-hidden="true">{{.Initial}}</span>{{end}}
{{if .HasNewChapter}}<span class="foot-rule"></span>
{{else if .Favorite}}<span class="foot-rule brass"></span>{{end}}
+10 -1
View File
@@ -50,6 +50,10 @@ type Handler struct {
// httpClient is the plain stdlib client that talks to Discord. It is not
// an injected interface: tests point APIBase at a stub server instead.
httpClient *http.Client
// covers proxies kagane cover images, which no browser can load directly.
// Nil disables the endpoint — see CoverFetcher.
covers CoverFetcher
coverCache coverCache
}
// listView is what every list-rendering template receives.
@@ -111,7 +115,7 @@ type loginView struct {
// New parses every template up front so a broken one kills the process at
// startup rather than the first request that touches it.
func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, novelPath string) (*Handler, error) {
func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, novelPath string, covers CoverFetcher) (*Handler, error) {
tmpl, err := template.ParseFS(templateFS, "templates/*.html")
if err != nil {
return nil, err
@@ -126,6 +130,7 @@ func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, nove
states: newOAuthStates(),
limiter: session.NewLoginLimiter(),
httpClient: &http.Client{Timeout: discordTimeout},
covers: covers,
}, nil
}
@@ -142,6 +147,10 @@ func (h *Handler) Register(mux *http.ServeMux) {
mux.HandleFunc("POST /ui/bookmarks/{key}/chapter", h.requireSession(h.uiChapter))
mux.HandleFunc("DELETE /ui/bookmarks/{key}", h.requireSession(h.uiDelete))
// Session-gated like every other UI route: the deployment proxies kagane's
// images for its own Readers, not for the internet.
mux.HandleFunc("GET /img/kagane/{id}", h.requireSession(h.kaganeCover))
// Install endpoints render the script directly under the session: the
// credential travels inside the served bytes, never in the address bar or
// the page markup. Updates after install use the credential-bearing /u/
+35 -24
View File
@@ -47,6 +47,10 @@ type Config struct {
NovelUserscriptPath string
// LatestPoll configures the background latest-chapter fetcher.
LatestPoll LatestPoll
// Covers proxies kagane cover images for the web UI. Not from the
// environment: it is the shared headless browser, wired in main once it
// connects, and nil in every test router.
Covers web.CoverFetcher
}
// LatestPoll configures the background latest-chapter poller.
@@ -206,7 +210,7 @@ func newRouter(s *store.Store, cfg Config) http.Handler {
// The browser UI is always registered; signing in is Discord OAuth, so
// there is no password to forget and no gate to leave unset.
wh, err := web.New(s, cfg.Discord, []byte(cfg.TokenKey),
cfg.UserscriptPath, cfg.NovelUserscriptPath)
cfg.UserscriptPath, cfg.NovelUserscriptPath, cfg.Covers)
if err != nil {
log.Fatalf("web handler: %v", err)
}
@@ -270,9 +274,26 @@ func main() {
// The poller is off the request path entirely: if it cannot start, the
// service still serves bookmarks and the userscript still captures latest
// chapters on its own.
//
// One headless browser serves both consumers that need a Cloudflare
// challenge cleared: the poller's kagane/novelfull fetches and the web
// UI's kagane cover proxy. Optional — unset leaves both degraded to what
// they were before the sidecar existed.
var browser latest.Fetcher
pollCtx, stopPoll := context.WithCancel(context.Background())
defer stopPoll()
startLatestPoller(pollCtx, s, cfg.LatestPoll)
if ws := strings.TrimSpace(os.Getenv("BROWSER_WS_URL")); ws != "" {
bf, err := latest.NewBrowserFetcher(ws)
if err != nil {
log.Printf("browser fetcher disabled: %v", err)
} else {
browser = bf
cfg.Covers = bf
context.AfterFunc(pollCtx, bf.Close)
log.Printf("browser fetcher at %s", ws)
}
}
startLatestPoller(pollCtx, s, cfg.LatestPoll, browser)
srv := &http.Server{
Addr: ":" + cfg.Port,
@@ -307,7 +328,7 @@ func main() {
// HTTP client cannot be built. Any problem here is logged and skipped: this
// feature going missing degrades the service to userscript-only latest-chapter
// tracking, which is exactly how it behaved before.
func startLatestPoller(ctx context.Context, s *store.Store, cfg LatestPoll) {
func startLatestPoller(ctx context.Context, s *store.Store, cfg LatestPoll, browser latest.Fetcher) {
if !cfg.Enabled {
log.Println("latest-chapter poller: disabled by config")
return
@@ -317,28 +338,18 @@ func startLatestPoller(ctx context.Context, s *store.Store, cfg LatestPoll) {
log.Printf("latest-chapter poller: disabled, cannot build client: %v", err)
return
}
// Nil browser: sites behind a JavaScript challenge are simply not polled,
// and their latest_chapter comes from the userscript alone — which is how
// the service behaved before the sidecar existed.
p := &latest.Poller{
Store: s,
Fetch: f,
Now: time.Now,
Cooldown: cfg.Cooldown,
Interval: cfg.Interval,
Stagger: cfg.Stagger,
Batch: cfg.Batch,
}
// Optional: without it, sites behind a JavaScript challenge are simply not
// polled, and their latest_chapter comes from the userscript alone — which
// is how the service behaved before the sidecar existed.
if ws := strings.TrimSpace(os.Getenv("BROWSER_WS_URL")); ws != "" {
bf, err := latest.NewBrowserFetcher(ws)
if err != nil {
log.Printf("latest-chapter poller: browser fetcher disabled: %v", err)
} else {
p.BrowserFetch = bf
context.AfterFunc(ctx, bf.Close)
log.Printf("latest-chapter poller: browser fetcher at %s", ws)
}
Store: s,
Fetch: f,
BrowserFetch: browser,
Now: time.Now,
Cooldown: cfg.Cooldown,
Interval: cfg.Interval,
Stagger: cfg.Stagger,
Batch: cfg.Batch,
}
go p.Run(ctx)
+39
View File
@@ -0,0 +1,39 @@
# syntax=docker/dockerfile:1
# Real Google Chrome for the latest-chapter poller and the kagane cover proxy.
#
# Not chromedp/headless-shell, which this replaces. headless-shell is a stripped
# Chrome build and Cloudflare's managed challenge on kagane.to never clears for
# it: measured 2026-08-08, 60s of a held-open tab still served the interstitial,
# while stock Chrome from the same IP cleared in ~4s. The tells are structural
# rather than a header — navigator.webdriver true, an empty plugin list, and
# Chromium- rather than Chrome-branded client hints. Overriding webdriver alone
# was tried and did not move it, so the browser build itself is the fix.
#
# zenika/alpine-chrome was also tried: its Chrome is 124 (2024), old enough that
# Cloudflare refuses it outright and old enough to break chromedp's CDP structs.
FROM debian:trixie-slim
# Chrome is deliberately unpinned, against the usual rule. A pinned build goes
# stale, and a stale browser is exactly what Cloudflare turns away — the 124 in
# alpine-chrome is the worked example. Rebuild is the upgrade path.
RUN apt-get update \
&& apt-get install -y --no-install-recommends ca-certificates wget gnupg \
&& wget -qO- https://dl.google.com/linux/linux_signing_key.pub \
| gpg --dearmor -o /usr/share/keyrings/google-chrome.gpg \
&& echo "deb [arch=amd64 signed-by=/usr/share/keyrings/google-chrome.gpg] https://dl.google.com/linux/chrome/deb/ stable main" \
> /etc/apt/sources.list.d/google-chrome.list \
&& apt-get update \
&& apt-get install -y --no-install-recommends google-chrome-stable socat \
&& rm -rf /var/lib/apt/lists/*
# Unprivileged: Chrome refuses to run as root, and the CDP endpoint is a shell
# on whatever user owns it.
RUN useradd --create-home --shell /usr/sbin/nologin chrome
USER chrome
WORKDIR /home/chrome
COPY entrypoint.sh /entrypoint.sh
EXPOSE 9222
ENTRYPOINT ["/entrypoint.sh"]
+61
View File
@@ -0,0 +1,61 @@
#!/bin/sh
set -eu
# A UTC clock is itself the bot signal — Cloudflare treats it as the datacenter
# default — and kagane's challenge then never clears. Measured 2026-08-08 with
# an identical container on one Indonesian egress IP: UTC never cleared in 60s
# (twice), while Asia/Jakarta and America/New_York both cleared in 4s. Any real
# zone will do; the zone does not have to match the IP's country, it just must
# not be UTC. It does have to be right the way Chrome reads it.
#
# TZ must carry the zone *name*. Chrome resolves the zone through ICU, which
# takes the name from /etc/localtime's symlink target and ignores the file's
# contents; bind-mounting the host's /etc/localtime therefore lands on the
# image's own symlink to Etc/UTC and leaves glibc reporting +07 while Chrome
# still reports UTC. /etc/timezone, mounted by docker-compose.yml, is the name.
[ -n "${TZ:-}" ] || TZ=$(cat /etc/timezone 2>/dev/null || echo UTC)
export TZ
# Chrome's own UA advertises "HeadlessChrome" under --headless=new, and that one
# token is the difference between kagane.to's challenge clearing in ~4s and
# never clearing at all (measured 2026-08-08, same host, same Chrome, only the
# UA changed). Overriding it does not touch the Sec-CH-UA client hints, which
# report the real version, so the version is read back out of the binary rather
# than hardcoded: a hardcoded one would drift out of step with the hints on the
# next Chrome update and become a fresh tell.
major=$(google-chrome-stable --version | sed -E 's/[^0-9]*([0-9]+)\..*/\1/')
ua="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/${major}.0.0.0 Safari/537.36"
# Chrome binds its DevTools port to loopback and silently ignores
# --remote-debugging-address (verified 2026-08-08: Chrome 151 with
# --remote-debugging-address=0.0.0.0 still listened on 127.0.0.1 only), so the
# caller — another container — cannot reach it directly. socat fronting the
# loopback port is how chromedp/headless-shell solved the same problem and is
# why this image is a drop-in for it.
#
# Nothing publishes 9222; reachability is the `browser` network in
# docker-compose.yml, and an exposed CDP endpoint is remote code execution.
socat TCP-LISTEN:9222,fork,reuseaddr TCP:127.0.0.1:9223 &
# Chrome stays in the foreground so that its death takes the container down and
# compose's restart policy applies; a backgrounded browser behind a live socat
# would leave the sidecar looking healthy while answering nothing.
#
# No --enable-automation: it sets navigator.webdriver, the first thing a bot
# check reads.
#
# --no-sandbox because Chrome's zygote wants user namespaces, which Docker's
# default profile does not hand out; the alternative is --cap-add=SYS_ADMIN,
# which gives the container strictly more than it takes away. Containment here
# is the unprivileged user, the isolated network, and the fact that this
# browser only ever navigates to kagane.to and novelfull.com.
exec google-chrome-stable \
--headless=new \
--no-sandbox \
--remote-debugging-port=9223 \
--user-agent="$ua" \
--user-data-dir=/home/chrome/profile \
--no-first-run \
--no-default-browser-check \
--disable-gpu \
about:blank
+27 -12
View File
@@ -24,6 +24,13 @@ services:
# comes from .env so it is never committed.
DATABASE_URL: ${DATABASE_URL:-postgres://bookmarks:${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD in .env}@postgres:5432/bookmarks?sslmode=disable}
PORT: "8080"
# Log timestamps only. Go's `log` stamps lines in local time, and this
# service has no other use for a zone: bookmark timestamps are unix ms
# and the two real time columns are timestamptz, both absolute instants.
# Purely so these lines read on the same clock as the sidecar's. Named
# API_TZ rather than TZ so an operator's exported shell TZ cannot leak
# in; distroless already carries tzdata, so the name just resolves.
TZ: ${API_TZ:-Asia/Jakarta}
# Discord OAuth for the browser UI (ADR-0002). The first four are
# required; DISCORD_REQUIRED_ROLE is optional and empty by default.
# Guild membership is the whole gate: any member becomes a Reader.
@@ -95,8 +102,25 @@ services:
- db
headless-shell:
image: chromedp/headless-shell:stable
# Real Google Chrome, not chromedp/headless-shell — see chrome/Dockerfile.
# The service name is kept so existing overrides and BROWSER_WS_URL stay put.
build: ./chrome
image: bookmarkmanager-chrome:latest
restart: unless-stopped
environment:
# A UTC clock is itself the bot signal: Cloudflare treats it as the
# datacenter default, and kagane's challenge then never clears. Measured
# 2026-08-08, identical container, one Indonesian egress IP: UTC never
# cleared in 60s (twice); Asia/Jakarta and America/New_York both cleared
# in 4s. So any real zone works and it need not match the IP's country —
# only UTC fails. Unset falls back to the host's /etc/timezone below,
# which is a real zone whenever the host clock is set to local time; set
# BROWSER_TZ when the host runs UTC.
TZ: ${BROWSER_TZ:-}
volumes:
# The zone *name*, which is what Chrome's ICU needs — see chrome/entrypoint.sh.
# Absent on a non-Debian host, which the entrypoint handles by falling back to UTC.
- /etc/timezone:/etc/timezone:ro
# Chrome allocates shared memory per tab and dies on Docker's 64MB default.
shm_size: '1gb'
# Reaps zombie renderer processes, which otherwise accumulate for the
@@ -104,17 +128,8 @@ services:
init: true
# Deliberately no `ports:` — an exposed CDP endpoint is remote code
# execution. Only bookmark-api, via the `browser` network below, may reach it.
# Don't pass --remote-debugging-address/--remote-debugging-port here: the
# image's own entrypoint (/headless-shell/run.sh) already starts Chrome on
# 127.0.0.1:9223 and fronts it with a socat proxy listening on 0.0.0.0:9222.
# Redeclaring the port flag here overrides Chrome's, so it binds 9222
# directly (IPv6 loopback only) instead of 9223 — collides with socat's own
# bind on 9222 and leaves nothing listening on 9223, so every external
# connection to headless-shell:9222 fails with EOF. Only pass flags the
# entrypoint doesn't already set.
command:
- --disable-gpu
- --no-sandbox
# No `command:` either: every flag this browser needs is in its entrypoint,
# and the UA override there is load-bearing for the challenge.
networks:
browser:
# Pinned so BROWSER_WS_URL can name an IP (required, see above) that
+19
View File
@@ -44,6 +44,25 @@ Guidance for OpenCode (and Claude Code) working under `userscript/`. See root `A
Encodings (incl. triple-encoded punctuation like `%25252D`) identical
on /manga/ and /title/ pages, so decode-once seriesIds match — verified
2026-07-28.
- **comix.to**: series `/title/<id>-<slug>`, chapter
`/title/<id>-<slug>/<uploadId>-chapter-<n>`. Only the leading `<id>` is
identity — the slug re-renders when a series is renamed (`comixSeriesId`).
An SPA that **never rewrites `og:title`**: the server-rendered head keeps
whatever document loaded first, so on a cold load `og:title` is the homepage's
"Comix — Read Comics online for free" and after an in-page hop it is the
*previous* series' name. `document.title` is the one thing client routing does
update, so titles come from there, with the chapter page's `" · Ch.<n>"` tail
stripped. Covers likewise: `og:image` is absent, so the cover is the `img`
whose `alt` matches the cleaned title — verified live 2026-08-08.
- **kagane.to**: series `/series/<uuid>`, reader
`/series/<uuid>/reader/<bookUuid>`. Reader URLs carry no chapter number, so
the number comes out of `og:title`. Two shapes exist: `"<Series> - Chapter
<n>[ - Episode <n>]"` and, for volume-numbered series, `"<Series> - Volume <v>
Chapter <n>"` with no episode name — both must yield a bare series title, or
the volume tail lands in the bookmark's title. Its covers are challenge- and
CORP-protected, so the web UI proxies them; the userscript still stores the
raw `og:image`. Behind a Cloudflare JS challenge, so the backend polls it
through the headless browser.
- **novelfull.com** (novel script): series `/<slug>.html`, chapter
`/<slug>/chapter-<n>[-<title-slug>].html`. No `og:*` tags at all — title from
`h3.title` (series) or `a.truyen-title` (chapter), cover from
+39 -13
View File
@@ -242,6 +242,12 @@
matches: (loc) => /(^|\.)comix\.to$/.test(loc.hostname),
detect(loc) {
const path = loc.pathname;
// comix client-routes without ever rewriting og:title — the head keeps
// whatever the first server-rendered document carried, so a bookmark
// taken after a client route got the homepage's title, then the
// previous series'. document.title is the one thing its router does
// update. Verified live 2026-08-08; do not "restore" meta("og:title").
const pageTitle = cleanTitle(document.title);
// /title/<id>-<slug>/<uploadId>-chapter-<n>. Several uploads (different
// groups or languages) share one chapter number; the number is the
// progress identity, the upload id is not.
@@ -252,8 +258,8 @@
type: "chapter",
site: this.site,
seriesId: comixSeriesId(m[1]),
title: cleanTitle(meta("og:title")),
cover: coverFromPage(),
title: pageTitle,
cover: coverFromPage(pageTitle),
seriesUrl: loc.origin + "/title/" + m[1],
chapterLabel: "Chapter " + m[2],
chapterNum: isNaN(num) ? null : num,
@@ -267,8 +273,8 @@
type: "series",
site: this.site,
seriesId: comixSeriesId(m[1]),
title: cleanTitle(meta("og:title")),
cover: coverFromPage(),
title: pageTitle,
cover: coverFromPage(pageTitle),
seriesUrl: loc.origin + "/title/" + m[1],
chapterLabel: null,
chapterNum: null,
@@ -277,7 +283,7 @@
}
return { type: "other" };
// comix chapter og:title is "<Title> · Ch.<n>"; series is clean.
// comix chapter document.title is "<Title> · Ch.<n>"; series is clean.
function cleanTitle(t) {
if (!t) return "";
return t.replace(/\s*·\s*Ch\.[\d.]+\s*$/i, "").trim();
@@ -287,8 +293,7 @@
// the DOM for a cover. Matching on alt rather than a class keeps it off
// the site's styling: the cover is the image whose alt is the title.
// Do not "simplify" this into meta("og:image") — that returns null.
function coverFromPage() {
const title = cleanTitle(meta("og:title"));
function coverFromPage(title) {
if (!title || !document.querySelectorAll) return "";
for (const img of document.querySelectorAll("img[alt]")) {
if (img.getAttribute("alt") === title) return img.getAttribute("src") || "";
@@ -313,6 +318,14 @@
},
};
// Kagane builds the reader og:title suffix out of the book's metadata, so
// every combination occurs: the volume part appears only when the book has a
// volume_no, the episode part only when it has a non-empty title. All four
// shapes captured live 2026-08-08 — "SP Baby - Volume 1 Chapter 1" is the one
// the old trailing-space regex missed, which left both the number and the
// series title wrong.
const KAGANE_CHAPTER_SUFFIX = /\s-\s(?:Volume\s[\d.]+\s)?Chapter\s([\d.]+)(?:\s-\s.*)?$/i;
const kagane = {
site: "kagane",
matches: (loc) => /(^|\.)kagane\.to$/.test(loc.hostname),
@@ -353,9 +366,8 @@
}
return { type: "other" };
// Reader og:title is "<Title> - Chapter <n> - <episode name>".
function chapterNumFromTitle(t) {
const m = t && t.match(/\s-\sChapter\s([\d.]+)\s/);
const m = t && t.match(KAGANE_CHAPTER_SUFFIX);
if (!m) return null;
const num = parseFloat(m[1]);
return isNaN(num) ? null : num;
@@ -363,7 +375,7 @@
function cleanTitle(t) {
if (!t) return "";
return t.replace(/\s-\sChapter\s[\d.]+\s-\s.*$/i, "").trim();
return t.replace(KAGANE_CHAPTER_SUFFIX, "").trim();
}
},
// Reader hrefs are uuids with no number in them, so no maximum can be taken
@@ -1450,8 +1462,15 @@
}
let lastUrl = location.href;
function onNavigate() {
let lastPageSig = "";
function setPage() {
state.page = detect();
lastPageSig = JSON.stringify(state.page);
}
function onNavigate() {
setPage();
render();
maybeAutoUpdate();
maybeCaptureLatestOnSeriesPage();
@@ -1484,7 +1503,14 @@
wrap("pushState");
wrap("replaceState");
window.addEventListener("popstate", fire);
setInterval(fire, 1500); // catch routes that bypass history
// Also catches routes that bypass history — and comix, which fills
// document.title a beat after the route changes, so the 300ms snapshot
// above can still hold the previous page's title. Re-detect whenever what
// we would read has changed, not only when the URL has.
setInterval(() => {
fire();
if (JSON.stringify(detect()) !== lastPageSig) onNavigate();
}, 1500);
}
// ============================================================
@@ -1563,7 +1589,7 @@
function init() {
buildUI();
state.page = detect();
setPage();
render();
installNavWatcher();
installLongPress();
+57 -8
View File
@@ -31,6 +31,9 @@ globalThis.location = {
// og: meta tags the adapters read through meta(). Reassigned per test.
let metaTags = {};
// document.title. comix's SPA rewrites this on client routing but never
// og:title, so the comix adapter reads it instead. Reassigned per test.
let docTitle = "";
// img[alt] elements comix's coverFromPage() scans. Reassigned per test; each
// entry is {alt, src}.
let pageImages = [];
@@ -48,6 +51,9 @@ globalThis.document = {
}));
},
addEventListener() {},
get title() {
return docTitle;
},
body: undefined,
};
@@ -193,8 +199,15 @@ test("comixSeriesId leaves a bare id untouched", () => {
assert.equal(comixSeriesId("n8we"), "n8we");
});
// comix is an SPA that rewrites document.title on client routing but leaves the
// server-rendered og:title untouched, so every test here pins og:title to a
// STALE value — the homepage title on first hop, the previous series after
// that. Captured live 2026-08-08.
const COMIX_STALE_HOME = "Comix - Read Comics online for free";
test("comix detects a series page", () => {
metaTags = { "og:title": "Dungeons and Crayons" };
metaTags = { "og:title": COMIX_STALE_HOME };
docTitle = "Dungeons and Crayons";
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
assert.equal(p.type, "series");
assert.equal(p.site, "comix");
@@ -204,8 +217,16 @@ test("comix detects a series page", () => {
assert.equal(p.chapterNum, null);
});
test("comix ignores a previous series' stale og:title", () => {
metaTags = { "og:title": "Full-Time Awakening" };
docTitle = "Dungeons and Crayons";
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
assert.equal(p.title, "Dungeons and Crayons");
});
test("comix detects a chapter page and strips the Ch. suffix from the title", () => {
metaTags = { "og:title": "Dungeons and Crayons · Ch.80" };
metaTags = { "og:title": COMIX_STALE_HOME };
docTitle = "Dungeons and Crayons · Ch.80";
const p = comix.detect(
loc("https://comix.to/title/n8we-dungeons-and-crayons/11139891-chapter-80")
);
@@ -218,7 +239,8 @@ test("comix detects a chapter page and strips the Ch. suffix from the title", ()
});
test("comix parses decimal chapter numbers", () => {
metaTags = { "og:title": "Dungeons and Crayons · Ch.80.5" };
metaTags = {};
docTitle = "Dungeons and Crayons · Ch.80.5";
const p = comix.detect(
loc("https://comix.to/title/n8we-dungeons-and-crayons/11139891-chapter-80.5")
);
@@ -226,19 +248,19 @@ test("comix parses decimal chapter numbers", () => {
});
test("comix.detect reads the cover from an img whose alt matches the cleaned title", () => {
metaTags = { "og:title": "Dungeons and Crayons · Ch.80" };
metaTags = { "og:title": COMIX_STALE_HOME };
docTitle = "Dungeons and Crayons";
pageImages = [
{ alt: "Some Other Series", src: "https://cdn.example/other.jpg" },
{ alt: "Dungeons and Crayons", src: "https://cdn.example/cover.jpg" },
];
const p = comix.detect(
loc("https://comix.to/title/n8we-dungeons-and-crayons/11139891-chapter-80")
);
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
assert.equal(p.cover, "https://cdn.example/cover.jpg");
});
test("comix.detect leaves cover empty when no img alt matches the title", () => {
metaTags = { "og:title": "Dungeons and Crayons" };
metaTags = {};
docTitle = "Dungeons and Crayons";
pageImages = [{ alt: "Some Other Series", src: "https://cdn.example/other.jpg" }];
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
assert.equal(p.cover, "");
@@ -315,6 +337,33 @@ test("kagane reads the chapter number out of og:title", () => {
assert.equal(p.seriesUrl, "https://kagane.to/series/" + KAGANE_SERIES);
});
// Volume-numbered series render the suffix as "- Volume <v> Chapter <n>" with
// no episode name, because the book carries volume_no and an empty title.
// Captured live 2026-08-08 from SP Baby.
test("kagane reads through a Volume-numbered chapter suffix", () => {
metaTags = {
"og:title": "SP Baby - Volume 1 Chapter 1",
"og:image": "https://kagane.to/api/v2/image/abc/compressed",
};
const p = kagane.detect(
loc("https://kagane.to/series/" + KAGANE_SERIES + "/reader/" + KAGANE_BOOK)
);
assert.equal(p.title, "SP Baby");
assert.equal(p.chapterNum, 1);
assert.equal(p.chapterLabel, "Chapter 1");
});
// A book with neither a volume nor an episode name ends the title right after
// the number, which the old trailing-\s regex could not match.
test("kagane reads a chapter suffix with no episode name", () => {
metaTags = { "og:title": "Some Series - Chapter 7.5" };
const p = kagane.detect(
loc("https://kagane.to/series/" + KAGANE_SERIES + "/reader/" + KAGANE_BOOK)
);
assert.equal(p.title, "Some Series");
assert.equal(p.chapterNum, 7.5);
});
test("kagane yields a null chapterNum when og:title has no chapter", () => {
metaTags = { "og:title": "Infinite Decryption: The Strongest Level 0" };
const p = kagane.detect(