Fix comix titles and covers, kagane volume chapters, and kagane cover rendering #37

Merged
sulthan merged 6 commits from fix/comix-kagane-titles-and-covers into main 2026-08-08 23:27:33 +07:00
19 changed files with 886 additions and 94 deletions
+20
View File
@@ -90,3 +90,23 @@ DISCORD_REDIRECT_URI=
# HTTP handler 500s any /json/version request whose Host header isn't an IP or # HTTP handler 500s any /json/version request whose Host header isn't an IP or
# "localhost", which silently breaks every kagane poll. # "localhost", which silently breaks every kagane poll.
# BROWSER_WS_URL=ws://172.28.0.10:9222 # BROWSER_WS_URL=ws://172.28.0.10:9222
# Clock zone the headless browser reports. A UTC clock is itself the bot
# signal — Cloudflare treats it as the datacenter default — and kagane's
# challenge then never clears. Measured 2026-08-08, identical container, one
# Indonesian egress IP: UTC never cleared in 60s (twice); Asia/Jakarta and
# America/New_York both cleared in 4s. So any real zone works; it does not
# have to match the IP's country, it just must not be UTC.
#
# Unset falls back to the host's /etc/timezone, which is a real zone whenever
# the host clock is set to local time. Set this when the host runs UTC — a UTC
# server is exactly the case that fails. Only the browser sidecar reads it —
# the backend's own zone is API_TZ below, and is cosmetic.
# BROWSER_TZ=Asia/Jakarta
# Zone the backend stamps its log lines in. Cosmetic only — it exists so the
# API's logs read on the same clock as the browser sidecar's. Nothing else in
# the service has a zone: bookmark timestamps are unix ms, and the two real
# time columns are timestamptz. Defaults to Asia/Jakarta; set to UTC for the
# conventional server default.
# API_TZ=Asia/Jakarta
+9 -1
View File
@@ -20,6 +20,9 @@ Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free
- Userscript run in **isolated world**, so embedded API token safe from site's JS. - Userscript run in **isolated world**, so embedded API token safe from site's JS.
- Cloudflare's block on manga sites **IP-reputation-based, not universal — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — Cloudflare's bot scoring can flip previously-clean IP without notice. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test. - Cloudflare's block on manga sites **IP-reputation-based, not universal — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — Cloudflare's bot scoring can flip previously-clean IP without notice. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test.
- **kagane.to and novelfull.com are the exception to the above** — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`) and skips them entirely when that's unset. The four other sites poll fine over plain TLS. - **kagane.to and novelfull.com are the exception to the above** — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`) and skips them entirely when that's unset. The four other sites poll fine over plain TLS.
- **The CDP sidecar must look like a real browser, and stock headless images don't.** Measured 2026-08-08 against kagane.to, all from the same IP: `chromedp/headless-shell:stable` never cleared the challenge in 90s (`navigator.webdriver` true, empty plugin list, Chromium-branded client hints — suppressing `webdriver` alone changed nothing); `zenika/alpine-chrome` ships Chrome 124, refused outright; real Chrome with the default `--headless=new` UA never cleared, because the UA says `HeadlessChrome`; real Chrome with a stock UA **and** a non-UTC clock zone cleared in ~4s. Hence `chrome/` — a Debian image with `google-chrome-stable`, a version-derived UA, and `TZ`/`BROWSER_TZ`. Chrome reads the zone *name* through ICU from `/etc/localtime`'s symlink target, ignoring the file's contents, so mounting the host's `/etc/localtime` does **not** work; `/etc/timezone` is mounted instead.
- **UTC is the tell, not a country mismatch.** A UTC clock is the datacenter default, so Cloudflare scores it as one; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while `Asia/Jakarta` **and** `America/New_York` both cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (`Asia/Bangkok`) rather than the measured egress. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one.
- **A challenged page needs the tab kept open.** The interstitial takes seconds to solve and only then writes clearance into the browser's shared cookie jar. Navigate-read-close never clears anything; `BrowserFetcher.run` holds one tab and re-reads until the payload arrives.
## Architecture ## Architecture
@@ -37,7 +40,12 @@ Backend (`cd backend`):
- Single test: `go test -run TestName ./...` - Single test: `go test -run TestName ./...`
- Build static binary: `CGO_ENABLED=0 go build` - Build static binary: `CGO_ENABLED=0 go build`
Local stack: `docker compose up` (bookmark-api + postgres + headless-shell; `postgres-data` named volume, `restart: unless-stopped`). Local stack: `docker compose up` (bookmark-api + postgres + headless-shell; `postgres-data` named volume, `restart: unless-stopped`). The `headless-shell` service keeps its name but now builds `chrome/` — real Google Chrome, for the reason in the hard constraints above.
Live CDP proof (needs a sidecar and network, skipped otherwise):
`SMOKE_BROWSER_WS_URL=ws://<host>:<port> go test -run TestSmokeKagane ./internal/latest`
— fetches a real kagane cover and chapter list. A red run means the challenge is
not clearing from this IP, which is a live fact to re-check, not necessarily a defect.
Smoke test: `curl` endpoints with `Authorization: Bearer <token>`; confirm `OPTIONS` preflight return CORS headers and `/healthz` return 200. Smoke test: `curl` endpoints with `Authorization: Bearer <token>`; confirm `OPTIONS` preflight return CORS headers and `/healthz` return 200.
+14 -3
View File
@@ -126,9 +126,20 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
`/userscript/novel-bookmark.user.js`, both supplied by bindmount; the `/userscript/novel-bookmark.user.js`, both supplied by bindmount; the
`__API_TOKEN__` placeholder inside them is substituted with the requesting `__API_TOKEN__` placeholder inside them is substituted with the requesting
Reader's credential at serve time). Reader's credential at serve time).
`BROWSER_WS_URL` (headless-shell CDP endpoint for kagane and novelfull; `BROWSER_WS_URL` (CDP endpoint of the `chrome/` sidecar, used by the poller
unset disables browser polling and leaves those sites to the userscript for kagane and novelfull *and* by the web UI's kagane cover proxy; unset
alone). disables browser polling and serves 404 from the proxy, leaving those sites
to the userscript alone).
- **kagane covers are proxied, not hot-linked:** kagane serves cover images
behind the same challenge as its pages and with
`cross-origin-resource-policy: same-origin`, so no `<img>` on the web UI's
origin can load one — not even from a browser holding the clearance cookie
(verified 2026-08-08). `Bookmark.CoverURL` rewrites a stored kagane
`og:image` to `/img/kagane/{id}`, served by `internal/web/cover.go` through
`latest.BrowserFetcher.Image` and memoised in-process. The templates render
`.CoverURL`, never `.Cover`. The id is matched against a UUID regex before it
reaches the browser: the stored value is client-supplied, so an unchecked one
is an SSRF primitive pointed at the deployment's own network.
- **Web UI also owns:** session-gated `GET /install/{manga,novel}-bookmark.user.js` - **Web UI also owns:** session-gated `GET /install/{manga,novel}-bookmark.user.js`
(renders the bindmounted script with the acting Reader's derived credential (renders the bindmounted script with the acting Reader's derived credential
substituted in — the credential never appears in page markup, the address substituted in — the credential never appears in page markup, the address
+155
View File
@@ -0,0 +1,155 @@
package main
import (
"context"
"errors"
"net/http"
"net/http/httptest"
"sync/atomic"
"testing"
)
// fakeCovers stands in for the headless browser. It counts calls so the test
// can prove the cache spares the browser a second navigation.
type fakeCovers struct {
body []byte
contentType string
err error
calls atomic.Int32
lastID atomic.Value
}
func (f *fakeCovers) Image(_ context.Context, imageID string) ([]byte, string, error) {
f.calls.Add(1)
f.lastID.Store(imageID)
if f.err != nil {
return nil, "", f.err
}
return f.body, f.contentType, nil
}
const testCoverID = "019fe11a-84c3-7fc3-a84b-88787374b617"
func getCover(t *testing.T, srv http.Handler, path string, cookie *http.Cookie) *httptest.ResponseRecorder {
t.Helper()
req := httptest.NewRequest(http.MethodGet, path, nil)
if cookie != nil {
req.AddCookie(cookie)
}
rr := httptest.NewRecorder()
srv.ServeHTTP(rr, req)
return rr
}
// kagane serves its covers behind a Cloudflare challenge and with
// cross-origin-resource-policy: same-origin, so the UI can only show one by
// re-serving the bytes from its own origin.
func TestKaganeCoverProxiesAndCaches(t *testing.T) {
cf := &fakeCovers{body: []byte("\x00webp-bytes"), contentType: "image/webp"}
cfg := testConfig()
cfg.Covers = cf
srv, st := newWebTestServer(t, cfg)
cookie := sessionCookie(t, st)
for i := range 2 {
rr := getCover(t, srv, "/img/kagane/"+testCoverID, cookie)
if rr.Code != http.StatusOK {
t.Fatalf("request %d: status = %d, want 200", i, rr.Code)
}
if got := rr.Body.String(); got != string(cf.body) {
t.Fatalf("request %d: body = %q, want %q", i, got, cf.body)
}
if got := rr.Header().Get("Content-Type"); got != "image/webp" {
t.Fatalf("request %d: Content-Type = %q, want image/webp", i, got)
}
}
if got := cf.calls.Load(); got != 1 {
t.Fatalf("fetcher called %d times, want 1 — the second read must come from the cache", got)
}
if got := cf.lastID.Load(); got != testCoverID {
t.Fatalf("fetched image id = %v, want %s", got, testCoverID)
}
}
// The proxy reaches a headless browser, so it is not open to the internet.
func TestKaganeCoverRequiresSession(t *testing.T) {
cf := &fakeCovers{body: []byte("x"), contentType: "image/webp"}
cfg := testConfig()
cfg.Covers = cf
srv, _ := newWebTestServer(t, cfg)
rr := getCover(t, srv, "/img/kagane/"+testCoverID, nil)
if rr.Code != http.StatusUnauthorized {
t.Fatalf("status = %d, want 401", rr.Code)
}
if got := cf.calls.Load(); got != 0 {
t.Fatalf("fetcher called %d times for an unauthenticated request, want 0", got)
}
}
func TestKaganeCoverRejectsBadInput(t *testing.T) {
cases := []struct {
name string
id string
fetch *fakeCovers
}{
{
"an id that is not a uuid never reaches the browser",
"solo-leveling",
&fakeCovers{body: []byte("x"), contentType: "image/webp"},
},
{
"a uuid-shaped id with a trailing segment is rejected whole",
testCoverID + "x",
&fakeCovers{body: []byte("x"), contentType: "image/webp"},
},
{
"a challenged fetch is a missing cover",
testCoverID,
&fakeCovers{err: errors.New("challenge held")},
},
{
"a content type outside the image set is not echoed back",
testCoverID,
&fakeCovers{body: []byte("<script>"), contentType: "text/html"},
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
cfg := testConfig()
cfg.Covers = tc.fetch
srv, st := newWebTestServer(t, cfg)
rr := getCover(t, srv, "/img/kagane/"+tc.id, sessionCookie(t, st))
if rr.Code != http.StatusNotFound {
t.Fatalf("status = %d, want 404", rr.Code)
}
})
}
}
// ServeMux path-cleans a traversal into a redirect before the handler runs, so
// the guarantee to pin down is that no request shaped like one ever gets bytes.
func TestKaganeCoverTraversalServesNothing(t *testing.T) {
cf := &fakeCovers{body: []byte("secret"), contentType: "image/webp"}
cfg := testConfig()
cfg.Covers = cf
srv, st := newWebTestServer(t, cfg)
rr := getCover(t, srv, "/img/kagane/../../etc/passwd", sessionCookie(t, st))
if rr.Code == http.StatusOK {
t.Fatalf("status = 200, want anything but a served body")
}
if got := cf.calls.Load(); got != 0 {
t.Fatalf("fetcher called %d times for a traversal, want 0", got)
}
}
// Without BROWSER_WS_URL there is no fetcher, and the endpoint must answer
// rather than reach for a nil one.
func TestKaganeCoverWithoutFetcher(t *testing.T) {
srv, st := newWebTestServer(t, testConfig())
rr := getCover(t, srv, "/img/kagane/"+testCoverID, sessionCookie(t, st))
if rr.Code != http.StatusNotFound {
t.Fatalf("status = %d, want 404", rr.Code)
}
}
+131 -30
View File
@@ -2,7 +2,9 @@ package latest
import ( import (
"context" "context"
"encoding/base64"
"encoding/json" "encoding/json"
"errors"
"fmt" "fmt"
"net/url" "net/url"
"regexp" "regexp"
@@ -22,6 +24,12 @@ const challengeTimeout = 45 * time.Second
var kaganeSeriesRe = regexp.MustCompile(`^/series/([0-9a-f-]{36})/?$`) var kaganeSeriesRe = regexp.MustCompile(`^/series/([0-9a-f-]{36})/?$`)
// kaganeImageIDRe pins the only path segment Image interpolates into an
// outbound URL. The id arrives from a stored cover URL, which a client
// supplied, so it is matched rather than trusted: a headless browser is a
// strong SSRF primitive.
var kaganeImageIDRe = regexp.MustCompile(`^[0-9a-f-]{36}$`)
// BrowserFetcher retrieves pages through a remote headless Chrome over the // BrowserFetcher retrieves pages through a remote headless Chrome over the
// DevTools Protocol. // DevTools Protocol.
// //
@@ -91,23 +99,6 @@ func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int
return "", 0, fmt.Errorf("not a fetchable browser series url: %q", seriesURL) return "", 0, fmt.Errorf("not a fetchable browser series url: %q", seriesURL)
} }
f.mu.Lock()
defer f.mu.Unlock()
ctx, cancel := context.WithTimeout(ctx, challengeTimeout)
defer cancel()
// A fresh tab per fetch, closed on return, so one wedged page cannot
// poison later polls.
tabCtx, cancelTab := chromedp.NewContext(f.allocCtx)
defer cancelTab()
// Bind the caller's deadline to the tab.
tabCtx, cancelDeadline := context.WithCancel(tabCtx)
defer cancelDeadline()
go func() {
<-ctx.Done()
cancelDeadline()
}()
var body string var body string
// kagane's chapter list is only in its JSON API, which must be called from // kagane's chapter list is only in its JSON API, which must be called from
// inside the page so the request carries the clearance cookie. novelfull // inside the page so the request carries the clearance cookie. novelfull
@@ -123,24 +114,134 @@ func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int
) )
} }
err := chromedp.Run(tabCtx, // novelfull's payload is the DOM itself, and the interstitial has a DOM
chromedp.Navigate(seriesURL), // too, so "we have an answer" has to exclude it explicitly. kagane's
// The challenge reloads the page itself when it passes; waiting for the // in-page fetch just fails while challenged, which is already the signal.
// site's own root element is what tells us we are through it. done := func() bool { return body != "" && (isKagane || !isInterstitial(body)) }
chromedp.WaitReady("body", chromedp.ByQuery), if err := f.run(ctx, seriesURL, read, done); err != nil {
read, // Challenge never cleared, or the API refused. Indistinguishable from
) // here and handled identically by the caller.
if err != nil { if errors.Is(err, errChallengeHeld) {
return "", 403, nil
}
return "", 0, fmt.Errorf("browser fetch %q: %w", seriesURL, err) return "", 0, fmt.Errorf("browser fetch %q: %w", seriesURL, err)
} }
if body == "" {
// Challenge still up, or the API refused. Indistinguishable from here
// and handled identically by the caller.
return "", 403, nil
}
return body, 200, nil return body, 200, nil
} }
// Image retrieves one kagane cover as raw bytes and its content type.
//
// It exists because kagane serves covers behind the same challenge as its
// pages *and* with `cross-origin-resource-policy: same-origin`, so an <img> on
// the web UI's origin cannot load one even from a browser that already holds
// the clearance cookie (verified 2026-08-08). Proxying is the only route.
//
// The image URL is navigated to rather than fetched from some other kagane
// page: the challenge only runs on a top-level navigation, and once it clears
// the document *is* the image, so a same-origin fetch of location.href reads
// it straight back out of the cache.
//
// The challenge is not solved by the first read: WaitReady("body") is satisfied
// by the interstitial too. run holds the tab open until the in-page fetch
// succeeds, which is what gives the challenge script the seconds it needs.
func (f *BrowserFetcher) Image(ctx context.Context, imageID string) ([]byte, string, error) {
if !kaganeImageIDRe.MatchString(imageID) {
return nil, "", fmt.Errorf("not a kagane image id: %q", imageID)
}
var dataURL string
err := f.run(ctx, "https://kagane.to/api/v2/image/"+imageID+"/compressed",
chromedp.Evaluate(`fetch(location.href).then(r => r.ok
? r.blob().then(b => new Promise(res => {
const fr = new FileReader();
fr.onload = () => res(fr.result);
fr.readAsDataURL(b);
}))
: "")`, &dataURL, awaitPromise),
func() bool { return dataURL != "" })
if err != nil {
return nil, "", fmt.Errorf("browser image %s: %w", imageID, err)
}
// "data:image/webp;base64,<payload>".
head, payload, ok := strings.Cut(dataURL, ";base64,")
if !ok {
return nil, "", fmt.Errorf("browser image %s: not a data url", imageID)
}
raw, err := base64.StdEncoding.DecodeString(payload)
if err != nil {
return nil, "", fmt.Errorf("browser image %s: %w", imageID, err)
}
return raw, strings.TrimPrefix(head, "data:"), nil
}
// errChallengeHeld reports that the budget ran out with the interstitial still
// up. Distinct from a transport failure: it means "this site said no", which
// the poller answers with a 403 and its ordinary cooldown.
var errChallengeHeld = errors.New("challenge held")
// challengePollInterval paces re-reads while a challenge solves itself.
const challengePollInterval = 2 * time.Second
// isInterstitial reports whether html is Cloudflare's challenge page rather
// than the site's own. Matched on the challenge runtime's script path, which is
// stable across the interstitial's wording and locale — the visible "Just a
// moment..." title is neither.
func isInterstitial(html string) bool {
return strings.Contains(html, "/cdn-cgi/challenge-platform/")
}
// run navigates to target and re-reads until done reports an answer, bounded by
// challengeTimeout and by the caller's own deadline, in a tab that is closed on
// return so one wedged page cannot poison later calls.
//
// Holding the tab open across re-reads is the whole point. A Cloudflare
// interstitial needs several seconds of a live page to solve itself and write
// clearance into the browser's shared cookie jar; reading once and closing the
// tab — which is what this did before 2026-08-08 — never gives it that window,
// so every fetch lands on the interstitial and the clearance that would have
// unblocked all the later ones is never obtained.
func (f *BrowserFetcher) run(ctx context.Context, target string, read chromedp.Action, done func() bool) error {
f.mu.Lock()
defer f.mu.Unlock()
ctx, cancel := context.WithTimeout(ctx, challengeTimeout)
defer cancel()
tabCtx, cancelTab := chromedp.NewContext(f.allocCtx)
defer cancelTab()
// Bind the caller's deadline to the tab.
tabCtx, cancelDeadline := context.WithCancel(tabCtx)
defer cancelDeadline()
go func() {
<-ctx.Done()
cancelDeadline()
}()
if err := chromedp.Run(tabCtx,
chromedp.Navigate(target),
chromedp.WaitReady("body", chromedp.ByQuery),
); err != nil {
return err
}
var lastErr error
for {
// The challenge reloads the page when it passes, which tears down the
// execution context mid-read. That is a retry, not a failure.
if err := chromedp.Run(tabCtx, read); err != nil {
lastErr = err
} else if done() {
return nil
}
select {
case <-ctx.Done():
if lastErr != nil {
return fmt.Errorf("%w (last read: %v)", errChallengeHeld, lastErr)
}
return errChallengeHeld
case <-time.After(challengePollInterval):
}
}
}
// kaganeAPIURL maps a stored series_url to the JSON endpoint carrying its // kaganeAPIURL maps a stored series_url to the JSON endpoint carrying its
// chapter list. Returning false for anything else is a second line of defence // chapter list. Returning false for anything else is a second line of defence
// behind fetchableSeriesURL: a headless browser is a strong SSRF primitive and // behind fetchableSeriesURL: a headless browser is a strong SSRF primitive and
@@ -0,0 +1,91 @@
package latest
import (
"context"
"net/http"
"os"
"testing"
"time"
)
// TestSmokeKaganeImage is the live proof that the cover proxy's fetch actually
// clears Cloudflare and returns image bytes. It needs a real headless Chrome
// with outbound network, so it runs only when SMOKE_BROWSER_WS_URL is set:
//
// docker run --rm --shm-size=1gb -p 19222:9222 chromedp/headless-shell:stable
// SMOKE_BROWSER_WS_URL=ws://127.0.0.1:19222 go test -run TestSmokeKaganeImage ./internal/latest
func TestSmokeKaganeImage(t *testing.T) {
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
if ws == "" {
t.Skip("SMOKE_BROWSER_WS_URL unset")
}
const imageID = "019fe11a-84c3-7fc3-a84b-88787374b617" // SP Baby's cover
// The same URL through a plain client is what the web UI's <img> gets.
// Asserting on it keeps the test honest about why the browser is needed.
req, err := http.NewRequest(http.MethodGet,
"https://kagane.to/api/v2/image/"+imageID+"/compressed", nil)
if err != nil {
t.Fatal(err)
}
if res, err := (&http.Client{Timeout: 15 * time.Second}).Do(req); err == nil {
res.Body.Close()
if res.StatusCode == http.StatusOK {
t.Log("note: kagane answered a plain request 200 — the challenge is not up right now")
}
}
f, err := NewBrowserFetcher(ws)
if err != nil {
t.Fatalf("NewBrowserFetcher: %v", err)
}
defer f.Close()
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
defer cancel()
body, contentType, err := f.Image(ctx, imageID)
if err != nil {
t.Fatalf("Image: %v", err)
}
if len(body) < 1000 {
t.Fatalf("body is %d bytes, want a real image", len(body))
}
if contentType != "image/webp" {
t.Fatalf("content type = %q, want image/webp", contentType)
}
// WebP files start with "RIFF....WEBP".
if string(body[:4]) != "RIFF" || string(body[8:12]) != "WEBP" {
t.Fatalf("body is not a WebP: % x", body[:12])
}
t.Logf("fetched %d bytes of %s", len(body), contentType)
if _, _, err := f.Image(ctx, "not-a-uuid"); err == nil {
t.Fatal("Image accepted a non-uuid id")
}
}
// Control for the test above: the poller's own kagane path, same sidecar. If
// this fails too, the sidecar is not clearing the challenge at all and the
// image result says nothing about Image itself.
func TestSmokeKaganeGet(t *testing.T) {
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
if ws == "" {
t.Skip("SMOKE_BROWSER_WS_URL unset")
}
f, err := NewBrowserFetcher(ws)
if err != nil {
t.Fatalf("NewBrowserFetcher: %v", err)
}
defer f.Close()
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
defer cancel()
body, status, err := f.Get(ctx, "https://kagane.to/series/019fe11a-8670-7cf3-8343-0b02057d3787")
if err != nil {
t.Fatalf("Get: %v", err)
}
t.Logf("status=%d bytes=%d head=%.80q", status, len(body), body)
if status != 200 {
t.Fatalf("status = %d, want 200 — the sidecar is not clearing the challenge", status)
}
}
+16
View File
@@ -137,6 +137,22 @@ func (b Bookmark) Initial() string {
return "?" return "?"
} }
// kaganeCoverRe matches the cover URL kagane's og:image carries, which is what
// the userscript stores for that site.
var kaganeCoverRe = regexp.MustCompile(`^https://kagane\.to/api/v2/image/([0-9a-f-]{36})/compressed$`)
// CoverURL is the src the web UI puts in an <img>. For every site but kagane
// that is Cover as stored. kagane serves its images behind a Cloudflare
// challenge *and* with `cross-origin-resource-policy: same-origin`, so no page
// on another origin can load one however it asks (verified 2026-08-08); those
// go through the backend's own proxy instead.
func (b Bookmark) CoverURL() string {
if m := kaganeCoverRe.FindStringSubmatch(b.Cover); m != nil {
return "/img/kagane/" + m[1]
}
return b.Cover
}
// Library buckets. A bookmark is in exactly one. This cannot be derived from // Library buckets. A bookmark is in exactly one. This cannot be derived from
// Site: asurascans serves manga and novels from the same /comics/ path, so the // Site: asurascans serves manga and novels from the same /comics/ path, so the
// userscript that recorded the page is the only party that knows which. // userscript that recorded the page is the only party that knows which.
+32
View File
@@ -543,6 +543,38 @@ func TestDisplayChapter(t *testing.T) {
} }
} }
func TestCoverURL(t *testing.T) {
cases := []struct {
name string
cover string
want string
}{
{
"kagane routes through the proxy",
"https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
"/img/kagane/019fe11a-84c3-7fc3-a84b-88787374b617",
},
{
"another site is served as stored",
"https://gg.asuracomic.net/storage/media/1/conversions/cover.webp",
"https://gg.asuracomic.net/storage/media/1/conversions/cover.webp",
},
{
"a lookalike host is not rewritten",
"https://evil.example/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
"https://evil.example/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
},
{"no cover stays empty", "", ""},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
if got := (Bookmark{Cover: tc.cover}).CoverURL(); got != tc.want {
t.Errorf("CoverURL() = %q, want %q", got, tc.want)
}
})
}
}
func TestUpsertKindDefaultsToManga(t *testing.T) { func TestUpsertKindDefaultsToManga(t *testing.T) {
store := newTestStore(t) store := newTestStore(t)
got, err := store.Upsert(store.OwnerID(), Bookmark{ got, err := store.Upsert(store.OwnerID(), Bookmark{
+129
View File
@@ -0,0 +1,129 @@
package web
import (
"context"
"log"
"net/http"
"regexp"
"sync"
"time"
)
// CoverFetcher retrieves one kagane cover by image id. Satisfied by
// latest.BrowserFetcher, and nil when BROWSER_WS_URL is unset — which leaves
// kagane covers exactly as unavailable as they were before this endpoint
// existed, rather than hanging a request on a fetcher that cannot run.
type CoverFetcher interface {
Image(ctx context.Context, imageID string) (body []byte, contentType string, err error)
}
// coverIDRe matches the request path segment that becomes part of an outbound
// URL. The proxy is session-gated, but the id still reaches a headless browser,
// so it is validated at the boundary rather than passed through.
var coverIDRe = regexp.MustCompile(`^[0-9a-f-]{36}$`)
// coverTypes is the set of content types the proxy will echo back. A response
// header sourced from a third party is not repeated verbatim: anything outside
// this set is treated as "not a cover".
var coverTypes = map[string]bool{
"image/webp": true,
"image/jpeg": true,
"image/png": true,
"image/avif": true,
"image/gif": true,
}
// coverTimeout bounds one proxied cover. Shorter than the fetcher's own
// challenge budget on purpose: a browser page is waiting on this, and a cover
// that has not arrived by now is better left as a broken slot than as a request
// holding a connection open.
const coverTimeout = 20 * time.Second
// coverCacheMax caps the in-memory cover cache. Covers are immutable per image
// id and a library holds tens of series, so this is a ceiling that is never
// reached in practice; reaching it clears the map rather than evicting by age.
//
// ponytail: flush-on-full, not LRU. Swap it for an LRU if a library ever grows
// past this and the flush starts costing refetches.
const coverCacheMax = 500
type cachedCover struct {
body []byte
contentType string
}
type coverCache struct {
mu sync.Mutex
m map[string]cachedCover
}
func (c *coverCache) get(id string) (cachedCover, bool) {
c.mu.Lock()
defer c.mu.Unlock()
v, ok := c.m[id]
return v, ok
}
func (c *coverCache) put(id string, v cachedCover) {
c.mu.Lock()
defer c.mu.Unlock()
if c.m == nil || len(c.m) >= coverCacheMax {
c.m = make(map[string]cachedCover, coverCacheMax)
}
c.m[id] = v
}
// kaganeCover serves a kagane cover from the backend's own origin.
//
// kagane answers image requests with a Cloudflare challenge and
// `cross-origin-resource-policy: same-origin`, so the web UI cannot render one
// directly under any combination of referrer policy or crossorigin attribute
// (verified 2026-08-08). Fetching it through the headless browser that already
// clears the challenge, and re-serving it here, is what puts the bytes on an
// origin the page may load from.
//
// ponytail: covers are fetched on first view, one browser navigation at a time
// behind the fetcher's mutex, so a first load of a large kagane library
// trickles in over a few seconds. The cache makes it a one-off. Prefetching
// during the poll cycle is the upgrade if that ever grates.
func (h *Handler) kaganeCover(w http.ResponseWriter, r *http.Request) {
id := r.PathValue("id")
if !coverIDRe.MatchString(id) {
http.NotFound(w, r)
return
}
if h.covers == nil {
http.NotFound(w, r)
return
}
if v, ok := h.coverCache.get(id); ok {
writeCover(w, v)
return
}
ctx, cancel := context.WithTimeout(r.Context(), coverTimeout)
defer cancel()
body, contentType, err := h.covers.Image(ctx, id)
if err != nil {
log.Printf("kagane cover %s: %v", id, err)
http.NotFound(w, r)
return
}
if !coverTypes[contentType] {
log.Printf("kagane cover %s: unexpected content type %q", id, contentType)
http.NotFound(w, r)
return
}
v := cachedCover{body: body, contentType: contentType}
h.coverCache.put(id, v)
writeCover(w, v)
}
// writeCover sends the bytes with a long cache life: an image id names one
// immutable rendering, so a client that has it never needs to ask again.
func writeCover(w http.ResponseWriter, v cachedCover) {
w.Header().Set("Content-Type", v.contentType)
w.Header().Set("Cache-Control", "private, max-age=604800, immutable")
w.Write(v.body)
}
+1 -1
View File
@@ -6,7 +6,7 @@
<div class="row"> <div class="row">
<a class="cover" href="{{.ContinueURL}}" target="_blank" rel="noopener noreferrer" <a class="cover" href="{{.ContinueURL}}" target="_blank" rel="noopener noreferrer"
tabindex="-1" aria-hidden="true"> tabindex="-1" aria-hidden="true">
{{if .Cover}}<img src="{{.Cover}}" alt="" loading="lazy"> {{if .CoverURL}}<img src="{{.CoverURL}}" alt="" loading="lazy">
{{/* aria-hidden on the cover link is not enough — Chromium still exposes {{/* aria-hidden on the cover link is not enough — Chromium still exposes
the letter because the link is programmatically focusable — so the the letter because the link is programmatically focusable — so the
monogram carries its own, same as the recent strip's. */}} monogram carries its own, same as the recent strip's. */}}
+1 -1
View File
@@ -14,7 +14,7 @@
<a class="recent-card {{if .HasNewChapter}}is-new{{end}}" href="{{.ContinueURL}}" <a class="recent-card {{if .HasNewChapter}}is-new{{end}}" href="{{.ContinueURL}}"
target="_blank" rel="noopener noreferrer"> target="_blank" rel="noopener noreferrer">
<span class="recent-cover"> <span class="recent-cover">
{{if .Cover}}<img src="{{.Cover}}" alt="" loading="lazy"> {{if .CoverURL}}<img src="{{.CoverURL}}" alt="" loading="lazy">
{{else}}<span class="monogram" aria-hidden="true">{{.Initial}}</span>{{end}} {{else}}<span class="monogram" aria-hidden="true">{{.Initial}}</span>{{end}}
{{if .HasNewChapter}}<span class="foot-rule"></span> {{if .HasNewChapter}}<span class="foot-rule"></span>
{{else if .Favorite}}<span class="foot-rule brass"></span>{{end}} {{else if .Favorite}}<span class="foot-rule brass"></span>{{end}}
+10 -1
View File
@@ -50,6 +50,10 @@ type Handler struct {
// httpClient is the plain stdlib client that talks to Discord. It is not // httpClient is the plain stdlib client that talks to Discord. It is not
// an injected interface: tests point APIBase at a stub server instead. // an injected interface: tests point APIBase at a stub server instead.
httpClient *http.Client httpClient *http.Client
// covers proxies kagane cover images, which no browser can load directly.
// Nil disables the endpoint — see CoverFetcher.
covers CoverFetcher
coverCache coverCache
} }
// listView is what every list-rendering template receives. // listView is what every list-rendering template receives.
@@ -111,7 +115,7 @@ type loginView struct {
// New parses every template up front so a broken one kills the process at // New parses every template up front so a broken one kills the process at
// startup rather than the first request that touches it. // startup rather than the first request that touches it.
func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, novelPath string) (*Handler, error) { func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, novelPath string, covers CoverFetcher) (*Handler, error) {
tmpl, err := template.ParseFS(templateFS, "templates/*.html") tmpl, err := template.ParseFS(templateFS, "templates/*.html")
if err != nil { if err != nil {
return nil, err return nil, err
@@ -126,6 +130,7 @@ func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, nove
states: newOAuthStates(), states: newOAuthStates(),
limiter: session.NewLoginLimiter(), limiter: session.NewLoginLimiter(),
httpClient: &http.Client{Timeout: discordTimeout}, httpClient: &http.Client{Timeout: discordTimeout},
covers: covers,
}, nil }, nil
} }
@@ -142,6 +147,10 @@ func (h *Handler) Register(mux *http.ServeMux) {
mux.HandleFunc("POST /ui/bookmarks/{key}/chapter", h.requireSession(h.uiChapter)) mux.HandleFunc("POST /ui/bookmarks/{key}/chapter", h.requireSession(h.uiChapter))
mux.HandleFunc("DELETE /ui/bookmarks/{key}", h.requireSession(h.uiDelete)) mux.HandleFunc("DELETE /ui/bookmarks/{key}", h.requireSession(h.uiDelete))
// Session-gated like every other UI route: the deployment proxies kagane's
// images for its own Readers, not for the internet.
mux.HandleFunc("GET /img/kagane/{id}", h.requireSession(h.kaganeCover))
// Install endpoints render the script directly under the session: the // Install endpoints render the script directly under the session: the
// credential travels inside the served bytes, never in the address bar or // credential travels inside the served bytes, never in the address bar or
// the page markup. Updates after install use the credential-bearing /u/ // the page markup. Updates after install use the credential-bearing /u/
+35 -24
View File
@@ -47,6 +47,10 @@ type Config struct {
NovelUserscriptPath string NovelUserscriptPath string
// LatestPoll configures the background latest-chapter fetcher. // LatestPoll configures the background latest-chapter fetcher.
LatestPoll LatestPoll LatestPoll LatestPoll
// Covers proxies kagane cover images for the web UI. Not from the
// environment: it is the shared headless browser, wired in main once it
// connects, and nil in every test router.
Covers web.CoverFetcher
} }
// LatestPoll configures the background latest-chapter poller. // LatestPoll configures the background latest-chapter poller.
@@ -206,7 +210,7 @@ func newRouter(s *store.Store, cfg Config) http.Handler {
// The browser UI is always registered; signing in is Discord OAuth, so // The browser UI is always registered; signing in is Discord OAuth, so
// there is no password to forget and no gate to leave unset. // there is no password to forget and no gate to leave unset.
wh, err := web.New(s, cfg.Discord, []byte(cfg.TokenKey), wh, err := web.New(s, cfg.Discord, []byte(cfg.TokenKey),
cfg.UserscriptPath, cfg.NovelUserscriptPath) cfg.UserscriptPath, cfg.NovelUserscriptPath, cfg.Covers)
if err != nil { if err != nil {
log.Fatalf("web handler: %v", err) log.Fatalf("web handler: %v", err)
} }
@@ -270,9 +274,26 @@ func main() {
// The poller is off the request path entirely: if it cannot start, the // The poller is off the request path entirely: if it cannot start, the
// service still serves bookmarks and the userscript still captures latest // service still serves bookmarks and the userscript still captures latest
// chapters on its own. // chapters on its own.
//
// One headless browser serves both consumers that need a Cloudflare
// challenge cleared: the poller's kagane/novelfull fetches and the web
// UI's kagane cover proxy. Optional — unset leaves both degraded to what
// they were before the sidecar existed.
var browser latest.Fetcher
pollCtx, stopPoll := context.WithCancel(context.Background()) pollCtx, stopPoll := context.WithCancel(context.Background())
defer stopPoll() defer stopPoll()
startLatestPoller(pollCtx, s, cfg.LatestPoll) if ws := strings.TrimSpace(os.Getenv("BROWSER_WS_URL")); ws != "" {
bf, err := latest.NewBrowserFetcher(ws)
if err != nil {
log.Printf("browser fetcher disabled: %v", err)
} else {
browser = bf
cfg.Covers = bf
context.AfterFunc(pollCtx, bf.Close)
log.Printf("browser fetcher at %s", ws)
}
}
startLatestPoller(pollCtx, s, cfg.LatestPoll, browser)
srv := &http.Server{ srv := &http.Server{
Addr: ":" + cfg.Port, Addr: ":" + cfg.Port,
@@ -307,7 +328,7 @@ func main() {
// HTTP client cannot be built. Any problem here is logged and skipped: this // HTTP client cannot be built. Any problem here is logged and skipped: this
// feature going missing degrades the service to userscript-only latest-chapter // feature going missing degrades the service to userscript-only latest-chapter
// tracking, which is exactly how it behaved before. // tracking, which is exactly how it behaved before.
func startLatestPoller(ctx context.Context, s *store.Store, cfg LatestPoll) { func startLatestPoller(ctx context.Context, s *store.Store, cfg LatestPoll, browser latest.Fetcher) {
if !cfg.Enabled { if !cfg.Enabled {
log.Println("latest-chapter poller: disabled by config") log.Println("latest-chapter poller: disabled by config")
return return
@@ -317,28 +338,18 @@ func startLatestPoller(ctx context.Context, s *store.Store, cfg LatestPoll) {
log.Printf("latest-chapter poller: disabled, cannot build client: %v", err) log.Printf("latest-chapter poller: disabled, cannot build client: %v", err)
return return
} }
// Nil browser: sites behind a JavaScript challenge are simply not polled,
// and their latest_chapter comes from the userscript alone — which is how
// the service behaved before the sidecar existed.
p := &latest.Poller{ p := &latest.Poller{
Store: s, Store: s,
Fetch: f, Fetch: f,
Now: time.Now, BrowserFetch: browser,
Cooldown: cfg.Cooldown, Now: time.Now,
Interval: cfg.Interval, Cooldown: cfg.Cooldown,
Stagger: cfg.Stagger, Interval: cfg.Interval,
Batch: cfg.Batch, Stagger: cfg.Stagger,
} Batch: cfg.Batch,
// Optional: without it, sites behind a JavaScript challenge are simply not
// polled, and their latest_chapter comes from the userscript alone — which
// is how the service behaved before the sidecar existed.
if ws := strings.TrimSpace(os.Getenv("BROWSER_WS_URL")); ws != "" {
bf, err := latest.NewBrowserFetcher(ws)
if err != nil {
log.Printf("latest-chapter poller: browser fetcher disabled: %v", err)
} else {
p.BrowserFetch = bf
context.AfterFunc(ctx, bf.Close)
log.Printf("latest-chapter poller: browser fetcher at %s", ws)
}
} }
go p.Run(ctx) go p.Run(ctx)
+39
View File
@@ -0,0 +1,39 @@
# syntax=docker/dockerfile:1
# Real Google Chrome for the latest-chapter poller and the kagane cover proxy.
#
# Not chromedp/headless-shell, which this replaces. headless-shell is a stripped
# Chrome build and Cloudflare's managed challenge on kagane.to never clears for
# it: measured 2026-08-08, 60s of a held-open tab still served the interstitial,
# while stock Chrome from the same IP cleared in ~4s. The tells are structural
# rather than a header — navigator.webdriver true, an empty plugin list, and
# Chromium- rather than Chrome-branded client hints. Overriding webdriver alone
# was tried and did not move it, so the browser build itself is the fix.
#
# zenika/alpine-chrome was also tried: its Chrome is 124 (2024), old enough that
# Cloudflare refuses it outright and old enough to break chromedp's CDP structs.
FROM debian:trixie-slim
# Chrome is deliberately unpinned, against the usual rule. A pinned build goes
# stale, and a stale browser is exactly what Cloudflare turns away — the 124 in
# alpine-chrome is the worked example. Rebuild is the upgrade path.
RUN apt-get update \
&& apt-get install -y --no-install-recommends ca-certificates wget gnupg \
&& wget -qO- https://dl.google.com/linux/linux_signing_key.pub \
| gpg --dearmor -o /usr/share/keyrings/google-chrome.gpg \
&& echo "deb [arch=amd64 signed-by=/usr/share/keyrings/google-chrome.gpg] https://dl.google.com/linux/chrome/deb/ stable main" \
> /etc/apt/sources.list.d/google-chrome.list \
&& apt-get update \
&& apt-get install -y --no-install-recommends google-chrome-stable socat \
&& rm -rf /var/lib/apt/lists/*
# Unprivileged: Chrome refuses to run as root, and the CDP endpoint is a shell
# on whatever user owns it.
RUN useradd --create-home --shell /usr/sbin/nologin chrome
USER chrome
WORKDIR /home/chrome
COPY entrypoint.sh /entrypoint.sh
EXPOSE 9222
ENTRYPOINT ["/entrypoint.sh"]
+61
View File
@@ -0,0 +1,61 @@
#!/bin/sh
set -eu
# A UTC clock is itself the bot signal — Cloudflare treats it as the datacenter
# default — and kagane's challenge then never clears. Measured 2026-08-08 with
# an identical container on one Indonesian egress IP: UTC never cleared in 60s
# (twice), while Asia/Jakarta and America/New_York both cleared in 4s. Any real
# zone will do; the zone does not have to match the IP's country, it just must
# not be UTC. It does have to be right the way Chrome reads it.
#
# TZ must carry the zone *name*. Chrome resolves the zone through ICU, which
# takes the name from /etc/localtime's symlink target and ignores the file's
# contents; bind-mounting the host's /etc/localtime therefore lands on the
# image's own symlink to Etc/UTC and leaves glibc reporting +07 while Chrome
# still reports UTC. /etc/timezone, mounted by docker-compose.yml, is the name.
[ -n "${TZ:-}" ] || TZ=$(cat /etc/timezone 2>/dev/null || echo UTC)
export TZ
# Chrome's own UA advertises "HeadlessChrome" under --headless=new, and that one
# token is the difference between kagane.to's challenge clearing in ~4s and
# never clearing at all (measured 2026-08-08, same host, same Chrome, only the
# UA changed). Overriding it does not touch the Sec-CH-UA client hints, which
# report the real version, so the version is read back out of the binary rather
# than hardcoded: a hardcoded one would drift out of step with the hints on the
# next Chrome update and become a fresh tell.
major=$(google-chrome-stable --version | sed -E 's/[^0-9]*([0-9]+)\..*/\1/')
ua="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/${major}.0.0.0 Safari/537.36"
# Chrome binds its DevTools port to loopback and silently ignores
# --remote-debugging-address (verified 2026-08-08: Chrome 151 with
# --remote-debugging-address=0.0.0.0 still listened on 127.0.0.1 only), so the
# caller — another container — cannot reach it directly. socat fronting the
# loopback port is how chromedp/headless-shell solved the same problem and is
# why this image is a drop-in for it.
#
# Nothing publishes 9222; reachability is the `browser` network in
# docker-compose.yml, and an exposed CDP endpoint is remote code execution.
socat TCP-LISTEN:9222,fork,reuseaddr TCP:127.0.0.1:9223 &
# Chrome stays in the foreground so that its death takes the container down and
# compose's restart policy applies; a backgrounded browser behind a live socat
# would leave the sidecar looking healthy while answering nothing.
#
# No --enable-automation: it sets navigator.webdriver, the first thing a bot
# check reads.
#
# --no-sandbox because Chrome's zygote wants user namespaces, which Docker's
# default profile does not hand out; the alternative is --cap-add=SYS_ADMIN,
# which gives the container strictly more than it takes away. Containment here
# is the unprivileged user, the isolated network, and the fact that this
# browser only ever navigates to kagane.to and novelfull.com.
exec google-chrome-stable \
--headless=new \
--no-sandbox \
--remote-debugging-port=9223 \
--user-agent="$ua" \
--user-data-dir=/home/chrome/profile \
--no-first-run \
--no-default-browser-check \
--disable-gpu \
about:blank
+27 -12
View File
@@ -24,6 +24,13 @@ services:
# comes from .env so it is never committed. # comes from .env so it is never committed.
DATABASE_URL: ${DATABASE_URL:-postgres://bookmarks:${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD in .env}@postgres:5432/bookmarks?sslmode=disable} DATABASE_URL: ${DATABASE_URL:-postgres://bookmarks:${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD in .env}@postgres:5432/bookmarks?sslmode=disable}
PORT: "8080" PORT: "8080"
# Log timestamps only. Go's `log` stamps lines in local time, and this
# service has no other use for a zone: bookmark timestamps are unix ms
# and the two real time columns are timestamptz, both absolute instants.
# Purely so these lines read on the same clock as the sidecar's. Named
# API_TZ rather than TZ so an operator's exported shell TZ cannot leak
# in; distroless already carries tzdata, so the name just resolves.
TZ: ${API_TZ:-Asia/Jakarta}
# Discord OAuth for the browser UI (ADR-0002). The first four are # Discord OAuth for the browser UI (ADR-0002). The first four are
# required; DISCORD_REQUIRED_ROLE is optional and empty by default. # required; DISCORD_REQUIRED_ROLE is optional and empty by default.
# Guild membership is the whole gate: any member becomes a Reader. # Guild membership is the whole gate: any member becomes a Reader.
@@ -95,8 +102,25 @@ services:
- db - db
headless-shell: headless-shell:
image: chromedp/headless-shell:stable # Real Google Chrome, not chromedp/headless-shell — see chrome/Dockerfile.
# The service name is kept so existing overrides and BROWSER_WS_URL stay put.
build: ./chrome
image: bookmarkmanager-chrome:latest
restart: unless-stopped restart: unless-stopped
environment:
# A UTC clock is itself the bot signal: Cloudflare treats it as the
# datacenter default, and kagane's challenge then never clears. Measured
# 2026-08-08, identical container, one Indonesian egress IP: UTC never
# cleared in 60s (twice); Asia/Jakarta and America/New_York both cleared
# in 4s. So any real zone works and it need not match the IP's country —
# only UTC fails. Unset falls back to the host's /etc/timezone below,
# which is a real zone whenever the host clock is set to local time; set
# BROWSER_TZ when the host runs UTC.
TZ: ${BROWSER_TZ:-}
volumes:
# The zone *name*, which is what Chrome's ICU needs — see chrome/entrypoint.sh.
# Absent on a non-Debian host, which the entrypoint handles by falling back to UTC.
- /etc/timezone:/etc/timezone:ro
# Chrome allocates shared memory per tab and dies on Docker's 64MB default. # Chrome allocates shared memory per tab and dies on Docker's 64MB default.
shm_size: '1gb' shm_size: '1gb'
# Reaps zombie renderer processes, which otherwise accumulate for the # Reaps zombie renderer processes, which otherwise accumulate for the
@@ -104,17 +128,8 @@ services:
init: true init: true
# Deliberately no `ports:` — an exposed CDP endpoint is remote code # Deliberately no `ports:` — an exposed CDP endpoint is remote code
# execution. Only bookmark-api, via the `browser` network below, may reach it. # execution. Only bookmark-api, via the `browser` network below, may reach it.
# Don't pass --remote-debugging-address/--remote-debugging-port here: the # No `command:` either: every flag this browser needs is in its entrypoint,
# image's own entrypoint (/headless-shell/run.sh) already starts Chrome on # and the UA override there is load-bearing for the challenge.
# 127.0.0.1:9223 and fronts it with a socat proxy listening on 0.0.0.0:9222.
# Redeclaring the port flag here overrides Chrome's, so it binds 9222
# directly (IPv6 loopback only) instead of 9223 — collides with socat's own
# bind on 9222 and leaves nothing listening on 9223, so every external
# connection to headless-shell:9222 fails with EOF. Only pass flags the
# entrypoint doesn't already set.
command:
- --disable-gpu
- --no-sandbox
networks: networks:
browser: browser:
# Pinned so BROWSER_WS_URL can name an IP (required, see above) that # Pinned so BROWSER_WS_URL can name an IP (required, see above) that
+19
View File
@@ -44,6 +44,25 @@ Guidance for OpenCode (and Claude Code) working under `userscript/`. See root `A
Encodings (incl. triple-encoded punctuation like `%25252D`) identical Encodings (incl. triple-encoded punctuation like `%25252D`) identical
on /manga/ and /title/ pages, so decode-once seriesIds match — verified on /manga/ and /title/ pages, so decode-once seriesIds match — verified
2026-07-28. 2026-07-28.
- **comix.to**: series `/title/<id>-<slug>`, chapter
`/title/<id>-<slug>/<uploadId>-chapter-<n>`. Only the leading `<id>` is
identity — the slug re-renders when a series is renamed (`comixSeriesId`).
An SPA that **never rewrites `og:title`**: the server-rendered head keeps
whatever document loaded first, so on a cold load `og:title` is the homepage's
"Comix — Read Comics online for free" and after an in-page hop it is the
*previous* series' name. `document.title` is the one thing client routing does
update, so titles come from there, with the chapter page's `" · Ch.<n>"` tail
stripped. Covers likewise: `og:image` is absent, so the cover is the `img`
whose `alt` matches the cleaned title — verified live 2026-08-08.
- **kagane.to**: series `/series/<uuid>`, reader
`/series/<uuid>/reader/<bookUuid>`. Reader URLs carry no chapter number, so
the number comes out of `og:title`. Two shapes exist: `"<Series> - Chapter
<n>[ - Episode <n>]"` and, for volume-numbered series, `"<Series> - Volume <v>
Chapter <n>"` with no episode name — both must yield a bare series title, or
the volume tail lands in the bookmark's title. Its covers are challenge- and
CORP-protected, so the web UI proxies them; the userscript still stores the
raw `og:image`. Behind a Cloudflare JS challenge, so the backend polls it
through the headless browser.
- **novelfull.com** (novel script): series `/<slug>.html`, chapter - **novelfull.com** (novel script): series `/<slug>.html`, chapter
`/<slug>/chapter-<n>[-<title-slug>].html`. No `og:*` tags at all — title from `/<slug>/chapter-<n>[-<title-slug>].html`. No `og:*` tags at all — title from
`h3.title` (series) or `a.truyen-title` (chapter), cover from `h3.title` (series) or `a.truyen-title` (chapter), cover from
+39 -13
View File
@@ -242,6 +242,12 @@
matches: (loc) => /(^|\.)comix\.to$/.test(loc.hostname), matches: (loc) => /(^|\.)comix\.to$/.test(loc.hostname),
detect(loc) { detect(loc) {
const path = loc.pathname; const path = loc.pathname;
// comix client-routes without ever rewriting og:title — the head keeps
// whatever the first server-rendered document carried, so a bookmark
// taken after a client route got the homepage's title, then the
// previous series'. document.title is the one thing its router does
// update. Verified live 2026-08-08; do not "restore" meta("og:title").
const pageTitle = cleanTitle(document.title);
// /title/<id>-<slug>/<uploadId>-chapter-<n>. Several uploads (different // /title/<id>-<slug>/<uploadId>-chapter-<n>. Several uploads (different
// groups or languages) share one chapter number; the number is the // groups or languages) share one chapter number; the number is the
// progress identity, the upload id is not. // progress identity, the upload id is not.
@@ -252,8 +258,8 @@
type: "chapter", type: "chapter",
site: this.site, site: this.site,
seriesId: comixSeriesId(m[1]), seriesId: comixSeriesId(m[1]),
title: cleanTitle(meta("og:title")), title: pageTitle,
cover: coverFromPage(), cover: coverFromPage(pageTitle),
seriesUrl: loc.origin + "/title/" + m[1], seriesUrl: loc.origin + "/title/" + m[1],
chapterLabel: "Chapter " + m[2], chapterLabel: "Chapter " + m[2],
chapterNum: isNaN(num) ? null : num, chapterNum: isNaN(num) ? null : num,
@@ -267,8 +273,8 @@
type: "series", type: "series",
site: this.site, site: this.site,
seriesId: comixSeriesId(m[1]), seriesId: comixSeriesId(m[1]),
title: cleanTitle(meta("og:title")), title: pageTitle,
cover: coverFromPage(), cover: coverFromPage(pageTitle),
seriesUrl: loc.origin + "/title/" + m[1], seriesUrl: loc.origin + "/title/" + m[1],
chapterLabel: null, chapterLabel: null,
chapterNum: null, chapterNum: null,
@@ -277,7 +283,7 @@
} }
return { type: "other" }; return { type: "other" };
// comix chapter og:title is "<Title> · Ch.<n>"; series is clean. // comix chapter document.title is "<Title> · Ch.<n>"; series is clean.
function cleanTitle(t) { function cleanTitle(t) {
if (!t) return ""; if (!t) return "";
return t.replace(/\s*·\s*Ch\.[\d.]+\s*$/i, "").trim(); return t.replace(/\s*·\s*Ch\.[\d.]+\s*$/i, "").trim();
@@ -287,8 +293,7 @@
// the DOM for a cover. Matching on alt rather than a class keeps it off // the DOM for a cover. Matching on alt rather than a class keeps it off
// the site's styling: the cover is the image whose alt is the title. // the site's styling: the cover is the image whose alt is the title.
// Do not "simplify" this into meta("og:image") — that returns null. // Do not "simplify" this into meta("og:image") — that returns null.
function coverFromPage() { function coverFromPage(title) {
const title = cleanTitle(meta("og:title"));
if (!title || !document.querySelectorAll) return ""; if (!title || !document.querySelectorAll) return "";
for (const img of document.querySelectorAll("img[alt]")) { for (const img of document.querySelectorAll("img[alt]")) {
if (img.getAttribute("alt") === title) return img.getAttribute("src") || ""; if (img.getAttribute("alt") === title) return img.getAttribute("src") || "";
@@ -313,6 +318,14 @@
}, },
}; };
// Kagane builds the reader og:title suffix out of the book's metadata, so
// every combination occurs: the volume part appears only when the book has a
// volume_no, the episode part only when it has a non-empty title. All four
// shapes captured live 2026-08-08 — "SP Baby - Volume 1 Chapter 1" is the one
// the old trailing-space regex missed, which left both the number and the
// series title wrong.
const KAGANE_CHAPTER_SUFFIX = /\s-\s(?:Volume\s[\d.]+\s)?Chapter\s([\d.]+)(?:\s-\s.*)?$/i;
const kagane = { const kagane = {
site: "kagane", site: "kagane",
matches: (loc) => /(^|\.)kagane\.to$/.test(loc.hostname), matches: (loc) => /(^|\.)kagane\.to$/.test(loc.hostname),
@@ -353,9 +366,8 @@
} }
return { type: "other" }; return { type: "other" };
// Reader og:title is "<Title> - Chapter <n> - <episode name>".
function chapterNumFromTitle(t) { function chapterNumFromTitle(t) {
const m = t && t.match(/\s-\sChapter\s([\d.]+)\s/); const m = t && t.match(KAGANE_CHAPTER_SUFFIX);
if (!m) return null; if (!m) return null;
const num = parseFloat(m[1]); const num = parseFloat(m[1]);
return isNaN(num) ? null : num; return isNaN(num) ? null : num;
@@ -363,7 +375,7 @@
function cleanTitle(t) { function cleanTitle(t) {
if (!t) return ""; if (!t) return "";
return t.replace(/\s-\sChapter\s[\d.]+\s-\s.*$/i, "").trim(); return t.replace(KAGANE_CHAPTER_SUFFIX, "").trim();
} }
}, },
// Reader hrefs are uuids with no number in them, so no maximum can be taken // Reader hrefs are uuids with no number in them, so no maximum can be taken
@@ -1450,8 +1462,15 @@
} }
let lastUrl = location.href; let lastUrl = location.href;
function onNavigate() { let lastPageSig = "";
function setPage() {
state.page = detect(); state.page = detect();
lastPageSig = JSON.stringify(state.page);
}
function onNavigate() {
setPage();
render(); render();
maybeAutoUpdate(); maybeAutoUpdate();
maybeCaptureLatestOnSeriesPage(); maybeCaptureLatestOnSeriesPage();
@@ -1484,7 +1503,14 @@
wrap("pushState"); wrap("pushState");
wrap("replaceState"); wrap("replaceState");
window.addEventListener("popstate", fire); window.addEventListener("popstate", fire);
setInterval(fire, 1500); // catch routes that bypass history // Also catches routes that bypass history — and comix, which fills
// document.title a beat after the route changes, so the 300ms snapshot
// above can still hold the previous page's title. Re-detect whenever what
// we would read has changed, not only when the URL has.
setInterval(() => {
fire();
if (JSON.stringify(detect()) !== lastPageSig) onNavigate();
}, 1500);
} }
// ============================================================ // ============================================================
@@ -1563,7 +1589,7 @@
function init() { function init() {
buildUI(); buildUI();
state.page = detect(); setPage();
render(); render();
installNavWatcher(); installNavWatcher();
installLongPress(); installLongPress();
+57 -8
View File
@@ -31,6 +31,9 @@ globalThis.location = {
// og: meta tags the adapters read through meta(). Reassigned per test. // og: meta tags the adapters read through meta(). Reassigned per test.
let metaTags = {}; let metaTags = {};
// document.title. comix's SPA rewrites this on client routing but never
// og:title, so the comix adapter reads it instead. Reassigned per test.
let docTitle = "";
// img[alt] elements comix's coverFromPage() scans. Reassigned per test; each // img[alt] elements comix's coverFromPage() scans. Reassigned per test; each
// entry is {alt, src}. // entry is {alt, src}.
let pageImages = []; let pageImages = [];
@@ -48,6 +51,9 @@ globalThis.document = {
})); }));
}, },
addEventListener() {}, addEventListener() {},
get title() {
return docTitle;
},
body: undefined, body: undefined,
}; };
@@ -193,8 +199,15 @@ test("comixSeriesId leaves a bare id untouched", () => {
assert.equal(comixSeriesId("n8we"), "n8we"); assert.equal(comixSeriesId("n8we"), "n8we");
}); });
// comix is an SPA that rewrites document.title on client routing but leaves the
// server-rendered og:title untouched, so every test here pins og:title to a
// STALE value — the homepage title on first hop, the previous series after
// that. Captured live 2026-08-08.
const COMIX_STALE_HOME = "Comix - Read Comics online for free";
test("comix detects a series page", () => { test("comix detects a series page", () => {
metaTags = { "og:title": "Dungeons and Crayons" }; metaTags = { "og:title": COMIX_STALE_HOME };
docTitle = "Dungeons and Crayons";
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons")); const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
assert.equal(p.type, "series"); assert.equal(p.type, "series");
assert.equal(p.site, "comix"); assert.equal(p.site, "comix");
@@ -204,8 +217,16 @@ test("comix detects a series page", () => {
assert.equal(p.chapterNum, null); assert.equal(p.chapterNum, null);
}); });
test("comix ignores a previous series' stale og:title", () => {
metaTags = { "og:title": "Full-Time Awakening" };
docTitle = "Dungeons and Crayons";
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
assert.equal(p.title, "Dungeons and Crayons");
});
test("comix detects a chapter page and strips the Ch. suffix from the title", () => { test("comix detects a chapter page and strips the Ch. suffix from the title", () => {
metaTags = { "og:title": "Dungeons and Crayons · Ch.80" }; metaTags = { "og:title": COMIX_STALE_HOME };
docTitle = "Dungeons and Crayons · Ch.80";
const p = comix.detect( const p = comix.detect(
loc("https://comix.to/title/n8we-dungeons-and-crayons/11139891-chapter-80") loc("https://comix.to/title/n8we-dungeons-and-crayons/11139891-chapter-80")
); );
@@ -218,7 +239,8 @@ test("comix detects a chapter page and strips the Ch. suffix from the title", ()
}); });
test("comix parses decimal chapter numbers", () => { test("comix parses decimal chapter numbers", () => {
metaTags = { "og:title": "Dungeons and Crayons · Ch.80.5" }; metaTags = {};
docTitle = "Dungeons and Crayons · Ch.80.5";
const p = comix.detect( const p = comix.detect(
loc("https://comix.to/title/n8we-dungeons-and-crayons/11139891-chapter-80.5") loc("https://comix.to/title/n8we-dungeons-and-crayons/11139891-chapter-80.5")
); );
@@ -226,19 +248,19 @@ test("comix parses decimal chapter numbers", () => {
}); });
test("comix.detect reads the cover from an img whose alt matches the cleaned title", () => { test("comix.detect reads the cover from an img whose alt matches the cleaned title", () => {
metaTags = { "og:title": "Dungeons and Crayons · Ch.80" }; metaTags = { "og:title": COMIX_STALE_HOME };
docTitle = "Dungeons and Crayons";
pageImages = [ pageImages = [
{ alt: "Some Other Series", src: "https://cdn.example/other.jpg" }, { alt: "Some Other Series", src: "https://cdn.example/other.jpg" },
{ alt: "Dungeons and Crayons", src: "https://cdn.example/cover.jpg" }, { alt: "Dungeons and Crayons", src: "https://cdn.example/cover.jpg" },
]; ];
const p = comix.detect( const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
loc("https://comix.to/title/n8we-dungeons-and-crayons/11139891-chapter-80")
);
assert.equal(p.cover, "https://cdn.example/cover.jpg"); assert.equal(p.cover, "https://cdn.example/cover.jpg");
}); });
test("comix.detect leaves cover empty when no img alt matches the title", () => { test("comix.detect leaves cover empty when no img alt matches the title", () => {
metaTags = { "og:title": "Dungeons and Crayons" }; metaTags = {};
docTitle = "Dungeons and Crayons";
pageImages = [{ alt: "Some Other Series", src: "https://cdn.example/other.jpg" }]; pageImages = [{ alt: "Some Other Series", src: "https://cdn.example/other.jpg" }];
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons")); const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
assert.equal(p.cover, ""); assert.equal(p.cover, "");
@@ -315,6 +337,33 @@ test("kagane reads the chapter number out of og:title", () => {
assert.equal(p.seriesUrl, "https://kagane.to/series/" + KAGANE_SERIES); assert.equal(p.seriesUrl, "https://kagane.to/series/" + KAGANE_SERIES);
}); });
// Volume-numbered series render the suffix as "- Volume <v> Chapter <n>" with
// no episode name, because the book carries volume_no and an empty title.
// Captured live 2026-08-08 from SP Baby.
test("kagane reads through a Volume-numbered chapter suffix", () => {
metaTags = {
"og:title": "SP Baby - Volume 1 Chapter 1",
"og:image": "https://kagane.to/api/v2/image/abc/compressed",
};
const p = kagane.detect(
loc("https://kagane.to/series/" + KAGANE_SERIES + "/reader/" + KAGANE_BOOK)
);
assert.equal(p.title, "SP Baby");
assert.equal(p.chapterNum, 1);
assert.equal(p.chapterLabel, "Chapter 1");
});
// A book with neither a volume nor an episode name ends the title right after
// the number, which the old trailing-\s regex could not match.
test("kagane reads a chapter suffix with no episode name", () => {
metaTags = { "og:title": "Some Series - Chapter 7.5" };
const p = kagane.detect(
loc("https://kagane.to/series/" + KAGANE_SERIES + "/reader/" + KAGANE_BOOK)
);
assert.equal(p.title, "Some Series");
assert.equal(p.chapterNum, 7.5);
});
test("kagane yields a null chapterNum when og:title has no chapter", () => { test("kagane yields a null chapterNum when og:title has no chapter", () => {
metaTags = { "og:title": "Infinite Decryption: The Strongest Level 0" }; metaTags = { "og:title": "Infinite Decryption: The Strongest Level 0" };
const p = kagane.detect( const p = kagane.detect(