Fix comix titles and covers, kagane volume chapters, and kagane cover rendering #37
@@ -90,3 +90,23 @@ DISCORD_REDIRECT_URI=
|
||||
# HTTP handler 500s any /json/version request whose Host header isn't an IP or
|
||||
# "localhost", which silently breaks every kagane poll.
|
||||
# BROWSER_WS_URL=ws://172.28.0.10:9222
|
||||
|
||||
# Clock zone the headless browser reports. A UTC clock is itself the bot
|
||||
# signal — Cloudflare treats it as the datacenter default — and kagane's
|
||||
# challenge then never clears. Measured 2026-08-08, identical container, one
|
||||
# Indonesian egress IP: UTC never cleared in 60s (twice); Asia/Jakarta and
|
||||
# America/New_York both cleared in 4s. So any real zone works; it does not
|
||||
# have to match the IP's country, it just must not be UTC.
|
||||
#
|
||||
# Unset falls back to the host's /etc/timezone, which is a real zone whenever
|
||||
# the host clock is set to local time. Set this when the host runs UTC — a UTC
|
||||
# server is exactly the case that fails. Only the browser sidecar reads it —
|
||||
# the backend's own zone is API_TZ below, and is cosmetic.
|
||||
# BROWSER_TZ=Asia/Jakarta
|
||||
|
||||
# Zone the backend stamps its log lines in. Cosmetic only — it exists so the
|
||||
# API's logs read on the same clock as the browser sidecar's. Nothing else in
|
||||
# the service has a zone: bookmark timestamps are unix ms, and the two real
|
||||
# time columns are timestamptz. Defaults to Asia/Jakarta; set to UTC for the
|
||||
# conventional server default.
|
||||
# API_TZ=Asia/Jakarta
|
||||
|
||||
@@ -20,6 +20,9 @@ Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free
|
||||
- Userscript run in **isolated world**, so embedded API token safe from site's JS.
|
||||
- Cloudflare's block on manga sites **IP-reputation-based, not universal — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — Cloudflare's bot scoring can flip previously-clean IP without notice. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test.
|
||||
- **kagane.to and novelfull.com are the exception to the above** — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`) and skips them entirely when that's unset. The four other sites poll fine over plain TLS.
|
||||
- **The CDP sidecar must look like a real browser, and stock headless images don't.** Measured 2026-08-08 against kagane.to, all from the same IP: `chromedp/headless-shell:stable` never cleared the challenge in 90s (`navigator.webdriver` true, empty plugin list, Chromium-branded client hints — suppressing `webdriver` alone changed nothing); `zenika/alpine-chrome` ships Chrome 124, refused outright; real Chrome with the default `--headless=new` UA never cleared, because the UA says `HeadlessChrome`; real Chrome with a stock UA **and** a non-UTC clock zone cleared in ~4s. Hence `chrome/` — a Debian image with `google-chrome-stable`, a version-derived UA, and `TZ`/`BROWSER_TZ`. Chrome reads the zone *name* through ICU from `/etc/localtime`'s symlink target, ignoring the file's contents, so mounting the host's `/etc/localtime` does **not** work; `/etc/timezone` is mounted instead.
|
||||
- **UTC is the tell, not a country mismatch.** A UTC clock is the datacenter default, so Cloudflare scores it as one; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while `Asia/Jakarta` **and** `America/New_York` both cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (`Asia/Bangkok`) rather than the measured egress. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one.
|
||||
- **A challenged page needs the tab kept open.** The interstitial takes seconds to solve and only then writes clearance into the browser's shared cookie jar. Navigate-read-close never clears anything; `BrowserFetcher.run` holds one tab and re-reads until the payload arrives.
|
||||
|
||||
## Architecture
|
||||
|
||||
@@ -37,7 +40,12 @@ Backend (`cd backend`):
|
||||
- Single test: `go test -run TestName ./...`
|
||||
- Build static binary: `CGO_ENABLED=0 go build`
|
||||
|
||||
Local stack: `docker compose up` (bookmark-api + postgres + headless-shell; `postgres-data` named volume, `restart: unless-stopped`).
|
||||
Local stack: `docker compose up` (bookmark-api + postgres + headless-shell; `postgres-data` named volume, `restart: unless-stopped`). The `headless-shell` service keeps its name but now builds `chrome/` — real Google Chrome, for the reason in the hard constraints above.
|
||||
|
||||
Live CDP proof (needs a sidecar and network, skipped otherwise):
|
||||
`SMOKE_BROWSER_WS_URL=ws://<host>:<port> go test -run TestSmokeKagane ./internal/latest`
|
||||
— fetches a real kagane cover and chapter list. A red run means the challenge is
|
||||
not clearing from this IP, which is a live fact to re-check, not necessarily a defect.
|
||||
|
||||
Smoke test: `curl` endpoints with `Authorization: Bearer <token>`; confirm `OPTIONS` preflight return CORS headers and `/healthz` return 200.
|
||||
|
||||
|
||||
+14
-3
@@ -126,9 +126,20 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
|
||||
`/userscript/novel-bookmark.user.js`, both supplied by bindmount; the
|
||||
`__API_TOKEN__` placeholder inside them is substituted with the requesting
|
||||
Reader's credential at serve time).
|
||||
`BROWSER_WS_URL` (headless-shell CDP endpoint for kagane and novelfull;
|
||||
unset disables browser polling and leaves those sites to the userscript
|
||||
alone).
|
||||
`BROWSER_WS_URL` (CDP endpoint of the `chrome/` sidecar, used by the poller
|
||||
for kagane and novelfull *and* by the web UI's kagane cover proxy; unset
|
||||
disables browser polling and serves 404 from the proxy, leaving those sites
|
||||
to the userscript alone).
|
||||
- **kagane covers are proxied, not hot-linked:** kagane serves cover images
|
||||
behind the same challenge as its pages and with
|
||||
`cross-origin-resource-policy: same-origin`, so no `<img>` on the web UI's
|
||||
origin can load one — not even from a browser holding the clearance cookie
|
||||
(verified 2026-08-08). `Bookmark.CoverURL` rewrites a stored kagane
|
||||
`og:image` to `/img/kagane/{id}`, served by `internal/web/cover.go` through
|
||||
`latest.BrowserFetcher.Image` and memoised in-process. The templates render
|
||||
`.CoverURL`, never `.Cover`. The id is matched against a UUID regex before it
|
||||
reaches the browser: the stored value is client-supplied, so an unchecked one
|
||||
is an SSRF primitive pointed at the deployment's own network.
|
||||
- **Web UI also owns:** session-gated `GET /install/{manga,novel}-bookmark.user.js`
|
||||
(renders the bindmounted script with the acting Reader's derived credential
|
||||
substituted in — the credential never appears in page markup, the address
|
||||
|
||||
@@ -0,0 +1,155 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"sync/atomic"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// fakeCovers stands in for the headless browser. It counts calls so the test
|
||||
// can prove the cache spares the browser a second navigation.
|
||||
type fakeCovers struct {
|
||||
body []byte
|
||||
contentType string
|
||||
err error
|
||||
calls atomic.Int32
|
||||
lastID atomic.Value
|
||||
}
|
||||
|
||||
func (f *fakeCovers) Image(_ context.Context, imageID string) ([]byte, string, error) {
|
||||
f.calls.Add(1)
|
||||
f.lastID.Store(imageID)
|
||||
if f.err != nil {
|
||||
return nil, "", f.err
|
||||
}
|
||||
return f.body, f.contentType, nil
|
||||
}
|
||||
|
||||
const testCoverID = "019fe11a-84c3-7fc3-a84b-88787374b617"
|
||||
|
||||
func getCover(t *testing.T, srv http.Handler, path string, cookie *http.Cookie) *httptest.ResponseRecorder {
|
||||
t.Helper()
|
||||
req := httptest.NewRequest(http.MethodGet, path, nil)
|
||||
if cookie != nil {
|
||||
req.AddCookie(cookie)
|
||||
}
|
||||
rr := httptest.NewRecorder()
|
||||
srv.ServeHTTP(rr, req)
|
||||
return rr
|
||||
}
|
||||
|
||||
// kagane serves its covers behind a Cloudflare challenge and with
|
||||
// cross-origin-resource-policy: same-origin, so the UI can only show one by
|
||||
// re-serving the bytes from its own origin.
|
||||
func TestKaganeCoverProxiesAndCaches(t *testing.T) {
|
||||
cf := &fakeCovers{body: []byte("\x00webp-bytes"), contentType: "image/webp"}
|
||||
cfg := testConfig()
|
||||
cfg.Covers = cf
|
||||
srv, st := newWebTestServer(t, cfg)
|
||||
cookie := sessionCookie(t, st)
|
||||
|
||||
for i := range 2 {
|
||||
rr := getCover(t, srv, "/img/kagane/"+testCoverID, cookie)
|
||||
if rr.Code != http.StatusOK {
|
||||
t.Fatalf("request %d: status = %d, want 200", i, rr.Code)
|
||||
}
|
||||
if got := rr.Body.String(); got != string(cf.body) {
|
||||
t.Fatalf("request %d: body = %q, want %q", i, got, cf.body)
|
||||
}
|
||||
if got := rr.Header().Get("Content-Type"); got != "image/webp" {
|
||||
t.Fatalf("request %d: Content-Type = %q, want image/webp", i, got)
|
||||
}
|
||||
}
|
||||
if got := cf.calls.Load(); got != 1 {
|
||||
t.Fatalf("fetcher called %d times, want 1 — the second read must come from the cache", got)
|
||||
}
|
||||
if got := cf.lastID.Load(); got != testCoverID {
|
||||
t.Fatalf("fetched image id = %v, want %s", got, testCoverID)
|
||||
}
|
||||
}
|
||||
|
||||
// The proxy reaches a headless browser, so it is not open to the internet.
|
||||
func TestKaganeCoverRequiresSession(t *testing.T) {
|
||||
cf := &fakeCovers{body: []byte("x"), contentType: "image/webp"}
|
||||
cfg := testConfig()
|
||||
cfg.Covers = cf
|
||||
srv, _ := newWebTestServer(t, cfg)
|
||||
|
||||
rr := getCover(t, srv, "/img/kagane/"+testCoverID, nil)
|
||||
if rr.Code != http.StatusUnauthorized {
|
||||
t.Fatalf("status = %d, want 401", rr.Code)
|
||||
}
|
||||
if got := cf.calls.Load(); got != 0 {
|
||||
t.Fatalf("fetcher called %d times for an unauthenticated request, want 0", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestKaganeCoverRejectsBadInput(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
id string
|
||||
fetch *fakeCovers
|
||||
}{
|
||||
{
|
||||
"an id that is not a uuid never reaches the browser",
|
||||
"solo-leveling",
|
||||
&fakeCovers{body: []byte("x"), contentType: "image/webp"},
|
||||
},
|
||||
{
|
||||
"a uuid-shaped id with a trailing segment is rejected whole",
|
||||
testCoverID + "x",
|
||||
&fakeCovers{body: []byte("x"), contentType: "image/webp"},
|
||||
},
|
||||
{
|
||||
"a challenged fetch is a missing cover",
|
||||
testCoverID,
|
||||
&fakeCovers{err: errors.New("challenge held")},
|
||||
},
|
||||
{
|
||||
"a content type outside the image set is not echoed back",
|
||||
testCoverID,
|
||||
&fakeCovers{body: []byte("<script>"), contentType: "text/html"},
|
||||
},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
cfg := testConfig()
|
||||
cfg.Covers = tc.fetch
|
||||
srv, st := newWebTestServer(t, cfg)
|
||||
rr := getCover(t, srv, "/img/kagane/"+tc.id, sessionCookie(t, st))
|
||||
if rr.Code != http.StatusNotFound {
|
||||
t.Fatalf("status = %d, want 404", rr.Code)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// ServeMux path-cleans a traversal into a redirect before the handler runs, so
|
||||
// the guarantee to pin down is that no request shaped like one ever gets bytes.
|
||||
func TestKaganeCoverTraversalServesNothing(t *testing.T) {
|
||||
cf := &fakeCovers{body: []byte("secret"), contentType: "image/webp"}
|
||||
cfg := testConfig()
|
||||
cfg.Covers = cf
|
||||
srv, st := newWebTestServer(t, cfg)
|
||||
|
||||
rr := getCover(t, srv, "/img/kagane/../../etc/passwd", sessionCookie(t, st))
|
||||
if rr.Code == http.StatusOK {
|
||||
t.Fatalf("status = 200, want anything but a served body")
|
||||
}
|
||||
if got := cf.calls.Load(); got != 0 {
|
||||
t.Fatalf("fetcher called %d times for a traversal, want 0", got)
|
||||
}
|
||||
}
|
||||
|
||||
// Without BROWSER_WS_URL there is no fetcher, and the endpoint must answer
|
||||
// rather than reach for a nil one.
|
||||
func TestKaganeCoverWithoutFetcher(t *testing.T) {
|
||||
srv, st := newWebTestServer(t, testConfig())
|
||||
rr := getCover(t, srv, "/img/kagane/"+testCoverID, sessionCookie(t, st))
|
||||
if rr.Code != http.StatusNotFound {
|
||||
t.Fatalf("status = %d, want 404", rr.Code)
|
||||
}
|
||||
}
|
||||
@@ -2,7 +2,9 @@ package latest
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/base64"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"fmt"
|
||||
"net/url"
|
||||
"regexp"
|
||||
@@ -22,6 +24,12 @@ const challengeTimeout = 45 * time.Second
|
||||
|
||||
var kaganeSeriesRe = regexp.MustCompile(`^/series/([0-9a-f-]{36})/?$`)
|
||||
|
||||
// kaganeImageIDRe pins the only path segment Image interpolates into an
|
||||
// outbound URL. The id arrives from a stored cover URL, which a client
|
||||
// supplied, so it is matched rather than trusted: a headless browser is a
|
||||
// strong SSRF primitive.
|
||||
var kaganeImageIDRe = regexp.MustCompile(`^[0-9a-f-]{36}$`)
|
||||
|
||||
// BrowserFetcher retrieves pages through a remote headless Chrome over the
|
||||
// DevTools Protocol.
|
||||
//
|
||||
@@ -91,23 +99,6 @@ func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int
|
||||
return "", 0, fmt.Errorf("not a fetchable browser series url: %q", seriesURL)
|
||||
}
|
||||
|
||||
f.mu.Lock()
|
||||
defer f.mu.Unlock()
|
||||
|
||||
ctx, cancel := context.WithTimeout(ctx, challengeTimeout)
|
||||
defer cancel()
|
||||
// A fresh tab per fetch, closed on return, so one wedged page cannot
|
||||
// poison later polls.
|
||||
tabCtx, cancelTab := chromedp.NewContext(f.allocCtx)
|
||||
defer cancelTab()
|
||||
// Bind the caller's deadline to the tab.
|
||||
tabCtx, cancelDeadline := context.WithCancel(tabCtx)
|
||||
defer cancelDeadline()
|
||||
go func() {
|
||||
<-ctx.Done()
|
||||
cancelDeadline()
|
||||
}()
|
||||
|
||||
var body string
|
||||
// kagane's chapter list is only in its JSON API, which must be called from
|
||||
// inside the page so the request carries the clearance cookie. novelfull
|
||||
@@ -123,24 +114,134 @@ func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int
|
||||
)
|
||||
}
|
||||
|
||||
err := chromedp.Run(tabCtx,
|
||||
chromedp.Navigate(seriesURL),
|
||||
// The challenge reloads the page itself when it passes; waiting for the
|
||||
// site's own root element is what tells us we are through it.
|
||||
chromedp.WaitReady("body", chromedp.ByQuery),
|
||||
read,
|
||||
)
|
||||
if err != nil {
|
||||
return "", 0, fmt.Errorf("browser fetch %q: %w", seriesURL, err)
|
||||
}
|
||||
if body == "" {
|
||||
// Challenge still up, or the API refused. Indistinguishable from here
|
||||
// and handled identically by the caller.
|
||||
// novelfull's payload is the DOM itself, and the interstitial has a DOM
|
||||
// too, so "we have an answer" has to exclude it explicitly. kagane's
|
||||
// in-page fetch just fails while challenged, which is already the signal.
|
||||
done := func() bool { return body != "" && (isKagane || !isInterstitial(body)) }
|
||||
if err := f.run(ctx, seriesURL, read, done); err != nil {
|
||||
// Challenge never cleared, or the API refused. Indistinguishable from
|
||||
// here and handled identically by the caller.
|
||||
if errors.Is(err, errChallengeHeld) {
|
||||
return "", 403, nil
|
||||
}
|
||||
return "", 0, fmt.Errorf("browser fetch %q: %w", seriesURL, err)
|
||||
}
|
||||
return body, 200, nil
|
||||
}
|
||||
|
||||
// Image retrieves one kagane cover as raw bytes and its content type.
|
||||
//
|
||||
// It exists because kagane serves covers behind the same challenge as its
|
||||
// pages *and* with `cross-origin-resource-policy: same-origin`, so an <img> on
|
||||
// the web UI's origin cannot load one even from a browser that already holds
|
||||
// the clearance cookie (verified 2026-08-08). Proxying is the only route.
|
||||
//
|
||||
// The image URL is navigated to rather than fetched from some other kagane
|
||||
// page: the challenge only runs on a top-level navigation, and once it clears
|
||||
// the document *is* the image, so a same-origin fetch of location.href reads
|
||||
// it straight back out of the cache.
|
||||
//
|
||||
// The challenge is not solved by the first read: WaitReady("body") is satisfied
|
||||
// by the interstitial too. run holds the tab open until the in-page fetch
|
||||
// succeeds, which is what gives the challenge script the seconds it needs.
|
||||
func (f *BrowserFetcher) Image(ctx context.Context, imageID string) ([]byte, string, error) {
|
||||
if !kaganeImageIDRe.MatchString(imageID) {
|
||||
return nil, "", fmt.Errorf("not a kagane image id: %q", imageID)
|
||||
}
|
||||
var dataURL string
|
||||
err := f.run(ctx, "https://kagane.to/api/v2/image/"+imageID+"/compressed",
|
||||
chromedp.Evaluate(`fetch(location.href).then(r => r.ok
|
||||
? r.blob().then(b => new Promise(res => {
|
||||
const fr = new FileReader();
|
||||
fr.onload = () => res(fr.result);
|
||||
fr.readAsDataURL(b);
|
||||
}))
|
||||
: "")`, &dataURL, awaitPromise),
|
||||
func() bool { return dataURL != "" })
|
||||
if err != nil {
|
||||
return nil, "", fmt.Errorf("browser image %s: %w", imageID, err)
|
||||
}
|
||||
// "data:image/webp;base64,<payload>".
|
||||
head, payload, ok := strings.Cut(dataURL, ";base64,")
|
||||
if !ok {
|
||||
return nil, "", fmt.Errorf("browser image %s: not a data url", imageID)
|
||||
}
|
||||
raw, err := base64.StdEncoding.DecodeString(payload)
|
||||
if err != nil {
|
||||
return nil, "", fmt.Errorf("browser image %s: %w", imageID, err)
|
||||
}
|
||||
return raw, strings.TrimPrefix(head, "data:"), nil
|
||||
}
|
||||
|
||||
// errChallengeHeld reports that the budget ran out with the interstitial still
|
||||
// up. Distinct from a transport failure: it means "this site said no", which
|
||||
// the poller answers with a 403 and its ordinary cooldown.
|
||||
var errChallengeHeld = errors.New("challenge held")
|
||||
|
||||
// challengePollInterval paces re-reads while a challenge solves itself.
|
||||
const challengePollInterval = 2 * time.Second
|
||||
|
||||
// isInterstitial reports whether html is Cloudflare's challenge page rather
|
||||
// than the site's own. Matched on the challenge runtime's script path, which is
|
||||
// stable across the interstitial's wording and locale — the visible "Just a
|
||||
// moment..." title is neither.
|
||||
func isInterstitial(html string) bool {
|
||||
return strings.Contains(html, "/cdn-cgi/challenge-platform/")
|
||||
}
|
||||
|
||||
// run navigates to target and re-reads until done reports an answer, bounded by
|
||||
// challengeTimeout and by the caller's own deadline, in a tab that is closed on
|
||||
// return so one wedged page cannot poison later calls.
|
||||
//
|
||||
// Holding the tab open across re-reads is the whole point. A Cloudflare
|
||||
// interstitial needs several seconds of a live page to solve itself and write
|
||||
// clearance into the browser's shared cookie jar; reading once and closing the
|
||||
// tab — which is what this did before 2026-08-08 — never gives it that window,
|
||||
// so every fetch lands on the interstitial and the clearance that would have
|
||||
// unblocked all the later ones is never obtained.
|
||||
func (f *BrowserFetcher) run(ctx context.Context, target string, read chromedp.Action, done func() bool) error {
|
||||
f.mu.Lock()
|
||||
defer f.mu.Unlock()
|
||||
|
||||
ctx, cancel := context.WithTimeout(ctx, challengeTimeout)
|
||||
defer cancel()
|
||||
tabCtx, cancelTab := chromedp.NewContext(f.allocCtx)
|
||||
defer cancelTab()
|
||||
// Bind the caller's deadline to the tab.
|
||||
tabCtx, cancelDeadline := context.WithCancel(tabCtx)
|
||||
defer cancelDeadline()
|
||||
go func() {
|
||||
<-ctx.Done()
|
||||
cancelDeadline()
|
||||
}()
|
||||
|
||||
if err := chromedp.Run(tabCtx,
|
||||
chromedp.Navigate(target),
|
||||
chromedp.WaitReady("body", chromedp.ByQuery),
|
||||
); err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
var lastErr error
|
||||
for {
|
||||
// The challenge reloads the page when it passes, which tears down the
|
||||
// execution context mid-read. That is a retry, not a failure.
|
||||
if err := chromedp.Run(tabCtx, read); err != nil {
|
||||
lastErr = err
|
||||
} else if done() {
|
||||
return nil
|
||||
}
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
if lastErr != nil {
|
||||
return fmt.Errorf("%w (last read: %v)", errChallengeHeld, lastErr)
|
||||
}
|
||||
return errChallengeHeld
|
||||
case <-time.After(challengePollInterval):
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// kaganeAPIURL maps a stored series_url to the JSON endpoint carrying its
|
||||
// chapter list. Returning false for anything else is a second line of defence
|
||||
// behind fetchableSeriesURL: a headless browser is a strong SSRF primitive and
|
||||
|
||||
@@ -0,0 +1,91 @@
|
||||
package latest
|
||||
|
||||
import (
|
||||
"context"
|
||||
"net/http"
|
||||
"os"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// TestSmokeKaganeImage is the live proof that the cover proxy's fetch actually
|
||||
// clears Cloudflare and returns image bytes. It needs a real headless Chrome
|
||||
// with outbound network, so it runs only when SMOKE_BROWSER_WS_URL is set:
|
||||
//
|
||||
// docker run --rm --shm-size=1gb -p 19222:9222 chromedp/headless-shell:stable
|
||||
// SMOKE_BROWSER_WS_URL=ws://127.0.0.1:19222 go test -run TestSmokeKaganeImage ./internal/latest
|
||||
func TestSmokeKaganeImage(t *testing.T) {
|
||||
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
|
||||
if ws == "" {
|
||||
t.Skip("SMOKE_BROWSER_WS_URL unset")
|
||||
}
|
||||
const imageID = "019fe11a-84c3-7fc3-a84b-88787374b617" // SP Baby's cover
|
||||
|
||||
// The same URL through a plain client is what the web UI's <img> gets.
|
||||
// Asserting on it keeps the test honest about why the browser is needed.
|
||||
req, err := http.NewRequest(http.MethodGet,
|
||||
"https://kagane.to/api/v2/image/"+imageID+"/compressed", nil)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if res, err := (&http.Client{Timeout: 15 * time.Second}).Do(req); err == nil {
|
||||
res.Body.Close()
|
||||
if res.StatusCode == http.StatusOK {
|
||||
t.Log("note: kagane answered a plain request 200 — the challenge is not up right now")
|
||||
}
|
||||
}
|
||||
|
||||
f, err := NewBrowserFetcher(ws)
|
||||
if err != nil {
|
||||
t.Fatalf("NewBrowserFetcher: %v", err)
|
||||
}
|
||||
defer f.Close()
|
||||
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
|
||||
defer cancel()
|
||||
body, contentType, err := f.Image(ctx, imageID)
|
||||
if err != nil {
|
||||
t.Fatalf("Image: %v", err)
|
||||
}
|
||||
if len(body) < 1000 {
|
||||
t.Fatalf("body is %d bytes, want a real image", len(body))
|
||||
}
|
||||
if contentType != "image/webp" {
|
||||
t.Fatalf("content type = %q, want image/webp", contentType)
|
||||
}
|
||||
// WebP files start with "RIFF....WEBP".
|
||||
if string(body[:4]) != "RIFF" || string(body[8:12]) != "WEBP" {
|
||||
t.Fatalf("body is not a WebP: % x", body[:12])
|
||||
}
|
||||
t.Logf("fetched %d bytes of %s", len(body), contentType)
|
||||
|
||||
if _, _, err := f.Image(ctx, "not-a-uuid"); err == nil {
|
||||
t.Fatal("Image accepted a non-uuid id")
|
||||
}
|
||||
}
|
||||
|
||||
// Control for the test above: the poller's own kagane path, same sidecar. If
|
||||
// this fails too, the sidecar is not clearing the challenge at all and the
|
||||
// image result says nothing about Image itself.
|
||||
func TestSmokeKaganeGet(t *testing.T) {
|
||||
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
|
||||
if ws == "" {
|
||||
t.Skip("SMOKE_BROWSER_WS_URL unset")
|
||||
}
|
||||
f, err := NewBrowserFetcher(ws)
|
||||
if err != nil {
|
||||
t.Fatalf("NewBrowserFetcher: %v", err)
|
||||
}
|
||||
defer f.Close()
|
||||
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
|
||||
defer cancel()
|
||||
body, status, err := f.Get(ctx, "https://kagane.to/series/019fe11a-8670-7cf3-8343-0b02057d3787")
|
||||
if err != nil {
|
||||
t.Fatalf("Get: %v", err)
|
||||
}
|
||||
t.Logf("status=%d bytes=%d head=%.80q", status, len(body), body)
|
||||
if status != 200 {
|
||||
t.Fatalf("status = %d, want 200 — the sidecar is not clearing the challenge", status)
|
||||
}
|
||||
}
|
||||
@@ -137,6 +137,22 @@ func (b Bookmark) Initial() string {
|
||||
return "?"
|
||||
}
|
||||
|
||||
// kaganeCoverRe matches the cover URL kagane's og:image carries, which is what
|
||||
// the userscript stores for that site.
|
||||
var kaganeCoverRe = regexp.MustCompile(`^https://kagane\.to/api/v2/image/([0-9a-f-]{36})/compressed$`)
|
||||
|
||||
// CoverURL is the src the web UI puts in an <img>. For every site but kagane
|
||||
// that is Cover as stored. kagane serves its images behind a Cloudflare
|
||||
// challenge *and* with `cross-origin-resource-policy: same-origin`, so no page
|
||||
// on another origin can load one however it asks (verified 2026-08-08); those
|
||||
// go through the backend's own proxy instead.
|
||||
func (b Bookmark) CoverURL() string {
|
||||
if m := kaganeCoverRe.FindStringSubmatch(b.Cover); m != nil {
|
||||
return "/img/kagane/" + m[1]
|
||||
}
|
||||
return b.Cover
|
||||
}
|
||||
|
||||
// Library buckets. A bookmark is in exactly one. This cannot be derived from
|
||||
// Site: asurascans serves manga and novels from the same /comics/ path, so the
|
||||
// userscript that recorded the page is the only party that knows which.
|
||||
|
||||
@@ -543,6 +543,38 @@ func TestDisplayChapter(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestCoverURL(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
cover string
|
||||
want string
|
||||
}{
|
||||
{
|
||||
"kagane routes through the proxy",
|
||||
"https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
|
||||
"/img/kagane/019fe11a-84c3-7fc3-a84b-88787374b617",
|
||||
},
|
||||
{
|
||||
"another site is served as stored",
|
||||
"https://gg.asuracomic.net/storage/media/1/conversions/cover.webp",
|
||||
"https://gg.asuracomic.net/storage/media/1/conversions/cover.webp",
|
||||
},
|
||||
{
|
||||
"a lookalike host is not rewritten",
|
||||
"https://evil.example/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
|
||||
"https://evil.example/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
|
||||
},
|
||||
{"no cover stays empty", "", ""},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
if got := (Bookmark{Cover: tc.cover}).CoverURL(); got != tc.want {
|
||||
t.Errorf("CoverURL() = %q, want %q", got, tc.want)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpsertKindDefaultsToManga(t *testing.T) {
|
||||
store := newTestStore(t)
|
||||
got, err := store.Upsert(store.OwnerID(), Bookmark{
|
||||
|
||||
@@ -0,0 +1,129 @@
|
||||
package web
|
||||
|
||||
import (
|
||||
"context"
|
||||
"log"
|
||||
"net/http"
|
||||
"regexp"
|
||||
"sync"
|
||||
"time"
|
||||
)
|
||||
|
||||
// CoverFetcher retrieves one kagane cover by image id. Satisfied by
|
||||
// latest.BrowserFetcher, and nil when BROWSER_WS_URL is unset — which leaves
|
||||
// kagane covers exactly as unavailable as they were before this endpoint
|
||||
// existed, rather than hanging a request on a fetcher that cannot run.
|
||||
type CoverFetcher interface {
|
||||
Image(ctx context.Context, imageID string) (body []byte, contentType string, err error)
|
||||
}
|
||||
|
||||
// coverIDRe matches the request path segment that becomes part of an outbound
|
||||
// URL. The proxy is session-gated, but the id still reaches a headless browser,
|
||||
// so it is validated at the boundary rather than passed through.
|
||||
var coverIDRe = regexp.MustCompile(`^[0-9a-f-]{36}$`)
|
||||
|
||||
// coverTypes is the set of content types the proxy will echo back. A response
|
||||
// header sourced from a third party is not repeated verbatim: anything outside
|
||||
// this set is treated as "not a cover".
|
||||
var coverTypes = map[string]bool{
|
||||
"image/webp": true,
|
||||
"image/jpeg": true,
|
||||
"image/png": true,
|
||||
"image/avif": true,
|
||||
"image/gif": true,
|
||||
}
|
||||
|
||||
// coverTimeout bounds one proxied cover. Shorter than the fetcher's own
|
||||
// challenge budget on purpose: a browser page is waiting on this, and a cover
|
||||
// that has not arrived by now is better left as a broken slot than as a request
|
||||
// holding a connection open.
|
||||
const coverTimeout = 20 * time.Second
|
||||
|
||||
// coverCacheMax caps the in-memory cover cache. Covers are immutable per image
|
||||
// id and a library holds tens of series, so this is a ceiling that is never
|
||||
// reached in practice; reaching it clears the map rather than evicting by age.
|
||||
//
|
||||
// ponytail: flush-on-full, not LRU. Swap it for an LRU if a library ever grows
|
||||
// past this and the flush starts costing refetches.
|
||||
const coverCacheMax = 500
|
||||
|
||||
type cachedCover struct {
|
||||
body []byte
|
||||
contentType string
|
||||
}
|
||||
|
||||
type coverCache struct {
|
||||
mu sync.Mutex
|
||||
m map[string]cachedCover
|
||||
}
|
||||
|
||||
func (c *coverCache) get(id string) (cachedCover, bool) {
|
||||
c.mu.Lock()
|
||||
defer c.mu.Unlock()
|
||||
v, ok := c.m[id]
|
||||
return v, ok
|
||||
}
|
||||
|
||||
func (c *coverCache) put(id string, v cachedCover) {
|
||||
c.mu.Lock()
|
||||
defer c.mu.Unlock()
|
||||
if c.m == nil || len(c.m) >= coverCacheMax {
|
||||
c.m = make(map[string]cachedCover, coverCacheMax)
|
||||
}
|
||||
c.m[id] = v
|
||||
}
|
||||
|
||||
// kaganeCover serves a kagane cover from the backend's own origin.
|
||||
//
|
||||
// kagane answers image requests with a Cloudflare challenge and
|
||||
// `cross-origin-resource-policy: same-origin`, so the web UI cannot render one
|
||||
// directly under any combination of referrer policy or crossorigin attribute
|
||||
// (verified 2026-08-08). Fetching it through the headless browser that already
|
||||
// clears the challenge, and re-serving it here, is what puts the bytes on an
|
||||
// origin the page may load from.
|
||||
//
|
||||
// ponytail: covers are fetched on first view, one browser navigation at a time
|
||||
// behind the fetcher's mutex, so a first load of a large kagane library
|
||||
// trickles in over a few seconds. The cache makes it a one-off. Prefetching
|
||||
// during the poll cycle is the upgrade if that ever grates.
|
||||
func (h *Handler) kaganeCover(w http.ResponseWriter, r *http.Request) {
|
||||
id := r.PathValue("id")
|
||||
if !coverIDRe.MatchString(id) {
|
||||
http.NotFound(w, r)
|
||||
return
|
||||
}
|
||||
if h.covers == nil {
|
||||
http.NotFound(w, r)
|
||||
return
|
||||
}
|
||||
if v, ok := h.coverCache.get(id); ok {
|
||||
writeCover(w, v)
|
||||
return
|
||||
}
|
||||
|
||||
ctx, cancel := context.WithTimeout(r.Context(), coverTimeout)
|
||||
defer cancel()
|
||||
body, contentType, err := h.covers.Image(ctx, id)
|
||||
if err != nil {
|
||||
log.Printf("kagane cover %s: %v", id, err)
|
||||
http.NotFound(w, r)
|
||||
return
|
||||
}
|
||||
if !coverTypes[contentType] {
|
||||
log.Printf("kagane cover %s: unexpected content type %q", id, contentType)
|
||||
http.NotFound(w, r)
|
||||
return
|
||||
}
|
||||
|
||||
v := cachedCover{body: body, contentType: contentType}
|
||||
h.coverCache.put(id, v)
|
||||
writeCover(w, v)
|
||||
}
|
||||
|
||||
// writeCover sends the bytes with a long cache life: an image id names one
|
||||
// immutable rendering, so a client that has it never needs to ask again.
|
||||
func writeCover(w http.ResponseWriter, v cachedCover) {
|
||||
w.Header().Set("Content-Type", v.contentType)
|
||||
w.Header().Set("Cache-Control", "private, max-age=604800, immutable")
|
||||
w.Write(v.body)
|
||||
}
|
||||
@@ -6,7 +6,7 @@
|
||||
<div class="row">
|
||||
<a class="cover" href="{{.ContinueURL}}" target="_blank" rel="noopener noreferrer"
|
||||
tabindex="-1" aria-hidden="true">
|
||||
{{if .Cover}}<img src="{{.Cover}}" alt="" loading="lazy">
|
||||
{{if .CoverURL}}<img src="{{.CoverURL}}" alt="" loading="lazy">
|
||||
{{/* aria-hidden on the cover link is not enough — Chromium still exposes
|
||||
the letter because the link is programmatically focusable — so the
|
||||
monogram carries its own, same as the recent strip's. */}}
|
||||
|
||||
@@ -14,7 +14,7 @@
|
||||
<a class="recent-card {{if .HasNewChapter}}is-new{{end}}" href="{{.ContinueURL}}"
|
||||
target="_blank" rel="noopener noreferrer">
|
||||
<span class="recent-cover">
|
||||
{{if .Cover}}<img src="{{.Cover}}" alt="" loading="lazy">
|
||||
{{if .CoverURL}}<img src="{{.CoverURL}}" alt="" loading="lazy">
|
||||
{{else}}<span class="monogram" aria-hidden="true">{{.Initial}}</span>{{end}}
|
||||
{{if .HasNewChapter}}<span class="foot-rule"></span>
|
||||
{{else if .Favorite}}<span class="foot-rule brass"></span>{{end}}
|
||||
|
||||
@@ -50,6 +50,10 @@ type Handler struct {
|
||||
// httpClient is the plain stdlib client that talks to Discord. It is not
|
||||
// an injected interface: tests point APIBase at a stub server instead.
|
||||
httpClient *http.Client
|
||||
// covers proxies kagane cover images, which no browser can load directly.
|
||||
// Nil disables the endpoint — see CoverFetcher.
|
||||
covers CoverFetcher
|
||||
coverCache coverCache
|
||||
}
|
||||
|
||||
// listView is what every list-rendering template receives.
|
||||
@@ -111,7 +115,7 @@ type loginView struct {
|
||||
|
||||
// New parses every template up front so a broken one kills the process at
|
||||
// startup rather than the first request that touches it.
|
||||
func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, novelPath string) (*Handler, error) {
|
||||
func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, novelPath string, covers CoverFetcher) (*Handler, error) {
|
||||
tmpl, err := template.ParseFS(templateFS, "templates/*.html")
|
||||
if err != nil {
|
||||
return nil, err
|
||||
@@ -126,6 +130,7 @@ func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, nove
|
||||
states: newOAuthStates(),
|
||||
limiter: session.NewLoginLimiter(),
|
||||
httpClient: &http.Client{Timeout: discordTimeout},
|
||||
covers: covers,
|
||||
}, nil
|
||||
}
|
||||
|
||||
@@ -142,6 +147,10 @@ func (h *Handler) Register(mux *http.ServeMux) {
|
||||
mux.HandleFunc("POST /ui/bookmarks/{key}/chapter", h.requireSession(h.uiChapter))
|
||||
mux.HandleFunc("DELETE /ui/bookmarks/{key}", h.requireSession(h.uiDelete))
|
||||
|
||||
// Session-gated like every other UI route: the deployment proxies kagane's
|
||||
// images for its own Readers, not for the internet.
|
||||
mux.HandleFunc("GET /img/kagane/{id}", h.requireSession(h.kaganeCover))
|
||||
|
||||
// Install endpoints render the script directly under the session: the
|
||||
// credential travels inside the served bytes, never in the address bar or
|
||||
// the page markup. Updates after install use the credential-bearing /u/
|
||||
|
||||
+28
-17
@@ -47,6 +47,10 @@ type Config struct {
|
||||
NovelUserscriptPath string
|
||||
// LatestPoll configures the background latest-chapter fetcher.
|
||||
LatestPoll LatestPoll
|
||||
// Covers proxies kagane cover images for the web UI. Not from the
|
||||
// environment: it is the shared headless browser, wired in main once it
|
||||
// connects, and nil in every test router.
|
||||
Covers web.CoverFetcher
|
||||
}
|
||||
|
||||
// LatestPoll configures the background latest-chapter poller.
|
||||
@@ -206,7 +210,7 @@ func newRouter(s *store.Store, cfg Config) http.Handler {
|
||||
// The browser UI is always registered; signing in is Discord OAuth, so
|
||||
// there is no password to forget and no gate to leave unset.
|
||||
wh, err := web.New(s, cfg.Discord, []byte(cfg.TokenKey),
|
||||
cfg.UserscriptPath, cfg.NovelUserscriptPath)
|
||||
cfg.UserscriptPath, cfg.NovelUserscriptPath, cfg.Covers)
|
||||
if err != nil {
|
||||
log.Fatalf("web handler: %v", err)
|
||||
}
|
||||
@@ -270,9 +274,26 @@ func main() {
|
||||
// The poller is off the request path entirely: if it cannot start, the
|
||||
// service still serves bookmarks and the userscript still captures latest
|
||||
// chapters on its own.
|
||||
//
|
||||
// One headless browser serves both consumers that need a Cloudflare
|
||||
// challenge cleared: the poller's kagane/novelfull fetches and the web
|
||||
// UI's kagane cover proxy. Optional — unset leaves both degraded to what
|
||||
// they were before the sidecar existed.
|
||||
var browser latest.Fetcher
|
||||
pollCtx, stopPoll := context.WithCancel(context.Background())
|
||||
defer stopPoll()
|
||||
startLatestPoller(pollCtx, s, cfg.LatestPoll)
|
||||
if ws := strings.TrimSpace(os.Getenv("BROWSER_WS_URL")); ws != "" {
|
||||
bf, err := latest.NewBrowserFetcher(ws)
|
||||
if err != nil {
|
||||
log.Printf("browser fetcher disabled: %v", err)
|
||||
} else {
|
||||
browser = bf
|
||||
cfg.Covers = bf
|
||||
context.AfterFunc(pollCtx, bf.Close)
|
||||
log.Printf("browser fetcher at %s", ws)
|
||||
}
|
||||
}
|
||||
startLatestPoller(pollCtx, s, cfg.LatestPoll, browser)
|
||||
|
||||
srv := &http.Server{
|
||||
Addr: ":" + cfg.Port,
|
||||
@@ -307,7 +328,7 @@ func main() {
|
||||
// HTTP client cannot be built. Any problem here is logged and skipped: this
|
||||
// feature going missing degrades the service to userscript-only latest-chapter
|
||||
// tracking, which is exactly how it behaved before.
|
||||
func startLatestPoller(ctx context.Context, s *store.Store, cfg LatestPoll) {
|
||||
func startLatestPoller(ctx context.Context, s *store.Store, cfg LatestPoll, browser latest.Fetcher) {
|
||||
if !cfg.Enabled {
|
||||
log.Println("latest-chapter poller: disabled by config")
|
||||
return
|
||||
@@ -317,9 +338,13 @@ func startLatestPoller(ctx context.Context, s *store.Store, cfg LatestPoll) {
|
||||
log.Printf("latest-chapter poller: disabled, cannot build client: %v", err)
|
||||
return
|
||||
}
|
||||
// Nil browser: sites behind a JavaScript challenge are simply not polled,
|
||||
// and their latest_chapter comes from the userscript alone — which is how
|
||||
// the service behaved before the sidecar existed.
|
||||
p := &latest.Poller{
|
||||
Store: s,
|
||||
Fetch: f,
|
||||
BrowserFetch: browser,
|
||||
Now: time.Now,
|
||||
Cooldown: cfg.Cooldown,
|
||||
Interval: cfg.Interval,
|
||||
@@ -327,19 +352,5 @@ func startLatestPoller(ctx context.Context, s *store.Store, cfg LatestPoll) {
|
||||
Batch: cfg.Batch,
|
||||
}
|
||||
|
||||
// Optional: without it, sites behind a JavaScript challenge are simply not
|
||||
// polled, and their latest_chapter comes from the userscript alone — which
|
||||
// is how the service behaved before the sidecar existed.
|
||||
if ws := strings.TrimSpace(os.Getenv("BROWSER_WS_URL")); ws != "" {
|
||||
bf, err := latest.NewBrowserFetcher(ws)
|
||||
if err != nil {
|
||||
log.Printf("latest-chapter poller: browser fetcher disabled: %v", err)
|
||||
} else {
|
||||
p.BrowserFetch = bf
|
||||
context.AfterFunc(ctx, bf.Close)
|
||||
log.Printf("latest-chapter poller: browser fetcher at %s", ws)
|
||||
}
|
||||
}
|
||||
|
||||
go p.Run(ctx)
|
||||
}
|
||||
|
||||
@@ -0,0 +1,39 @@
|
||||
# syntax=docker/dockerfile:1
|
||||
|
||||
# Real Google Chrome for the latest-chapter poller and the kagane cover proxy.
|
||||
#
|
||||
# Not chromedp/headless-shell, which this replaces. headless-shell is a stripped
|
||||
# Chrome build and Cloudflare's managed challenge on kagane.to never clears for
|
||||
# it: measured 2026-08-08, 60s of a held-open tab still served the interstitial,
|
||||
# while stock Chrome from the same IP cleared in ~4s. The tells are structural
|
||||
# rather than a header — navigator.webdriver true, an empty plugin list, and
|
||||
# Chromium- rather than Chrome-branded client hints. Overriding webdriver alone
|
||||
# was tried and did not move it, so the browser build itself is the fix.
|
||||
#
|
||||
# zenika/alpine-chrome was also tried: its Chrome is 124 (2024), old enough that
|
||||
# Cloudflare refuses it outright and old enough to break chromedp's CDP structs.
|
||||
FROM debian:trixie-slim
|
||||
|
||||
# Chrome is deliberately unpinned, against the usual rule. A pinned build goes
|
||||
# stale, and a stale browser is exactly what Cloudflare turns away — the 124 in
|
||||
# alpine-chrome is the worked example. Rebuild is the upgrade path.
|
||||
RUN apt-get update \
|
||||
&& apt-get install -y --no-install-recommends ca-certificates wget gnupg \
|
||||
&& wget -qO- https://dl.google.com/linux/linux_signing_key.pub \
|
||||
| gpg --dearmor -o /usr/share/keyrings/google-chrome.gpg \
|
||||
&& echo "deb [arch=amd64 signed-by=/usr/share/keyrings/google-chrome.gpg] https://dl.google.com/linux/chrome/deb/ stable main" \
|
||||
> /etc/apt/sources.list.d/google-chrome.list \
|
||||
&& apt-get update \
|
||||
&& apt-get install -y --no-install-recommends google-chrome-stable socat \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
# Unprivileged: Chrome refuses to run as root, and the CDP endpoint is a shell
|
||||
# on whatever user owns it.
|
||||
RUN useradd --create-home --shell /usr/sbin/nologin chrome
|
||||
USER chrome
|
||||
WORKDIR /home/chrome
|
||||
|
||||
COPY entrypoint.sh /entrypoint.sh
|
||||
|
||||
EXPOSE 9222
|
||||
ENTRYPOINT ["/entrypoint.sh"]
|
||||
Executable
+61
@@ -0,0 +1,61 @@
|
||||
#!/bin/sh
|
||||
set -eu
|
||||
|
||||
# A UTC clock is itself the bot signal — Cloudflare treats it as the datacenter
|
||||
# default — and kagane's challenge then never clears. Measured 2026-08-08 with
|
||||
# an identical container on one Indonesian egress IP: UTC never cleared in 60s
|
||||
# (twice), while Asia/Jakarta and America/New_York both cleared in 4s. Any real
|
||||
# zone will do; the zone does not have to match the IP's country, it just must
|
||||
# not be UTC. It does have to be right the way Chrome reads it.
|
||||
#
|
||||
# TZ must carry the zone *name*. Chrome resolves the zone through ICU, which
|
||||
# takes the name from /etc/localtime's symlink target and ignores the file's
|
||||
# contents; bind-mounting the host's /etc/localtime therefore lands on the
|
||||
# image's own symlink to Etc/UTC and leaves glibc reporting +07 while Chrome
|
||||
# still reports UTC. /etc/timezone, mounted by docker-compose.yml, is the name.
|
||||
[ -n "${TZ:-}" ] || TZ=$(cat /etc/timezone 2>/dev/null || echo UTC)
|
||||
export TZ
|
||||
|
||||
# Chrome's own UA advertises "HeadlessChrome" under --headless=new, and that one
|
||||
# token is the difference between kagane.to's challenge clearing in ~4s and
|
||||
# never clearing at all (measured 2026-08-08, same host, same Chrome, only the
|
||||
# UA changed). Overriding it does not touch the Sec-CH-UA client hints, which
|
||||
# report the real version, so the version is read back out of the binary rather
|
||||
# than hardcoded: a hardcoded one would drift out of step with the hints on the
|
||||
# next Chrome update and become a fresh tell.
|
||||
major=$(google-chrome-stable --version | sed -E 's/[^0-9]*([0-9]+)\..*/\1/')
|
||||
ua="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/${major}.0.0.0 Safari/537.36"
|
||||
|
||||
# Chrome binds its DevTools port to loopback and silently ignores
|
||||
# --remote-debugging-address (verified 2026-08-08: Chrome 151 with
|
||||
# --remote-debugging-address=0.0.0.0 still listened on 127.0.0.1 only), so the
|
||||
# caller — another container — cannot reach it directly. socat fronting the
|
||||
# loopback port is how chromedp/headless-shell solved the same problem and is
|
||||
# why this image is a drop-in for it.
|
||||
#
|
||||
# Nothing publishes 9222; reachability is the `browser` network in
|
||||
# docker-compose.yml, and an exposed CDP endpoint is remote code execution.
|
||||
socat TCP-LISTEN:9222,fork,reuseaddr TCP:127.0.0.1:9223 &
|
||||
|
||||
# Chrome stays in the foreground so that its death takes the container down and
|
||||
# compose's restart policy applies; a backgrounded browser behind a live socat
|
||||
# would leave the sidecar looking healthy while answering nothing.
|
||||
#
|
||||
# No --enable-automation: it sets navigator.webdriver, the first thing a bot
|
||||
# check reads.
|
||||
#
|
||||
# --no-sandbox because Chrome's zygote wants user namespaces, which Docker's
|
||||
# default profile does not hand out; the alternative is --cap-add=SYS_ADMIN,
|
||||
# which gives the container strictly more than it takes away. Containment here
|
||||
# is the unprivileged user, the isolated network, and the fact that this
|
||||
# browser only ever navigates to kagane.to and novelfull.com.
|
||||
exec google-chrome-stable \
|
||||
--headless=new \
|
||||
--no-sandbox \
|
||||
--remote-debugging-port=9223 \
|
||||
--user-agent="$ua" \
|
||||
--user-data-dir=/home/chrome/profile \
|
||||
--no-first-run \
|
||||
--no-default-browser-check \
|
||||
--disable-gpu \
|
||||
about:blank
|
||||
+27
-12
@@ -24,6 +24,13 @@ services:
|
||||
# comes from .env so it is never committed.
|
||||
DATABASE_URL: ${DATABASE_URL:-postgres://bookmarks:${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD in .env}@postgres:5432/bookmarks?sslmode=disable}
|
||||
PORT: "8080"
|
||||
# Log timestamps only. Go's `log` stamps lines in local time, and this
|
||||
# service has no other use for a zone: bookmark timestamps are unix ms
|
||||
# and the two real time columns are timestamptz, both absolute instants.
|
||||
# Purely so these lines read on the same clock as the sidecar's. Named
|
||||
# API_TZ rather than TZ so an operator's exported shell TZ cannot leak
|
||||
# in; distroless already carries tzdata, so the name just resolves.
|
||||
TZ: ${API_TZ:-Asia/Jakarta}
|
||||
# Discord OAuth for the browser UI (ADR-0002). The first four are
|
||||
# required; DISCORD_REQUIRED_ROLE is optional and empty by default.
|
||||
# Guild membership is the whole gate: any member becomes a Reader.
|
||||
@@ -95,8 +102,25 @@ services:
|
||||
- db
|
||||
|
||||
headless-shell:
|
||||
image: chromedp/headless-shell:stable
|
||||
# Real Google Chrome, not chromedp/headless-shell — see chrome/Dockerfile.
|
||||
# The service name is kept so existing overrides and BROWSER_WS_URL stay put.
|
||||
build: ./chrome
|
||||
image: bookmarkmanager-chrome:latest
|
||||
restart: unless-stopped
|
||||
environment:
|
||||
# A UTC clock is itself the bot signal: Cloudflare treats it as the
|
||||
# datacenter default, and kagane's challenge then never clears. Measured
|
||||
# 2026-08-08, identical container, one Indonesian egress IP: UTC never
|
||||
# cleared in 60s (twice); Asia/Jakarta and America/New_York both cleared
|
||||
# in 4s. So any real zone works and it need not match the IP's country —
|
||||
# only UTC fails. Unset falls back to the host's /etc/timezone below,
|
||||
# which is a real zone whenever the host clock is set to local time; set
|
||||
# BROWSER_TZ when the host runs UTC.
|
||||
TZ: ${BROWSER_TZ:-}
|
||||
volumes:
|
||||
# The zone *name*, which is what Chrome's ICU needs — see chrome/entrypoint.sh.
|
||||
# Absent on a non-Debian host, which the entrypoint handles by falling back to UTC.
|
||||
- /etc/timezone:/etc/timezone:ro
|
||||
# Chrome allocates shared memory per tab and dies on Docker's 64MB default.
|
||||
shm_size: '1gb'
|
||||
# Reaps zombie renderer processes, which otherwise accumulate for the
|
||||
@@ -104,17 +128,8 @@ services:
|
||||
init: true
|
||||
# Deliberately no `ports:` — an exposed CDP endpoint is remote code
|
||||
# execution. Only bookmark-api, via the `browser` network below, may reach it.
|
||||
# Don't pass --remote-debugging-address/--remote-debugging-port here: the
|
||||
# image's own entrypoint (/headless-shell/run.sh) already starts Chrome on
|
||||
# 127.0.0.1:9223 and fronts it with a socat proxy listening on 0.0.0.0:9222.
|
||||
# Redeclaring the port flag here overrides Chrome's, so it binds 9222
|
||||
# directly (IPv6 loopback only) instead of 9223 — collides with socat's own
|
||||
# bind on 9222 and leaves nothing listening on 9223, so every external
|
||||
# connection to headless-shell:9222 fails with EOF. Only pass flags the
|
||||
# entrypoint doesn't already set.
|
||||
command:
|
||||
- --disable-gpu
|
||||
- --no-sandbox
|
||||
# No `command:` either: every flag this browser needs is in its entrypoint,
|
||||
# and the UA override there is load-bearing for the challenge.
|
||||
networks:
|
||||
browser:
|
||||
# Pinned so BROWSER_WS_URL can name an IP (required, see above) that
|
||||
|
||||
@@ -44,6 +44,25 @@ Guidance for OpenCode (and Claude Code) working under `userscript/`. See root `A
|
||||
Encodings (incl. triple-encoded punctuation like `%25252D`) identical
|
||||
on /manga/ and /title/ pages, so decode-once seriesIds match — verified
|
||||
2026-07-28.
|
||||
- **comix.to**: series `/title/<id>-<slug>`, chapter
|
||||
`/title/<id>-<slug>/<uploadId>-chapter-<n>`. Only the leading `<id>` is
|
||||
identity — the slug re-renders when a series is renamed (`comixSeriesId`).
|
||||
An SPA that **never rewrites `og:title`**: the server-rendered head keeps
|
||||
whatever document loaded first, so on a cold load `og:title` is the homepage's
|
||||
"Comix — Read Comics online for free" and after an in-page hop it is the
|
||||
*previous* series' name. `document.title` is the one thing client routing does
|
||||
update, so titles come from there, with the chapter page's `" · Ch.<n>"` tail
|
||||
stripped. Covers likewise: `og:image` is absent, so the cover is the `img`
|
||||
whose `alt` matches the cleaned title — verified live 2026-08-08.
|
||||
- **kagane.to**: series `/series/<uuid>`, reader
|
||||
`/series/<uuid>/reader/<bookUuid>`. Reader URLs carry no chapter number, so
|
||||
the number comes out of `og:title`. Two shapes exist: `"<Series> - Chapter
|
||||
<n>[ - Episode <n>]"` and, for volume-numbered series, `"<Series> - Volume <v>
|
||||
Chapter <n>"` with no episode name — both must yield a bare series title, or
|
||||
the volume tail lands in the bookmark's title. Its covers are challenge- and
|
||||
CORP-protected, so the web UI proxies them; the userscript still stores the
|
||||
raw `og:image`. Behind a Cloudflare JS challenge, so the backend polls it
|
||||
through the headless browser.
|
||||
- **novelfull.com** (novel script): series `/<slug>.html`, chapter
|
||||
`/<slug>/chapter-<n>[-<title-slug>].html`. No `og:*` tags at all — title from
|
||||
`h3.title` (series) or `a.truyen-title` (chapter), cover from
|
||||
|
||||
@@ -242,6 +242,12 @@
|
||||
matches: (loc) => /(^|\.)comix\.to$/.test(loc.hostname),
|
||||
detect(loc) {
|
||||
const path = loc.pathname;
|
||||
// comix client-routes without ever rewriting og:title — the head keeps
|
||||
// whatever the first server-rendered document carried, so a bookmark
|
||||
// taken after a client route got the homepage's title, then the
|
||||
// previous series'. document.title is the one thing its router does
|
||||
// update. Verified live 2026-08-08; do not "restore" meta("og:title").
|
||||
const pageTitle = cleanTitle(document.title);
|
||||
// /title/<id>-<slug>/<uploadId>-chapter-<n>. Several uploads (different
|
||||
// groups or languages) share one chapter number; the number is the
|
||||
// progress identity, the upload id is not.
|
||||
@@ -252,8 +258,8 @@
|
||||
type: "chapter",
|
||||
site: this.site,
|
||||
seriesId: comixSeriesId(m[1]),
|
||||
title: cleanTitle(meta("og:title")),
|
||||
cover: coverFromPage(),
|
||||
title: pageTitle,
|
||||
cover: coverFromPage(pageTitle),
|
||||
seriesUrl: loc.origin + "/title/" + m[1],
|
||||
chapterLabel: "Chapter " + m[2],
|
||||
chapterNum: isNaN(num) ? null : num,
|
||||
@@ -267,8 +273,8 @@
|
||||
type: "series",
|
||||
site: this.site,
|
||||
seriesId: comixSeriesId(m[1]),
|
||||
title: cleanTitle(meta("og:title")),
|
||||
cover: coverFromPage(),
|
||||
title: pageTitle,
|
||||
cover: coverFromPage(pageTitle),
|
||||
seriesUrl: loc.origin + "/title/" + m[1],
|
||||
chapterLabel: null,
|
||||
chapterNum: null,
|
||||
@@ -277,7 +283,7 @@
|
||||
}
|
||||
return { type: "other" };
|
||||
|
||||
// comix chapter og:title is "<Title> · Ch.<n>"; series is clean.
|
||||
// comix chapter document.title is "<Title> · Ch.<n>"; series is clean.
|
||||
function cleanTitle(t) {
|
||||
if (!t) return "";
|
||||
return t.replace(/\s*·\s*Ch\.[\d.]+\s*$/i, "").trim();
|
||||
@@ -287,8 +293,7 @@
|
||||
// the DOM for a cover. Matching on alt rather than a class keeps it off
|
||||
// the site's styling: the cover is the image whose alt is the title.
|
||||
// Do not "simplify" this into meta("og:image") — that returns null.
|
||||
function coverFromPage() {
|
||||
const title = cleanTitle(meta("og:title"));
|
||||
function coverFromPage(title) {
|
||||
if (!title || !document.querySelectorAll) return "";
|
||||
for (const img of document.querySelectorAll("img[alt]")) {
|
||||
if (img.getAttribute("alt") === title) return img.getAttribute("src") || "";
|
||||
@@ -313,6 +318,14 @@
|
||||
},
|
||||
};
|
||||
|
||||
// Kagane builds the reader og:title suffix out of the book's metadata, so
|
||||
// every combination occurs: the volume part appears only when the book has a
|
||||
// volume_no, the episode part only when it has a non-empty title. All four
|
||||
// shapes captured live 2026-08-08 — "SP Baby - Volume 1 Chapter 1" is the one
|
||||
// the old trailing-space regex missed, which left both the number and the
|
||||
// series title wrong.
|
||||
const KAGANE_CHAPTER_SUFFIX = /\s-\s(?:Volume\s[\d.]+\s)?Chapter\s([\d.]+)(?:\s-\s.*)?$/i;
|
||||
|
||||
const kagane = {
|
||||
site: "kagane",
|
||||
matches: (loc) => /(^|\.)kagane\.to$/.test(loc.hostname),
|
||||
@@ -353,9 +366,8 @@
|
||||
}
|
||||
return { type: "other" };
|
||||
|
||||
// Reader og:title is "<Title> - Chapter <n> - <episode name>".
|
||||
function chapterNumFromTitle(t) {
|
||||
const m = t && t.match(/\s-\sChapter\s([\d.]+)\s/);
|
||||
const m = t && t.match(KAGANE_CHAPTER_SUFFIX);
|
||||
if (!m) return null;
|
||||
const num = parseFloat(m[1]);
|
||||
return isNaN(num) ? null : num;
|
||||
@@ -363,7 +375,7 @@
|
||||
|
||||
function cleanTitle(t) {
|
||||
if (!t) return "";
|
||||
return t.replace(/\s-\sChapter\s[\d.]+\s-\s.*$/i, "").trim();
|
||||
return t.replace(KAGANE_CHAPTER_SUFFIX, "").trim();
|
||||
}
|
||||
},
|
||||
// Reader hrefs are uuids with no number in them, so no maximum can be taken
|
||||
@@ -1450,8 +1462,15 @@
|
||||
}
|
||||
|
||||
let lastUrl = location.href;
|
||||
function onNavigate() {
|
||||
let lastPageSig = "";
|
||||
|
||||
function setPage() {
|
||||
state.page = detect();
|
||||
lastPageSig = JSON.stringify(state.page);
|
||||
}
|
||||
|
||||
function onNavigate() {
|
||||
setPage();
|
||||
render();
|
||||
maybeAutoUpdate();
|
||||
maybeCaptureLatestOnSeriesPage();
|
||||
@@ -1484,7 +1503,14 @@
|
||||
wrap("pushState");
|
||||
wrap("replaceState");
|
||||
window.addEventListener("popstate", fire);
|
||||
setInterval(fire, 1500); // catch routes that bypass history
|
||||
// Also catches routes that bypass history — and comix, which fills
|
||||
// document.title a beat after the route changes, so the 300ms snapshot
|
||||
// above can still hold the previous page's title. Re-detect whenever what
|
||||
// we would read has changed, not only when the URL has.
|
||||
setInterval(() => {
|
||||
fire();
|
||||
if (JSON.stringify(detect()) !== lastPageSig) onNavigate();
|
||||
}, 1500);
|
||||
}
|
||||
|
||||
// ============================================================
|
||||
@@ -1563,7 +1589,7 @@
|
||||
|
||||
function init() {
|
||||
buildUI();
|
||||
state.page = detect();
|
||||
setPage();
|
||||
render();
|
||||
installNavWatcher();
|
||||
installLongPress();
|
||||
|
||||
@@ -31,6 +31,9 @@ globalThis.location = {
|
||||
|
||||
// og: meta tags the adapters read through meta(). Reassigned per test.
|
||||
let metaTags = {};
|
||||
// document.title. comix's SPA rewrites this on client routing but never
|
||||
// og:title, so the comix adapter reads it instead. Reassigned per test.
|
||||
let docTitle = "";
|
||||
// img[alt] elements comix's coverFromPage() scans. Reassigned per test; each
|
||||
// entry is {alt, src}.
|
||||
let pageImages = [];
|
||||
@@ -48,6 +51,9 @@ globalThis.document = {
|
||||
}));
|
||||
},
|
||||
addEventListener() {},
|
||||
get title() {
|
||||
return docTitle;
|
||||
},
|
||||
body: undefined,
|
||||
};
|
||||
|
||||
@@ -193,8 +199,15 @@ test("comixSeriesId leaves a bare id untouched", () => {
|
||||
assert.equal(comixSeriesId("n8we"), "n8we");
|
||||
});
|
||||
|
||||
// comix is an SPA that rewrites document.title on client routing but leaves the
|
||||
// server-rendered og:title untouched, so every test here pins og:title to a
|
||||
// STALE value — the homepage title on first hop, the previous series after
|
||||
// that. Captured live 2026-08-08.
|
||||
const COMIX_STALE_HOME = "Comix - Read Comics online for free";
|
||||
|
||||
test("comix detects a series page", () => {
|
||||
metaTags = { "og:title": "Dungeons and Crayons" };
|
||||
metaTags = { "og:title": COMIX_STALE_HOME };
|
||||
docTitle = "Dungeons and Crayons";
|
||||
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
|
||||
assert.equal(p.type, "series");
|
||||
assert.equal(p.site, "comix");
|
||||
@@ -204,8 +217,16 @@ test("comix detects a series page", () => {
|
||||
assert.equal(p.chapterNum, null);
|
||||
});
|
||||
|
||||
test("comix ignores a previous series' stale og:title", () => {
|
||||
metaTags = { "og:title": "Full-Time Awakening" };
|
||||
docTitle = "Dungeons and Crayons";
|
||||
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
|
||||
assert.equal(p.title, "Dungeons and Crayons");
|
||||
});
|
||||
|
||||
test("comix detects a chapter page and strips the Ch. suffix from the title", () => {
|
||||
metaTags = { "og:title": "Dungeons and Crayons · Ch.80" };
|
||||
metaTags = { "og:title": COMIX_STALE_HOME };
|
||||
docTitle = "Dungeons and Crayons · Ch.80";
|
||||
const p = comix.detect(
|
||||
loc("https://comix.to/title/n8we-dungeons-and-crayons/11139891-chapter-80")
|
||||
);
|
||||
@@ -218,7 +239,8 @@ test("comix detects a chapter page and strips the Ch. suffix from the title", ()
|
||||
});
|
||||
|
||||
test("comix parses decimal chapter numbers", () => {
|
||||
metaTags = { "og:title": "Dungeons and Crayons · Ch.80.5" };
|
||||
metaTags = {};
|
||||
docTitle = "Dungeons and Crayons · Ch.80.5";
|
||||
const p = comix.detect(
|
||||
loc("https://comix.to/title/n8we-dungeons-and-crayons/11139891-chapter-80.5")
|
||||
);
|
||||
@@ -226,19 +248,19 @@ test("comix parses decimal chapter numbers", () => {
|
||||
});
|
||||
|
||||
test("comix.detect reads the cover from an img whose alt matches the cleaned title", () => {
|
||||
metaTags = { "og:title": "Dungeons and Crayons · Ch.80" };
|
||||
metaTags = { "og:title": COMIX_STALE_HOME };
|
||||
docTitle = "Dungeons and Crayons";
|
||||
pageImages = [
|
||||
{ alt: "Some Other Series", src: "https://cdn.example/other.jpg" },
|
||||
{ alt: "Dungeons and Crayons", src: "https://cdn.example/cover.jpg" },
|
||||
];
|
||||
const p = comix.detect(
|
||||
loc("https://comix.to/title/n8we-dungeons-and-crayons/11139891-chapter-80")
|
||||
);
|
||||
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
|
||||
assert.equal(p.cover, "https://cdn.example/cover.jpg");
|
||||
});
|
||||
|
||||
test("comix.detect leaves cover empty when no img alt matches the title", () => {
|
||||
metaTags = { "og:title": "Dungeons and Crayons" };
|
||||
metaTags = {};
|
||||
docTitle = "Dungeons and Crayons";
|
||||
pageImages = [{ alt: "Some Other Series", src: "https://cdn.example/other.jpg" }];
|
||||
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
|
||||
assert.equal(p.cover, "");
|
||||
@@ -315,6 +337,33 @@ test("kagane reads the chapter number out of og:title", () => {
|
||||
assert.equal(p.seriesUrl, "https://kagane.to/series/" + KAGANE_SERIES);
|
||||
});
|
||||
|
||||
// Volume-numbered series render the suffix as "- Volume <v> Chapter <n>" with
|
||||
// no episode name, because the book carries volume_no and an empty title.
|
||||
// Captured live 2026-08-08 from SP Baby.
|
||||
test("kagane reads through a Volume-numbered chapter suffix", () => {
|
||||
metaTags = {
|
||||
"og:title": "SP Baby - Volume 1 Chapter 1",
|
||||
"og:image": "https://kagane.to/api/v2/image/abc/compressed",
|
||||
};
|
||||
const p = kagane.detect(
|
||||
loc("https://kagane.to/series/" + KAGANE_SERIES + "/reader/" + KAGANE_BOOK)
|
||||
);
|
||||
assert.equal(p.title, "SP Baby");
|
||||
assert.equal(p.chapterNum, 1);
|
||||
assert.equal(p.chapterLabel, "Chapter 1");
|
||||
});
|
||||
|
||||
// A book with neither a volume nor an episode name ends the title right after
|
||||
// the number, which the old trailing-\s regex could not match.
|
||||
test("kagane reads a chapter suffix with no episode name", () => {
|
||||
metaTags = { "og:title": "Some Series - Chapter 7.5" };
|
||||
const p = kagane.detect(
|
||||
loc("https://kagane.to/series/" + KAGANE_SERIES + "/reader/" + KAGANE_BOOK)
|
||||
);
|
||||
assert.equal(p.title, "Some Series");
|
||||
assert.equal(p.chapterNum, 7.5);
|
||||
});
|
||||
|
||||
test("kagane yields a null chapterNum when og:title has no chapter", () => {
|
||||
metaTags = { "og:title": "Infinite Decryption: The Strongest Level 0" };
|
||||
const p = kagane.detect(
|
||||
|
||||
Reference in New Issue
Block a user