Fix comix titles and covers, kagane volume chapters, and kagane cover rendering #37
@@ -90,3 +90,23 @@ DISCORD_REDIRECT_URI=
|
|||||||
# HTTP handler 500s any /json/version request whose Host header isn't an IP or
|
# HTTP handler 500s any /json/version request whose Host header isn't an IP or
|
||||||
# "localhost", which silently breaks every kagane poll.
|
# "localhost", which silently breaks every kagane poll.
|
||||||
# BROWSER_WS_URL=ws://172.28.0.10:9222
|
# BROWSER_WS_URL=ws://172.28.0.10:9222
|
||||||
|
|
||||||
|
# Clock zone the headless browser reports. A UTC clock is itself the bot
|
||||||
|
# signal — Cloudflare treats it as the datacenter default — and kagane's
|
||||||
|
# challenge then never clears. Measured 2026-08-08, identical container, one
|
||||||
|
# Indonesian egress IP: UTC never cleared in 60s (twice); Asia/Jakarta and
|
||||||
|
# America/New_York both cleared in 4s. So any real zone works; it does not
|
||||||
|
# have to match the IP's country, it just must not be UTC.
|
||||||
|
#
|
||||||
|
# Unset falls back to the host's /etc/timezone, which is a real zone whenever
|
||||||
|
# the host clock is set to local time. Set this when the host runs UTC — a UTC
|
||||||
|
# server is exactly the case that fails. Only the browser sidecar reads it —
|
||||||
|
# the backend's own zone is API_TZ below, and is cosmetic.
|
||||||
|
# BROWSER_TZ=Asia/Jakarta
|
||||||
|
|
||||||
|
# Zone the backend stamps its log lines in. Cosmetic only — it exists so the
|
||||||
|
# API's logs read on the same clock as the browser sidecar's. Nothing else in
|
||||||
|
# the service has a zone: bookmark timestamps are unix ms, and the two real
|
||||||
|
# time columns are timestamptz. Defaults to Asia/Jakarta; set to UTC for the
|
||||||
|
# conventional server default.
|
||||||
|
# API_TZ=Asia/Jakarta
|
||||||
|
|||||||
@@ -20,6 +20,9 @@ Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free
|
|||||||
- Userscript run in **isolated world**, so embedded API token safe from site's JS.
|
- Userscript run in **isolated world**, so embedded API token safe from site's JS.
|
||||||
- Cloudflare's block on manga sites **IP-reputation-based, not universal — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — Cloudflare's bot scoring can flip previously-clean IP without notice. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test.
|
- Cloudflare's block on manga sites **IP-reputation-based, not universal — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — Cloudflare's bot scoring can flip previously-clean IP without notice. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test.
|
||||||
- **kagane.to and novelfull.com are the exception to the above** — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`) and skips them entirely when that's unset. The four other sites poll fine over plain TLS.
|
- **kagane.to and novelfull.com are the exception to the above** — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`) and skips them entirely when that's unset. The four other sites poll fine over plain TLS.
|
||||||
|
- **The CDP sidecar must look like a real browser, and stock headless images don't.** Measured 2026-08-08 against kagane.to, all from the same IP: `chromedp/headless-shell:stable` never cleared the challenge in 90s (`navigator.webdriver` true, empty plugin list, Chromium-branded client hints — suppressing `webdriver` alone changed nothing); `zenika/alpine-chrome` ships Chrome 124, refused outright; real Chrome with the default `--headless=new` UA never cleared, because the UA says `HeadlessChrome`; real Chrome with a stock UA **and** a non-UTC clock zone cleared in ~4s. Hence `chrome/` — a Debian image with `google-chrome-stable`, a version-derived UA, and `TZ`/`BROWSER_TZ`. Chrome reads the zone *name* through ICU from `/etc/localtime`'s symlink target, ignoring the file's contents, so mounting the host's `/etc/localtime` does **not** work; `/etc/timezone` is mounted instead.
|
||||||
|
- **UTC is the tell, not a country mismatch.** A UTC clock is the datacenter default, so Cloudflare scores it as one; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while `Asia/Jakarta` **and** `America/New_York` both cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (`Asia/Bangkok`) rather than the measured egress. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one.
|
||||||
|
- **A challenged page needs the tab kept open.** The interstitial takes seconds to solve and only then writes clearance into the browser's shared cookie jar. Navigate-read-close never clears anything; `BrowserFetcher.run` holds one tab and re-reads until the payload arrives.
|
||||||
|
|
||||||
## Architecture
|
## Architecture
|
||||||
|
|
||||||
@@ -37,7 +40,12 @@ Backend (`cd backend`):
|
|||||||
- Single test: `go test -run TestName ./...`
|
- Single test: `go test -run TestName ./...`
|
||||||
- Build static binary: `CGO_ENABLED=0 go build`
|
- Build static binary: `CGO_ENABLED=0 go build`
|
||||||
|
|
||||||
Local stack: `docker compose up` (bookmark-api + postgres + headless-shell; `postgres-data` named volume, `restart: unless-stopped`).
|
Local stack: `docker compose up` (bookmark-api + postgres + headless-shell; `postgres-data` named volume, `restart: unless-stopped`). The `headless-shell` service keeps its name but now builds `chrome/` — real Google Chrome, for the reason in the hard constraints above.
|
||||||
|
|
||||||
|
Live CDP proof (needs a sidecar and network, skipped otherwise):
|
||||||
|
`SMOKE_BROWSER_WS_URL=ws://<host>:<port> go test -run TestSmokeKagane ./internal/latest`
|
||||||
|
— fetches a real kagane cover and chapter list. A red run means the challenge is
|
||||||
|
not clearing from this IP, which is a live fact to re-check, not necessarily a defect.
|
||||||
|
|
||||||
Smoke test: `curl` endpoints with `Authorization: Bearer <token>`; confirm `OPTIONS` preflight return CORS headers and `/healthz` return 200.
|
Smoke test: `curl` endpoints with `Authorization: Bearer <token>`; confirm `OPTIONS` preflight return CORS headers and `/healthz` return 200.
|
||||||
|
|
||||||
|
|||||||
+14
-3
@@ -126,9 +126,20 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
|
|||||||
`/userscript/novel-bookmark.user.js`, both supplied by bindmount; the
|
`/userscript/novel-bookmark.user.js`, both supplied by bindmount; the
|
||||||
`__API_TOKEN__` placeholder inside them is substituted with the requesting
|
`__API_TOKEN__` placeholder inside them is substituted with the requesting
|
||||||
Reader's credential at serve time).
|
Reader's credential at serve time).
|
||||||
`BROWSER_WS_URL` (headless-shell CDP endpoint for kagane and novelfull;
|
`BROWSER_WS_URL` (CDP endpoint of the `chrome/` sidecar, used by the poller
|
||||||
unset disables browser polling and leaves those sites to the userscript
|
for kagane and novelfull *and* by the web UI's kagane cover proxy; unset
|
||||||
alone).
|
disables browser polling and serves 404 from the proxy, leaving those sites
|
||||||
|
to the userscript alone).
|
||||||
|
- **kagane covers are proxied, not hot-linked:** kagane serves cover images
|
||||||
|
behind the same challenge as its pages and with
|
||||||
|
`cross-origin-resource-policy: same-origin`, so no `<img>` on the web UI's
|
||||||
|
origin can load one — not even from a browser holding the clearance cookie
|
||||||
|
(verified 2026-08-08). `Bookmark.CoverURL` rewrites a stored kagane
|
||||||
|
`og:image` to `/img/kagane/{id}`, served by `internal/web/cover.go` through
|
||||||
|
`latest.BrowserFetcher.Image` and memoised in-process. The templates render
|
||||||
|
`.CoverURL`, never `.Cover`. The id is matched against a UUID regex before it
|
||||||
|
reaches the browser: the stored value is client-supplied, so an unchecked one
|
||||||
|
is an SSRF primitive pointed at the deployment's own network.
|
||||||
- **Web UI also owns:** session-gated `GET /install/{manga,novel}-bookmark.user.js`
|
- **Web UI also owns:** session-gated `GET /install/{manga,novel}-bookmark.user.js`
|
||||||
(renders the bindmounted script with the acting Reader's derived credential
|
(renders the bindmounted script with the acting Reader's derived credential
|
||||||
substituted in — the credential never appears in page markup, the address
|
substituted in — the credential never appears in page markup, the address
|
||||||
|
|||||||
@@ -0,0 +1,155 @@
|
|||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"errors"
|
||||||
|
"net/http"
|
||||||
|
"net/http/httptest"
|
||||||
|
"sync/atomic"
|
||||||
|
"testing"
|
||||||
|
)
|
||||||
|
|
||||||
|
// fakeCovers stands in for the headless browser. It counts calls so the test
|
||||||
|
// can prove the cache spares the browser a second navigation.
|
||||||
|
type fakeCovers struct {
|
||||||
|
body []byte
|
||||||
|
contentType string
|
||||||
|
err error
|
||||||
|
calls atomic.Int32
|
||||||
|
lastID atomic.Value
|
||||||
|
}
|
||||||
|
|
||||||
|
func (f *fakeCovers) Image(_ context.Context, imageID string) ([]byte, string, error) {
|
||||||
|
f.calls.Add(1)
|
||||||
|
f.lastID.Store(imageID)
|
||||||
|
if f.err != nil {
|
||||||
|
return nil, "", f.err
|
||||||
|
}
|
||||||
|
return f.body, f.contentType, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
const testCoverID = "019fe11a-84c3-7fc3-a84b-88787374b617"
|
||||||
|
|
||||||
|
func getCover(t *testing.T, srv http.Handler, path string, cookie *http.Cookie) *httptest.ResponseRecorder {
|
||||||
|
t.Helper()
|
||||||
|
req := httptest.NewRequest(http.MethodGet, path, nil)
|
||||||
|
if cookie != nil {
|
||||||
|
req.AddCookie(cookie)
|
||||||
|
}
|
||||||
|
rr := httptest.NewRecorder()
|
||||||
|
srv.ServeHTTP(rr, req)
|
||||||
|
return rr
|
||||||
|
}
|
||||||
|
|
||||||
|
// kagane serves its covers behind a Cloudflare challenge and with
|
||||||
|
// cross-origin-resource-policy: same-origin, so the UI can only show one by
|
||||||
|
// re-serving the bytes from its own origin.
|
||||||
|
func TestKaganeCoverProxiesAndCaches(t *testing.T) {
|
||||||
|
cf := &fakeCovers{body: []byte("\x00webp-bytes"), contentType: "image/webp"}
|
||||||
|
cfg := testConfig()
|
||||||
|
cfg.Covers = cf
|
||||||
|
srv, st := newWebTestServer(t, cfg)
|
||||||
|
cookie := sessionCookie(t, st)
|
||||||
|
|
||||||
|
for i := range 2 {
|
||||||
|
rr := getCover(t, srv, "/img/kagane/"+testCoverID, cookie)
|
||||||
|
if rr.Code != http.StatusOK {
|
||||||
|
t.Fatalf("request %d: status = %d, want 200", i, rr.Code)
|
||||||
|
}
|
||||||
|
if got := rr.Body.String(); got != string(cf.body) {
|
||||||
|
t.Fatalf("request %d: body = %q, want %q", i, got, cf.body)
|
||||||
|
}
|
||||||
|
if got := rr.Header().Get("Content-Type"); got != "image/webp" {
|
||||||
|
t.Fatalf("request %d: Content-Type = %q, want image/webp", i, got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if got := cf.calls.Load(); got != 1 {
|
||||||
|
t.Fatalf("fetcher called %d times, want 1 — the second read must come from the cache", got)
|
||||||
|
}
|
||||||
|
if got := cf.lastID.Load(); got != testCoverID {
|
||||||
|
t.Fatalf("fetched image id = %v, want %s", got, testCoverID)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The proxy reaches a headless browser, so it is not open to the internet.
|
||||||
|
func TestKaganeCoverRequiresSession(t *testing.T) {
|
||||||
|
cf := &fakeCovers{body: []byte("x"), contentType: "image/webp"}
|
||||||
|
cfg := testConfig()
|
||||||
|
cfg.Covers = cf
|
||||||
|
srv, _ := newWebTestServer(t, cfg)
|
||||||
|
|
||||||
|
rr := getCover(t, srv, "/img/kagane/"+testCoverID, nil)
|
||||||
|
if rr.Code != http.StatusUnauthorized {
|
||||||
|
t.Fatalf("status = %d, want 401", rr.Code)
|
||||||
|
}
|
||||||
|
if got := cf.calls.Load(); got != 0 {
|
||||||
|
t.Fatalf("fetcher called %d times for an unauthenticated request, want 0", got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestKaganeCoverRejectsBadInput(t *testing.T) {
|
||||||
|
cases := []struct {
|
||||||
|
name string
|
||||||
|
id string
|
||||||
|
fetch *fakeCovers
|
||||||
|
}{
|
||||||
|
{
|
||||||
|
"an id that is not a uuid never reaches the browser",
|
||||||
|
"solo-leveling",
|
||||||
|
&fakeCovers{body: []byte("x"), contentType: "image/webp"},
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"a uuid-shaped id with a trailing segment is rejected whole",
|
||||||
|
testCoverID + "x",
|
||||||
|
&fakeCovers{body: []byte("x"), contentType: "image/webp"},
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"a challenged fetch is a missing cover",
|
||||||
|
testCoverID,
|
||||||
|
&fakeCovers{err: errors.New("challenge held")},
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"a content type outside the image set is not echoed back",
|
||||||
|
testCoverID,
|
||||||
|
&fakeCovers{body: []byte("<script>"), contentType: "text/html"},
|
||||||
|
},
|
||||||
|
}
|
||||||
|
for _, tc := range cases {
|
||||||
|
t.Run(tc.name, func(t *testing.T) {
|
||||||
|
cfg := testConfig()
|
||||||
|
cfg.Covers = tc.fetch
|
||||||
|
srv, st := newWebTestServer(t, cfg)
|
||||||
|
rr := getCover(t, srv, "/img/kagane/"+tc.id, sessionCookie(t, st))
|
||||||
|
if rr.Code != http.StatusNotFound {
|
||||||
|
t.Fatalf("status = %d, want 404", rr.Code)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// ServeMux path-cleans a traversal into a redirect before the handler runs, so
|
||||||
|
// the guarantee to pin down is that no request shaped like one ever gets bytes.
|
||||||
|
func TestKaganeCoverTraversalServesNothing(t *testing.T) {
|
||||||
|
cf := &fakeCovers{body: []byte("secret"), contentType: "image/webp"}
|
||||||
|
cfg := testConfig()
|
||||||
|
cfg.Covers = cf
|
||||||
|
srv, st := newWebTestServer(t, cfg)
|
||||||
|
|
||||||
|
rr := getCover(t, srv, "/img/kagane/../../etc/passwd", sessionCookie(t, st))
|
||||||
|
if rr.Code == http.StatusOK {
|
||||||
|
t.Fatalf("status = 200, want anything but a served body")
|
||||||
|
}
|
||||||
|
if got := cf.calls.Load(); got != 0 {
|
||||||
|
t.Fatalf("fetcher called %d times for a traversal, want 0", got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Without BROWSER_WS_URL there is no fetcher, and the endpoint must answer
|
||||||
|
// rather than reach for a nil one.
|
||||||
|
func TestKaganeCoverWithoutFetcher(t *testing.T) {
|
||||||
|
srv, st := newWebTestServer(t, testConfig())
|
||||||
|
rr := getCover(t, srv, "/img/kagane/"+testCoverID, sessionCookie(t, st))
|
||||||
|
if rr.Code != http.StatusNotFound {
|
||||||
|
t.Fatalf("status = %d, want 404", rr.Code)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -2,7 +2,9 @@ package latest
|
|||||||
|
|
||||||
import (
|
import (
|
||||||
"context"
|
"context"
|
||||||
|
"encoding/base64"
|
||||||
"encoding/json"
|
"encoding/json"
|
||||||
|
"errors"
|
||||||
"fmt"
|
"fmt"
|
||||||
"net/url"
|
"net/url"
|
||||||
"regexp"
|
"regexp"
|
||||||
@@ -22,6 +24,12 @@ const challengeTimeout = 45 * time.Second
|
|||||||
|
|
||||||
var kaganeSeriesRe = regexp.MustCompile(`^/series/([0-9a-f-]{36})/?$`)
|
var kaganeSeriesRe = regexp.MustCompile(`^/series/([0-9a-f-]{36})/?$`)
|
||||||
|
|
||||||
|
// kaganeImageIDRe pins the only path segment Image interpolates into an
|
||||||
|
// outbound URL. The id arrives from a stored cover URL, which a client
|
||||||
|
// supplied, so it is matched rather than trusted: a headless browser is a
|
||||||
|
// strong SSRF primitive.
|
||||||
|
var kaganeImageIDRe = regexp.MustCompile(`^[0-9a-f-]{36}$`)
|
||||||
|
|
||||||
// BrowserFetcher retrieves pages through a remote headless Chrome over the
|
// BrowserFetcher retrieves pages through a remote headless Chrome over the
|
||||||
// DevTools Protocol.
|
// DevTools Protocol.
|
||||||
//
|
//
|
||||||
@@ -91,23 +99,6 @@ func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int
|
|||||||
return "", 0, fmt.Errorf("not a fetchable browser series url: %q", seriesURL)
|
return "", 0, fmt.Errorf("not a fetchable browser series url: %q", seriesURL)
|
||||||
}
|
}
|
||||||
|
|
||||||
f.mu.Lock()
|
|
||||||
defer f.mu.Unlock()
|
|
||||||
|
|
||||||
ctx, cancel := context.WithTimeout(ctx, challengeTimeout)
|
|
||||||
defer cancel()
|
|
||||||
// A fresh tab per fetch, closed on return, so one wedged page cannot
|
|
||||||
// poison later polls.
|
|
||||||
tabCtx, cancelTab := chromedp.NewContext(f.allocCtx)
|
|
||||||
defer cancelTab()
|
|
||||||
// Bind the caller's deadline to the tab.
|
|
||||||
tabCtx, cancelDeadline := context.WithCancel(tabCtx)
|
|
||||||
defer cancelDeadline()
|
|
||||||
go func() {
|
|
||||||
<-ctx.Done()
|
|
||||||
cancelDeadline()
|
|
||||||
}()
|
|
||||||
|
|
||||||
var body string
|
var body string
|
||||||
// kagane's chapter list is only in its JSON API, which must be called from
|
// kagane's chapter list is only in its JSON API, which must be called from
|
||||||
// inside the page so the request carries the clearance cookie. novelfull
|
// inside the page so the request carries the clearance cookie. novelfull
|
||||||
@@ -123,24 +114,134 @@ func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int
|
|||||||
)
|
)
|
||||||
}
|
}
|
||||||
|
|
||||||
err := chromedp.Run(tabCtx,
|
// novelfull's payload is the DOM itself, and the interstitial has a DOM
|
||||||
chromedp.Navigate(seriesURL),
|
// too, so "we have an answer" has to exclude it explicitly. kagane's
|
||||||
// The challenge reloads the page itself when it passes; waiting for the
|
// in-page fetch just fails while challenged, which is already the signal.
|
||||||
// site's own root element is what tells us we are through it.
|
done := func() bool { return body != "" && (isKagane || !isInterstitial(body)) }
|
||||||
chromedp.WaitReady("body", chromedp.ByQuery),
|
if err := f.run(ctx, seriesURL, read, done); err != nil {
|
||||||
read,
|
// Challenge never cleared, or the API refused. Indistinguishable from
|
||||||
)
|
// here and handled identically by the caller.
|
||||||
if err != nil {
|
if errors.Is(err, errChallengeHeld) {
|
||||||
return "", 0, fmt.Errorf("browser fetch %q: %w", seriesURL, err)
|
|
||||||
}
|
|
||||||
if body == "" {
|
|
||||||
// Challenge still up, or the API refused. Indistinguishable from here
|
|
||||||
// and handled identically by the caller.
|
|
||||||
return "", 403, nil
|
return "", 403, nil
|
||||||
}
|
}
|
||||||
|
return "", 0, fmt.Errorf("browser fetch %q: %w", seriesURL, err)
|
||||||
|
}
|
||||||
return body, 200, nil
|
return body, 200, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Image retrieves one kagane cover as raw bytes and its content type.
|
||||||
|
//
|
||||||
|
// It exists because kagane serves covers behind the same challenge as its
|
||||||
|
// pages *and* with `cross-origin-resource-policy: same-origin`, so an <img> on
|
||||||
|
// the web UI's origin cannot load one even from a browser that already holds
|
||||||
|
// the clearance cookie (verified 2026-08-08). Proxying is the only route.
|
||||||
|
//
|
||||||
|
// The image URL is navigated to rather than fetched from some other kagane
|
||||||
|
// page: the challenge only runs on a top-level navigation, and once it clears
|
||||||
|
// the document *is* the image, so a same-origin fetch of location.href reads
|
||||||
|
// it straight back out of the cache.
|
||||||
|
//
|
||||||
|
// The challenge is not solved by the first read: WaitReady("body") is satisfied
|
||||||
|
// by the interstitial too. run holds the tab open until the in-page fetch
|
||||||
|
// succeeds, which is what gives the challenge script the seconds it needs.
|
||||||
|
func (f *BrowserFetcher) Image(ctx context.Context, imageID string) ([]byte, string, error) {
|
||||||
|
if !kaganeImageIDRe.MatchString(imageID) {
|
||||||
|
return nil, "", fmt.Errorf("not a kagane image id: %q", imageID)
|
||||||
|
}
|
||||||
|
var dataURL string
|
||||||
|
err := f.run(ctx, "https://kagane.to/api/v2/image/"+imageID+"/compressed",
|
||||||
|
chromedp.Evaluate(`fetch(location.href).then(r => r.ok
|
||||||
|
? r.blob().then(b => new Promise(res => {
|
||||||
|
const fr = new FileReader();
|
||||||
|
fr.onload = () => res(fr.result);
|
||||||
|
fr.readAsDataURL(b);
|
||||||
|
}))
|
||||||
|
: "")`, &dataURL, awaitPromise),
|
||||||
|
func() bool { return dataURL != "" })
|
||||||
|
if err != nil {
|
||||||
|
return nil, "", fmt.Errorf("browser image %s: %w", imageID, err)
|
||||||
|
}
|
||||||
|
// "data:image/webp;base64,<payload>".
|
||||||
|
head, payload, ok := strings.Cut(dataURL, ";base64,")
|
||||||
|
if !ok {
|
||||||
|
return nil, "", fmt.Errorf("browser image %s: not a data url", imageID)
|
||||||
|
}
|
||||||
|
raw, err := base64.StdEncoding.DecodeString(payload)
|
||||||
|
if err != nil {
|
||||||
|
return nil, "", fmt.Errorf("browser image %s: %w", imageID, err)
|
||||||
|
}
|
||||||
|
return raw, strings.TrimPrefix(head, "data:"), nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// errChallengeHeld reports that the budget ran out with the interstitial still
|
||||||
|
// up. Distinct from a transport failure: it means "this site said no", which
|
||||||
|
// the poller answers with a 403 and its ordinary cooldown.
|
||||||
|
var errChallengeHeld = errors.New("challenge held")
|
||||||
|
|
||||||
|
// challengePollInterval paces re-reads while a challenge solves itself.
|
||||||
|
const challengePollInterval = 2 * time.Second
|
||||||
|
|
||||||
|
// isInterstitial reports whether html is Cloudflare's challenge page rather
|
||||||
|
// than the site's own. Matched on the challenge runtime's script path, which is
|
||||||
|
// stable across the interstitial's wording and locale — the visible "Just a
|
||||||
|
// moment..." title is neither.
|
||||||
|
func isInterstitial(html string) bool {
|
||||||
|
return strings.Contains(html, "/cdn-cgi/challenge-platform/")
|
||||||
|
}
|
||||||
|
|
||||||
|
// run navigates to target and re-reads until done reports an answer, bounded by
|
||||||
|
// challengeTimeout and by the caller's own deadline, in a tab that is closed on
|
||||||
|
// return so one wedged page cannot poison later calls.
|
||||||
|
//
|
||||||
|
// Holding the tab open across re-reads is the whole point. A Cloudflare
|
||||||
|
// interstitial needs several seconds of a live page to solve itself and write
|
||||||
|
// clearance into the browser's shared cookie jar; reading once and closing the
|
||||||
|
// tab — which is what this did before 2026-08-08 — never gives it that window,
|
||||||
|
// so every fetch lands on the interstitial and the clearance that would have
|
||||||
|
// unblocked all the later ones is never obtained.
|
||||||
|
func (f *BrowserFetcher) run(ctx context.Context, target string, read chromedp.Action, done func() bool) error {
|
||||||
|
f.mu.Lock()
|
||||||
|
defer f.mu.Unlock()
|
||||||
|
|
||||||
|
ctx, cancel := context.WithTimeout(ctx, challengeTimeout)
|
||||||
|
defer cancel()
|
||||||
|
tabCtx, cancelTab := chromedp.NewContext(f.allocCtx)
|
||||||
|
defer cancelTab()
|
||||||
|
// Bind the caller's deadline to the tab.
|
||||||
|
tabCtx, cancelDeadline := context.WithCancel(tabCtx)
|
||||||
|
defer cancelDeadline()
|
||||||
|
go func() {
|
||||||
|
<-ctx.Done()
|
||||||
|
cancelDeadline()
|
||||||
|
}()
|
||||||
|
|
||||||
|
if err := chromedp.Run(tabCtx,
|
||||||
|
chromedp.Navigate(target),
|
||||||
|
chromedp.WaitReady("body", chromedp.ByQuery),
|
||||||
|
); err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
|
||||||
|
var lastErr error
|
||||||
|
for {
|
||||||
|
// The challenge reloads the page when it passes, which tears down the
|
||||||
|
// execution context mid-read. That is a retry, not a failure.
|
||||||
|
if err := chromedp.Run(tabCtx, read); err != nil {
|
||||||
|
lastErr = err
|
||||||
|
} else if done() {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
select {
|
||||||
|
case <-ctx.Done():
|
||||||
|
if lastErr != nil {
|
||||||
|
return fmt.Errorf("%w (last read: %v)", errChallengeHeld, lastErr)
|
||||||
|
}
|
||||||
|
return errChallengeHeld
|
||||||
|
case <-time.After(challengePollInterval):
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
// kaganeAPIURL maps a stored series_url to the JSON endpoint carrying its
|
// kaganeAPIURL maps a stored series_url to the JSON endpoint carrying its
|
||||||
// chapter list. Returning false for anything else is a second line of defence
|
// chapter list. Returning false for anything else is a second line of defence
|
||||||
// behind fetchableSeriesURL: a headless browser is a strong SSRF primitive and
|
// behind fetchableSeriesURL: a headless browser is a strong SSRF primitive and
|
||||||
|
|||||||
@@ -0,0 +1,91 @@
|
|||||||
|
package latest
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"net/http"
|
||||||
|
"os"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
// TestSmokeKaganeImage is the live proof that the cover proxy's fetch actually
|
||||||
|
// clears Cloudflare and returns image bytes. It needs a real headless Chrome
|
||||||
|
// with outbound network, so it runs only when SMOKE_BROWSER_WS_URL is set:
|
||||||
|
//
|
||||||
|
// docker run --rm --shm-size=1gb -p 19222:9222 chromedp/headless-shell:stable
|
||||||
|
// SMOKE_BROWSER_WS_URL=ws://127.0.0.1:19222 go test -run TestSmokeKaganeImage ./internal/latest
|
||||||
|
func TestSmokeKaganeImage(t *testing.T) {
|
||||||
|
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
|
||||||
|
if ws == "" {
|
||||||
|
t.Skip("SMOKE_BROWSER_WS_URL unset")
|
||||||
|
}
|
||||||
|
const imageID = "019fe11a-84c3-7fc3-a84b-88787374b617" // SP Baby's cover
|
||||||
|
|
||||||
|
// The same URL through a plain client is what the web UI's <img> gets.
|
||||||
|
// Asserting on it keeps the test honest about why the browser is needed.
|
||||||
|
req, err := http.NewRequest(http.MethodGet,
|
||||||
|
"https://kagane.to/api/v2/image/"+imageID+"/compressed", nil)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
if res, err := (&http.Client{Timeout: 15 * time.Second}).Do(req); err == nil {
|
||||||
|
res.Body.Close()
|
||||||
|
if res.StatusCode == http.StatusOK {
|
||||||
|
t.Log("note: kagane answered a plain request 200 — the challenge is not up right now")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
f, err := NewBrowserFetcher(ws)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("NewBrowserFetcher: %v", err)
|
||||||
|
}
|
||||||
|
defer f.Close()
|
||||||
|
|
||||||
|
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
|
||||||
|
defer cancel()
|
||||||
|
body, contentType, err := f.Image(ctx, imageID)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("Image: %v", err)
|
||||||
|
}
|
||||||
|
if len(body) < 1000 {
|
||||||
|
t.Fatalf("body is %d bytes, want a real image", len(body))
|
||||||
|
}
|
||||||
|
if contentType != "image/webp" {
|
||||||
|
t.Fatalf("content type = %q, want image/webp", contentType)
|
||||||
|
}
|
||||||
|
// WebP files start with "RIFF....WEBP".
|
||||||
|
if string(body[:4]) != "RIFF" || string(body[8:12]) != "WEBP" {
|
||||||
|
t.Fatalf("body is not a WebP: % x", body[:12])
|
||||||
|
}
|
||||||
|
t.Logf("fetched %d bytes of %s", len(body), contentType)
|
||||||
|
|
||||||
|
if _, _, err := f.Image(ctx, "not-a-uuid"); err == nil {
|
||||||
|
t.Fatal("Image accepted a non-uuid id")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Control for the test above: the poller's own kagane path, same sidecar. If
|
||||||
|
// this fails too, the sidecar is not clearing the challenge at all and the
|
||||||
|
// image result says nothing about Image itself.
|
||||||
|
func TestSmokeKaganeGet(t *testing.T) {
|
||||||
|
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
|
||||||
|
if ws == "" {
|
||||||
|
t.Skip("SMOKE_BROWSER_WS_URL unset")
|
||||||
|
}
|
||||||
|
f, err := NewBrowserFetcher(ws)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("NewBrowserFetcher: %v", err)
|
||||||
|
}
|
||||||
|
defer f.Close()
|
||||||
|
|
||||||
|
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
|
||||||
|
defer cancel()
|
||||||
|
body, status, err := f.Get(ctx, "https://kagane.to/series/019fe11a-8670-7cf3-8343-0b02057d3787")
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("Get: %v", err)
|
||||||
|
}
|
||||||
|
t.Logf("status=%d bytes=%d head=%.80q", status, len(body), body)
|
||||||
|
if status != 200 {
|
||||||
|
t.Fatalf("status = %d, want 200 — the sidecar is not clearing the challenge", status)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -137,6 +137,22 @@ func (b Bookmark) Initial() string {
|
|||||||
return "?"
|
return "?"
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// kaganeCoverRe matches the cover URL kagane's og:image carries, which is what
|
||||||
|
// the userscript stores for that site.
|
||||||
|
var kaganeCoverRe = regexp.MustCompile(`^https://kagane\.to/api/v2/image/([0-9a-f-]{36})/compressed$`)
|
||||||
|
|
||||||
|
// CoverURL is the src the web UI puts in an <img>. For every site but kagane
|
||||||
|
// that is Cover as stored. kagane serves its images behind a Cloudflare
|
||||||
|
// challenge *and* with `cross-origin-resource-policy: same-origin`, so no page
|
||||||
|
// on another origin can load one however it asks (verified 2026-08-08); those
|
||||||
|
// go through the backend's own proxy instead.
|
||||||
|
func (b Bookmark) CoverURL() string {
|
||||||
|
if m := kaganeCoverRe.FindStringSubmatch(b.Cover); m != nil {
|
||||||
|
return "/img/kagane/" + m[1]
|
||||||
|
}
|
||||||
|
return b.Cover
|
||||||
|
}
|
||||||
|
|
||||||
// Library buckets. A bookmark is in exactly one. This cannot be derived from
|
// Library buckets. A bookmark is in exactly one. This cannot be derived from
|
||||||
// Site: asurascans serves manga and novels from the same /comics/ path, so the
|
// Site: asurascans serves manga and novels from the same /comics/ path, so the
|
||||||
// userscript that recorded the page is the only party that knows which.
|
// userscript that recorded the page is the only party that knows which.
|
||||||
|
|||||||
@@ -543,6 +543,38 @@ func TestDisplayChapter(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func TestCoverURL(t *testing.T) {
|
||||||
|
cases := []struct {
|
||||||
|
name string
|
||||||
|
cover string
|
||||||
|
want string
|
||||||
|
}{
|
||||||
|
{
|
||||||
|
"kagane routes through the proxy",
|
||||||
|
"https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
|
||||||
|
"/img/kagane/019fe11a-84c3-7fc3-a84b-88787374b617",
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"another site is served as stored",
|
||||||
|
"https://gg.asuracomic.net/storage/media/1/conversions/cover.webp",
|
||||||
|
"https://gg.asuracomic.net/storage/media/1/conversions/cover.webp",
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"a lookalike host is not rewritten",
|
||||||
|
"https://evil.example/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
|
||||||
|
"https://evil.example/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed",
|
||||||
|
},
|
||||||
|
{"no cover stays empty", "", ""},
|
||||||
|
}
|
||||||
|
for _, tc := range cases {
|
||||||
|
t.Run(tc.name, func(t *testing.T) {
|
||||||
|
if got := (Bookmark{Cover: tc.cover}).CoverURL(); got != tc.want {
|
||||||
|
t.Errorf("CoverURL() = %q, want %q", got, tc.want)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
func TestUpsertKindDefaultsToManga(t *testing.T) {
|
func TestUpsertKindDefaultsToManga(t *testing.T) {
|
||||||
store := newTestStore(t)
|
store := newTestStore(t)
|
||||||
got, err := store.Upsert(store.OwnerID(), Bookmark{
|
got, err := store.Upsert(store.OwnerID(), Bookmark{
|
||||||
|
|||||||
@@ -0,0 +1,129 @@
|
|||||||
|
package web
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"log"
|
||||||
|
"net/http"
|
||||||
|
"regexp"
|
||||||
|
"sync"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
// CoverFetcher retrieves one kagane cover by image id. Satisfied by
|
||||||
|
// latest.BrowserFetcher, and nil when BROWSER_WS_URL is unset — which leaves
|
||||||
|
// kagane covers exactly as unavailable as they were before this endpoint
|
||||||
|
// existed, rather than hanging a request on a fetcher that cannot run.
|
||||||
|
type CoverFetcher interface {
|
||||||
|
Image(ctx context.Context, imageID string) (body []byte, contentType string, err error)
|
||||||
|
}
|
||||||
|
|
||||||
|
// coverIDRe matches the request path segment that becomes part of an outbound
|
||||||
|
// URL. The proxy is session-gated, but the id still reaches a headless browser,
|
||||||
|
// so it is validated at the boundary rather than passed through.
|
||||||
|
var coverIDRe = regexp.MustCompile(`^[0-9a-f-]{36}$`)
|
||||||
|
|
||||||
|
// coverTypes is the set of content types the proxy will echo back. A response
|
||||||
|
// header sourced from a third party is not repeated verbatim: anything outside
|
||||||
|
// this set is treated as "not a cover".
|
||||||
|
var coverTypes = map[string]bool{
|
||||||
|
"image/webp": true,
|
||||||
|
"image/jpeg": true,
|
||||||
|
"image/png": true,
|
||||||
|
"image/avif": true,
|
||||||
|
"image/gif": true,
|
||||||
|
}
|
||||||
|
|
||||||
|
// coverTimeout bounds one proxied cover. Shorter than the fetcher's own
|
||||||
|
// challenge budget on purpose: a browser page is waiting on this, and a cover
|
||||||
|
// that has not arrived by now is better left as a broken slot than as a request
|
||||||
|
// holding a connection open.
|
||||||
|
const coverTimeout = 20 * time.Second
|
||||||
|
|
||||||
|
// coverCacheMax caps the in-memory cover cache. Covers are immutable per image
|
||||||
|
// id and a library holds tens of series, so this is a ceiling that is never
|
||||||
|
// reached in practice; reaching it clears the map rather than evicting by age.
|
||||||
|
//
|
||||||
|
// ponytail: flush-on-full, not LRU. Swap it for an LRU if a library ever grows
|
||||||
|
// past this and the flush starts costing refetches.
|
||||||
|
const coverCacheMax = 500
|
||||||
|
|
||||||
|
type cachedCover struct {
|
||||||
|
body []byte
|
||||||
|
contentType string
|
||||||
|
}
|
||||||
|
|
||||||
|
type coverCache struct {
|
||||||
|
mu sync.Mutex
|
||||||
|
m map[string]cachedCover
|
||||||
|
}
|
||||||
|
|
||||||
|
func (c *coverCache) get(id string) (cachedCover, bool) {
|
||||||
|
c.mu.Lock()
|
||||||
|
defer c.mu.Unlock()
|
||||||
|
v, ok := c.m[id]
|
||||||
|
return v, ok
|
||||||
|
}
|
||||||
|
|
||||||
|
func (c *coverCache) put(id string, v cachedCover) {
|
||||||
|
c.mu.Lock()
|
||||||
|
defer c.mu.Unlock()
|
||||||
|
if c.m == nil || len(c.m) >= coverCacheMax {
|
||||||
|
c.m = make(map[string]cachedCover, coverCacheMax)
|
||||||
|
}
|
||||||
|
c.m[id] = v
|
||||||
|
}
|
||||||
|
|
||||||
|
// kaganeCover serves a kagane cover from the backend's own origin.
|
||||||
|
//
|
||||||
|
// kagane answers image requests with a Cloudflare challenge and
|
||||||
|
// `cross-origin-resource-policy: same-origin`, so the web UI cannot render one
|
||||||
|
// directly under any combination of referrer policy or crossorigin attribute
|
||||||
|
// (verified 2026-08-08). Fetching it through the headless browser that already
|
||||||
|
// clears the challenge, and re-serving it here, is what puts the bytes on an
|
||||||
|
// origin the page may load from.
|
||||||
|
//
|
||||||
|
// ponytail: covers are fetched on first view, one browser navigation at a time
|
||||||
|
// behind the fetcher's mutex, so a first load of a large kagane library
|
||||||
|
// trickles in over a few seconds. The cache makes it a one-off. Prefetching
|
||||||
|
// during the poll cycle is the upgrade if that ever grates.
|
||||||
|
func (h *Handler) kaganeCover(w http.ResponseWriter, r *http.Request) {
|
||||||
|
id := r.PathValue("id")
|
||||||
|
if !coverIDRe.MatchString(id) {
|
||||||
|
http.NotFound(w, r)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
if h.covers == nil {
|
||||||
|
http.NotFound(w, r)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
if v, ok := h.coverCache.get(id); ok {
|
||||||
|
writeCover(w, v)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
|
||||||
|
ctx, cancel := context.WithTimeout(r.Context(), coverTimeout)
|
||||||
|
defer cancel()
|
||||||
|
body, contentType, err := h.covers.Image(ctx, id)
|
||||||
|
if err != nil {
|
||||||
|
log.Printf("kagane cover %s: %v", id, err)
|
||||||
|
http.NotFound(w, r)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
if !coverTypes[contentType] {
|
||||||
|
log.Printf("kagane cover %s: unexpected content type %q", id, contentType)
|
||||||
|
http.NotFound(w, r)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
|
||||||
|
v := cachedCover{body: body, contentType: contentType}
|
||||||
|
h.coverCache.put(id, v)
|
||||||
|
writeCover(w, v)
|
||||||
|
}
|
||||||
|
|
||||||
|
// writeCover sends the bytes with a long cache life: an image id names one
|
||||||
|
// immutable rendering, so a client that has it never needs to ask again.
|
||||||
|
func writeCover(w http.ResponseWriter, v cachedCover) {
|
||||||
|
w.Header().Set("Content-Type", v.contentType)
|
||||||
|
w.Header().Set("Cache-Control", "private, max-age=604800, immutable")
|
||||||
|
w.Write(v.body)
|
||||||
|
}
|
||||||
@@ -6,7 +6,7 @@
|
|||||||
<div class="row">
|
<div class="row">
|
||||||
<a class="cover" href="{{.ContinueURL}}" target="_blank" rel="noopener noreferrer"
|
<a class="cover" href="{{.ContinueURL}}" target="_blank" rel="noopener noreferrer"
|
||||||
tabindex="-1" aria-hidden="true">
|
tabindex="-1" aria-hidden="true">
|
||||||
{{if .Cover}}<img src="{{.Cover}}" alt="" loading="lazy">
|
{{if .CoverURL}}<img src="{{.CoverURL}}" alt="" loading="lazy">
|
||||||
{{/* aria-hidden on the cover link is not enough — Chromium still exposes
|
{{/* aria-hidden on the cover link is not enough — Chromium still exposes
|
||||||
the letter because the link is programmatically focusable — so the
|
the letter because the link is programmatically focusable — so the
|
||||||
monogram carries its own, same as the recent strip's. */}}
|
monogram carries its own, same as the recent strip's. */}}
|
||||||
|
|||||||
@@ -14,7 +14,7 @@
|
|||||||
<a class="recent-card {{if .HasNewChapter}}is-new{{end}}" href="{{.ContinueURL}}"
|
<a class="recent-card {{if .HasNewChapter}}is-new{{end}}" href="{{.ContinueURL}}"
|
||||||
target="_blank" rel="noopener noreferrer">
|
target="_blank" rel="noopener noreferrer">
|
||||||
<span class="recent-cover">
|
<span class="recent-cover">
|
||||||
{{if .Cover}}<img src="{{.Cover}}" alt="" loading="lazy">
|
{{if .CoverURL}}<img src="{{.CoverURL}}" alt="" loading="lazy">
|
||||||
{{else}}<span class="monogram" aria-hidden="true">{{.Initial}}</span>{{end}}
|
{{else}}<span class="monogram" aria-hidden="true">{{.Initial}}</span>{{end}}
|
||||||
{{if .HasNewChapter}}<span class="foot-rule"></span>
|
{{if .HasNewChapter}}<span class="foot-rule"></span>
|
||||||
{{else if .Favorite}}<span class="foot-rule brass"></span>{{end}}
|
{{else if .Favorite}}<span class="foot-rule brass"></span>{{end}}
|
||||||
|
|||||||
@@ -50,6 +50,10 @@ type Handler struct {
|
|||||||
// httpClient is the plain stdlib client that talks to Discord. It is not
|
// httpClient is the plain stdlib client that talks to Discord. It is not
|
||||||
// an injected interface: tests point APIBase at a stub server instead.
|
// an injected interface: tests point APIBase at a stub server instead.
|
||||||
httpClient *http.Client
|
httpClient *http.Client
|
||||||
|
// covers proxies kagane cover images, which no browser can load directly.
|
||||||
|
// Nil disables the endpoint — see CoverFetcher.
|
||||||
|
covers CoverFetcher
|
||||||
|
coverCache coverCache
|
||||||
}
|
}
|
||||||
|
|
||||||
// listView is what every list-rendering template receives.
|
// listView is what every list-rendering template receives.
|
||||||
@@ -111,7 +115,7 @@ type loginView struct {
|
|||||||
|
|
||||||
// New parses every template up front so a broken one kills the process at
|
// New parses every template up front so a broken one kills the process at
|
||||||
// startup rather than the first request that touches it.
|
// startup rather than the first request that touches it.
|
||||||
func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, novelPath string) (*Handler, error) {
|
func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, novelPath string, covers CoverFetcher) (*Handler, error) {
|
||||||
tmpl, err := template.ParseFS(templateFS, "templates/*.html")
|
tmpl, err := template.ParseFS(templateFS, "templates/*.html")
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return nil, err
|
return nil, err
|
||||||
@@ -126,6 +130,7 @@ func New(s *store.Store, discord DiscordConfig, tokenKey []byte, mangaPath, nove
|
|||||||
states: newOAuthStates(),
|
states: newOAuthStates(),
|
||||||
limiter: session.NewLoginLimiter(),
|
limiter: session.NewLoginLimiter(),
|
||||||
httpClient: &http.Client{Timeout: discordTimeout},
|
httpClient: &http.Client{Timeout: discordTimeout},
|
||||||
|
covers: covers,
|
||||||
}, nil
|
}, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -142,6 +147,10 @@ func (h *Handler) Register(mux *http.ServeMux) {
|
|||||||
mux.HandleFunc("POST /ui/bookmarks/{key}/chapter", h.requireSession(h.uiChapter))
|
mux.HandleFunc("POST /ui/bookmarks/{key}/chapter", h.requireSession(h.uiChapter))
|
||||||
mux.HandleFunc("DELETE /ui/bookmarks/{key}", h.requireSession(h.uiDelete))
|
mux.HandleFunc("DELETE /ui/bookmarks/{key}", h.requireSession(h.uiDelete))
|
||||||
|
|
||||||
|
// Session-gated like every other UI route: the deployment proxies kagane's
|
||||||
|
// images for its own Readers, not for the internet.
|
||||||
|
mux.HandleFunc("GET /img/kagane/{id}", h.requireSession(h.kaganeCover))
|
||||||
|
|
||||||
// Install endpoints render the script directly under the session: the
|
// Install endpoints render the script directly under the session: the
|
||||||
// credential travels inside the served bytes, never in the address bar or
|
// credential travels inside the served bytes, never in the address bar or
|
||||||
// the page markup. Updates after install use the credential-bearing /u/
|
// the page markup. Updates after install use the credential-bearing /u/
|
||||||
|
|||||||
+28
-17
@@ -47,6 +47,10 @@ type Config struct {
|
|||||||
NovelUserscriptPath string
|
NovelUserscriptPath string
|
||||||
// LatestPoll configures the background latest-chapter fetcher.
|
// LatestPoll configures the background latest-chapter fetcher.
|
||||||
LatestPoll LatestPoll
|
LatestPoll LatestPoll
|
||||||
|
// Covers proxies kagane cover images for the web UI. Not from the
|
||||||
|
// environment: it is the shared headless browser, wired in main once it
|
||||||
|
// connects, and nil in every test router.
|
||||||
|
Covers web.CoverFetcher
|
||||||
}
|
}
|
||||||
|
|
||||||
// LatestPoll configures the background latest-chapter poller.
|
// LatestPoll configures the background latest-chapter poller.
|
||||||
@@ -206,7 +210,7 @@ func newRouter(s *store.Store, cfg Config) http.Handler {
|
|||||||
// The browser UI is always registered; signing in is Discord OAuth, so
|
// The browser UI is always registered; signing in is Discord OAuth, so
|
||||||
// there is no password to forget and no gate to leave unset.
|
// there is no password to forget and no gate to leave unset.
|
||||||
wh, err := web.New(s, cfg.Discord, []byte(cfg.TokenKey),
|
wh, err := web.New(s, cfg.Discord, []byte(cfg.TokenKey),
|
||||||
cfg.UserscriptPath, cfg.NovelUserscriptPath)
|
cfg.UserscriptPath, cfg.NovelUserscriptPath, cfg.Covers)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
log.Fatalf("web handler: %v", err)
|
log.Fatalf("web handler: %v", err)
|
||||||
}
|
}
|
||||||
@@ -270,9 +274,26 @@ func main() {
|
|||||||
// The poller is off the request path entirely: if it cannot start, the
|
// The poller is off the request path entirely: if it cannot start, the
|
||||||
// service still serves bookmarks and the userscript still captures latest
|
// service still serves bookmarks and the userscript still captures latest
|
||||||
// chapters on its own.
|
// chapters on its own.
|
||||||
|
//
|
||||||
|
// One headless browser serves both consumers that need a Cloudflare
|
||||||
|
// challenge cleared: the poller's kagane/novelfull fetches and the web
|
||||||
|
// UI's kagane cover proxy. Optional — unset leaves both degraded to what
|
||||||
|
// they were before the sidecar existed.
|
||||||
|
var browser latest.Fetcher
|
||||||
pollCtx, stopPoll := context.WithCancel(context.Background())
|
pollCtx, stopPoll := context.WithCancel(context.Background())
|
||||||
defer stopPoll()
|
defer stopPoll()
|
||||||
startLatestPoller(pollCtx, s, cfg.LatestPoll)
|
if ws := strings.TrimSpace(os.Getenv("BROWSER_WS_URL")); ws != "" {
|
||||||
|
bf, err := latest.NewBrowserFetcher(ws)
|
||||||
|
if err != nil {
|
||||||
|
log.Printf("browser fetcher disabled: %v", err)
|
||||||
|
} else {
|
||||||
|
browser = bf
|
||||||
|
cfg.Covers = bf
|
||||||
|
context.AfterFunc(pollCtx, bf.Close)
|
||||||
|
log.Printf("browser fetcher at %s", ws)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
startLatestPoller(pollCtx, s, cfg.LatestPoll, browser)
|
||||||
|
|
||||||
srv := &http.Server{
|
srv := &http.Server{
|
||||||
Addr: ":" + cfg.Port,
|
Addr: ":" + cfg.Port,
|
||||||
@@ -307,7 +328,7 @@ func main() {
|
|||||||
// HTTP client cannot be built. Any problem here is logged and skipped: this
|
// HTTP client cannot be built. Any problem here is logged and skipped: this
|
||||||
// feature going missing degrades the service to userscript-only latest-chapter
|
// feature going missing degrades the service to userscript-only latest-chapter
|
||||||
// tracking, which is exactly how it behaved before.
|
// tracking, which is exactly how it behaved before.
|
||||||
func startLatestPoller(ctx context.Context, s *store.Store, cfg LatestPoll) {
|
func startLatestPoller(ctx context.Context, s *store.Store, cfg LatestPoll, browser latest.Fetcher) {
|
||||||
if !cfg.Enabled {
|
if !cfg.Enabled {
|
||||||
log.Println("latest-chapter poller: disabled by config")
|
log.Println("latest-chapter poller: disabled by config")
|
||||||
return
|
return
|
||||||
@@ -317,9 +338,13 @@ func startLatestPoller(ctx context.Context, s *store.Store, cfg LatestPoll) {
|
|||||||
log.Printf("latest-chapter poller: disabled, cannot build client: %v", err)
|
log.Printf("latest-chapter poller: disabled, cannot build client: %v", err)
|
||||||
return
|
return
|
||||||
}
|
}
|
||||||
|
// Nil browser: sites behind a JavaScript challenge are simply not polled,
|
||||||
|
// and their latest_chapter comes from the userscript alone — which is how
|
||||||
|
// the service behaved before the sidecar existed.
|
||||||
p := &latest.Poller{
|
p := &latest.Poller{
|
||||||
Store: s,
|
Store: s,
|
||||||
Fetch: f,
|
Fetch: f,
|
||||||
|
BrowserFetch: browser,
|
||||||
Now: time.Now,
|
Now: time.Now,
|
||||||
Cooldown: cfg.Cooldown,
|
Cooldown: cfg.Cooldown,
|
||||||
Interval: cfg.Interval,
|
Interval: cfg.Interval,
|
||||||
@@ -327,19 +352,5 @@ func startLatestPoller(ctx context.Context, s *store.Store, cfg LatestPoll) {
|
|||||||
Batch: cfg.Batch,
|
Batch: cfg.Batch,
|
||||||
}
|
}
|
||||||
|
|
||||||
// Optional: without it, sites behind a JavaScript challenge are simply not
|
|
||||||
// polled, and their latest_chapter comes from the userscript alone — which
|
|
||||||
// is how the service behaved before the sidecar existed.
|
|
||||||
if ws := strings.TrimSpace(os.Getenv("BROWSER_WS_URL")); ws != "" {
|
|
||||||
bf, err := latest.NewBrowserFetcher(ws)
|
|
||||||
if err != nil {
|
|
||||||
log.Printf("latest-chapter poller: browser fetcher disabled: %v", err)
|
|
||||||
} else {
|
|
||||||
p.BrowserFetch = bf
|
|
||||||
context.AfterFunc(ctx, bf.Close)
|
|
||||||
log.Printf("latest-chapter poller: browser fetcher at %s", ws)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
go p.Run(ctx)
|
go p.Run(ctx)
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,39 @@
|
|||||||
|
# syntax=docker/dockerfile:1
|
||||||
|
|
||||||
|
# Real Google Chrome for the latest-chapter poller and the kagane cover proxy.
|
||||||
|
#
|
||||||
|
# Not chromedp/headless-shell, which this replaces. headless-shell is a stripped
|
||||||
|
# Chrome build and Cloudflare's managed challenge on kagane.to never clears for
|
||||||
|
# it: measured 2026-08-08, 60s of a held-open tab still served the interstitial,
|
||||||
|
# while stock Chrome from the same IP cleared in ~4s. The tells are structural
|
||||||
|
# rather than a header — navigator.webdriver true, an empty plugin list, and
|
||||||
|
# Chromium- rather than Chrome-branded client hints. Overriding webdriver alone
|
||||||
|
# was tried and did not move it, so the browser build itself is the fix.
|
||||||
|
#
|
||||||
|
# zenika/alpine-chrome was also tried: its Chrome is 124 (2024), old enough that
|
||||||
|
# Cloudflare refuses it outright and old enough to break chromedp's CDP structs.
|
||||||
|
FROM debian:trixie-slim
|
||||||
|
|
||||||
|
# Chrome is deliberately unpinned, against the usual rule. A pinned build goes
|
||||||
|
# stale, and a stale browser is exactly what Cloudflare turns away — the 124 in
|
||||||
|
# alpine-chrome is the worked example. Rebuild is the upgrade path.
|
||||||
|
RUN apt-get update \
|
||||||
|
&& apt-get install -y --no-install-recommends ca-certificates wget gnupg \
|
||||||
|
&& wget -qO- https://dl.google.com/linux/linux_signing_key.pub \
|
||||||
|
| gpg --dearmor -o /usr/share/keyrings/google-chrome.gpg \
|
||||||
|
&& echo "deb [arch=amd64 signed-by=/usr/share/keyrings/google-chrome.gpg] https://dl.google.com/linux/chrome/deb/ stable main" \
|
||||||
|
> /etc/apt/sources.list.d/google-chrome.list \
|
||||||
|
&& apt-get update \
|
||||||
|
&& apt-get install -y --no-install-recommends google-chrome-stable socat \
|
||||||
|
&& rm -rf /var/lib/apt/lists/*
|
||||||
|
|
||||||
|
# Unprivileged: Chrome refuses to run as root, and the CDP endpoint is a shell
|
||||||
|
# on whatever user owns it.
|
||||||
|
RUN useradd --create-home --shell /usr/sbin/nologin chrome
|
||||||
|
USER chrome
|
||||||
|
WORKDIR /home/chrome
|
||||||
|
|
||||||
|
COPY entrypoint.sh /entrypoint.sh
|
||||||
|
|
||||||
|
EXPOSE 9222
|
||||||
|
ENTRYPOINT ["/entrypoint.sh"]
|
||||||
Executable
+61
@@ -0,0 +1,61 @@
|
|||||||
|
#!/bin/sh
|
||||||
|
set -eu
|
||||||
|
|
||||||
|
# A UTC clock is itself the bot signal — Cloudflare treats it as the datacenter
|
||||||
|
# default — and kagane's challenge then never clears. Measured 2026-08-08 with
|
||||||
|
# an identical container on one Indonesian egress IP: UTC never cleared in 60s
|
||||||
|
# (twice), while Asia/Jakarta and America/New_York both cleared in 4s. Any real
|
||||||
|
# zone will do; the zone does not have to match the IP's country, it just must
|
||||||
|
# not be UTC. It does have to be right the way Chrome reads it.
|
||||||
|
#
|
||||||
|
# TZ must carry the zone *name*. Chrome resolves the zone through ICU, which
|
||||||
|
# takes the name from /etc/localtime's symlink target and ignores the file's
|
||||||
|
# contents; bind-mounting the host's /etc/localtime therefore lands on the
|
||||||
|
# image's own symlink to Etc/UTC and leaves glibc reporting +07 while Chrome
|
||||||
|
# still reports UTC. /etc/timezone, mounted by docker-compose.yml, is the name.
|
||||||
|
[ -n "${TZ:-}" ] || TZ=$(cat /etc/timezone 2>/dev/null || echo UTC)
|
||||||
|
export TZ
|
||||||
|
|
||||||
|
# Chrome's own UA advertises "HeadlessChrome" under --headless=new, and that one
|
||||||
|
# token is the difference between kagane.to's challenge clearing in ~4s and
|
||||||
|
# never clearing at all (measured 2026-08-08, same host, same Chrome, only the
|
||||||
|
# UA changed). Overriding it does not touch the Sec-CH-UA client hints, which
|
||||||
|
# report the real version, so the version is read back out of the binary rather
|
||||||
|
# than hardcoded: a hardcoded one would drift out of step with the hints on the
|
||||||
|
# next Chrome update and become a fresh tell.
|
||||||
|
major=$(google-chrome-stable --version | sed -E 's/[^0-9]*([0-9]+)\..*/\1/')
|
||||||
|
ua="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/${major}.0.0.0 Safari/537.36"
|
||||||
|
|
||||||
|
# Chrome binds its DevTools port to loopback and silently ignores
|
||||||
|
# --remote-debugging-address (verified 2026-08-08: Chrome 151 with
|
||||||
|
# --remote-debugging-address=0.0.0.0 still listened on 127.0.0.1 only), so the
|
||||||
|
# caller — another container — cannot reach it directly. socat fronting the
|
||||||
|
# loopback port is how chromedp/headless-shell solved the same problem and is
|
||||||
|
# why this image is a drop-in for it.
|
||||||
|
#
|
||||||
|
# Nothing publishes 9222; reachability is the `browser` network in
|
||||||
|
# docker-compose.yml, and an exposed CDP endpoint is remote code execution.
|
||||||
|
socat TCP-LISTEN:9222,fork,reuseaddr TCP:127.0.0.1:9223 &
|
||||||
|
|
||||||
|
# Chrome stays in the foreground so that its death takes the container down and
|
||||||
|
# compose's restart policy applies; a backgrounded browser behind a live socat
|
||||||
|
# would leave the sidecar looking healthy while answering nothing.
|
||||||
|
#
|
||||||
|
# No --enable-automation: it sets navigator.webdriver, the first thing a bot
|
||||||
|
# check reads.
|
||||||
|
#
|
||||||
|
# --no-sandbox because Chrome's zygote wants user namespaces, which Docker's
|
||||||
|
# default profile does not hand out; the alternative is --cap-add=SYS_ADMIN,
|
||||||
|
# which gives the container strictly more than it takes away. Containment here
|
||||||
|
# is the unprivileged user, the isolated network, and the fact that this
|
||||||
|
# browser only ever navigates to kagane.to and novelfull.com.
|
||||||
|
exec google-chrome-stable \
|
||||||
|
--headless=new \
|
||||||
|
--no-sandbox \
|
||||||
|
--remote-debugging-port=9223 \
|
||||||
|
--user-agent="$ua" \
|
||||||
|
--user-data-dir=/home/chrome/profile \
|
||||||
|
--no-first-run \
|
||||||
|
--no-default-browser-check \
|
||||||
|
--disable-gpu \
|
||||||
|
about:blank
|
||||||
+27
-12
@@ -24,6 +24,13 @@ services:
|
|||||||
# comes from .env so it is never committed.
|
# comes from .env so it is never committed.
|
||||||
DATABASE_URL: ${DATABASE_URL:-postgres://bookmarks:${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD in .env}@postgres:5432/bookmarks?sslmode=disable}
|
DATABASE_URL: ${DATABASE_URL:-postgres://bookmarks:${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD in .env}@postgres:5432/bookmarks?sslmode=disable}
|
||||||
PORT: "8080"
|
PORT: "8080"
|
||||||
|
# Log timestamps only. Go's `log` stamps lines in local time, and this
|
||||||
|
# service has no other use for a zone: bookmark timestamps are unix ms
|
||||||
|
# and the two real time columns are timestamptz, both absolute instants.
|
||||||
|
# Purely so these lines read on the same clock as the sidecar's. Named
|
||||||
|
# API_TZ rather than TZ so an operator's exported shell TZ cannot leak
|
||||||
|
# in; distroless already carries tzdata, so the name just resolves.
|
||||||
|
TZ: ${API_TZ:-Asia/Jakarta}
|
||||||
# Discord OAuth for the browser UI (ADR-0002). The first four are
|
# Discord OAuth for the browser UI (ADR-0002). The first four are
|
||||||
# required; DISCORD_REQUIRED_ROLE is optional and empty by default.
|
# required; DISCORD_REQUIRED_ROLE is optional and empty by default.
|
||||||
# Guild membership is the whole gate: any member becomes a Reader.
|
# Guild membership is the whole gate: any member becomes a Reader.
|
||||||
@@ -95,8 +102,25 @@ services:
|
|||||||
- db
|
- db
|
||||||
|
|
||||||
headless-shell:
|
headless-shell:
|
||||||
image: chromedp/headless-shell:stable
|
# Real Google Chrome, not chromedp/headless-shell — see chrome/Dockerfile.
|
||||||
|
# The service name is kept so existing overrides and BROWSER_WS_URL stay put.
|
||||||
|
build: ./chrome
|
||||||
|
image: bookmarkmanager-chrome:latest
|
||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
|
environment:
|
||||||
|
# A UTC clock is itself the bot signal: Cloudflare treats it as the
|
||||||
|
# datacenter default, and kagane's challenge then never clears. Measured
|
||||||
|
# 2026-08-08, identical container, one Indonesian egress IP: UTC never
|
||||||
|
# cleared in 60s (twice); Asia/Jakarta and America/New_York both cleared
|
||||||
|
# in 4s. So any real zone works and it need not match the IP's country —
|
||||||
|
# only UTC fails. Unset falls back to the host's /etc/timezone below,
|
||||||
|
# which is a real zone whenever the host clock is set to local time; set
|
||||||
|
# BROWSER_TZ when the host runs UTC.
|
||||||
|
TZ: ${BROWSER_TZ:-}
|
||||||
|
volumes:
|
||||||
|
# The zone *name*, which is what Chrome's ICU needs — see chrome/entrypoint.sh.
|
||||||
|
# Absent on a non-Debian host, which the entrypoint handles by falling back to UTC.
|
||||||
|
- /etc/timezone:/etc/timezone:ro
|
||||||
# Chrome allocates shared memory per tab and dies on Docker's 64MB default.
|
# Chrome allocates shared memory per tab and dies on Docker's 64MB default.
|
||||||
shm_size: '1gb'
|
shm_size: '1gb'
|
||||||
# Reaps zombie renderer processes, which otherwise accumulate for the
|
# Reaps zombie renderer processes, which otherwise accumulate for the
|
||||||
@@ -104,17 +128,8 @@ services:
|
|||||||
init: true
|
init: true
|
||||||
# Deliberately no `ports:` — an exposed CDP endpoint is remote code
|
# Deliberately no `ports:` — an exposed CDP endpoint is remote code
|
||||||
# execution. Only bookmark-api, via the `browser` network below, may reach it.
|
# execution. Only bookmark-api, via the `browser` network below, may reach it.
|
||||||
# Don't pass --remote-debugging-address/--remote-debugging-port here: the
|
# No `command:` either: every flag this browser needs is in its entrypoint,
|
||||||
# image's own entrypoint (/headless-shell/run.sh) already starts Chrome on
|
# and the UA override there is load-bearing for the challenge.
|
||||||
# 127.0.0.1:9223 and fronts it with a socat proxy listening on 0.0.0.0:9222.
|
|
||||||
# Redeclaring the port flag here overrides Chrome's, so it binds 9222
|
|
||||||
# directly (IPv6 loopback only) instead of 9223 — collides with socat's own
|
|
||||||
# bind on 9222 and leaves nothing listening on 9223, so every external
|
|
||||||
# connection to headless-shell:9222 fails with EOF. Only pass flags the
|
|
||||||
# entrypoint doesn't already set.
|
|
||||||
command:
|
|
||||||
- --disable-gpu
|
|
||||||
- --no-sandbox
|
|
||||||
networks:
|
networks:
|
||||||
browser:
|
browser:
|
||||||
# Pinned so BROWSER_WS_URL can name an IP (required, see above) that
|
# Pinned so BROWSER_WS_URL can name an IP (required, see above) that
|
||||||
|
|||||||
@@ -44,6 +44,25 @@ Guidance for OpenCode (and Claude Code) working under `userscript/`. See root `A
|
|||||||
Encodings (incl. triple-encoded punctuation like `%25252D`) identical
|
Encodings (incl. triple-encoded punctuation like `%25252D`) identical
|
||||||
on /manga/ and /title/ pages, so decode-once seriesIds match — verified
|
on /manga/ and /title/ pages, so decode-once seriesIds match — verified
|
||||||
2026-07-28.
|
2026-07-28.
|
||||||
|
- **comix.to**: series `/title/<id>-<slug>`, chapter
|
||||||
|
`/title/<id>-<slug>/<uploadId>-chapter-<n>`. Only the leading `<id>` is
|
||||||
|
identity — the slug re-renders when a series is renamed (`comixSeriesId`).
|
||||||
|
An SPA that **never rewrites `og:title`**: the server-rendered head keeps
|
||||||
|
whatever document loaded first, so on a cold load `og:title` is the homepage's
|
||||||
|
"Comix — Read Comics online for free" and after an in-page hop it is the
|
||||||
|
*previous* series' name. `document.title` is the one thing client routing does
|
||||||
|
update, so titles come from there, with the chapter page's `" · Ch.<n>"` tail
|
||||||
|
stripped. Covers likewise: `og:image` is absent, so the cover is the `img`
|
||||||
|
whose `alt` matches the cleaned title — verified live 2026-08-08.
|
||||||
|
- **kagane.to**: series `/series/<uuid>`, reader
|
||||||
|
`/series/<uuid>/reader/<bookUuid>`. Reader URLs carry no chapter number, so
|
||||||
|
the number comes out of `og:title`. Two shapes exist: `"<Series> - Chapter
|
||||||
|
<n>[ - Episode <n>]"` and, for volume-numbered series, `"<Series> - Volume <v>
|
||||||
|
Chapter <n>"` with no episode name — both must yield a bare series title, or
|
||||||
|
the volume tail lands in the bookmark's title. Its covers are challenge- and
|
||||||
|
CORP-protected, so the web UI proxies them; the userscript still stores the
|
||||||
|
raw `og:image`. Behind a Cloudflare JS challenge, so the backend polls it
|
||||||
|
through the headless browser.
|
||||||
- **novelfull.com** (novel script): series `/<slug>.html`, chapter
|
- **novelfull.com** (novel script): series `/<slug>.html`, chapter
|
||||||
`/<slug>/chapter-<n>[-<title-slug>].html`. No `og:*` tags at all — title from
|
`/<slug>/chapter-<n>[-<title-slug>].html`. No `og:*` tags at all — title from
|
||||||
`h3.title` (series) or `a.truyen-title` (chapter), cover from
|
`h3.title` (series) or `a.truyen-title` (chapter), cover from
|
||||||
|
|||||||
@@ -242,6 +242,12 @@
|
|||||||
matches: (loc) => /(^|\.)comix\.to$/.test(loc.hostname),
|
matches: (loc) => /(^|\.)comix\.to$/.test(loc.hostname),
|
||||||
detect(loc) {
|
detect(loc) {
|
||||||
const path = loc.pathname;
|
const path = loc.pathname;
|
||||||
|
// comix client-routes without ever rewriting og:title — the head keeps
|
||||||
|
// whatever the first server-rendered document carried, so a bookmark
|
||||||
|
// taken after a client route got the homepage's title, then the
|
||||||
|
// previous series'. document.title is the one thing its router does
|
||||||
|
// update. Verified live 2026-08-08; do not "restore" meta("og:title").
|
||||||
|
const pageTitle = cleanTitle(document.title);
|
||||||
// /title/<id>-<slug>/<uploadId>-chapter-<n>. Several uploads (different
|
// /title/<id>-<slug>/<uploadId>-chapter-<n>. Several uploads (different
|
||||||
// groups or languages) share one chapter number; the number is the
|
// groups or languages) share one chapter number; the number is the
|
||||||
// progress identity, the upload id is not.
|
// progress identity, the upload id is not.
|
||||||
@@ -252,8 +258,8 @@
|
|||||||
type: "chapter",
|
type: "chapter",
|
||||||
site: this.site,
|
site: this.site,
|
||||||
seriesId: comixSeriesId(m[1]),
|
seriesId: comixSeriesId(m[1]),
|
||||||
title: cleanTitle(meta("og:title")),
|
title: pageTitle,
|
||||||
cover: coverFromPage(),
|
cover: coverFromPage(pageTitle),
|
||||||
seriesUrl: loc.origin + "/title/" + m[1],
|
seriesUrl: loc.origin + "/title/" + m[1],
|
||||||
chapterLabel: "Chapter " + m[2],
|
chapterLabel: "Chapter " + m[2],
|
||||||
chapterNum: isNaN(num) ? null : num,
|
chapterNum: isNaN(num) ? null : num,
|
||||||
@@ -267,8 +273,8 @@
|
|||||||
type: "series",
|
type: "series",
|
||||||
site: this.site,
|
site: this.site,
|
||||||
seriesId: comixSeriesId(m[1]),
|
seriesId: comixSeriesId(m[1]),
|
||||||
title: cleanTitle(meta("og:title")),
|
title: pageTitle,
|
||||||
cover: coverFromPage(),
|
cover: coverFromPage(pageTitle),
|
||||||
seriesUrl: loc.origin + "/title/" + m[1],
|
seriesUrl: loc.origin + "/title/" + m[1],
|
||||||
chapterLabel: null,
|
chapterLabel: null,
|
||||||
chapterNum: null,
|
chapterNum: null,
|
||||||
@@ -277,7 +283,7 @@
|
|||||||
}
|
}
|
||||||
return { type: "other" };
|
return { type: "other" };
|
||||||
|
|
||||||
// comix chapter og:title is "<Title> · Ch.<n>"; series is clean.
|
// comix chapter document.title is "<Title> · Ch.<n>"; series is clean.
|
||||||
function cleanTitle(t) {
|
function cleanTitle(t) {
|
||||||
if (!t) return "";
|
if (!t) return "";
|
||||||
return t.replace(/\s*·\s*Ch\.[\d.]+\s*$/i, "").trim();
|
return t.replace(/\s*·\s*Ch\.[\d.]+\s*$/i, "").trim();
|
||||||
@@ -287,8 +293,7 @@
|
|||||||
// the DOM for a cover. Matching on alt rather than a class keeps it off
|
// the DOM for a cover. Matching on alt rather than a class keeps it off
|
||||||
// the site's styling: the cover is the image whose alt is the title.
|
// the site's styling: the cover is the image whose alt is the title.
|
||||||
// Do not "simplify" this into meta("og:image") — that returns null.
|
// Do not "simplify" this into meta("og:image") — that returns null.
|
||||||
function coverFromPage() {
|
function coverFromPage(title) {
|
||||||
const title = cleanTitle(meta("og:title"));
|
|
||||||
if (!title || !document.querySelectorAll) return "";
|
if (!title || !document.querySelectorAll) return "";
|
||||||
for (const img of document.querySelectorAll("img[alt]")) {
|
for (const img of document.querySelectorAll("img[alt]")) {
|
||||||
if (img.getAttribute("alt") === title) return img.getAttribute("src") || "";
|
if (img.getAttribute("alt") === title) return img.getAttribute("src") || "";
|
||||||
@@ -313,6 +318,14 @@
|
|||||||
},
|
},
|
||||||
};
|
};
|
||||||
|
|
||||||
|
// Kagane builds the reader og:title suffix out of the book's metadata, so
|
||||||
|
// every combination occurs: the volume part appears only when the book has a
|
||||||
|
// volume_no, the episode part only when it has a non-empty title. All four
|
||||||
|
// shapes captured live 2026-08-08 — "SP Baby - Volume 1 Chapter 1" is the one
|
||||||
|
// the old trailing-space regex missed, which left both the number and the
|
||||||
|
// series title wrong.
|
||||||
|
const KAGANE_CHAPTER_SUFFIX = /\s-\s(?:Volume\s[\d.]+\s)?Chapter\s([\d.]+)(?:\s-\s.*)?$/i;
|
||||||
|
|
||||||
const kagane = {
|
const kagane = {
|
||||||
site: "kagane",
|
site: "kagane",
|
||||||
matches: (loc) => /(^|\.)kagane\.to$/.test(loc.hostname),
|
matches: (loc) => /(^|\.)kagane\.to$/.test(loc.hostname),
|
||||||
@@ -353,9 +366,8 @@
|
|||||||
}
|
}
|
||||||
return { type: "other" };
|
return { type: "other" };
|
||||||
|
|
||||||
// Reader og:title is "<Title> - Chapter <n> - <episode name>".
|
|
||||||
function chapterNumFromTitle(t) {
|
function chapterNumFromTitle(t) {
|
||||||
const m = t && t.match(/\s-\sChapter\s([\d.]+)\s/);
|
const m = t && t.match(KAGANE_CHAPTER_SUFFIX);
|
||||||
if (!m) return null;
|
if (!m) return null;
|
||||||
const num = parseFloat(m[1]);
|
const num = parseFloat(m[1]);
|
||||||
return isNaN(num) ? null : num;
|
return isNaN(num) ? null : num;
|
||||||
@@ -363,7 +375,7 @@
|
|||||||
|
|
||||||
function cleanTitle(t) {
|
function cleanTitle(t) {
|
||||||
if (!t) return "";
|
if (!t) return "";
|
||||||
return t.replace(/\s-\sChapter\s[\d.]+\s-\s.*$/i, "").trim();
|
return t.replace(KAGANE_CHAPTER_SUFFIX, "").trim();
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
// Reader hrefs are uuids with no number in them, so no maximum can be taken
|
// Reader hrefs are uuids with no number in them, so no maximum can be taken
|
||||||
@@ -1450,8 +1462,15 @@
|
|||||||
}
|
}
|
||||||
|
|
||||||
let lastUrl = location.href;
|
let lastUrl = location.href;
|
||||||
function onNavigate() {
|
let lastPageSig = "";
|
||||||
|
|
||||||
|
function setPage() {
|
||||||
state.page = detect();
|
state.page = detect();
|
||||||
|
lastPageSig = JSON.stringify(state.page);
|
||||||
|
}
|
||||||
|
|
||||||
|
function onNavigate() {
|
||||||
|
setPage();
|
||||||
render();
|
render();
|
||||||
maybeAutoUpdate();
|
maybeAutoUpdate();
|
||||||
maybeCaptureLatestOnSeriesPage();
|
maybeCaptureLatestOnSeriesPage();
|
||||||
@@ -1484,7 +1503,14 @@
|
|||||||
wrap("pushState");
|
wrap("pushState");
|
||||||
wrap("replaceState");
|
wrap("replaceState");
|
||||||
window.addEventListener("popstate", fire);
|
window.addEventListener("popstate", fire);
|
||||||
setInterval(fire, 1500); // catch routes that bypass history
|
// Also catches routes that bypass history — and comix, which fills
|
||||||
|
// document.title a beat after the route changes, so the 300ms snapshot
|
||||||
|
// above can still hold the previous page's title. Re-detect whenever what
|
||||||
|
// we would read has changed, not only when the URL has.
|
||||||
|
setInterval(() => {
|
||||||
|
fire();
|
||||||
|
if (JSON.stringify(detect()) !== lastPageSig) onNavigate();
|
||||||
|
}, 1500);
|
||||||
}
|
}
|
||||||
|
|
||||||
// ============================================================
|
// ============================================================
|
||||||
@@ -1563,7 +1589,7 @@
|
|||||||
|
|
||||||
function init() {
|
function init() {
|
||||||
buildUI();
|
buildUI();
|
||||||
state.page = detect();
|
setPage();
|
||||||
render();
|
render();
|
||||||
installNavWatcher();
|
installNavWatcher();
|
||||||
installLongPress();
|
installLongPress();
|
||||||
|
|||||||
@@ -31,6 +31,9 @@ globalThis.location = {
|
|||||||
|
|
||||||
// og: meta tags the adapters read through meta(). Reassigned per test.
|
// og: meta tags the adapters read through meta(). Reassigned per test.
|
||||||
let metaTags = {};
|
let metaTags = {};
|
||||||
|
// document.title. comix's SPA rewrites this on client routing but never
|
||||||
|
// og:title, so the comix adapter reads it instead. Reassigned per test.
|
||||||
|
let docTitle = "";
|
||||||
// img[alt] elements comix's coverFromPage() scans. Reassigned per test; each
|
// img[alt] elements comix's coverFromPage() scans. Reassigned per test; each
|
||||||
// entry is {alt, src}.
|
// entry is {alt, src}.
|
||||||
let pageImages = [];
|
let pageImages = [];
|
||||||
@@ -48,6 +51,9 @@ globalThis.document = {
|
|||||||
}));
|
}));
|
||||||
},
|
},
|
||||||
addEventListener() {},
|
addEventListener() {},
|
||||||
|
get title() {
|
||||||
|
return docTitle;
|
||||||
|
},
|
||||||
body: undefined,
|
body: undefined,
|
||||||
};
|
};
|
||||||
|
|
||||||
@@ -193,8 +199,15 @@ test("comixSeriesId leaves a bare id untouched", () => {
|
|||||||
assert.equal(comixSeriesId("n8we"), "n8we");
|
assert.equal(comixSeriesId("n8we"), "n8we");
|
||||||
});
|
});
|
||||||
|
|
||||||
|
// comix is an SPA that rewrites document.title on client routing but leaves the
|
||||||
|
// server-rendered og:title untouched, so every test here pins og:title to a
|
||||||
|
// STALE value — the homepage title on first hop, the previous series after
|
||||||
|
// that. Captured live 2026-08-08.
|
||||||
|
const COMIX_STALE_HOME = "Comix - Read Comics online for free";
|
||||||
|
|
||||||
test("comix detects a series page", () => {
|
test("comix detects a series page", () => {
|
||||||
metaTags = { "og:title": "Dungeons and Crayons" };
|
metaTags = { "og:title": COMIX_STALE_HOME };
|
||||||
|
docTitle = "Dungeons and Crayons";
|
||||||
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
|
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
|
||||||
assert.equal(p.type, "series");
|
assert.equal(p.type, "series");
|
||||||
assert.equal(p.site, "comix");
|
assert.equal(p.site, "comix");
|
||||||
@@ -204,8 +217,16 @@ test("comix detects a series page", () => {
|
|||||||
assert.equal(p.chapterNum, null);
|
assert.equal(p.chapterNum, null);
|
||||||
});
|
});
|
||||||
|
|
||||||
|
test("comix ignores a previous series' stale og:title", () => {
|
||||||
|
metaTags = { "og:title": "Full-Time Awakening" };
|
||||||
|
docTitle = "Dungeons and Crayons";
|
||||||
|
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
|
||||||
|
assert.equal(p.title, "Dungeons and Crayons");
|
||||||
|
});
|
||||||
|
|
||||||
test("comix detects a chapter page and strips the Ch. suffix from the title", () => {
|
test("comix detects a chapter page and strips the Ch. suffix from the title", () => {
|
||||||
metaTags = { "og:title": "Dungeons and Crayons · Ch.80" };
|
metaTags = { "og:title": COMIX_STALE_HOME };
|
||||||
|
docTitle = "Dungeons and Crayons · Ch.80";
|
||||||
const p = comix.detect(
|
const p = comix.detect(
|
||||||
loc("https://comix.to/title/n8we-dungeons-and-crayons/11139891-chapter-80")
|
loc("https://comix.to/title/n8we-dungeons-and-crayons/11139891-chapter-80")
|
||||||
);
|
);
|
||||||
@@ -218,7 +239,8 @@ test("comix detects a chapter page and strips the Ch. suffix from the title", ()
|
|||||||
});
|
});
|
||||||
|
|
||||||
test("comix parses decimal chapter numbers", () => {
|
test("comix parses decimal chapter numbers", () => {
|
||||||
metaTags = { "og:title": "Dungeons and Crayons · Ch.80.5" };
|
metaTags = {};
|
||||||
|
docTitle = "Dungeons and Crayons · Ch.80.5";
|
||||||
const p = comix.detect(
|
const p = comix.detect(
|
||||||
loc("https://comix.to/title/n8we-dungeons-and-crayons/11139891-chapter-80.5")
|
loc("https://comix.to/title/n8we-dungeons-and-crayons/11139891-chapter-80.5")
|
||||||
);
|
);
|
||||||
@@ -226,19 +248,19 @@ test("comix parses decimal chapter numbers", () => {
|
|||||||
});
|
});
|
||||||
|
|
||||||
test("comix.detect reads the cover from an img whose alt matches the cleaned title", () => {
|
test("comix.detect reads the cover from an img whose alt matches the cleaned title", () => {
|
||||||
metaTags = { "og:title": "Dungeons and Crayons · Ch.80" };
|
metaTags = { "og:title": COMIX_STALE_HOME };
|
||||||
|
docTitle = "Dungeons and Crayons";
|
||||||
pageImages = [
|
pageImages = [
|
||||||
{ alt: "Some Other Series", src: "https://cdn.example/other.jpg" },
|
{ alt: "Some Other Series", src: "https://cdn.example/other.jpg" },
|
||||||
{ alt: "Dungeons and Crayons", src: "https://cdn.example/cover.jpg" },
|
{ alt: "Dungeons and Crayons", src: "https://cdn.example/cover.jpg" },
|
||||||
];
|
];
|
||||||
const p = comix.detect(
|
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
|
||||||
loc("https://comix.to/title/n8we-dungeons-and-crayons/11139891-chapter-80")
|
|
||||||
);
|
|
||||||
assert.equal(p.cover, "https://cdn.example/cover.jpg");
|
assert.equal(p.cover, "https://cdn.example/cover.jpg");
|
||||||
});
|
});
|
||||||
|
|
||||||
test("comix.detect leaves cover empty when no img alt matches the title", () => {
|
test("comix.detect leaves cover empty when no img alt matches the title", () => {
|
||||||
metaTags = { "og:title": "Dungeons and Crayons" };
|
metaTags = {};
|
||||||
|
docTitle = "Dungeons and Crayons";
|
||||||
pageImages = [{ alt: "Some Other Series", src: "https://cdn.example/other.jpg" }];
|
pageImages = [{ alt: "Some Other Series", src: "https://cdn.example/other.jpg" }];
|
||||||
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
|
const p = comix.detect(loc("https://comix.to/title/n8we-dungeons-and-crayons"));
|
||||||
assert.equal(p.cover, "");
|
assert.equal(p.cover, "");
|
||||||
@@ -315,6 +337,33 @@ test("kagane reads the chapter number out of og:title", () => {
|
|||||||
assert.equal(p.seriesUrl, "https://kagane.to/series/" + KAGANE_SERIES);
|
assert.equal(p.seriesUrl, "https://kagane.to/series/" + KAGANE_SERIES);
|
||||||
});
|
});
|
||||||
|
|
||||||
|
// Volume-numbered series render the suffix as "- Volume <v> Chapter <n>" with
|
||||||
|
// no episode name, because the book carries volume_no and an empty title.
|
||||||
|
// Captured live 2026-08-08 from SP Baby.
|
||||||
|
test("kagane reads through a Volume-numbered chapter suffix", () => {
|
||||||
|
metaTags = {
|
||||||
|
"og:title": "SP Baby - Volume 1 Chapter 1",
|
||||||
|
"og:image": "https://kagane.to/api/v2/image/abc/compressed",
|
||||||
|
};
|
||||||
|
const p = kagane.detect(
|
||||||
|
loc("https://kagane.to/series/" + KAGANE_SERIES + "/reader/" + KAGANE_BOOK)
|
||||||
|
);
|
||||||
|
assert.equal(p.title, "SP Baby");
|
||||||
|
assert.equal(p.chapterNum, 1);
|
||||||
|
assert.equal(p.chapterLabel, "Chapter 1");
|
||||||
|
});
|
||||||
|
|
||||||
|
// A book with neither a volume nor an episode name ends the title right after
|
||||||
|
// the number, which the old trailing-\s regex could not match.
|
||||||
|
test("kagane reads a chapter suffix with no episode name", () => {
|
||||||
|
metaTags = { "og:title": "Some Series - Chapter 7.5" };
|
||||||
|
const p = kagane.detect(
|
||||||
|
loc("https://kagane.to/series/" + KAGANE_SERIES + "/reader/" + KAGANE_BOOK)
|
||||||
|
);
|
||||||
|
assert.equal(p.title, "Some Series");
|
||||||
|
assert.equal(p.chapterNum, 7.5);
|
||||||
|
});
|
||||||
|
|
||||||
test("kagane yields a null chapterNum when og:title has no chapter", () => {
|
test("kagane yields a null chapterNum when og:title has no chapter", () => {
|
||||||
metaTags = { "og:title": "Infinite Decryption: The Strongest Level 0" };
|
metaTags = { "og:title": "Infinite Decryption: The Strongest Level 0" };
|
||||||
const p = kagane.detect(
|
const p = kagane.detect(
|
||||||
|
|||||||
Reference in New Issue
Block a user