Browser-backed Sites join the Cover pipeline (#62) (#72)

Fixes #62

Browser-backed Sites join the Cover pipeline: kagane and novelfull Series now get their Covers at creation, through the same acquisition path as every other Site, instead of waiting for a poll pass.

## What changed

`latest.Acquirer` (creation-time acquisition, fired by the first Bookmark of a Series) previously skipped kagane and novelfull entirely — their pages only yield a Cloudflare challenge to the TLS client, so the request was spent for nothing. It now routes them like the poller does, with the two Sites split exactly as the issue demands:

- **kagane** — page fetched through the browser sidecar, cover URL extracted from the API JSON, bytes fetched through the browser sidecar (the only path that clears the challenge) into the content-addressed store. With no `BROWSER_WS_URL` configured, acquisition is skipped entirely and nothing falls back to a plain fetch.
- **novelfull** — page fetched through the browser sidecar, cover URL extracted from the HTML, bytes fetched over plain TLS through the ordinary gated fetcher (its image paths answer 200 with `access-control-allow-origin: *`, measured 2026-08-09). With no browser configured, the page fetch falls back to the TLS client — novelfull's challenge is a live time-varying fact (AGENTS.md), so when the page body answers, the Cover still lands; when it is challenged, nothing happens.

The byte-routing rule (kagane → browser, every other Site → TLS) is now one shared function (`latest.fetchCoverBytes`) used by both the Poller and the Acquirer, so the two cannot drift apart.

## Acceptance criteria

- [x] kagane cover bytes are fetched through the browser sidecar and stored in the content-addressed store — `TestAcquireKaganeCoverThroughBrowser`
- [x] novelfull cover URLs are extracted from the browser-fetched HTML, and its bytes are fetched over plain TLS — `TestAcquireNovelfullCoverOverPlainTLS`
- [x] With no browser sidecar configured, kagane Covers are absent and nothing falls back to a plain fetch — `TestAcquireKaganeSkippedWithoutBrowser`
- [x] With no browser sidecar configured, novelfull Covers still work if its page body is available — `TestAcquireNovelfullCoverWithoutBrowser`
- [x] Manually verified on-device: a kagane Series shows its Cover in the panel, not a broken-image glyph — being run by a separate manual-verification agent against a mocked scenario (no prod data); not part of this PR
- [x] `go test ./...` is green, with live-network checks gated behind `SMOKE_BROWSER_WS_URL` like the existing kagane image smoke test — new `TestSmokeAcquireKaganeCover` proves the end-to-end acquire path against the real browser when the env var is set

## Verification

- `go test ./...` green across all packages
- New unit tests exercise every routing decision with fakes — no network in the default suite
- Smoke test gated behind `SMOKE_BROWSER_WS_URL`, skipped by default

## Post-review changes (a66491a)

- **One routing rule for pages too** — `fetcherFor` is now a shared function used by both the Poller and the Acquirer; novelfull falls back to the plain-TLS fetcher in *both* when no browser is configured, so pre-existing (client-scraped) novelfull rows get healed by the poll as well, not just Series created after this change (`TestNovelfullUsesTLSWhenNoBrowserFetcher`).
- **Byte-level no-fallback proof** — `TestAcquireKaganeBytesNeverFallBackToPlainTLS` pins that kagane cover bytes never route to the TLS fetcher even when the page came through a browser.
- **Acquirer wired independent of the TLS client** — if `NewTLSFetcher` fails, kagane/novelfull acquisition still works via the sidecar (`main.go`).
- AGENTS.md (root + backend) updated for the novelfull plain-TLS fallback.

Reviewed-on: #72
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
This commit was merged in pull request #72.
This commit is contained in:
2026-08-10 11:06:36 +07:00
committed by sulthan
parent b9220b3dfc
commit 78234f3c19
9 changed files with 406 additions and 53 deletions
+1 -1
View File
@@ -19,7 +19,7 @@ Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free
- Every site is its **own origin with its own `localStorage`** — a shared remote store is the only way to unify bookmarks. Cloud sync required, not optional.
- Userscript run in **isolated world**, so embedded API token safe from site's JS.
- Cloudflare's block on manga sites **IP-reputation-based, not universal — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — Cloudflare's bot scoring can flip previously-clean IP without notice. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test.
- **kagane.to and novelfull.com are the exception to the above** — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`) and skips them entirely when that's unset. The four other sites poll fine over plain TLS.
- **kagane.to and novelfull.com are the exception to the above** — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`). When that's unset, kagane is skipped entirely (a plain fetch would only retrieve a challenge page) while novelfull pages are still attempted over plain TLS — its challenge is a live time-varying fact and its cover bytes never need the browser. The four other sites poll fine over plain TLS.
- **The CDP browser must look like a real browser, and stock headless images don't.** Measured 2026-08-08 against kagane.to, all from the same IP: `chromedp/headless-shell:stable` never cleared the challenge in 90s (`navigator.webdriver` true, empty plugin list, Chromium-branded client hints — suppressing `webdriver` alone changed nothing); `zenika/alpine-chrome` ships Chrome 124, refused outright; real Chrome with the default `--headless=new` UA never cleared, because the UA says `HeadlessChrome`; real Chrome with a stock UA **and** a non-UTC clock zone cleared in ~4s. Hence `chrome/` — a Debian image with `google-chrome-stable`, a version-derived UA, and `TZ`/`BROWSER_TZ`. Chrome reads the zone *name* through ICU from `/etc/localtime`'s symlink target, ignoring the file's contents, so mounting the host's `/etc/localtime` does **not** work; `/etc/timezone` is mounted instead.
- **The browser is not in the API stack and must not be put back.** It's its own compose unit (`chrome/docker-compose.yml`) on a second machine, reached over the tailnet — it held 471 MiB on a 1974 MiB swapless VPS, and a residential egress scores better with Cloudflare anyway (ADR-0006). Consequences that constrain code: `BROWSER_WS_URL` must be a tailnet **IP** (a MagicDNS name 500s at `/json/version`, same trap as the old Docker service name); the CDP port binds to the tailnet address only, since CDP authenticates nothing and that host has a real LAN; and the browser is on-demand (ADR-0005), so an unreachable or asleep one must degrade exactly as an unset `BROWSER_WS_URL` — plain-TLS libraries unaffected, kagane/novelfull logged and skipped, stored covers still served. Never add `chromedp.NoModifyURL`: discovery per fetch is what makes a restarted Chrome invisible.
- **UTC is the tell, not a country mismatch.** A UTC clock is the datacenter default, so Cloudflare scores it as one; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while `Asia/Jakarta` **and** `America/New_York` both cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (`Asia/Bangkok`) rather than the measured egress. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one.
+12 -2
View File
@@ -91,8 +91,10 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
Fetches use `bogdanfinn/tls-client` with Chrome profile as defence in depth
against fingerprint-based blocking; any failure log and skip. kagane and
novelfull sit behind Cloudflare JavaScript challenges the TLS client can't
clear, so they are browser-only: fetched over CDP via `BROWSER_WS_URL`, and
simply not polled when that's unset. See
clear, so they are fetched over CDP via `BROWSER_WS_URL`; kagane is simply
not polled when that's unset, while novelfull falls back to a plain-TLS
attempt — its challenge is a live time-varying fact, and its cover bytes
never need the browser. See
`docs/superpowers/specs/2026-07-26-server-latest-chapter-polling-design.md`.
The poller's series write is a single-column UPDATE
(`Store.SetLatestChapter`), not a read-modify-write of the whole bookmark:
@@ -110,6 +112,14 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
is public and uncredentialed: the userscript renders it on a Site's origin,
where no cookie or token of ours travels. A client-sent `cover` is decoded
and discarded, permanently (ADR-0004 compatibility).
Browser-backed Sites join the same pipeline (issue #62): kagane pages *and*
cover bytes go through the browser sidecar (nothing falls back to a plain
fetch, which would only retrieve a challenge page), while novelfull needs
the browser only for its HTML — the cover URL comes out of the
browser-fetched page and the bytes go over plain TLS. With no browser
configured, kagane Covers are simply absent; novelfull still gets one — at
creation and on the poll — when its page body happens to answer a plain
request (the challenge is a live time-varying fact).
- **`updated_at` drives list order, so moves only on real reading progress:** server apply its timestamp when row new or `last_chapter_num` changes, else keep stored value — favouriting series or recording newly published chapter must not reorder list. `PUT` therefore returns row **as stored**, clients must adopt that response rather than own payload. See `plans/2026-07-25-bookmark-list-favorites-design.md` §4.
- **Lifecycle buckets:** `status` on each bookmark is `reading` | `archived` |
`finished`, orthogonal to `favorite`. Archived and finished appear only in
+20 -9
View File
@@ -3,7 +3,6 @@ package latest
import (
"context"
"log"
"slices"
"sync"
"time"
@@ -32,11 +31,21 @@ const acquireTimeout = 45 * time.Second
// left blank until the poll's own cover pass (#61) fills it.
type Acquirer struct {
Store *store.Store
// Fetch retrieves the series page. Nil disables acquisition entirely.
// Fetch retrieves the series page over plain TLS. Nil with a nil
// BrowserFetch disables acquisition entirely.
Fetch Fetcher
// BrowserFetch retrieves kagane and novelfull pages through the browser
// sidecar, the only thing that clears their Cloudflare challenge. The
// per-site fallback policy lives in fetcherFor. Nil leaves those Sites
// unacquired when no fallback applies.
BrowserFetch Fetcher
// Covers retrieves the cover bytes. Nil leaves the Cover blank and the
// chapter half working.
Covers CoverBytesFetcher
// BrowserCoverFetch retrieves kagane cover bytes through the browser
// sidecar. Nil leaves kagane Covers blank; nothing falls back to a plain
// fetch, which would only ever retrieve a challenge page.
BrowserCoverFetch BrowserCoverFetcher
// Ctx cancels in-flight acquisitions at shutdown. A hook signature has
// nowhere to pass one, so it lives here; nil means context.Background.
Ctx context.Context
@@ -85,10 +94,7 @@ func (a *Acquirer) Acquire(sr store.Series) {
func (a *Acquirer) Wait() { a.inflight.Wait() }
func (a *Acquirer) acquire(ctx context.Context, sr store.Series) {
// Browser-backed Sites are deliberately not acquired here: their pages
// only yield a Cloudflare challenge to the TLS client, so the request
// would be spent for nothing.
if a.Fetch == nil || slices.Contains(browserBackedSites, sr.Site) {
if a.Fetch == nil && a.BrowserFetch == nil {
return
}
// series_url arrives in a client-supplied PUT body, so the same gate the
@@ -99,7 +105,12 @@ func (a *Acquirer) acquire(ctx context.Context, sr store.Series) {
return
}
body, status, err := a.Fetch.Get(ctx, sr.SeriesURL)
f := fetcherFor(sr.Site, a.BrowserFetch, a.Fetch)
if f == nil {
log.Printf("acquire %q: no fetcher for site %q", sr.Key(), sr.Site)
return
}
body, status, err := f.Get(ctx, sr.SeriesURL)
if err != nil {
log.Printf("acquire %q: fetch %s: %v", sr.Key(), sr.SeriesURL, err)
return
@@ -122,10 +133,10 @@ func (a *Acquirer) acquire(ctx context.Context, sr store.Series) {
}
cover, ok := coverFrom(sr.Site, sr.SeriesURL, body)
if !ok || a.Covers == nil {
if !ok {
return
}
bytes, contentType, err := a.Covers.Fetch(ctx, cover)
bytes, contentType, err := fetchCoverBytes(ctx, sr.Site, cover, a.BrowserCoverFetch, a.Covers)
if err != nil {
log.Printf("acquire %q: fetch cover %s: %v", sr.Key(), cover, err)
return
+191
View File
@@ -20,6 +20,29 @@ const (
acquireCoverURL = "https://cdn.asurascans.com/asura-images/covers/chronicles-of-the-demon-faction.d4dcb8.webp"
)
const (
kaganeKey = "kagane:019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
kaganeSeriesID = "019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
kaganeSeriesURL = "https://kagane.to/series/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b"
kaganeImageID = "019fe11a-84c3-7fc3-a84b-88787374b617"
kaganeCoverSrc = "https://kagane.to/api/v2/image/" + kaganeImageID + "/compressed"
)
// kagane's browser-fetched body is one JSON object carrying both the chapter
// list (series_books) and the cover image ids (series_covers), so the single
// acquisition fetch yields both facts.
const kaganeSeriesAndCoverFixture = `{"series_id":"019f84bc-9ba0-7ed9-86f5-8b905ec7c28b",` +
`"series_books":[{"book_id":"b","title":"Episode 41","chapter_no":"41","sort_no":41}],` +
`"series_covers":[{"cover_id":"019fe11a-84d1-714b-9cf4-2827f277f3c0","language":"en",` +
`"image_id":"019fe11a-84c3-7fc3-a84b-88787374b617"}]}`
const (
novelfullKey = "novelfull:reverend-insanity"
novelfullSeriesID = "reverend-insanity"
novelfullSeriesURI = "https://novelfull.com/reverend-insanity.html"
novelfullCoverURL = "https://novelfull.com/uploads/webp/novel/reverend-insanity-82661d911a.webp"
)
// newAcquirer wires an acquirer onto the store's creation hook, which is how
// main wires it: the write path is what starts an acquisition.
func newAcquirer(s *store.Store, page *fakeFetcher, covers *fakeBytesCoverFetcher) *Acquirer {
@@ -50,6 +73,30 @@ func readBookmark(t *testing.T, s *store.Store, key string) store.Bookmark {
return b
}
func bookmarkNewKaganeSeries(t *testing.T, s *store.Store) store.Bookmark {
t.Helper()
stored, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: kaganeKey, Site: "kagane", SeriesID: kaganeSeriesID,
Title: "Infinite Decryption", SeriesURL: kaganeSeriesURL, UpdatedAt: 1000,
})
if err != nil {
t.Fatalf("Upsert: %v", err)
}
return stored
}
func bookmarkNewNovelfullSeries(t *testing.T, s *store.Store) store.Bookmark {
t.Helper()
stored, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: novelfullKey, Site: "novelfull", SeriesID: novelfullSeriesID,
Title: "Reverend Insanity", SeriesURL: novelfullSeriesURI, UpdatedAt: 1000,
})
if err != nil {
t.Fatalf("Upsert: %v", err)
}
return stored
}
// The reported bug: a Reader bookmarks a Series nobody holds and expects the
// Cover, not a broken image. Both facts come from the one series-page fetch.
func TestAcquireFillsChapterAndCoverFromOneFetch(t *testing.T) {
@@ -237,3 +284,147 @@ func TestAcquireDoesNotBlockTheWrite(t *testing.T) {
close(release)
acq.Wait()
}
// The second symptom of #47: a kagane Series bookmarked from a chapter page
// gets its Cover at creation, with the bytes fetched through the browser
// sidecar — the only path that clears the challenge — into the
// content-addressed store.
func TestAcquireKaganeCoverThroughBrowser(t *testing.T) {
s, _ := newTestStore(t)
tlsPage := &fakeFetcher{body: "", status: 403}
browserPage := &fakeFetcher{body: kaganeSeriesAndCoverFixture, status: 200}
covers := &fakeCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
acq := &Acquirer{
Store: s, Fetch: tlsPage, BrowserFetch: browserPage,
BrowserCoverFetch: covers, Covers: &fakeBytesCoverFetcher{},
}
s.OnSeriesCreated = acq.Acquire
bookmarkNewKaganeSeries(t, s)
acq.Wait()
if got := tlsPage.callCount(); got != 0 {
t.Fatalf("plain-TLS page fetches = %d, want 0 — kagane pages are browser-only", got)
}
if got := browserPage.callCount(); got != 1 {
t.Fatalf("browser page fetches = %d, want 1", got)
}
if got := covers.callCount(); got != 1 {
t.Fatalf("browser cover fetches = %d, want 1", got)
}
if got := covers.calls[0]; got != kaganeImageID {
t.Fatalf("browser cover fetched image id %q, want %q", got, kaganeImageID)
}
got := readBookmark(t, s, kaganeKey)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(kaganeCoverSrc); got.Cover != want {
t.Fatalf("Cover = %q, want the content-addressed URL %q", got.Cover, want)
}
body, contentType, ok, err := s.CoverByAddress(store.CoverAddress(kaganeCoverSrc))
if err != nil || !ok {
t.Fatalf("CoverByAddress = %v, %v", ok, err)
}
if string(body) != "cover-bytes" || contentType != "image/webp" {
t.Fatalf("stored cover = (%q, %q), want the browser-fetched bytes", body, contentType)
}
}
// novelfull needs the browser only for its HTML: the cover URL comes out of
// the browser-fetched page, but the bytes go over plain TLS through the
// ordinary gated fetcher, never through the browser (issue #62).
func TestAcquireNovelfullCoverOverPlainTLS(t *testing.T) {
s, _ := newTestStore(t)
browserPage := &fakeFetcher{body: novelfullSeriesFixture + novelfullCoverFixture, status: 200}
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
acq := &Acquirer{
Store: s, Fetch: &fakeFetcher{body: "", status: 403},
BrowserFetch: browserPage, Covers: covers,
}
s.OnSeriesCreated = acq.Acquire
bookmarkNewNovelfullSeries(t, s)
acq.Wait()
if got := browserPage.callCount(); got != 1 {
t.Fatalf("browser page fetches = %d, want 1", got)
}
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetches = %d, want 1 — novelfull bytes never touch the browser", got)
}
if got := covers.calls[0]; got != novelfullCoverURL {
t.Fatalf("cover fetched from %q, want %q", got, novelfullCoverURL)
}
got := readBookmark(t, s, novelfullKey)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(novelfullCoverURL); got.Cover != want {
t.Fatalf("Cover = %q, want %q", got.Cover, want)
}
}
// With no browser sidecar configured, kagane is simply not acquired: no
// request is spent on a page that could only ever answer with a challenge,
// and nothing falls back to a plain fetch.
func TestAcquireKaganeSkippedWithoutBrowser(t *testing.T) {
s, _ := newTestStore(t)
tlsPage := &fakeFetcher{body: kaganeSeriesAndCoverFixture, status: 200}
acq := &Acquirer{
Store: s, Fetch: tlsPage,
Covers: &fakeBytesCoverFetcher{body: []byte("x"), contentType: "image/webp"},
}
s.OnSeriesCreated = acq.Acquire
bookmarkNewKaganeSeries(t, s)
acq.Wait()
if got := tlsPage.callCount(); got != 0 {
t.Fatalf("plain-TLS fetches for kagane = %d, want 0", got)
}
if got := readBookmark(t, s, kaganeKey); got.Cover != "" {
t.Fatalf("Cover = %q, want empty without a browser", got.Cover)
}
}
// The byte half of "nothing falls back to a plain fetch": with a browser for
// the page but none for the bytes, a kagane Cover stays absent and the TLS
// cover fetcher is never consulted.
func TestAcquireKaganeBytesNeverFallBackToPlainTLS(t *testing.T) {
s, _ := newTestStore(t)
browserPage := &fakeFetcher{body: kaganeSeriesAndCoverFixture, status: 200}
tlsCovers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
acq := &Acquirer{
Store: s, Fetch: &fakeFetcher{body: "", status: 403},
BrowserFetch: browserPage, Covers: tlsCovers,
}
s.OnSeriesCreated = acq.Acquire
bookmarkNewKaganeSeries(t, s)
acq.Wait()
if got := tlsCovers.callCount(); got != 0 {
t.Fatalf("plain-TLS cover fetches = %d, want 0 — kagane bytes are browser-only", got)
}
if got := readBookmark(t, s, kaganeKey); got.Cover != "" {
t.Fatalf("Cover = %q, want empty without a browser cover fetcher", got.Cover)
}
}
// novelfull's no-browser degradation differs from kagane's: only its HTML
// needs the sidecar, so when the page body is available — the challenge is a
// live time-varying fact that sometimes answers a plain request — the Cover
// still lands, bytes over plain TLS.
func TestAcquireNovelfullCoverWithoutBrowser(t *testing.T) {
s, _ := newTestStore(t)
tlsPage := &fakeFetcher{body: novelfullSeriesFixture + novelfullCoverFixture, status: 200}
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
acq := &Acquirer{Store: s, Fetch: tlsPage, Covers: covers}
s.OnSeriesCreated = acq.Acquire
bookmarkNewNovelfullSeries(t, s)
acq.Wait()
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetches = %d, want 1", got)
}
got := readBookmark(t, s, novelfullKey)
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(novelfullCoverURL); got.Cover != want {
t.Fatalf("Cover = %q, want %q", got.Cover, want)
}
}
+23
View File
@@ -22,6 +22,29 @@ type CoverBytesFetcher interface {
Fetch(ctx context.Context, sourceURL string) (body []byte, contentType string, err error)
}
// fetchCoverBytes routes a cover's byte retrieval by Site: only kagane needs
// the browser for image bytes — its covers answer a plain fetch with a
// challenge and `cross-origin-resource-policy: same-origin` — while every
// other Site's CDN answers plain TLS. Missing fetchers degrade to an error the
// caller logs, never a fallback onto a path that cannot succeed. One routing
// rule for the poll and the acquirer, so the two cannot drift apart.
func fetchCoverBytes(ctx context.Context, site, cover string, browser BrowserCoverFetcher, tls CoverBytesFetcher) ([]byte, string, error) {
if site == "kagane" {
if browser == nil {
return nil, "", errors.New("no cover fetcher")
}
imageID, ok := store.KaganeImageID(cover)
if !ok {
return nil, "", errors.New("invalid kagane cover URL")
}
return browser.Image(ctx, imageID)
}
if tls == nil {
return nil, "", errors.New("no cover fetcher")
}
return tls.Fetch(ctx, cover)
}
// CoverResolver resolves a host before any connection is attempted. Tests
// inject it to exercise hostile DNS results without touching the live network.
type CoverResolver func(context.Context, string) ([]netip.Addr, error)
+22 -31
View File
@@ -2,7 +2,6 @@ package latest
import (
"context"
"errors"
"log"
"net/url"
"slices"
@@ -107,7 +106,7 @@ func (p *Poller) prefetchCover(ctx context.Context, sr store.Series) {
// failure is logged against the Series and swallowed so the chapter poll
// cannot see it.
func (p *Poller) storeCover(ctx context.Context, sr store.Series, sourceURL string) {
bytes, contentType, err := p.fetchCoverBytes(ctx, sr, sourceURL)
bytes, contentType, err := fetchCoverBytes(ctx, sr.Site, sourceURL, p.CoverFetch, p.CoverBytesFetch)
if err != nil {
log.Printf("latest poll %q: fetch cover %s: %v", sr.Key(), sourceURL, err)
return
@@ -117,36 +116,28 @@ func (p *Poller) storeCover(ctx context.Context, sr store.Series, sourceURL stri
}
}
// fetchCoverBytes routes by Site: only kagane needs the browser for image
// bytes; every other Site's CDN answers plain TLS. Missing fetchers degrade to
// a blank Cover rather than falling back onto a path that cannot succeed.
func (p *Poller) fetchCoverBytes(ctx context.Context, sr store.Series, cover string) ([]byte, string, error) {
if sr.Site == "kagane" {
if p.CoverFetch == nil {
return nil, "", errors.New("no cover fetcher")
// fetcherFor returns the fetcher a site's page needs, or nil when the site
// cannot be fetched at all right now. kagane and novelfull pages sit behind a
// Cloudflare JavaScript challenge that no TLS fingerprint clears (kagane
// verified 2026-08-03, novelfull verified 2026-08-05, both against the same
// Chrome_133 profile TLSFetcher uses), so both prefer the browser; novelfull
// alone falls back to the plain-TLS fetcher when no browser is configured,
// because its challenge is a live time-varying fact (AGENTS.md) and its cover
// bytes never need the browser. kagane never falls back: a plain fetch of a
// kagane page or cover would only ever retrieve a challenge page. One routing
// rule for the poll and the acquirer, so the two cannot drift apart.
func fetcherFor(site string, browser, tls Fetcher) Fetcher {
switch {
case site == "kagane":
return browser
case slices.Contains(browserBackedSites, site): // novelfull
if browser != nil {
return browser
}
imageID, ok := store.KaganeImageID(cover)
if !ok {
return nil, "", errors.New("invalid kagane cover URL")
return tls
default:
return tls
}
return p.CoverFetch.Image(ctx, imageID)
}
if p.CoverBytesFetch == nil {
return nil, "", errors.New("no cover fetcher")
}
return p.CoverBytesFetch.Fetch(ctx, cover)
}
// fetcherFor returns the fetcher a site needs, or nil when the site cannot be
// fetched at all right now. kagane and novelfull both sit behind a Cloudflare
// JavaScript challenge that no TLS fingerprint clears — kagane verified
// 2026-08-03, novelfull verified 2026-08-05, both against the same Chrome_133
// profile TLSFetcher uses — so they are browser-only or nothing.
func (p *Poller) fetcherFor(site string) Fetcher {
if slices.Contains(browserBackedSites, site) {
return p.BrowserFetch
}
return p.Fetch
}
// Run polls until ctx is cancelled.
@@ -242,7 +233,7 @@ func (p *Poller) checkOne(ctx context.Context, sr store.Series) {
return
}
f := p.fetcherFor(sr.Site)
f := fetcherFor(sr.Site, p.BrowserFetch, p.Fetch)
if f == nil {
log.Printf("latest poll %q: no fetcher for site %q", sr.Key(), sr.Site)
return
+55 -6
View File
@@ -671,6 +671,51 @@ func TestKaganeSkippedWhenNoBrowserFetcher(t *testing.T) {
}
}
// novelfull without a browser is not skipped outright: its challenge is a
// live time-varying fact, so the plain-TLS page fetch is attempted and — when
// the body answers — fills both the chapter and the Cover, exactly the
// client-scraped rows #62 wants healed.
func TestNovelfullUsesTLSWhenNoBrowserFetcher(t *testing.T) {
s, _ := newTestStore(t)
key := "novelfull:reverend-insanity"
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: key,
Site: "novelfull",
SeriesID: "reverend-insanity",
SeriesURL: "https://novelfull.com/reverend-insanity.html",
UpdatedAt: 1000,
}); err != nil {
t.Fatalf("seed: %v", err)
}
tlsF := &fakeFetcher{body: novelfullSeriesFixture + novelfullCoverFixture, status: 200}
covers := &fakeBytesCoverFetcher{body: []byte("cover-bytes"), contentType: "image/webp"}
p := &Poller{
Store: s, Fetch: tlsF, CoverBytesFetch: covers,
Now: func() time.Time { return time.UnixMilli(5_000_000) },
Cooldown: time.Hour, BrowserCooldown: time.Hour,
Interval: time.Hour, Batch: 10,
}
p.runOnce(context.Background())
if len(tlsF.calls) != 1 {
t.Fatalf("TLS fetcher calls = %d, want 1", len(tlsF.calls))
}
if got := covers.callCount(); got != 1 {
t.Fatalf("cover fetches = %d, want 1", got)
}
got, found, err := s.Get(s.OwnerID(), key)
if err != nil || !found {
t.Fatalf("Get: %v found=%v", err, found)
}
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(novelfullCoverURL); got.Cover != want {
t.Fatalf("Cover = %q, want %q", got.Cover, want)
}
if got.LatestChapterNum == nil || *got.LatestChapterNum != 2334 {
t.Fatalf("LatestChapterNum = %v, want 2334", got.LatestChapterNum)
}
}
// With a browser fetcher wired up, kagane goes to it and not to the TLS one.
func TestKaganeUsesBrowserFetcher(t *testing.T) {
s, _ := newTestStore(t)
@@ -893,20 +938,24 @@ func TestRunOnceRoutesNonKaganeCoverToPublicFetcher(t *testing.T) {
func TestFetcherForRoutesNovelSites(t *testing.T) {
tls := &fakeFetcher{}
browser := &fakeFetcher{}
p := &Poller{Fetch: tls, BrowserFetch: browser}
cases := []struct {
site string
browser Fetcher
tls Fetcher
want Fetcher
}{
{"asura", tls},
{"lightnovelworld", tls},
{"kagane", browser},
{"novelfull", browser},
{"asura", browser, tls, tls},
{"lightnovelworld", browser, tls, tls},
{"kagane", browser, tls, browser},
{"novelfull", browser, tls, browser},
// browser-less deployment: kagane is nothing, novelfull degrades to TLS
{"kagane", nil, tls, nil},
{"novelfull", nil, tls, tls},
}
for _, tc := range cases {
t.Run(tc.site, func(t *testing.T) {
if got := p.fetcherFor(tc.site); got != tc.want {
if got := fetcherFor(tc.site, tc.browser, tc.tls); got != tc.want {
t.Fatalf("fetcherFor(%q) = %v, want %v", tc.site, got, tc.want)
}
})
@@ -6,6 +6,8 @@ import (
"os"
"testing"
"time"
"bookmarkmanager/backend/internal/store"
)
// TestSmokeKaganeImage is the live proof that the cover proxy's fetch actually
@@ -92,3 +94,62 @@ func TestSmokeKaganeGet(t *testing.T) {
t.Fatalf("status = %d, want 200 — the sidecar is not clearing the challenge", status)
}
}
// TestSmokeAcquireKaganeCover proves the #62 acquisition path end to end
// against the real browser: a kagane Series bookmarked at creation gets its
// Cover, bytes fetched through the sidecar into the content-addressed store.
// Same SMOKE_BROWSER_WS_URL gate as the tests above; a red run means the
// challenge is not clearing from this IP (a live fact to re-check), not
// necessarily a defect in the pipeline.
func TestSmokeAcquireKaganeCover(t *testing.T) {
ws := os.Getenv("SMOKE_BROWSER_WS_URL")
if ws == "" {
t.Skip("SMOKE_BROWSER_WS_URL unset")
}
const (
seriesID = "019fe11a-8670-7cf3-8343-0b02057d3787"
coverURL = "https://kagane.to/api/v2/image/019fe11a-84c3-7fc3-a84b-88787374b617/compressed"
)
s, _ := newTestStore(t)
bf, err := NewBrowserFetcher(ws)
if err != nil {
t.Fatalf("NewBrowserFetcher: %v", err)
}
defer bf.Close()
tlsF, err := NewTLSFetcher()
if err != nil {
t.Fatalf("NewTLSFetcher: %v", err)
}
acq := &Acquirer{
Store: s, Fetch: tlsF, BrowserFetch: bf,
BrowserCoverFetch: bf, Covers: NewCoverFetcher(),
}
s.OnSeriesCreated = acq.Acquire
if _, err := s.Upsert(s.OwnerID(), store.Bookmark{
Key: "kagane:" + seriesID, Site: "kagane", SeriesID: seriesID,
Title: "smoke", SeriesURL: "https://kagane.to/series/" + seriesID, UpdatedAt: 1000,
}); err != nil {
t.Fatalf("Upsert: %v", err)
}
acq.Wait()
got, found, err := s.Get(s.OwnerID(), "kagane:"+seriesID)
if err != nil || !found {
t.Fatalf("Get: %v found=%v", err, found)
}
if want := testCoverBaseURL + "/covers/" + store.CoverAddress(coverURL); got.Cover != want {
t.Fatalf("Cover = %q, want %q — the acquire path did not store the browser-fetched bytes", got.Cover, want)
}
body, contentType, ok, err := s.CoverByAddress(store.CoverAddress(coverURL))
if err != nil || !ok {
t.Fatalf("CoverByAddress: %v found=%v", err, ok)
}
if len(body) < 1000 {
t.Fatalf("stored cover is %d bytes, want a real image", len(body))
}
if contentType != "image/webp" {
t.Fatalf("content type = %q, want image/webp", contentType)
}
t.Logf("stored %d bytes of %s", len(body), contentType)
}
+19 -2
View File
@@ -328,10 +328,27 @@ func main() {
// Cover from one fetch, at creation, instead of waiting out a poll queue
// ordered by Reader count. Off the write path: the hook returns as soon
// as the goroutine is started.
var tlsFetch latest.Fetcher
if f, err := latest.NewTLSFetcher(); err != nil {
log.Printf("creation-time acquisition disabled, cannot build client: %v", err)
log.Printf("creation-time acquisition: plain-TLS Sites disabled, cannot build client: %v", err)
} else {
acq := &latest.Acquirer{Store: s, Fetch: f, Covers: latest.NewCoverFetcher(), Ctx: pollCtx}
tlsFetch = f
}
var browserCover latest.BrowserCoverFetcher
if b, ok := browser.(latest.BrowserCoverFetcher); ok {
browserCover = b
}
// The Acquirer must survive a TLS client failure: kagane needs only the
// sidecar, and novelfull degrades to whatever is left.
if tlsFetch != nil || browser != nil {
acq := &latest.Acquirer{
Store: s,
Fetch: tlsFetch,
BrowserFetch: browser,
BrowserCoverFetch: browserCover,
Covers: latest.NewCoverFetcher(),
Ctx: pollCtx,
}
s.OnSeriesCreated = acq.Acquire
}
startLatestPoller(pollCtx, s, cfg.LatestPoll, browser)