Browser-backed Sites join the Cover pipeline (#62) (#72)

Fixes #62

Browser-backed Sites join the Cover pipeline: kagane and novelfull Series now get their Covers at creation, through the same acquisition path as every other Site, instead of waiting for a poll pass.

## What changed

`latest.Acquirer` (creation-time acquisition, fired by the first Bookmark of a Series) previously skipped kagane and novelfull entirely — their pages only yield a Cloudflare challenge to the TLS client, so the request was spent for nothing. It now routes them like the poller does, with the two Sites split exactly as the issue demands:

- **kagane** — page fetched through the browser sidecar, cover URL extracted from the API JSON, bytes fetched through the browser sidecar (the only path that clears the challenge) into the content-addressed store. With no `BROWSER_WS_URL` configured, acquisition is skipped entirely and nothing falls back to a plain fetch.
- **novelfull** — page fetched through the browser sidecar, cover URL extracted from the HTML, bytes fetched over plain TLS through the ordinary gated fetcher (its image paths answer 200 with `access-control-allow-origin: *`, measured 2026-08-09). With no browser configured, the page fetch falls back to the TLS client — novelfull's challenge is a live time-varying fact (AGENTS.md), so when the page body answers, the Cover still lands; when it is challenged, nothing happens.

The byte-routing rule (kagane → browser, every other Site → TLS) is now one shared function (`latest.fetchCoverBytes`) used by both the Poller and the Acquirer, so the two cannot drift apart.

## Acceptance criteria

- [x] kagane cover bytes are fetched through the browser sidecar and stored in the content-addressed store — `TestAcquireKaganeCoverThroughBrowser`
- [x] novelfull cover URLs are extracted from the browser-fetched HTML, and its bytes are fetched over plain TLS — `TestAcquireNovelfullCoverOverPlainTLS`
- [x] With no browser sidecar configured, kagane Covers are absent and nothing falls back to a plain fetch — `TestAcquireKaganeSkippedWithoutBrowser`
- [x] With no browser sidecar configured, novelfull Covers still work if its page body is available — `TestAcquireNovelfullCoverWithoutBrowser`
- [x] Manually verified on-device: a kagane Series shows its Cover in the panel, not a broken-image glyph — being run by a separate manual-verification agent against a mocked scenario (no prod data); not part of this PR
- [x] `go test ./...` is green, with live-network checks gated behind `SMOKE_BROWSER_WS_URL` like the existing kagane image smoke test — new `TestSmokeAcquireKaganeCover` proves the end-to-end acquire path against the real browser when the env var is set

## Verification

- `go test ./...` green across all packages
- New unit tests exercise every routing decision with fakes — no network in the default suite
- Smoke test gated behind `SMOKE_BROWSER_WS_URL`, skipped by default

## Post-review changes (a66491a)

- **One routing rule for pages too** — `fetcherFor` is now a shared function used by both the Poller and the Acquirer; novelfull falls back to the plain-TLS fetcher in *both* when no browser is configured, so pre-existing (client-scraped) novelfull rows get healed by the poll as well, not just Series created after this change (`TestNovelfullUsesTLSWhenNoBrowserFetcher`).
- **Byte-level no-fallback proof** — `TestAcquireKaganeBytesNeverFallBackToPlainTLS` pins that kagane cover bytes never route to the TLS fetcher even when the page came through a browser.
- **Acquirer wired independent of the TLS client** — if `NewTLSFetcher` fails, kagane/novelfull acquisition still works via the sidecar (`main.go`).
- AGENTS.md (root + backend) updated for the novelfull plain-TLS fallback.

Reviewed-on: #72
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
This commit was merged in pull request #72.
This commit is contained in:
2026-08-10 11:06:36 +07:00
committed by sulthan
parent b9220b3dfc
commit 78234f3c19
9 changed files with 406 additions and 53 deletions
+20 -9
View File
@@ -3,7 +3,6 @@ package latest
import (
"context"
"log"
"slices"
"sync"
"time"
@@ -32,11 +31,21 @@ const acquireTimeout = 45 * time.Second
// left blank until the poll's own cover pass (#61) fills it.
type Acquirer struct {
Store *store.Store
// Fetch retrieves the series page. Nil disables acquisition entirely.
// Fetch retrieves the series page over plain TLS. Nil with a nil
// BrowserFetch disables acquisition entirely.
Fetch Fetcher
// BrowserFetch retrieves kagane and novelfull pages through the browser
// sidecar, the only thing that clears their Cloudflare challenge. The
// per-site fallback policy lives in fetcherFor. Nil leaves those Sites
// unacquired when no fallback applies.
BrowserFetch Fetcher
// Covers retrieves the cover bytes. Nil leaves the Cover blank and the
// chapter half working.
Covers CoverBytesFetcher
// BrowserCoverFetch retrieves kagane cover bytes through the browser
// sidecar. Nil leaves kagane Covers blank; nothing falls back to a plain
// fetch, which would only ever retrieve a challenge page.
BrowserCoverFetch BrowserCoverFetcher
// Ctx cancels in-flight acquisitions at shutdown. A hook signature has
// nowhere to pass one, so it lives here; nil means context.Background.
Ctx context.Context
@@ -85,10 +94,7 @@ func (a *Acquirer) Acquire(sr store.Series) {
func (a *Acquirer) Wait() { a.inflight.Wait() }
func (a *Acquirer) acquire(ctx context.Context, sr store.Series) {
// Browser-backed Sites are deliberately not acquired here: their pages
// only yield a Cloudflare challenge to the TLS client, so the request
// would be spent for nothing.
if a.Fetch == nil || slices.Contains(browserBackedSites, sr.Site) {
if a.Fetch == nil && a.BrowserFetch == nil {
return
}
// series_url arrives in a client-supplied PUT body, so the same gate the
@@ -99,7 +105,12 @@ func (a *Acquirer) acquire(ctx context.Context, sr store.Series) {
return
}
body, status, err := a.Fetch.Get(ctx, sr.SeriesURL)
f := fetcherFor(sr.Site, a.BrowserFetch, a.Fetch)
if f == nil {
log.Printf("acquire %q: no fetcher for site %q", sr.Key(), sr.Site)
return
}
body, status, err := f.Get(ctx, sr.SeriesURL)
if err != nil {
log.Printf("acquire %q: fetch %s: %v", sr.Key(), sr.SeriesURL, err)
return
@@ -122,10 +133,10 @@ func (a *Acquirer) acquire(ctx context.Context, sr store.Series) {
}
cover, ok := coverFrom(sr.Site, sr.SeriesURL, body)
if !ok || a.Covers == nil {
if !ok {
return
}
bytes, contentType, err := a.Covers.Fetch(ctx, cover)
bytes, contentType, err := fetchCoverBytes(ctx, sr.Site, cover, a.BrowserCoverFetch, a.Covers)
if err != nil {
log.Printf("acquire %q: fetch cover %s: %v", sr.Key(), cover, err)
return