7c7d597019
Closes #63 Deletes the second way to reach a Cover. Since #62, every Site's cover bytes land in the content-addressed store at creation or on the poll, and the one public route serves them all — nothing needs the kagane proxy anymore. ## What went - **Template-level rewrite:** `Bookmark.CoverURL()` and both templates' use of it. Cards and chrome now render `.Cover` — the wire value — and nothing else. `Bookmark.CoverSource` was dead once `CoverURL` went, so it and its `bookmarkColumns` entry are gone too. - **Kagane-only cover route and its identifier validation:** `GET /img/kagane/{id}`, `web.CoverFetcher`, `coverIDRe`, and the whole `internal/web/cover.go`. - **The proxy's persistence:** `store.KaganeImageID`, `GetKaganeCover`, `PutKaganeCover`, `kaganeCoverSourceURL`, `kaganeCoverRe`. - **The kagane-shaped branch in the byte-fetch routing:** `fetchCoverBytes` no longer takes a `site` argument and no longer names a Site. The URL shape kagane's API publishes is claimed by the browser module itself — `kaganeImageURLRe` + `browserCoverURL` live in `latest/browser.go` with the rest of the per-Site knowledge — and `BrowserFetcher.Image` is now URL-driven (it validates the URL it will navigate to, same SSRF discipline as before). The no-plain-TLS-fallback rule for a claimed URL is preserved: a claimed address with no browser is an error, never a challenge-page fetch. ## What stayed (deliberately) - `BrowserFetcher.Image` and the browser-backed acquisition path: kagane genuinely serves cover bytes behind the challenge + `cross-origin-resource-policy: same-origin`, so the sidecar remains the only fetcher for them — it just routes by URL claim now instead of by Site name. - `fetcherFor`'s per-Site page routing (kagane/novelfull page fetches) — that is the page path, not a cover path. ## Acceptance criteria - [x] Template-level kagane cover rewrite gone - [x] Kagane-only cover route and its identifier validation gone - [x] Tests removed/rewritten against the general route, guarantees kept: unstored + traversal-shaped addresses serve nothing (`TestPublicCoverRejectsUnknownAddress`), non-image content types never echoed (`TestPublicCoverNeverEchoesNonImage` — new; the store-side gate was already pinned by `TestCoverStoreAcceptsAnySourceURL`). Store reopen-persistence and filesystem content-addressing tests rewritten against `PutCover`/`GetCover`, no guarantee lost. - [x] No Site name in a cover code path outside the acquisition module (`grep kagane backend`: store/web/templates/api are clean; remaining hits are `latest/browser.go` + `latest/sites.go`, tests, docs) - [x] Web UI and panel render Covers for all six Sites (templates render the wire address; panel renders `b.cover` — untouched, it never had a kagane path) - [x] `go test ./...` green ## Verification - `go vet ./...` clean - `go test ./...` — all packages pass (root 16.9s, latest 12.7s, store 12.7s, web 0.004s) - `CGO_ENABLED=0 go build` produces the static binary - Cover-path tests run verbosely: `TestPublicCoverServesStoredBytesUnauthenticated`, `TestPublicCoverRejectsUnknownAddress` (unknown/malformed/traversal/empty), `TestPublicCoverNeverEchoesNonImage`, `TestListRendersAcquiredCover`, `TestAcquireKaganeCoverThroughBrowser`, `TestRunOncePrefetchesKaganeCover`, `TestRunOnceRoutesNonKaganeCoverToPublicFetcher` all pass; the three `SMOKE_*` tests skip without the browser sidecar, as designed Live browser verification of the "web UI and panel render Covers for all six Sites" criterion is being run separately with Playwright against real Site pages and a locally mocked backend. Reviewed-on: #73 Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com> Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
328 lines
12 KiB
Go
328 lines
12 KiB
Go
package latest
|
|
|
|
import (
|
|
"context"
|
|
"log"
|
|
"net/url"
|
|
"slices"
|
|
"time"
|
|
|
|
"bookmarkmanager/backend/internal/store"
|
|
)
|
|
|
|
// Fetcher retrieves a series page. It exists as an interface so tests can inject
|
|
// a fake: nothing in the test suite may touch the network or the TLS client.
|
|
type Fetcher interface {
|
|
Get(ctx context.Context, url string) (body string, status int, err error)
|
|
}
|
|
|
|
// BrowserCoverFetcher retrieves one cover's bytes through the browser-backed
|
|
// path — the only route that clears the challenge kagane's image URLs answer
|
|
// a plain fetch with. Satisfied by BrowserFetcher.
|
|
type BrowserCoverFetcher interface {
|
|
Image(ctx context.Context, imageURL string) (body []byte, contentType string, err error)
|
|
}
|
|
|
|
// Poller re-checks each bookmarked series' newest published chapter on a
|
|
// schedule, independent of the userscript's own in-browser checks. The two run
|
|
// in parallel and report the same observable fact, so whichever writes last wins
|
|
// and neither needs to know about the other.
|
|
//
|
|
// Two clocks, deliberately independent:
|
|
//
|
|
// - Interval is how often this goroutine wakes up and looks.
|
|
// - Cooldowns are how long a series rests since its own last check. Browser-
|
|
// backed sites use the longer BrowserCooldown.
|
|
//
|
|
// Cooldowns are enforced by the WHERE clause in DueForLatestCheck rather than
|
|
// by any timer. Shortening Interval therefore cannot shorten anyone's cooldown;
|
|
// it only makes the poller wake up and find nothing due more often.
|
|
type Poller struct {
|
|
Store *store.Store
|
|
Fetch Fetcher
|
|
// BrowserFetch handles sites behind a JavaScript challenge that Fetch
|
|
// cannot clear. Nil disables those sites entirely rather than falling back
|
|
// to Fetch, which would only ever retrieve a challenge page.
|
|
BrowserFetch Fetcher
|
|
// CoverFetch is optional; failures are logged and never affect the chapter poll.
|
|
CoverFetch BrowserCoverFetcher
|
|
// CoverBytesFetch is optional; it handles plain-TLS sources through the
|
|
// same failure-isolated prefetch path.
|
|
CoverBytesFetch CoverBytesFetcher
|
|
Now func() time.Time // injected so tests can freeze it
|
|
Cooldown time.Duration
|
|
BrowserCooldown time.Duration
|
|
Interval time.Duration
|
|
Stagger time.Duration
|
|
Batch int
|
|
}
|
|
|
|
var browserBackedSites = []string{"kagane", "novelfull"}
|
|
|
|
// fillBlankCover gives a Series its Cover when it has none. The blank state is
|
|
// what "no Cover yet" means on the wire (ADR-0007): permanently-blank rows
|
|
// created before acquisition existed, and rows whose creation-time fetch
|
|
// failed, both heal here. A non-blank CoverAddress is left alone — refetching
|
|
// would add a request per Series per cycle and change artwork under the Reader
|
|
// for no visible reason. A row that already carries a source URL is owned by
|
|
// prefetchCover instead; this path only extracts from the series page.
|
|
//
|
|
// Failures are logged against the Series and never returned: the chapter poll
|
|
// must not notice. A failed fill is retried the next time this Series is due;
|
|
// there is no separate retry queue.
|
|
func (p *Poller) fillBlankCover(ctx context.Context, sr store.Series, body string) {
|
|
if sr.CoverAddress != "" || sr.Cover != "" {
|
|
return
|
|
}
|
|
cover, ok := coverFrom(sr.Site, sr.SeriesURL, body)
|
|
if !ok {
|
|
return
|
|
}
|
|
p.storeCover(ctx, sr, cover)
|
|
}
|
|
|
|
// prefetchCover heals Series that already carry a third-party source URL but
|
|
// no stored address — the state left by client-supplied covers before
|
|
// acquisition moved server-side. Every Site takes the same path; fetchCoverBytes
|
|
// routes by URL shape, so browser-claimed URLs still need the sidecar. New
|
|
// blanks have no source URL and go through fillBlankCover from the series page
|
|
// instead.
|
|
func (p *Poller) prefetchCover(ctx context.Context, sr store.Series) {
|
|
if sr.Cover == "" || sr.CoverAddress != "" {
|
|
return
|
|
}
|
|
body, contentType, found, err := p.Store.GetCover(sr.Cover)
|
|
if err != nil {
|
|
log.Printf("latest poll %q: read cover: %v", sr.Key(), err)
|
|
return
|
|
}
|
|
if found {
|
|
if err := p.Store.SetSeriesCover(sr.Site, sr.SeriesID, sr.Cover, body, contentType); err != nil {
|
|
log.Printf("latest poll %q: persist cover: %v", sr.Key(), err)
|
|
}
|
|
return
|
|
}
|
|
p.storeCover(ctx, sr, sr.Cover)
|
|
}
|
|
|
|
// storeCover fetches bytes for sourceURL and points the Series at them. Every
|
|
// failure is logged against the Series and swallowed so the chapter poll
|
|
// cannot see it.
|
|
func (p *Poller) storeCover(ctx context.Context, sr store.Series, sourceURL string) {
|
|
bytes, contentType, err := fetchCoverBytes(ctx, sourceURL, p.CoverFetch, p.CoverBytesFetch)
|
|
if err != nil {
|
|
log.Printf("latest poll %q: fetch cover %s: %v", sr.Key(), sourceURL, err)
|
|
return
|
|
}
|
|
if err := p.Store.SetSeriesCover(sr.Site, sr.SeriesID, sourceURL, bytes, contentType); err != nil {
|
|
log.Printf("latest poll %q: persist cover: %v", sr.Key(), err)
|
|
}
|
|
}
|
|
|
|
// fetcherFor returns the fetcher a site's page needs, or nil when the site
|
|
// cannot be fetched at all right now. kagane and novelfull pages sit behind a
|
|
// Cloudflare JavaScript challenge that no TLS fingerprint clears (kagane
|
|
// verified 2026-08-03, novelfull verified 2026-08-05, both against the same
|
|
// Chrome_133 profile TLSFetcher uses), so both prefer the browser; novelfull
|
|
// alone falls back to the plain-TLS fetcher when no browser is configured,
|
|
// because its challenge is a live time-varying fact (AGENTS.md) and its cover
|
|
// bytes never need the browser. kagane never falls back: a plain fetch of a
|
|
// kagane page or cover would only ever retrieve a challenge page. One routing
|
|
// rule for the poll and the acquirer, so the two cannot drift apart.
|
|
func fetcherFor(site string, browser, tls Fetcher) Fetcher {
|
|
switch {
|
|
case site == "kagane":
|
|
return browser
|
|
case slices.Contains(browserBackedSites, site): // novelfull
|
|
if browser != nil {
|
|
return browser
|
|
}
|
|
return tls
|
|
default:
|
|
return tls
|
|
}
|
|
}
|
|
|
|
// Run polls until ctx is cancelled.
|
|
//
|
|
// runOnce is called synchronously, so a batch that overruns the tick delays the
|
|
// next one instead of stacking a second batch on top of it. That is the intended
|
|
// failure mode for a misconfigured batch x stagger: a slower cadence, never
|
|
// concurrent fetch storms.
|
|
func (p *Poller) Run(ctx context.Context) {
|
|
log.Printf("latest-chapter poller: interval=%s cooldown=%s browser-cooldown=%s batch=%d stagger=%s",
|
|
p.Interval, p.Cooldown, p.BrowserCooldown, p.Batch, p.Stagger)
|
|
t := time.NewTicker(p.Interval)
|
|
defer t.Stop()
|
|
for {
|
|
select {
|
|
case <-ctx.Done():
|
|
log.Println("latest-chapter poller: stopped")
|
|
return
|
|
case <-t.C:
|
|
p.runOnce(ctx)
|
|
}
|
|
}
|
|
}
|
|
|
|
// runOnce processes one batch of due series.
|
|
func (p *Poller) runOnce(ctx context.Context) {
|
|
now := p.Now()
|
|
cutoff := now.Add(-p.Cooldown).UnixMilli()
|
|
browserCutoff := now.Add(-p.BrowserCooldown).UnixMilli()
|
|
due, err := p.Store.DueForLatestCheck(cutoff, browserCutoff, browserBackedSites, p.Batch)
|
|
if err != nil {
|
|
log.Printf("latest poll: due query: %v", err)
|
|
return
|
|
}
|
|
|
|
checked := 0
|
|
for i, sr := range due {
|
|
if ctx.Err() != nil {
|
|
break
|
|
}
|
|
// Staggered rather than fired together: a burst of simultaneous requests
|
|
// from one server IP is the traffic shape most likely to move that IP's
|
|
// bot score. This is the server-side analogue of the userscript's "one
|
|
// series per navigation ... indistinguishable from browsing" (L455-456).
|
|
stopped := false
|
|
if i > 0 && p.Stagger > 0 {
|
|
select {
|
|
case <-ctx.Done():
|
|
stopped = true
|
|
case <-time.After(p.Stagger):
|
|
}
|
|
}
|
|
if stopped {
|
|
break
|
|
}
|
|
p.checkOne(ctx, sr)
|
|
checked++
|
|
}
|
|
// due vs checked is how you tell which constraint is binding: ticks that
|
|
// report due=0 mean the cooldown is the limit, ticks that report due==batch
|
|
// every time mean throughput is.
|
|
log.Printf("latest poll: due=%d checked=%d", len(due), checked)
|
|
}
|
|
|
|
// checkOne re-checks one series. Every failure path here is "log and move on":
|
|
// the poller is a best-effort enhancement, and no single bad series may stall a
|
|
// batch or take down the process.
|
|
func (p *Poller) checkOne(ctx context.Context, sr store.Series) {
|
|
defer func() {
|
|
if r := recover(); r != nil {
|
|
log.Printf("latest poll %q: recovered from panic: %v", sr.Key(), r)
|
|
}
|
|
}()
|
|
|
|
// Stamped before the fetch, not after, so an error, a timeout, or a shutdown
|
|
// mid-request still consumes the cooldown. Otherwise a renamed or deleted
|
|
// series would be retried on every single tick forever. The userscript
|
|
// stamps in the same order and for the same reason (L471-473).
|
|
if err := p.Store.MarkLatestChecked(sr.Site, sr.SeriesID, p.Now().UnixMilli()); err != nil {
|
|
log.Printf("latest poll %q: mark checked: %v", sr.Key(), err)
|
|
return
|
|
}
|
|
|
|
// series_url is client-supplied (PUT /bookmarks/{key} accepts any string),
|
|
// so this is not just an optimisation against burning a request on an
|
|
// unknown site: without it, the server would issue a GET from its own
|
|
// network position to whatever URL a token-holder writes, including
|
|
// link-local/internal addresses or non-https schemes. The cooldown above
|
|
// is already consumed, so a row that never passes this check is retried at
|
|
// cooldown pace rather than hot-looping.
|
|
if !fetchableSeriesURL(sr.Site, sr.SeriesURL) {
|
|
log.Printf("latest poll %q: not fetchable: site=%q url=%q", sr.Key(), sr.Site, sr.SeriesURL)
|
|
return
|
|
}
|
|
|
|
f := fetcherFor(sr.Site, p.BrowserFetch, p.Fetch)
|
|
if f == nil {
|
|
log.Printf("latest poll %q: no fetcher for site %q", sr.Key(), sr.Site)
|
|
return
|
|
}
|
|
p.prefetchCover(ctx, sr)
|
|
|
|
body, status, err := f.Get(ctx, sr.SeriesURL)
|
|
if err != nil {
|
|
log.Printf("latest poll %q: fetch %s: %v", sr.Key(), sr.SeriesURL, err)
|
|
return
|
|
}
|
|
if status != 200 {
|
|
log.Printf("latest poll %q: fetch %s: status %d", sr.Key(), sr.SeriesURL, status)
|
|
return
|
|
}
|
|
|
|
latest, ok := latestChapterFrom(sr.Site, sr.SeriesURL, body)
|
|
// Cover fill is independent of the chapter signal: a page that lost its
|
|
// chapter list may keep its og:image, and a blank Series heals either way.
|
|
p.fillBlankCover(ctx, sr, body)
|
|
if !ok {
|
|
// Most likely a challenge page or a layout change. Either way the row is
|
|
// already stamped, so this waits out a cooldown instead of hot-looping.
|
|
log.Printf("latest poll %q: no chapter links in %d bytes", sr.Key(), len(body))
|
|
return
|
|
}
|
|
|
|
// Equality, not >, mirroring the userscript (L427): a site that retracts a
|
|
// chapter should correct the stored number downward. The comparison is
|
|
// against the due-query snapshot; a concurrent write in between only costs
|
|
// one redundant UPDATE of the same absolute value, never a wrong one.
|
|
if sr.LatestChapterNum != nil && *sr.LatestChapterNum == latest.Num {
|
|
return
|
|
}
|
|
|
|
// Series-level write: the row is shared, so one update refreshes every
|
|
// bookmark joining to it, and the bookmark's updated_at is never touched —
|
|
// a newly published chapter is not reading progress and must not reorder
|
|
// the list.
|
|
if err := p.Store.SetLatestChapter(sr.Site, sr.SeriesID, latest.Label, latest.Num); err != nil {
|
|
log.Printf("latest poll %q: set latest chapter: %v", sr.Key(), err)
|
|
return
|
|
}
|
|
log.Printf("latest poll %q: latest is now %s", sr.Key(), latest.Label)
|
|
}
|
|
|
|
// fetchableSeriesURL reports whether site is a site latestChapterFrom knows how
|
|
// to parse and seriesURL is safe to hand to a fetcher: an https URL with a
|
|
// non-empty host. series_url comes from client-supplied PUT bodies, so this is
|
|
// a defence against the poller being used to probe arbitrary hosts from the
|
|
// server's own network position, not just a check against wasted requests.
|
|
//
|
|
// Three sites are held to a stricter rule, each for a different reason:
|
|
//
|
|
// - kagane and novelfull are fetched by a headless browser, which executes
|
|
// JavaScript and carries cookies, and is therefore a far stronger SSRF
|
|
// primitive than an HTTP GET. Their hosts must match exactly, not merely
|
|
// be non-empty.
|
|
// - lightnovelworld's parser regex hardcodes its host, so a URL anywhere
|
|
// else could never yield a match — reject it here rather than burn the
|
|
// request.
|
|
func fetchableSeriesURL(site, seriesURL string) bool {
|
|
switch site {
|
|
case "asura", "demonic", "comix", "kagane", "novelfull", "lightnovelworld":
|
|
default:
|
|
return false
|
|
}
|
|
u, err := url.Parse(seriesURL)
|
|
if err != nil {
|
|
return false
|
|
}
|
|
if u.Scheme != "https" || u.Host == "" {
|
|
return false
|
|
}
|
|
switch site {
|
|
case "kagane":
|
|
return u.Hostname() == "kagane.to"
|
|
case "novelfull":
|
|
// Fetched by a real browser, same as kagane, so the host is pinned
|
|
// rather than merely non-empty.
|
|
return u.Hostname() == "novelfull.com"
|
|
case "lightnovelworld":
|
|
// Its parser regex hardcodes this host, so a URL anywhere else could
|
|
// never yield a match — reject it here rather than burn the request.
|
|
return u.Hostname() == "lightnovelworld.net"
|
|
}
|
|
return true
|
|
}
|