feat: one registry entry per Site, one shared Series-page read (#95)
Closes #94. ## What Two phases per the spec, in three feature commits plus two review-fix commits: **Phase one — one registry entry per Site** (`2d134fb`) The six per-site comparison points that used to live across three files collapse into one `sites` map in `backend/internal/latest/sites.go`: Latest Chapter parse, Cover parse, browser-backed list, fetcher route, host pins, and the browser payload read all become lookups into it. `browserBackedSites()` is derived from the registry (sorted, deterministic); `fetcherFor` and `fetchableSeriesURL` keep their signatures and become lookups; `BrowserFetcher.Get` dispatches through the entries' `Read`/`Done` while the tab lifecycle stays in `BrowserFetcher.run`. **Phase two — one shared Series-page read** (`f215130`) `readSeriesPage` (new `read.go`) performs the read the Poll and the Acquisition have in common: gate, route, fetch, parse Latest Chapter, parse Cover address. It returns facts only — polling and persistence policies (stamp order, cooldowns, cover policy) stay with the callers; `acquire.go` gained the comment naming the deliberate post-fetch stamp order. The poll's legacy cover heal and the no-chapter byte-count diagnostic were restored after review (`d998f87`) so the claims "the Poll keeps its own Cover policy" and "pinning is the only behavioural change" both hold. ## Behaviour - All six Sites now pin their host exactly; asura/demonic/comix previously accepted any https host. For asura this is a strict improvement: its dead old domain redirects deep links to the site root and would parse the wrong document. - Everything else is unchanged: existing parse tables, the challenge-body table and the gate table pass unmodified except the one deliberate exception — the gate table gains the three new pin cases. ## Security invariants preserved - The address gate is recognisably the same rule, now a single registry lookup: `https` + exact hostname match, all callers route through it. No fetch path was widened; asura/demonic/comix were narrowed. - The second host pin inside each browser entry's Read is retained deliberately (browser = strong SSRF primitive, `series_url` is client-supplied) and is not deduplicated against the shared gate. - Review hardening: `fetcherFor` now fails closed for unknown site strings (previously fell through to the TLS fetcher on an unreachable path), and the browser dispatch iterates a sorted list so outcomes cannot depend on map order. - The security review's log-injection finding was checked against Go's `url.Parse` and does not hold: control characters are rejected anywhere in a URL, so a client-supplied value in a log line cannot carry a newline. ## Review Reviewed on three axes (spec, standards, security) by read-only subagents over `672c16f..f1b26f4`. No blocking findings; all minor/nit findings addressed in `d998f87` and `700de20`. Verified end to end with `go test ./...` (Docker Postgres per test package) on every commit. ## Out of scope (tracked separately) - Dropping asuracomic.net (CORS allowlist, userscript match, API fixtures, live env) — separate issue, per spec. Reviewed-on: #95 Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com> Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
This commit was merged in pull request #95.
This commit is contained in:
@@ -2,6 +2,7 @@ package latest
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"log"
|
||||
"sync"
|
||||
"time"
|
||||
@@ -97,51 +98,49 @@ func (a *Acquirer) acquire(ctx context.Context, sr store.Series) {
|
||||
if a.Fetch == nil && a.BrowserFetch == nil {
|
||||
return
|
||||
}
|
||||
// series_url arrives in a client-supplied PUT body, so the same gate the
|
||||
// poller uses applies here — without it a token-holder chooses what the
|
||||
// server fetches from its own network position.
|
||||
if !fetchableSeriesURL(sr.Site, sr.SeriesURL) {
|
||||
log.Printf("acquire %q: not fetchable: site=%q url=%q", sr.Key(), sr.Site, sr.SeriesURL)
|
||||
return
|
||||
}
|
||||
|
||||
f := fetcherFor(sr.Site, a.BrowserFetch, a.Fetch)
|
||||
if f == nil {
|
||||
log.Printf("acquire %q: no fetcher for site %q", sr.Key(), sr.Site)
|
||||
return
|
||||
}
|
||||
body, status, err := f.Get(ctx, sr.SeriesURL)
|
||||
facts, err := readSeriesPage(ctx, sr.Site, sr.SeriesURL, a.BrowserFetch, a.Fetch)
|
||||
if err != nil {
|
||||
log.Printf("acquire %q: fetch %s: %v", sr.Key(), sr.SeriesURL, err)
|
||||
return
|
||||
}
|
||||
if status != 200 {
|
||||
log.Printf("acquire %q: fetch %s: status %d", sr.Key(), sr.SeriesURL, status)
|
||||
switch {
|
||||
case errors.Is(err, errNotFetchable):
|
||||
// series_url arrives in a client-supplied PUT body, so without the
|
||||
// gate a token-holder chooses what the server fetches from its own
|
||||
// network position.
|
||||
log.Printf("acquire %q: not fetchable: site=%q url=%q", sr.Key(), sr.Site, sr.SeriesURL)
|
||||
case errors.Is(err, errNoFetcher):
|
||||
log.Printf("acquire %q: no fetcher for site %q", sr.Key(), sr.Site)
|
||||
default:
|
||||
log.Printf("acquire %q: %v", sr.Key(), err)
|
||||
}
|
||||
return
|
||||
}
|
||||
|
||||
// This page just served the same purpose a poll tick would have; without
|
||||
// the stamp the row stays due and the poller refetches it immediately.
|
||||
//
|
||||
// Stamped after success — the reverse of the poller, which stamps before
|
||||
// the fetch: the Reader is here, watching the Series they just created, so
|
||||
// a failed acquisition must leave the row due for a fast retry rather than
|
||||
// consuming the cooldown. The stamp happens even when the page read
|
||||
// succeeded but produced no facts to persist.
|
||||
if err := a.Store.MarkLatestChecked(sr.Site, sr.SeriesID, time.Now().UnixMilli()); err != nil {
|
||||
log.Printf("acquire %q: mark checked: %v", sr.Key(), err)
|
||||
}
|
||||
|
||||
if latest, ok := latestChapterFrom(sr.Site, sr.SeriesURL, body); ok {
|
||||
if err := a.Store.SetLatestChapter(sr.Site, sr.SeriesID, latest.Label, latest.Num); err != nil {
|
||||
if facts.HasLatest {
|
||||
if err := a.Store.SetLatestChapter(sr.Site, sr.SeriesID, facts.Latest.Label, facts.Latest.Num); err != nil {
|
||||
log.Printf("acquire %q: set latest chapter: %v", sr.Key(), err)
|
||||
}
|
||||
}
|
||||
|
||||
cover, ok := coverFrom(sr.Site, sr.SeriesURL, body)
|
||||
if !ok {
|
||||
if !facts.HasCover {
|
||||
return
|
||||
}
|
||||
bytes, contentType, err := fetchCoverBytes(ctx, cover, a.BrowserCoverFetch, a.Covers)
|
||||
bytes, contentType, err := fetchCoverBytes(ctx, facts.Cover, a.BrowserCoverFetch, a.Covers)
|
||||
if err != nil {
|
||||
log.Printf("acquire %q: fetch cover %s: %v", sr.Key(), cover, err)
|
||||
log.Printf("acquire %q: fetch cover %s: %v", sr.Key(), facts.Cover, err)
|
||||
return
|
||||
}
|
||||
if err := a.Store.SetSeriesCover(sr.Site, sr.SeriesID, cover, bytes, contentType); err != nil {
|
||||
if err := a.Store.SetSeriesCover(sr.Site, sr.SeriesID, facts.Cover, bytes, contentType); err != nil {
|
||||
log.Printf("acquire %q: persist cover: %v", sr.Key(), err)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -86,46 +86,59 @@ func (f *BrowserFetcher) Close() {
|
||||
f.cancel()
|
||||
}
|
||||
|
||||
// Get navigates to seriesURL, lets any challenge resolve, then reads either the
|
||||
// site's JSON API (kagane) from inside the page so the request carries the
|
||||
// clearance cookie, or the served HTML itself (novelfull) — see
|
||||
// novelfullSeriesURL for the latter case. The returned body is whatever the
|
||||
// site's chapter list lives in, which is what latestChapterFrom's per-site
|
||||
// switch expects.
|
||||
// Get navigates to seriesURL, lets any challenge resolve, then reads the
|
||||
// payload the Site's registry entry describes — kagane's chapter-list API from
|
||||
// inside the page so the request carries the clearance cookie, novelfull's
|
||||
// served HTML. The returned body is whatever the Site's chapter list lives in,
|
||||
// which is what the entry's LatestChapter parse expects.
|
||||
func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int, error) {
|
||||
apiURL, isKagane := kaganeAPIURL(seriesURL)
|
||||
if !isKagane && !novelfullSeriesURL(seriesURL) {
|
||||
return "", 0, fmt.Errorf("not a fetchable browser series url: %q", seriesURL)
|
||||
}
|
||||
|
||||
var body string
|
||||
// kagane's chapter list is only in its JSON API, which must be called from
|
||||
// inside the page so the request carries the clearance cookie. novelfull
|
||||
// renders its chapters into the HTML, so the cleared DOM is the answer.
|
||||
// chromedp.OuterHTML returns a QueryAction and chromedp.Evaluate an
|
||||
// EvaluateAction, so the variable has to be the interface both implement.
|
||||
var read chromedp.Action = chromedp.OuterHTML("html", &body, chromedp.ByQuery)
|
||||
if isKagane {
|
||||
read = chromedp.Evaluate(
|
||||
`fetch(`+jsString(apiURL)+`).then(r => r.ok ? r.text() : "")`,
|
||||
&body,
|
||||
awaitPromise,
|
||||
)
|
||||
}
|
||||
|
||||
// novelfull's payload is the DOM itself, and the interstitial has a DOM
|
||||
// too, so "we have an answer" has to exclude it explicitly. kagane's
|
||||
// in-page fetch just fails while challenged, which is already the signal.
|
||||
done := func() bool { return body != "" && (isKagane || !isInterstitial(body)) }
|
||||
if err := f.run(ctx, seriesURL, read, done); err != nil {
|
||||
// Challenge never cleared, or the API refused. Indistinguishable from
|
||||
// here and handled identically by the caller.
|
||||
if errors.Is(err, errChallengeHeld) {
|
||||
return "", 403, nil
|
||||
// Sorted order (browserBackedSites sorts) makes dispatch deterministic:
|
||||
// entries' Read funcs are expected to refuse any address owned by another
|
||||
// Site, and the loop must not depend on that staying true.
|
||||
for _, name := range browserBackedSites() {
|
||||
s := sites[name]
|
||||
read, ok := s.Browser.Read(seriesURL, &body)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
return "", 0, fmt.Errorf("browser fetch %q: %w", seriesURL, err)
|
||||
if err := f.run(ctx, seriesURL, read,
|
||||
func() bool { return s.Browser.Done(body) }); err != nil {
|
||||
// Challenge never cleared, or the payload was refused.
|
||||
// Indistinguishable from here and handled identically by the caller.
|
||||
if errors.Is(err, errChallengeHeld) {
|
||||
return "", 403, nil
|
||||
}
|
||||
return "", 0, fmt.Errorf("browser fetch %q: %w", seriesURL, err)
|
||||
}
|
||||
return body, 200, nil
|
||||
}
|
||||
return body, 200, nil
|
||||
return "", 0, fmt.Errorf("not a fetchable browser series url: %q", seriesURL)
|
||||
}
|
||||
|
||||
// kaganeRead builds the in-tab fetch of kagane's chapter-list API: the
|
||||
// request must be made from inside the page so it carries the clearance
|
||||
// cookie, and the API is the only place the list exists. Refusing any other
|
||||
// address is the per-Site half of the SSRF gate, kept deliberately behind
|
||||
// fetchableSeriesURL (see browserRead.Read).
|
||||
func kaganeRead(seriesURL string, out *string) (chromedp.Action, bool) {
|
||||
apiURL, ok := kaganeAPIURL(seriesURL)
|
||||
if !ok {
|
||||
return nil, false
|
||||
}
|
||||
return chromedp.Evaluate(
|
||||
`fetch(`+jsString(apiURL)+`).then(r => r.ok ? r.text() : "")`,
|
||||
out, awaitPromise), true
|
||||
}
|
||||
|
||||
// novelfullRead reads the cleared DOM. novelfull renders its chapter list
|
||||
// into the served HTML, so there is no API to call from inside the page — the
|
||||
// challenge-cleared DOM is the payload.
|
||||
func novelfullRead(seriesURL string, out *string) (chromedp.Action, bool) {
|
||||
if !novelfullSeriesURL(seriesURL) {
|
||||
return nil, false
|
||||
}
|
||||
return chromedp.OuterHTML("html", out, chromedp.ByQuery), true
|
||||
}
|
||||
|
||||
// Image retrieves one cover's bytes through the browser sidecar, and its
|
||||
|
||||
@@ -2,9 +2,9 @@ package latest
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"log"
|
||||
"net/url"
|
||||
"slices"
|
||||
"time"
|
||||
|
||||
"bookmarkmanager/backend/internal/store"
|
||||
@@ -57,25 +57,23 @@ type Poller struct {
|
||||
Batch int
|
||||
}
|
||||
|
||||
var browserBackedSites = []string{"kagane", "novelfull"}
|
||||
|
||||
// fillBlankCover gives a Series its Cover when it has none. The blank state is
|
||||
// what "no Cover yet" means on the wire (ADR-0007): permanently-blank rows
|
||||
// created before acquisition existed, and rows whose creation-time fetch
|
||||
// failed, both heal here. A non-blank CoverAddress is left alone — refetching
|
||||
// would add a request per Series per cycle and change artwork under the Reader
|
||||
// for no visible reason. A row that already carries a source URL is owned by
|
||||
// prefetchCover instead; this path only extracts from the series page.
|
||||
// prefetchCover instead; this path only records a Cover address already
|
||||
// extracted from the series page.
|
||||
//
|
||||
// Failures are logged against the Series and never returned: the chapter poll
|
||||
// must not notice. A failed fill is retried the next time this Series is due;
|
||||
// there is no separate retry queue.
|
||||
func (p *Poller) fillBlankCover(ctx context.Context, sr store.Series, body string) {
|
||||
func (p *Poller) fillBlankCover(ctx context.Context, sr store.Series, cover string) {
|
||||
if sr.CoverAddress != "" || sr.Cover != "" {
|
||||
return
|
||||
}
|
||||
cover, ok := coverFrom(sr.Site, sr.SeriesURL, body)
|
||||
if !ok {
|
||||
if cover == "" {
|
||||
return
|
||||
}
|
||||
p.storeCover(ctx, sr, cover)
|
||||
@@ -120,27 +118,29 @@ func (p *Poller) storeCover(ctx context.Context, sr store.Series, sourceURL stri
|
||||
}
|
||||
|
||||
// fetcherFor returns the fetcher a site's page needs, or nil when the site
|
||||
// cannot be fetched at all right now. kagane and novelfull pages sit behind a
|
||||
// Cloudflare JavaScript challenge that no TLS fingerprint clears (kagane
|
||||
// verified 2026-08-03, novelfull verified 2026-08-05, both against the same
|
||||
// Chrome_133 profile TLSFetcher uses), so both prefer the browser; novelfull
|
||||
// alone falls back to the plain-TLS fetcher when no browser is configured,
|
||||
// because its challenge is a live time-varying fact (AGENTS.md) and its cover
|
||||
// bytes never need the browser. kagane never falls back: a plain fetch of a
|
||||
// kagane page or cover would only ever retrieve a challenge page. One routing
|
||||
// rule for the poll and the acquirer, so the two cannot drift apart.
|
||||
// cannot be fetched at all right now. A Site whose registry entry carries a
|
||||
// Browser read — kagane and novelfull, both behind a Cloudflare JavaScript
|
||||
// challenge no TLS fingerprint clears — prefers the browser; when it is
|
||||
// absent, the entry's Fallback decides whether plain TLS may take over. One
|
||||
// routing rule for the poll and the acquirer, so the two cannot drift apart.
|
||||
func fetcherFor(site string, browser, tls Fetcher) Fetcher {
|
||||
switch {
|
||||
case site == "kagane":
|
||||
return browser
|
||||
case slices.Contains(browserBackedSites, site): // novelfull
|
||||
if browser != nil {
|
||||
return browser
|
||||
}
|
||||
return tls
|
||||
default:
|
||||
s, known := sites[site]
|
||||
if !known {
|
||||
// No registry entry means nothing to fetch or parse; fail closed even
|
||||
// though the only caller gates first, so a future caller that skips
|
||||
// the gate cannot hand an arbitrary https URL to the TLS fetcher.
|
||||
return nil
|
||||
}
|
||||
if s.Browser == nil {
|
||||
return tls
|
||||
}
|
||||
if browser != nil {
|
||||
return browser
|
||||
}
|
||||
if s.Browser.Fallback {
|
||||
return tls
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// Run polls until ctx is cancelled.
|
||||
@@ -170,7 +170,7 @@ func (p *Poller) runOnce(ctx context.Context) {
|
||||
now := p.Now()
|
||||
cutoff := now.Add(-p.Cooldown).UnixMilli()
|
||||
browserCutoff := now.Add(-p.BrowserCooldown).UnixMilli()
|
||||
due, err := p.Store.DueForLatestCheck(cutoff, browserCutoff, browserBackedSites, p.Batch)
|
||||
due, err := p.Store.DueForLatestCheck(cutoff, browserCutoff, browserBackedSites(), p.Batch)
|
||||
if err != nil {
|
||||
log.Printf("latest poll: due query: %v", err)
|
||||
return
|
||||
@@ -224,43 +224,35 @@ func (p *Poller) checkOne(ctx context.Context, sr store.Series) {
|
||||
return
|
||||
}
|
||||
|
||||
// series_url is client-supplied (PUT /bookmarks/{key} accepts any string),
|
||||
// so this is not just an optimisation against burning a request on an
|
||||
// unknown site: without it, the server would issue a GET from its own
|
||||
// network position to whatever URL a token-holder writes, including
|
||||
// link-local/internal addresses or non-https schemes. The cooldown above
|
||||
// is already consumed, so a row that never passes this check is retried at
|
||||
// cooldown pace rather than hot-looping.
|
||||
if !fetchableSeriesURL(sr.Site, sr.SeriesURL) {
|
||||
log.Printf("latest poll %q: not fetchable: site=%q url=%q", sr.Key(), sr.Site, sr.SeriesURL)
|
||||
return
|
||||
}
|
||||
|
||||
f := fetcherFor(sr.Site, p.BrowserFetch, p.Fetch)
|
||||
if f == nil {
|
||||
log.Printf("latest poll %q: no fetcher for site %q", sr.Key(), sr.Site)
|
||||
return
|
||||
}
|
||||
p.prefetchCover(ctx, sr)
|
||||
|
||||
body, status, err := f.Get(ctx, sr.SeriesURL)
|
||||
facts, err := readSeriesPage(ctx, sr.Site, sr.SeriesURL, p.BrowserFetch, p.Fetch)
|
||||
if err != nil {
|
||||
log.Printf("latest poll %q: fetch %s: %v", sr.Key(), sr.SeriesURL, err)
|
||||
switch {
|
||||
case errors.Is(err, errNotFetchable):
|
||||
// The cooldown above is already consumed, so a row that never
|
||||
// passes the gate is retried at cooldown pace rather than
|
||||
// hot-looping.
|
||||
log.Printf("latest poll %q: not fetchable: site=%q url=%q", sr.Key(), sr.Site, sr.SeriesURL)
|
||||
return
|
||||
case errors.Is(err, errNoFetcher):
|
||||
log.Printf("latest poll %q: no fetcher for site %q", sr.Key(), sr.Site)
|
||||
return
|
||||
}
|
||||
// A legacy cover heals independently of the page read: its source may
|
||||
// answer — a CDN — while the origin does not, so a fetch failure does
|
||||
// not skip the heal, matching the order the shared read replaced.
|
||||
p.prefetchCover(ctx, sr)
|
||||
log.Printf("latest poll %q: %v", sr.Key(), err)
|
||||
return
|
||||
}
|
||||
if status != 200 {
|
||||
log.Printf("latest poll %q: fetch %s: status %d", sr.Key(), sr.SeriesURL, status)
|
||||
return
|
||||
}
|
||||
|
||||
latest, ok := latestChapterFrom(sr.Site, sr.SeriesURL, body)
|
||||
// A legacy cover source is healed independently of the page read.
|
||||
p.prefetchCover(ctx, sr)
|
||||
// Cover fill is independent of the chapter signal: a page that lost its
|
||||
// chapter list may keep its og:image, and a blank Series heals either way.
|
||||
p.fillBlankCover(ctx, sr, body)
|
||||
if !ok {
|
||||
p.fillBlankCover(ctx, sr, facts.Cover)
|
||||
if !facts.HasLatest {
|
||||
// Most likely a challenge page or a layout change. Either way the row is
|
||||
// already stamped, so this waits out a cooldown instead of hot-looping.
|
||||
log.Printf("latest poll %q: no chapter links in %d bytes", sr.Key(), len(body))
|
||||
log.Printf("latest poll %q: no chapter links in %d bytes", sr.Key(), facts.BodyLen)
|
||||
return
|
||||
}
|
||||
|
||||
@@ -268,7 +260,7 @@ func (p *Poller) checkOne(ctx context.Context, sr store.Series) {
|
||||
// chapter should correct the stored number downward. The comparison is
|
||||
// against the due-query snapshot; a concurrent write in between only costs
|
||||
// one redundant UPDATE of the same absolute value, never a wrong one.
|
||||
if sr.LatestChapterNum != nil && *sr.LatestChapterNum == latest.Num {
|
||||
if sr.LatestChapterNum != nil && *sr.LatestChapterNum == facts.Latest.Num {
|
||||
return
|
||||
}
|
||||
|
||||
@@ -276,52 +268,30 @@ func (p *Poller) checkOne(ctx context.Context, sr store.Series) {
|
||||
// bookmark joining to it, and the bookmark's updated_at is never touched —
|
||||
// a newly published chapter is not reading progress and must not reorder
|
||||
// the list.
|
||||
if err := p.Store.SetLatestChapter(sr.Site, sr.SeriesID, latest.Label, latest.Num); err != nil {
|
||||
if err := p.Store.SetLatestChapter(sr.Site, sr.SeriesID, facts.Latest.Label, facts.Latest.Num); err != nil {
|
||||
log.Printf("latest poll %q: set latest chapter: %v", sr.Key(), err)
|
||||
return
|
||||
}
|
||||
log.Printf("latest poll %q: latest is now %s", sr.Key(), latest.Label)
|
||||
log.Printf("latest poll %q: latest is now %s", sr.Key(), facts.Latest.Label)
|
||||
}
|
||||
|
||||
// fetchableSeriesURL reports whether site is a site latestChapterFrom knows how
|
||||
// to parse and seriesURL is safe to hand to a fetcher: an https URL with a
|
||||
// non-empty host. series_url comes from client-supplied PUT bodies, so this is
|
||||
// a defence against the poller being used to probe arbitrary hosts from the
|
||||
// server's own network position, not just a check against wasted requests.
|
||||
//
|
||||
// Three sites are held to a stricter rule, each for a different reason:
|
||||
//
|
||||
// - kagane and novelfull are fetched by a headless browser, which executes
|
||||
// JavaScript and carries cookies, and is therefore a far stronger SSRF
|
||||
// primitive than an HTTP GET. Their hosts must match exactly, not merely
|
||||
// be non-empty.
|
||||
// - lightnovelworld's parser regex hardcodes its host, so a URL anywhere
|
||||
// else could never yield a match — reject it here rather than burn the
|
||||
// request.
|
||||
// fetchableSeriesURL reports whether site is a Site the registry knows and
|
||||
// seriesURL is safe to hand to a fetcher: an https URL whose host matches the
|
||||
// Site's pinned hostname exactly. series_url comes from client-supplied PUT
|
||||
// bodies, so this is a defence against the poller being used to probe
|
||||
// arbitrary hosts from the server's own network position, not just a check
|
||||
// against wasted requests. The pin guards different things per Site — a
|
||||
// browser Site guards a control that executes JavaScript and carries cookies,
|
||||
// a parser Site guards a wasted request — but the rule is one rule, from the
|
||||
// registry.
|
||||
func fetchableSeriesURL(site, seriesURL string) bool {
|
||||
switch site {
|
||||
case "asura", "demonic", "comix", "kagane", "novelfull", "lightnovelworld":
|
||||
default:
|
||||
s, known := sites[site]
|
||||
if !known {
|
||||
return false
|
||||
}
|
||||
u, err := url.Parse(seriesURL)
|
||||
if err != nil {
|
||||
return false
|
||||
}
|
||||
if u.Scheme != "https" || u.Host == "" {
|
||||
return false
|
||||
}
|
||||
switch site {
|
||||
case "kagane":
|
||||
return u.Hostname() == "kagane.to"
|
||||
case "novelfull":
|
||||
// Fetched by a real browser, same as kagane, so the host is pinned
|
||||
// rather than merely non-empty.
|
||||
return u.Hostname() == "novelfull.com"
|
||||
case "lightnovelworld":
|
||||
// Its parser regex hardcodes this host, so a URL anywhere else could
|
||||
// never yield a match — reject it here rather than burn the request.
|
||||
return u.Hostname() == "lightnovelworld.net"
|
||||
}
|
||||
return true
|
||||
return u.Scheme == "https" && u.Hostname() == s.Host
|
||||
}
|
||||
|
||||
@@ -623,6 +623,13 @@ func TestFetchableSeriesURL(t *testing.T) {
|
||||
{"demonic https", "demonic", "https://demonicscans.org/manga/X", true},
|
||||
{"comix https", "comix", "https://comix.to/title/n8we-dungeons-and-crayons", true},
|
||||
{"kagane on its own host", "kagane", "https://kagane.to/series/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b", true},
|
||||
// The three plain-TLS Sites are pinned too: a client-supplied
|
||||
// series_url must not aim a fetcher at a lookalike host, even when
|
||||
// the fetcher is only an HTTP GET.
|
||||
{"asura on a foreign host", "asura", "https://asurascans.com.evil.example/comics/x", false},
|
||||
{"asura on the dead old domain", "asura", "https://asuracomic.net/comics/x", false},
|
||||
{"demonic on a lookalike host", "demonic", "https://demonicscans.org.evil.example/manga/X", false},
|
||||
{"comix on a foreign host", "comix", "https://evil.example/title/x", false},
|
||||
// The browser fetcher runs JavaScript and carries cookies, so a
|
||||
// client-supplied series_url must not be able to aim it anywhere else.
|
||||
{"kagane on a foreign host", "kagane", "https://evil.example/series/x", false},
|
||||
@@ -952,6 +959,8 @@ func TestFetcherForRoutesNovelSites(t *testing.T) {
|
||||
// browser-less deployment: kagane is nothing, novelfull degrades to TLS
|
||||
{"kagane", nil, tls, nil},
|
||||
{"novelfull", nil, tls, tls},
|
||||
// unknown site: fail closed — nothing to fetch or parse
|
||||
{"mangadex", browser, tls, nil},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.site, func(t *testing.T) {
|
||||
|
||||
@@ -0,0 +1,58 @@
|
||||
package latest
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
)
|
||||
|
||||
// seriesRead carries the two facts the poll and the acquirer both extract
|
||||
// from a series page. Persistence, stamps and scheduling stay with the
|
||||
// callers, so the policies that keep the two flows distinct (stamp order,
|
||||
// cooldowns) are not swallowed by the module.
|
||||
type seriesRead struct {
|
||||
Latest latestChapter
|
||||
HasLatest bool
|
||||
Cover string
|
||||
HasCover bool
|
||||
// BodyLen is the fetched body's length, surfaced because the no-chapter
|
||||
// log uses it to tell a markup change from a body the size cap cut short.
|
||||
BodyLen int
|
||||
}
|
||||
|
||||
// errNotFetchable and errNoFetcher separate the gate and the route from fetch
|
||||
// failures so each caller keeps its own distinct log line for all three.
|
||||
var (
|
||||
errNotFetchable = errors.New("series url not fetchable")
|
||||
errNoFetcher = errors.New("no fetcher for site")
|
||||
)
|
||||
|
||||
// readSeriesPage performs the series-page read the poll and the acquirer have
|
||||
// in common: gate the address, choose the route, fetch the page, extract the
|
||||
// Latest Chapter and the Cover address. It persists nothing and stamps
|
||||
// nothing.
|
||||
//
|
||||
// series_url arrives in a client-supplied PUT body (PUT /bookmarks/{key}
|
||||
// accepts any string), so the gate is not an optimisation against burning a
|
||||
// request on an unknown site: without it, the server would issue a GET from
|
||||
// its own network position to whatever URL a token-holder writes, including
|
||||
// link-local/internal addresses or non-https schemes.
|
||||
func readSeriesPage(ctx context.Context, site, seriesURL string, browser, tls Fetcher) (seriesRead, error) {
|
||||
if !fetchableSeriesURL(site, seriesURL) {
|
||||
return seriesRead{}, fmt.Errorf("%w: site=%q url=%q", errNotFetchable, site, seriesURL)
|
||||
}
|
||||
f := fetcherFor(site, browser, tls)
|
||||
if f == nil {
|
||||
return seriesRead{}, fmt.Errorf("%w: site %q", errNoFetcher, site)
|
||||
}
|
||||
body, status, err := f.Get(ctx, seriesURL)
|
||||
if err != nil {
|
||||
return seriesRead{}, fmt.Errorf("fetch %s: %w", seriesURL, err)
|
||||
}
|
||||
if status != 200 {
|
||||
return seriesRead{}, fmt.Errorf("fetch %s: status %d", seriesURL, status)
|
||||
}
|
||||
latest, hasLatest := latestChapterFrom(site, seriesURL, body)
|
||||
cover, hasCover := coverFrom(site, seriesURL, body)
|
||||
return seriesRead{Latest: latest, HasLatest: hasLatest, Cover: cover, HasCover: hasCover, BodyLen: len(body)}, nil
|
||||
}
|
||||
@@ -6,8 +6,11 @@ import (
|
||||
"log"
|
||||
"net/url"
|
||||
"regexp"
|
||||
"sort"
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
"github.com/chromedp/chromedp"
|
||||
)
|
||||
|
||||
// latestChapter is the newest chapter a series page advertises.
|
||||
@@ -16,6 +19,40 @@ type latestChapter struct {
|
||||
Label string
|
||||
}
|
||||
|
||||
// site answers the fixed questions every series-page read asks of its Site
|
||||
// (ADR-0009): the host its addresses must carry, how to find the Latest
|
||||
// Chapter and the Cover address in a body, and — for a Site behind a
|
||||
// JavaScript challenge — how to read its payload from a cleared tab. One
|
||||
// entry describes everything about one Site, and nowhere else gets to compare
|
||||
// the site string.
|
||||
type site struct {
|
||||
// Host is the exact hostname a series_url for this Site must carry.
|
||||
Host string
|
||||
// LatestChapter finds the newest chapter in a fetched body.
|
||||
LatestChapter func(seriesURL, body string) (latestChapter, bool)
|
||||
// Cover finds the Cover address in a fetched body.
|
||||
Cover func(seriesURL, body string) (string, bool)
|
||||
// Browser reads this Site's payload from a cleared browser tab; nil
|
||||
// means the page is fetched over plain TLS.
|
||||
Browser *browserRead
|
||||
}
|
||||
|
||||
type browserRead struct {
|
||||
// Read builds the tab read for seriesURL, refusing (false) an address
|
||||
// this Site will not open in a browser — the per-Site half of the SSRF
|
||||
// gate, kept deliberately behind fetchableSeriesURL: a headless browser
|
||||
// executes JavaScript and carries cookies, and series_url is
|
||||
// client-supplied.
|
||||
Read func(seriesURL string, out *string) (chromedp.Action, bool)
|
||||
// Done reports whether the payload arrived.
|
||||
Done func(body string) bool
|
||||
// Fallback allows the plain-TLS fetcher when no browser is configured.
|
||||
// False skips the Site instead. kagane is false — a plain fetch would
|
||||
// only ever retrieve a challenge page — and novelfull is true, because
|
||||
// its challenge is a live time-varying fact (AGENTS.md).
|
||||
Fallback bool
|
||||
}
|
||||
|
||||
// asuraSlugRe pulls the series slug out of a stored series_url.
|
||||
// Shape verified live 2026-07-26: https://asurascans.com/comics/<slug>, where
|
||||
// the slug carries a trailing build-hash suffix (e.g. "-f886a8af") that
|
||||
@@ -66,7 +103,7 @@ var novelfullSlugRe = regexp.MustCompile(`^/([^/?#]+)\.html$`)
|
||||
// under several Chapter Slugs (measured 2026-08-11: a sampled novel serves
|
||||
// 1-99 under one slug and 100-423 under another), so no stored-slug pattern can
|
||||
// cover a Series' whole list. An unscoped match is safe because
|
||||
// latestChapterFrom truncates the body at the comment thread before scanning
|
||||
// lnwLatestChapter truncates the body at the comment thread before scanning
|
||||
// (lnwCommentMarker); without that, a visitor's comment could set the Latest
|
||||
// Chapter on the shared Series row.
|
||||
var lnwChapterRe = regexp.MustCompile(`lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/`)
|
||||
@@ -80,86 +117,16 @@ var lnwChapterRe = regexp.MustCompile(`lightnovelworld\.net/[a-z0-9-]+-chapter-(
|
||||
// body is skipped, never scanned whole.
|
||||
const lnwCommentMarker = "wpd-threads"
|
||||
|
||||
// latestChapterFrom returns the highest chapter number body advertises for this
|
||||
// series. ok is false when the body yields nothing usable — an unknown site, an
|
||||
// empty body, a Cloudflare challenge page, and a site redesign all land here,
|
||||
// and the caller treats all four identically.
|
||||
//
|
||||
// Ported from the userscript's latestChapterFromAnchors (asura L123-133,
|
||||
// demonic L183-193), including its reason for taking a maximum rather than a
|
||||
// first or last: neither site lists chapters in a dependable order.
|
||||
// maxChapter returns the highest chapter number the regex finds in body. A
|
||||
// maximum rather than a first or last, ported from the userscript's
|
||||
// latestChapterFromAnchors (asura L123-133, demonic L183-193): neither site
|
||||
// lists chapters in a dependable order.
|
||||
//
|
||||
// The userscript's asura rule additionally requires the anchor text to match
|
||||
// /Chapter\s+[\d.]+/i. That check exists only to skip the "First Chapter"
|
||||
// shortcut, which points at chapter/1 and therefore can never win a maximum, so
|
||||
// it is redundant here. For asura, scoping the pattern to this series' own slug
|
||||
// replaces it with a stronger guarantee: a chapter link belonging to some other
|
||||
// series cannot contribute even if the page starts carrying them. demonic has no
|
||||
// such guarantee — demonicChapterRe matches any chaptered.php?manga=<id> anchor
|
||||
// with no per-series scoping, because the stored series_id for demonic is a
|
||||
// slug, not the numeric id the URL carries, so it cannot easily be scoped.
|
||||
// lightnovelworld is unscoped and body-truncated instead — see lnwChapterRe.
|
||||
func latestChapterFrom(site, seriesURL, body string) (latestChapter, bool) {
|
||||
var re *regexp.Regexp
|
||||
switch site {
|
||||
case "asura":
|
||||
m := asuraSlugRe.FindStringSubmatch(seriesURL)
|
||||
if m == nil {
|
||||
return latestChapter{}, false
|
||||
}
|
||||
// Stored URLs predating a redeploy may carry a stale build hash;
|
||||
// chapter hrefs in the fetched body carry the current one. Strip to
|
||||
// the stable ID and make the hash optional in the pattern, so scoping
|
||||
// survives rotations.
|
||||
slug := asuraBuildHash.ReplaceAllString(m[1], "")
|
||||
// Compiled per call rather than cached: this runs once per fetch, which
|
||||
// is at most a few times a minute, and the slug varies per series.
|
||||
re = regexp.MustCompile(`/comics/` + regexp.QuoteMeta(slug) + `(?:-[0-9a-f]{8})?/chapter/([0-9.]+)`)
|
||||
case "demonic":
|
||||
re = demonicChapterRe
|
||||
case "comix":
|
||||
id, ok := comixSeriesID(seriesURL)
|
||||
if !ok {
|
||||
return latestChapter{}, false
|
||||
}
|
||||
// comix ships an SPA: the served HTML carries a JSON state blob instead
|
||||
// of chapter anchors, and latestChapterUrl is the only place the newest
|
||||
// chapter appears. Scoping to this series' id prefix keeps a
|
||||
// "recommended" strip's entries from winning the maximum.
|
||||
re = regexp.MustCompile(`"latestChapterUrl":"/title/` + regexp.QuoteMeta(id) + `-[^"]*-chapter-([0-9.]+)"`)
|
||||
case "kagane":
|
||||
re = kaganeChapterRe
|
||||
case "novelfull":
|
||||
u, err := url.Parse(seriesURL)
|
||||
if err != nil {
|
||||
return latestChapter{}, false
|
||||
}
|
||||
m := novelfullSlugRe.FindStringSubmatch(u.Path)
|
||||
if m == nil {
|
||||
return latestChapter{}, false
|
||||
}
|
||||
// Scoped to this series' slug for the same reason asura is: page 1
|
||||
// carries a "latest chapters" widget and a "you may also like" strip,
|
||||
// and neither may contribute to the maximum.
|
||||
re = regexp.MustCompile(`/` + regexp.QuoteMeta(m[1]) + `/chapter-([0-9.]+)`)
|
||||
case "lightnovelworld":
|
||||
// The comment thread below the chapter list is the one region of the
|
||||
// page any visitor can write to, so the scan never reads past it (see
|
||||
// lnwChapterRe). A body without the marker is skipped, never scanned
|
||||
// whole — a redesign must degrade into staleness, not into a wrong
|
||||
// shared value; the logged body length tells a markup change from a
|
||||
// body the size cap cut short.
|
||||
i := strings.Index(body, lnwCommentMarker)
|
||||
if i < 0 {
|
||||
log.Printf("latest poll %q: no %s marker in %d bytes", seriesURL, lnwCommentMarker, len(body))
|
||||
return latestChapter{}, false
|
||||
}
|
||||
body = body[:i]
|
||||
re = lnwChapterRe
|
||||
default:
|
||||
return latestChapter{}, false
|
||||
}
|
||||
|
||||
// shortcut, which points at chapter/1 and therefore can never win a maximum,
|
||||
// so it is redundant once a maximum is taken.
|
||||
func maxChapter(re *regexp.Regexp, body string) (latestChapter, bool) {
|
||||
var best latestChapter
|
||||
found := false
|
||||
for _, m := range re.FindAllStringSubmatch(body, -1) {
|
||||
@@ -178,6 +145,94 @@ func latestChapterFrom(site, seriesURL, body string) (latestChapter, bool) {
|
||||
return best, found
|
||||
}
|
||||
|
||||
// asuraLatestChapter scopes chapter links to this series' own slug, which
|
||||
// replaces the userscript's anchor-text check with a stronger guarantee: a
|
||||
// chapter link belonging to some other series cannot contribute even if the
|
||||
// page starts carrying them.
|
||||
func asuraLatestChapter(seriesURL, body string) (latestChapter, bool) {
|
||||
m := asuraSlugRe.FindStringSubmatch(seriesURL)
|
||||
if m == nil {
|
||||
return latestChapter{}, false
|
||||
}
|
||||
// Stored URLs predating a redeploy may carry a stale build hash; chapter
|
||||
// hrefs in the fetched body carry the current one. Strip to the stable ID
|
||||
// and make the hash optional in the pattern, so scoping survives
|
||||
// rotations.
|
||||
slug := asuraBuildHash.ReplaceAllString(m[1], "")
|
||||
// Compiled per call rather than cached: this runs once per fetch, which is
|
||||
// at most a few times a minute, and the slug varies per series.
|
||||
re := regexp.MustCompile(`/comics/` + regexp.QuoteMeta(slug) + `(?:-[0-9a-f]{8})?/chapter/([0-9.]+)`)
|
||||
return maxChapter(re, body)
|
||||
}
|
||||
|
||||
// demonicLatestChapter is not scoped: demonicChapterRe matches any
|
||||
// chaptered.php?manga=<id> anchor, because the stored series_id is a slug,
|
||||
// not the numeric id the URL carries, so it cannot be scoped.
|
||||
func demonicLatestChapter(_, body string) (latestChapter, bool) {
|
||||
return maxChapter(demonicChapterRe, body)
|
||||
}
|
||||
|
||||
// comixLatestChapter reads comix's SPA: the served HTML carries a JSON state
|
||||
// blob instead of chapter anchors, and latestChapterUrl is the only place the
|
||||
// newest chapter appears. Scoping to this series' id prefix keeps a
|
||||
// "recommended" strip's entries from winning the maximum.
|
||||
func comixLatestChapter(seriesURL, body string) (latestChapter, bool) {
|
||||
id, ok := comixSeriesID(seriesURL)
|
||||
if !ok {
|
||||
return latestChapter{}, false
|
||||
}
|
||||
re := regexp.MustCompile(`"latestChapterUrl":"/title/` + regexp.QuoteMeta(id) + `-[^"]*-chapter-([0-9.]+)"`)
|
||||
return maxChapter(re, body)
|
||||
}
|
||||
|
||||
// kaganeLatestChapter scans the kagane series API JSON that the browser read
|
||||
// fetched from inside the page; the match rides on the property name,
|
||||
// regardless of the surrounding JSON shape.
|
||||
func kaganeLatestChapter(_, body string) (latestChapter, bool) {
|
||||
return maxChapter(kaganeChapterRe, body)
|
||||
}
|
||||
|
||||
// novelfullLatestChapter is scoped to this series' slug for the same reason
|
||||
// asura is: page 1 carries a "latest chapters" widget and a "you may also
|
||||
// like" strip, and neither may contribute to the maximum.
|
||||
func novelfullLatestChapter(seriesURL, body string) (latestChapter, bool) {
|
||||
u, err := url.Parse(seriesURL)
|
||||
if err != nil {
|
||||
return latestChapter{}, false
|
||||
}
|
||||
m := novelfullSlugRe.FindStringSubmatch(u.Path)
|
||||
if m == nil {
|
||||
return latestChapter{}, false
|
||||
}
|
||||
re := regexp.MustCompile(`/` + regexp.QuoteMeta(m[1]) + `/chapter-([0-9.]+)`)
|
||||
return maxChapter(re, body)
|
||||
}
|
||||
|
||||
// lnwLatestChapter truncates the body at the comment thread before scanning:
|
||||
// it is the one region of the page any visitor can write to (see lnwChapterRe).
|
||||
// A body without the marker is skipped, never scanned whole — a redesign must
|
||||
// degrade into staleness, not into a wrong shared value; the logged body length
|
||||
// tells a markup change from a body the size cap cut short.
|
||||
func lnwLatestChapter(seriesURL, body string) (latestChapter, bool) {
|
||||
i := strings.Index(body, lnwCommentMarker)
|
||||
if i < 0 {
|
||||
log.Printf("latest poll %q: no %s marker in %d bytes", seriesURL, lnwCommentMarker, len(body))
|
||||
return latestChapter{}, false
|
||||
}
|
||||
return maxChapter(lnwChapterRe, body[:i])
|
||||
}
|
||||
|
||||
// latestChapterFrom returns the highest chapter number body advertises for this
|
||||
// series, via the Site's registry entry. ok is false when the body yields
|
||||
// nothing usable — an unknown site, an empty body, a Cloudflare challenge page,
|
||||
// and a site redesign all land here, and the caller treats all four identically.
|
||||
func latestChapterFrom(site, seriesURL, body string) (latestChapter, bool) {
|
||||
if fn := sites[site].LatestChapter; fn != nil {
|
||||
return fn(seriesURL, body)
|
||||
}
|
||||
return latestChapter{}, false
|
||||
}
|
||||
|
||||
var metaTagRe = regexp.MustCompile(`(?is)<meta\b[^>]*>`)
|
||||
var doubleQuotedMetaAttrRe = regexp.MustCompile(`(?is)([a-z][a-z0-9:_-]*)\s*=\s*"([^"]*)"`)
|
||||
var singleQuotedMetaAttrRe = regexp.MustCompile(`(?is)([a-z][a-z0-9:_-]*)\s*=\s*'([^']*)'`)
|
||||
@@ -257,24 +312,40 @@ func comixCoverURL(seriesURL, body string) string {
|
||||
return publishedCoverURL(detail.Poster.Medium)
|
||||
}
|
||||
|
||||
// coverFrom reports false for unknown sites, challenge bodies, and pages with
|
||||
// no usable cover. Metadata extraction keeps scanning after an empty match so
|
||||
// a later published cover is not hidden by an empty tag.
|
||||
func coverFrom(site, seriesURL, body string) (string, bool) {
|
||||
var cover string
|
||||
switch site {
|
||||
case "asura", "demonic", "lightnovelworld":
|
||||
cover = metaContent(body, "property", "og:image")
|
||||
case "novelfull":
|
||||
cover = metaContent(body, "name", "image")
|
||||
case "comix":
|
||||
cover = comixCoverURL(seriesURL, body)
|
||||
case "kagane":
|
||||
cover = kaganeCoverURL(body)
|
||||
}
|
||||
// ogImageCover reads the og:image metadata shared by asura, demonic and
|
||||
// lightnovelworld.
|
||||
func ogImageCover(_, body string) (string, bool) {
|
||||
cover := metaContent(body, "property", "og:image")
|
||||
return cover, cover != ""
|
||||
}
|
||||
|
||||
func novelfullCoverEntry(_, body string) (string, bool) {
|
||||
cover := metaContent(body, "name", "image")
|
||||
return cover, cover != ""
|
||||
}
|
||||
|
||||
func comixCoverEntry(seriesURL, body string) (string, bool) {
|
||||
cover := comixCoverURL(seriesURL, body)
|
||||
return cover, cover != ""
|
||||
}
|
||||
|
||||
func kaganeCoverEntry(_, body string) (string, bool) {
|
||||
cover := kaganeCoverURL(body)
|
||||
return cover, cover != ""
|
||||
}
|
||||
|
||||
// coverFrom reports false for unknown sites, challenge bodies, and pages with
|
||||
// no usable cover, via the Site's registry entry.
|
||||
func coverFrom(site, seriesURL, body string) (string, bool) {
|
||||
if fn := sites[site].Cover; fn != nil {
|
||||
return fn(seriesURL, body)
|
||||
}
|
||||
return "", false
|
||||
}
|
||||
|
||||
// metaContent returns the content of the first <meta> whose attrName is
|
||||
// attrValue. It keeps scanning after an empty match so a later published cover
|
||||
// is not hidden by an empty tag.
|
||||
func metaContent(body, attrName, attrValue string) string {
|
||||
for _, tag := range metaTagRe.FindAllString(body, -1) {
|
||||
attrs := make(map[string]string)
|
||||
@@ -297,3 +368,70 @@ func publishedCoverURL(value string) string {
|
||||
value = strings.TrimSpace(html.UnescapeString(value))
|
||||
return strings.ReplaceAll(value, " ", "%20")
|
||||
}
|
||||
|
||||
// sites is the registry: one entry per Site, keyed by the stored site string.
|
||||
// Adding a Site means adding an entry here and nowhere else — the dispatch
|
||||
// functions above and the poller's route list are lookups into this map. An
|
||||
// unknown site string resolves to the zero entry, which fails the existing
|
||||
// not-fetchable and no-fetcher paths unchanged.
|
||||
var sites = map[string]site{
|
||||
"asura": {
|
||||
Host: "asurascans.com",
|
||||
LatestChapter: asuraLatestChapter,
|
||||
Cover: ogImageCover,
|
||||
},
|
||||
"demonic": {
|
||||
Host: "demonicscans.org",
|
||||
LatestChapter: demonicLatestChapter,
|
||||
Cover: ogImageCover,
|
||||
},
|
||||
"comix": {
|
||||
Host: "comix.to",
|
||||
LatestChapter: comixLatestChapter,
|
||||
Cover: comixCoverEntry,
|
||||
},
|
||||
"kagane": {
|
||||
Host: "kagane.to",
|
||||
LatestChapter: kaganeLatestChapter,
|
||||
Cover: kaganeCoverEntry,
|
||||
Browser: &browserRead{
|
||||
Read: kaganeRead,
|
||||
Done: func(body string) bool { return body != "" },
|
||||
// Never falls back: a plain fetch of a kagane page or cover would
|
||||
// only ever retrieve a challenge page (verified 2026-08-03).
|
||||
Fallback: false,
|
||||
},
|
||||
},
|
||||
"novelfull": {
|
||||
Host: "novelfull.com",
|
||||
LatestChapter: novelfullLatestChapter,
|
||||
Cover: novelfullCoverEntry,
|
||||
Browser: &browserRead{
|
||||
Read: novelfullRead,
|
||||
// The interstitial has a DOM too, so "the payload arrived" has to
|
||||
// exclude it explicitly.
|
||||
Done: func(body string) bool { return body != "" && !isInterstitial(body) },
|
||||
Fallback: true,
|
||||
},
|
||||
},
|
||||
"lightnovelworld": {
|
||||
Host: "lightnovelworld.net",
|
||||
LatestChapter: lnwLatestChapter,
|
||||
Cover: ogImageCover,
|
||||
},
|
||||
}
|
||||
|
||||
// browserBackedSites is derived from the registry: the Sites whose pages are
|
||||
// read through the browser sidecar, which are also the ones granted the longer
|
||||
// cooldown. Sorted so callers that range it (the due query, the browser
|
||||
// fetcher's dispatch) see a stable order instead of map-iteration noise.
|
||||
func browserBackedSites() []string {
|
||||
out := make([]string, 0, len(sites))
|
||||
for name, s := range sites {
|
||||
if s.Browser != nil {
|
||||
out = append(out, name)
|
||||
}
|
||||
}
|
||||
sort.Strings(out)
|
||||
return out
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user