62772e1eaa
Adds a background goroutine to the backend that re-checks each bookmarked series' newest published chapter on its own schedule, so `latest_chapter` stays fresh even when the manga sites are never opened in a browser.
This is a *second, parallel* signal, not a replacement: the userscript keeps its own `maybeCaptureLatestOnSeriesPage` / `backgroundRefreshLatest` logic, unchanged. `userscript/manga-bookmark.user.js` is byte-identical to `main`.
## How it works
One ticker goroutine in the same binary. Each wake it asks SQLite for bookmarks whose `latest_checked_at` has aged past a per-bookmark cooldown, fetches those series pages through a Chrome-fingerprinted HTTP client, extracts the max chapter number with a per-site regex, and writes it back through `Store.Get` + `Store.Upsert`. Every failure path logs and moves on.
Two independent clocks:
- **cooldown** — how long one bookmark rests between checks, enforced by the `WHERE` clause in `Store.DueForLatestCheck`, not by a timer.
- **interval** — how often the goroutine wakes and looks.
Shortening the interval therefore cannot shorten anyone's cooldown; it only makes the poller wake and find nothing due more often.
The row is stamped **before** the fetch, so an error, a timeout, or a shutdown mid-request still consumes the cooldown — a renamed or challenged series waits out a full cooldown instead of being retried every tick.
## Design decisions worth reviewing
**`latest_checked_at` is deliberately absent from the `Bookmark` struct and from `bookmarkColumns`.** `PUT /bookmarks/{key}` decodes a whole `Bookmark` and `Upsert` writes every column it knows about, so a userscript PUT — which has no idea this field exists — would write a zero and reset the cooldown, making the poller re-fetch that series on every tick for as long as the user kept reading it. Two tests guard this: `TestUpsertPreservesLatestCheckedAt` and `TestPutDoesNotClobberLatestCheckedAt`, the latter driving a real router PUT with a userscript-shaped body.
**`updated_at` never moves on a latest-chapter bump.** All chapter writes go through `Store.Get` + `Store.Upsert`, so the existing `CASE` keeps the stored timestamp when only `latest_chapter_num` changes and the bookmark list does not reorder. `TestRunOnceDoesNotReorderList` asserts both the timestamp and the `List()` head position.
**Fetches use `bogdanfinn/tls-client` with a Chrome profile.** Plain `net/http` was verified working against both sites on 2026-07-26, so this is not fixing an observed block — it is deliberate defence-in-depth against a future fingerprint-based one. The library is pure Go, so `CGO_ENABLED=0`, the static binary, and the distroless image are all unaffected. It does require the Go floor to move 1.23 → 1.24.
**`checkOne` validates before spending a request.** `series_url` is entirely client-supplied through `PUT /bookmarks/{key}`, so without a guard the poller would issue GETs from the server's own network position to any URL a token holder writes. The check requires a known site and an `https` URL with a non-empty host, and sits *after* the cooldown stamp so an unfetchable row is retried at cooldown pace rather than hot-looping.
## Config
Five new env vars, all with defaults sized for this deployment, all wired through `docker-compose.yml`:
| Variable | Default | Meaning |
| --- | --- | --- |
| `LATEST_CHAPTER_POLL_ENABLED` | `1` | Kill switch |
| `LATEST_CHAPTER_POLL_COOLDOWN` | `1h` | Per series, floored at `15m` |
| `LATEST_CHAPTER_POLL_INTERVAL` | `10m` | How often to wake |
| `LATEST_CHAPTER_POLL_BATCH` | `14` | Series per wake |
| `LATEST_CHAPTER_POLL_STAGGER` | `20s` | Delay between fetches in a batch |
`batch × (cooldown / interval)` = 84 series hold a true cooldown cadence at these defaults. Past that nothing breaks: the cadence stretches uniformly and the oldest-checked-first ordering keeps it fair. Bad values log and fall back rather than failing startup — the poller is an enhancement, and a typo in one of its knobs must not stop bookmark sync.
## Known limitation (accepted, documented)
The poller's `Store.Get` + `Store.Upsert` is not wrapped in a single transaction. If a userscript `PUT` commits in the sub-millisecond window between the two, the poller writes back its stale re-read — reverting that progress and, since the stored `last_chapter_num` now differs, tripping the `updated_at` `CASE` and reordering the list.
Accepted rather than fixed for a single-user deployment: the window is one SELECT wide, the poller only writes when a chapter number actually changed, and the next read self-heals it. The alternative — a transactional read-modify-write — means moving or duplicating the `updated_at` `CASE` that four tests and the whole list-ordering invariant depend on. Recorded in `CLAUDE.md` next to the poller's architecture bullet so it is not a silent trap.
## Testing
- Full suite green, including `-race`; `go vet` clean; `CGO_ENABLED=0` static build and `docker compose build` both pass on the bumped `golang:1.24-alpine`.
- No test touches the network: the `fetcher` interface exists so tests inject a fake, and no test imports `tls-client` or reaches either manga site.
- Extraction is fixture-driven against markup trimmed from real pages (2026-07-26), including a Cloudflare challenge page, cross-series chapter links, decimal chapters, and both raw `&` and `&` forms.
- Poller tests cover the no-reorder invariant, cooldown enforcement across passes, batch limiting, one bad series not stalling a batch, downward correction on a retracted chapter, cancelled contexts, and all four failure shapes still consuming the cooldown.
- Migration from a pre-column database has its own test — `newTestStore` takes the `CREATE TABLE` path, so the `ALTER TABLE` path would otherwise be untested.
- **Live smoke test:** real server, real fetch of asurascans.com. Log showed `latest is now Chapter 181` and `due=1 checked=1`; `GET /bookmarks` returned `latest_chapter_num: 181` with `updated_at` byte-identical to the PUT that created the row — the no-reorder invariant confirmed against a live site, not just a fake.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Reviewed-on: #2
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
195 lines
6.6 KiB
Go
195 lines
6.6 KiB
Go
package main
|
|
|
|
import (
|
|
"context"
|
|
"log"
|
|
"net/url"
|
|
"time"
|
|
)
|
|
|
|
// fetcher retrieves a series page. It exists as an interface so tests can inject
|
|
// a fake: nothing in the test suite may touch the network or the TLS client.
|
|
type fetcher interface {
|
|
Get(ctx context.Context, url string) (body string, status int, err error)
|
|
}
|
|
|
|
// latestPoller re-checks each bookmarked series' newest published chapter on a
|
|
// schedule, independent of the userscript's own in-browser checks. The two run
|
|
// in parallel and report the same observable fact, so whichever writes last wins
|
|
// and neither needs to know about the other.
|
|
//
|
|
// Two clocks, deliberately independent:
|
|
//
|
|
// - interval is how often this goroutine wakes up and looks.
|
|
// - cooldown is how long one bookmark rests since its own last check.
|
|
//
|
|
// Only the cooldown is per bookmark, and it is enforced by the WHERE clause in
|
|
// DueForLatestCheck rather than by any timer. Shortening interval therefore
|
|
// cannot shorten anyone's cooldown; it only makes the poller wake up and find
|
|
// nothing due more often.
|
|
type latestPoller struct {
|
|
store *Store
|
|
fetch fetcher
|
|
now func() time.Time // injected so tests can freeze it
|
|
cooldown time.Duration
|
|
interval time.Duration
|
|
stagger time.Duration
|
|
batch int
|
|
}
|
|
|
|
// Run polls until ctx is cancelled.
|
|
//
|
|
// runOnce is called synchronously, so a batch that overruns the tick delays the
|
|
// next one instead of stacking a second batch on top of it. That is the intended
|
|
// failure mode for a misconfigured batch x stagger: a slower cadence, never
|
|
// concurrent fetch storms.
|
|
func (p *latestPoller) Run(ctx context.Context) {
|
|
log.Printf("latest-chapter poller: interval=%s cooldown=%s batch=%d stagger=%s",
|
|
p.interval, p.cooldown, p.batch, p.stagger)
|
|
t := time.NewTicker(p.interval)
|
|
defer t.Stop()
|
|
for {
|
|
select {
|
|
case <-ctx.Done():
|
|
log.Println("latest-chapter poller: stopped")
|
|
return
|
|
case <-t.C:
|
|
p.runOnce(ctx)
|
|
}
|
|
}
|
|
}
|
|
|
|
// runOnce processes one batch of due bookmarks.
|
|
func (p *latestPoller) runOnce(ctx context.Context) {
|
|
cutoff := p.now().Add(-p.cooldown).UnixMilli()
|
|
due, err := p.store.DueForLatestCheck(cutoff, p.batch)
|
|
if err != nil {
|
|
log.Printf("latest poll: due query: %v", err)
|
|
return
|
|
}
|
|
|
|
checked := 0
|
|
for i, b := range due {
|
|
if ctx.Err() != nil {
|
|
break
|
|
}
|
|
// Staggered rather than fired together: a burst of simultaneous requests
|
|
// from one server IP is the traffic shape most likely to move that IP's
|
|
// bot score. This is the server-side analogue of the userscript's "one
|
|
// series per navigation ... indistinguishable from browsing" (L455-456).
|
|
stopped := false
|
|
if i > 0 && p.stagger > 0 {
|
|
select {
|
|
case <-ctx.Done():
|
|
stopped = true
|
|
case <-time.After(p.stagger):
|
|
}
|
|
}
|
|
if stopped {
|
|
break
|
|
}
|
|
p.checkOne(ctx, b)
|
|
checked++
|
|
}
|
|
// due vs checked is how you tell which constraint is binding: ticks that
|
|
// report due=0 mean the cooldown is the limit, ticks that report due==batch
|
|
// every time mean throughput is.
|
|
log.Printf("latest poll: due=%d checked=%d", len(due), checked)
|
|
}
|
|
|
|
// checkOne re-checks one series. Every failure path here is "log and move on":
|
|
// the poller is a best-effort enhancement, and no single bad series may stall a
|
|
// batch or take down the process.
|
|
func (p *latestPoller) checkOne(ctx context.Context, b Bookmark) {
|
|
defer func() {
|
|
if r := recover(); r != nil {
|
|
log.Printf("latest poll %q: recovered from panic: %v", b.Key, r)
|
|
}
|
|
}()
|
|
|
|
// Stamped before the fetch, not after, so an error, a timeout, or a shutdown
|
|
// mid-request still consumes the cooldown. Otherwise a renamed or deleted
|
|
// series would be retried on every single tick forever. The userscript
|
|
// stamps in the same order and for the same reason (L471-473).
|
|
if err := p.store.MarkLatestChecked(b.Key, p.now().UnixMilli()); err != nil {
|
|
log.Printf("latest poll %q: mark checked: %v", b.Key, err)
|
|
return
|
|
}
|
|
|
|
// series_url is client-supplied (PUT /bookmarks/{key} accepts any string),
|
|
// so this is not just an optimisation against burning a request on an
|
|
// unknown site: without it, the server would issue a GET from its own
|
|
// network position to whatever URL a token-holder writes, including
|
|
// link-local/internal addresses or non-https schemes. The cooldown above
|
|
// is already consumed, so a row that never passes this check is retried at
|
|
// cooldown pace rather than hot-looping.
|
|
if !fetchableSeriesURL(b.Site, b.SeriesURL) {
|
|
log.Printf("latest poll %q: not fetchable: site=%q url=%q", b.Key, b.Site, b.SeriesURL)
|
|
return
|
|
}
|
|
|
|
body, status, err := p.fetch.Get(ctx, b.SeriesURL)
|
|
if err != nil {
|
|
log.Printf("latest poll %q: fetch %s: %v", b.Key, b.SeriesURL, err)
|
|
return
|
|
}
|
|
if status != 200 {
|
|
log.Printf("latest poll %q: fetch %s: status %d", b.Key, b.SeriesURL, status)
|
|
return
|
|
}
|
|
|
|
latest, ok := latestChapterFrom(b.Site, b.SeriesURL, body)
|
|
if !ok {
|
|
// Most likely a challenge page or a layout change. Either way the row is
|
|
// already stamped, so this waits out a cooldown instead of hot-looping.
|
|
log.Printf("latest poll %q: no chapter links in %d bytes", b.Key, len(body))
|
|
return
|
|
}
|
|
|
|
// Re-read: the row may have been updated or deleted while the fetch was in
|
|
// flight, and writing b back wholesale would undo that.
|
|
cur, found, err := p.store.Get(b.Key)
|
|
if err != nil {
|
|
log.Printf("latest poll %q: reread: %v", b.Key, err)
|
|
return
|
|
}
|
|
if !found {
|
|
return
|
|
}
|
|
// Equality, not >, mirroring the userscript (L427): a site that retracts a
|
|
// chapter should correct the stored number downward.
|
|
if cur.LatestChapterNum != nil && *cur.LatestChapterNum == latest.Num {
|
|
return
|
|
}
|
|
|
|
num := latest.Num
|
|
cur.LatestChapter = latest.Label
|
|
cur.LatestChapterNum = &num
|
|
// A candidate only. last_chapter_num is untouched, so the CASE in Upsert
|
|
// keeps the stored updated_at and the bookmark list does not reorder.
|
|
cur.UpdatedAt = p.now().UnixMilli()
|
|
if _, err := p.store.Upsert(cur); err != nil {
|
|
log.Printf("latest poll %q: upsert: %v", b.Key, err)
|
|
return
|
|
}
|
|
log.Printf("latest poll %q: latest is now %s", b.Key, latest.Label)
|
|
}
|
|
|
|
// fetchableSeriesURL reports whether site is a site latestChapterFrom knows how
|
|
// to parse and seriesURL is safe to hand to the fetcher: an https URL with a
|
|
// non-empty host. series_url comes from client-supplied PUT bodies, so this is
|
|
// a defence against the poller being used to probe arbitrary hosts from the
|
|
// server's own network position, not just a check against wasted requests.
|
|
func fetchableSeriesURL(site, seriesURL string) bool {
|
|
switch site {
|
|
case "asura", "demonic":
|
|
default:
|
|
return false
|
|
}
|
|
u, err := url.Parse(seriesURL)
|
|
if err != nil {
|
|
return false
|
|
}
|
|
return u.Scheme == "https" && u.Host != ""
|
|
}
|