Each Site runs its own poll goroutine, paced by Rest and Gap from the registry (sites.go) instead of the shared cooldown/interval/stagger/batch config. Rest is enforced by the due query's WHERE clause; the Lane sleeps its effective gap between fetches (hour/eligible, floored at 1s). Lane-local failures: two refusals stop that Site for 15m, a lost browser stops only the round's remaining browser Lanes, and cover heals moved to background goroutines so a slow CDN cannot consume a Lane's gap. The five LATEST_CHAPTER_POLL_* pace env vars are gone; only the kill switch remains.
4.8 KiB
ADR-0010: Poll Lanes — one independent Poll stream per Site
Date: 2026-08-16 Status: accepted
Decision
Replace the single shared polling pace with one Poll Lane per Site: an
independent goroutine that polls only that Site's Series, paced by that Site's
registry entry. Pace moves out of config and into the Site registry
(internal/latest/sites.go): every entry carries a Rest (how long a Series
rests between Polls) and a Gap (how long the Lane waits between fetches).
Rest is enforced by the due query's WHERE clause (latest_checked_at <= now - Rest), never by a timer — the same mechanism that enforced the old cooldown.
The Lane enforces its own gap by sleeping between fetches. effectiveGap is
the registry gap, or one hour divided by the Site's eligible Series count when
that is smaller, never below one second.
The five environment settings that used to size the shared pace —
LATEST_CHAPTER_POLL_COOLDOWN, _BROWSER_COOLDOWN, _INTERVAL, _BATCH,
_STAGGER — are deleted. Only the kill switch LATEST_CHAPTER_POLL_ENABLED
remains. No deployed .env may carry the deleted knobs.
Why
The shared pace capped the whole backend at roughly 180 Polls an hour (one 20-second stagger across one queue). ~60 Series today, scaling to hundreds or thousands, would stretch the hour beyond what the New Chapter signal can tolerate. Worse, the queue mixed Sites with very different costs: kagane and comix pay seconds of a serialized single-tab Chrome per Poll (a challenged page, ADR-0005), and one hostile Site burning its challenge timeout made every other Site's Series wait — "one hostile Site can eat most of an hour".
Lanes fix both at once:
- Throughput scales per Site. The six Lanes fetch concurrently; a Lane's own gap paces it. The browser Lanes' combined ceiling stays about 360 Polls an hour (one tab), and when they cannot keep up the wait past Rest grows and is logged every pass — the "behind by X" measurement, so the decision to give browser Sites more pages is made from data.
- Hostility is contained. A refusal (two challenge-held reads in one
pass) stops only that Site's Lane for
refuseBackoff(15m); the rest of that Lane's Series stay unstamped and due. A lost browser gates the other browser Lanes' passes for the same window — the flag is shared Poller state, so the loss is noticed once instead of once per Lane per pass, and decays after 15m so the Lanes probe again. One Site can no longer tax the others.
Tradeoffs and rejections
- Per-Site env knobs (e.g.
KAGANE_POLL_GAP) rejected: the registry is the single place pace lives, testable and reviewable; config knobs would recreate the shared-pace sprawl with six times the surface. All six entries are deliberately uniform at first — rest an hour, gap ten seconds — so the structure exists to differ without inventing numbers for Sites that have not earned them. - Dynamic gap (
rest / eligible) is the one knob that stays automatic: a Site with more Series than one per ten seconds would otherwise back up behind its own gap, and the per-Series share of the hour is the natural pace. The ten-second default is not arbitrary: one request per ten seconds is the strictest rate rule a free-plan Site can even express (per-zone rate limiting, as documented indocs/research/cloudflare-bot-scoring-and-poll-cadence.md), so the default pace is exactly what the most restrictive Site would demand of us. The computed gap never goes below one second and logs loudly when the floor engages. - Timer-based pacing rejected: the old ticker made the poller's rate a function of wall clock rather than of what was due. The due-query cutoff is the only rate authority; the Lane sleep just prevents hammering.
- Batch size (the old
_BATCHcap) is gone with the shared pace: a Lane processes everything due, paced by its gap. There is no global queue left to bound.
Constraints preserved
- Stamp-before-fetch ("attempted" semantics): an untried Series stays due, so a browser that appears after a restart finds its full queue waiting.
- Browser wake gate (ADR-0005): a browser Lane leaves Chrome asleep below five due Series and 15 minutes of wait, per Lane.
- The browser is not in the API stack (ADR-0006): an unreachable browser
degrades a Lane exactly as an unset
BROWSER_WS_URL— browser-only Sites skipped, plain-TLS unaffected, stored covers still served. - Cover heals moved to background goroutines (joined by the test suite via
waitCovers) so a slow cover CDN cannot consume a Lane's gap.
Supersedes the pace mechanics of ADR-0003's "raise throughput instead" note
(the stagger cut it rejected is what the per-Lane gap replaces) and the
6-hour browser cooldown introduced with the browser-backed Sites; the
1-hour browser rest was already cleared as safe by
docs/research/cloudflare-bot-scoring-and-poll-cadence.md.