Implements spec #137 (spec 4 of 4 from wayfinder map #114). Closes #137. Tickets: #164, #165, #166, #167, #168, #169, #170, #171, #172, #173 — all closed, landed on this branch. ## Summary - #164/#168: sixth outcome word `not_found`; per-Site completed marker predicate. - #165/#169: `poll_failures` row is the failure state; the pass remembers a Site-reported completion. - #166/#171: two new admin filters (`failing`, `unverified`); outbound owner notification + stall condition. - #167/#170: a failure names itself on the Series page; the completion hint reaches the owner and decides nothing. - #172: the other three fault conditions (no-browser-route, sidecar-down, adapter-broken) feeding the notifier. - #173: the landing verdict line shares the same `latest.FaultsFrom` judgement the notifier uses, so the page and the push cannot disagree. `cd backend && go test ./...` green on the merged branch (8 packages). Reviewed-on: #174 Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com> Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
8.0 KiB
ADR-0016: The failure row is the state
Date: 2026-08-22 Status: accepted
Decision
A Series that used to work and has stopped is recorded in one table,
poll_failures, keyed by the same (site, series_id) composite the rest of
the system uses. One row per failing Series, carrying the outcome word and a
failing_since stamp. The row's existence is the failure state: there is
no success sentinel, no counter, no history, and nothing added to the Series
row. A correct read deletes the row, and a repeated identical failure writes
nothing — the stamp ages the run of failures, not the current word.
The write is one upsert, from the poll pass loop:
INSERT INTO poll_failures (site, series_id, outcome, failing_since)
VALUES ($1, $2, $3, $4)
ON CONFLICT (site, series_id) DO UPDATE SET outcome = EXCLUDED.outcome
WHERE poll_failures.outcome <> EXCLUDED.outcome
failing_since is simply absent from the SET arm, so it is never
overwritten, and the WHERE makes an unchanged word write nothing at all —
one statement, no read-then-write race. The clear is a plain DELETE; a row
that is not there is not an error, because the poller clears on every
successful read and most clears find nothing. The composite foreign key
cascades: deleting a Series takes its failure row, so orphan removal stays a
single statement.
Why a future reader will find this surprising
The failure state lives in a table of its own, not on the Series row and not
in the pass log. A flag on series would be the familiar shape, and the
pass log already records outcomes per pass — why a third place?
Because the failure must outlive the pass that observed it and mean
something the pass row cannot say. The pass log is a per-Lane stream: it
records that a Site returned N errors and M not_found on a given run,
but "this exact Series has been failing for three months" is not a question
any single pass row answers — it is a question across rows, and the Lanes
page is a live view of the last 14 days, not a Series index. A Series-level
fact needs Series-level storage, and a flag on the Series row is the wrong
shape too: the failure is transient by definition (a correct read ends it)
and repeatable (the same Series can fail again next year), so the honest
record of "failing since" is a stamp that moves, not a column that flips.
A row that is created and deleted by the read's outcome is that stamp, and
nothing else: the moment a read succeeds the row is gone, so "is this Series
failing?" is answered by one indexed existence check — no success sentinel
to keep consistent with the failure, no counter to reset, no history to
prune. The state cannot drift out of step with the reads that maintain it,
because the reads are the maintenance.
The two no-write outcomes are the second surprise. refused (the Site is
holding a challenge) and unreachable (the browser sidecar was lost
mid-loop) issue no statement at all — neither an upsert nor a delete.
The direction matters, and both directions are wrong to touch:
- Writing would condemn a whole library for one Site's bad day. A refusal is the Site's mood, not a fact about any particular Series: when Cloudflare turns a zone's JS detection on, every Series on that Site reads refused in the same pass. Writing those rows would stamp every Series in the library as failing on the evidence of one Site's configuration, and #166's worklist would present a Site-wide outage as thousands of broken Series.
- Deleting would claim a recovery nothing read. A lost sidecar tells us
nothing about the page — the read never happened. Deleting the row would
report the Series healthy, and worse, it would reset
failing_since: the age that makes a three-month failure findable would start over on a browser restart, erasing evidence no page read contradicted.
So a refusal or an unreachable pass leaves the table exactly as it found
it — the pass neither adds rows that would condemn nor deletes rows that
would claim recovery. The four outcomes that are real evidence about the
page (not_found, no_chapter, unfetchable, errors) write; the two
that are not write nothing; success deletes. There is no third case.
Considered options
A failing_since column on the Series row, alongside the failure word.
Rejected: it is two more columns on a table every due query and every
bookmark join already touches, for a state that is transient and repeatable.
The column would need its own "clear on success" writer anyway — the same
maintenance the row has — while permanently widening the hottest table in
the system for a value that is usually absent. And the worklist's join
would have to distinguish "never failed" from "currently healthy", which
is exactly the sentinel problem the row avoids: an empty column means both.
A counter or last-outcome timestamp instead of (or beside) the row. Rejected: nothing in the system consumes a failure count or a last-outcome time — the worklist needs "failing and since when". A counter invites "three failures = something" thresholds that the ticket explicitly keeps out of scope, and a last-outcome timestamp conflates "the word changed" with "the run restarted", which the age is deliberately defined against. The stamp is the only number that means anything, and the row carries exactly that.
Write into poll_passes and derive the failure state from the pass
stream.
Rejected: a pass row is per-Lane and per-run; deriving "this Series is
failing since" means scanning 14 days of outcome counts per Series and
guessing at continuity across retention boundaries. The pass log answers
"what did this Site's lane do recently"; the failure table answers "which
Series are broken right now". Two questions, two tables — the pass log is
already pruned on a fixed cutoff, which would silently reset every
failure's age the moment the evidence aged out.
Refused/unreachable write nothing, but the delete happens anyway (or the upsert happens, with no delete). Rejected in both directions, above: writing condemns a library for one Site's mood; deleting claims a recovery nothing read and resets the age that makes a long failure findable. The asymmetry is the point — the two outcomes are evidence about the environment, never about a page.
Consequences
- The poller writes from the pass loop, one call after
checkOne, and clears through the ordinary success path: a Forced Poll that reads the page deletes the row with no forced branch of its own, and a Forced Poll request — which only stamps the request — clears nothing. poll_failuresis a poll-write table likepoll_lanesandpoll_passes: the store's two methods live next to the Lane-pass writers, and neither runs on a request path. A store failure is logged and the poll continues — a failure row is best-effort, and no single bad Series may stall a Lane.- The due query does not join the table. A failing Series is polled at the same pace as any other, because the moment the query started pacing by failure state, the age would stop meaning what the worklist reads it as.
- Orphan removal stays a single statement: the composite foreign key's
ON DELETE CASCADEis what takes the failure row with the Series. - The table starts empty at deploy, so nothing is findable for the first twelve hours after deploy and a pre-existing breakage reads as new. Accepted; the alternative is inventing history.
Cost of reversing
The failure state is derived, not stored: a rollback drops the table and every in-progress failure age with it, leaving the pass log's outcome counts as the only trace — which is exactly the "log line nobody reads" the ticket set out to replace. Recreating the table later starts the ages over, so a reversal that is later reversed loses the evidence of the intervening failures. The failure state itself, however, is the one thing that is not lost by reversing: it is re-derived from the next pass, in the same direction the original design derives it — the row is recreated by the next failing read and deleted by the next good one.