148 lines
8.0 KiB
Markdown
148 lines
8.0 KiB
Markdown
# ADR-0016: The failure row is the state
|
|
|
|
Date: 2026-08-22
|
|
Status: accepted
|
|
|
|
## Decision
|
|
|
|
A Series that used to work and has stopped is recorded in one table,
|
|
`poll_failures`, keyed by the same `(site, series_id)` composite the rest of
|
|
the system uses. One row per failing Series, carrying the outcome word and a
|
|
`failing_since` stamp. The row's **existence** is the failure state: there is
|
|
no success sentinel, no counter, no history, and nothing added to the Series
|
|
row. A correct read deletes the row, and a repeated identical failure writes
|
|
nothing — the stamp ages the run of failures, not the current word.
|
|
|
|
The write is one upsert, from the poll pass loop:
|
|
|
|
```sql
|
|
INSERT INTO poll_failures (site, series_id, outcome, failing_since)
|
|
VALUES ($1, $2, $3, $4)
|
|
ON CONFLICT (site, series_id) DO UPDATE SET outcome = EXCLUDED.outcome
|
|
WHERE poll_failures.outcome <> EXCLUDED.outcome
|
|
```
|
|
|
|
`failing_since` is simply absent from the `SET` arm, so it is never
|
|
overwritten, and the `WHERE` makes an unchanged word write nothing at all —
|
|
one statement, no read-then-write race. The clear is a plain `DELETE`; a row
|
|
that is not there is not an error, because the poller clears on every
|
|
successful read and most clears find nothing. The composite foreign key
|
|
cascades: deleting a Series takes its failure row, so orphan removal stays a
|
|
single statement.
|
|
|
|
## Why a future reader will find this surprising
|
|
|
|
The failure state lives in a table of its own, not on the Series row and not
|
|
in the pass log. A flag on `series` would be the familiar shape, and the
|
|
pass log already records outcomes per pass — why a third place?
|
|
|
|
Because the failure must outlive the pass that observed it and mean
|
|
something the pass row cannot say. The pass log is a per-Lane stream: it
|
|
records that a Site returned N `errors` and M `not_found` on a given run,
|
|
but "this exact Series has been failing for three months" is not a question
|
|
any single pass row answers — it is a question across rows, and the Lanes
|
|
page is a live view of the last 14 days, not a Series index. A Series-level
|
|
fact needs Series-level storage, and a flag on the Series row is the wrong
|
|
shape too: the failure is *transient by definition* (a correct read ends it)
|
|
and *repeatable* (the same Series can fail again next year), so the honest
|
|
record of "failing since" is a stamp that moves, not a column that flips.
|
|
A row that is created and deleted by the read's outcome is that stamp, and
|
|
nothing else: the moment a read succeeds the row is gone, so "is this Series
|
|
failing?" is answered by one indexed existence check — no success sentinel
|
|
to keep consistent with the failure, no counter to reset, no history to
|
|
prune. The state cannot drift out of step with the reads that maintain it,
|
|
because the reads *are* the maintenance.
|
|
|
|
The two no-write outcomes are the second surprise. `refused` (the Site is
|
|
holding a challenge) and `unreachable` (the browser sidecar was lost
|
|
mid-loop) issue **no statement at all** — neither an upsert nor a delete.
|
|
The direction matters, and both directions are wrong to touch:
|
|
|
|
- **Writing would condemn a whole library for one Site's bad day.** A
|
|
refusal is the Site's mood, not a fact about any particular Series: when
|
|
Cloudflare turns a zone's JS detection on, every Series on that Site reads
|
|
refused in the same pass. Writing those rows would stamp every Series in
|
|
the library as failing on the evidence of one Site's configuration, and
|
|
#166's worklist would present a Site-wide outage as thousands of broken
|
|
Series.
|
|
- **Deleting would claim a recovery nothing read.** A lost sidecar tells us
|
|
nothing about the page — the read never happened. Deleting the row would
|
|
report the Series healthy, and worse, it would reset `failing_since`: the
|
|
age that makes a three-month failure findable would start over on a
|
|
browser restart, erasing evidence no page read contradicted.
|
|
|
|
So a refusal or an unreachable pass leaves the table exactly as it found
|
|
it — the pass neither adds rows that would condemn nor deletes rows that
|
|
would claim recovery. The four outcomes that are real evidence about the
|
|
page (`not_found`, `no_chapter`, `unfetchable`, `errors`) write; the two
|
|
that are not write nothing; success deletes. There is no third case.
|
|
|
|
## Considered options
|
|
|
|
**A `failing_since` column on the Series row, alongside the failure word.**
|
|
Rejected: it is two more columns on a table every due query and every
|
|
bookmark join already touches, for a state that is transient and repeatable.
|
|
The column would need its own "clear on success" writer anyway — the same
|
|
maintenance the row has — while permanently widening the hottest table in
|
|
the system for a value that is usually absent. And the worklist's join
|
|
would have to distinguish "never failed" from "currently healthy", which
|
|
is exactly the sentinel problem the row avoids: an empty column means both.
|
|
|
|
**A counter or last-outcome timestamp instead of (or beside) the row.**
|
|
Rejected: nothing in the system consumes a failure *count* or a
|
|
last-outcome time — the worklist needs "failing and since when". A counter
|
|
invites "three failures = something" thresholds that the ticket explicitly
|
|
keeps out of scope, and a last-outcome timestamp conflates "the word
|
|
changed" with "the run restarted", which the age is deliberately defined
|
|
against. The stamp is the only number that means anything, and the row
|
|
carries exactly that.
|
|
|
|
**Write into `poll_passes` and derive the failure state from the pass
|
|
stream.**
|
|
Rejected: a pass row is per-Lane and per-run; deriving "this Series is
|
|
failing since" means scanning 14 days of outcome counts per Series and
|
|
guessing at continuity across retention boundaries. The pass log answers
|
|
"what did this Site's lane do recently"; the failure table answers "which
|
|
Series are broken right now". Two questions, two tables — the pass log is
|
|
already pruned on a fixed cutoff, which would silently reset every
|
|
failure's age the moment the evidence aged out.
|
|
|
|
**Refused/unreachable write nothing, but the delete happens anyway (or the
|
|
upsert happens, with no delete).**
|
|
Rejected in both directions, above: writing condemns a library for one
|
|
Site's mood; deleting claims a recovery nothing read and resets the age
|
|
that makes a long failure findable. The asymmetry is the point — the two
|
|
outcomes are evidence about the *environment*, never about a page.
|
|
|
|
## Consequences
|
|
|
|
- The poller writes from the pass loop, one call after `checkOne`, and
|
|
clears through the ordinary success path: a Forced Poll that reads the
|
|
page deletes the row with no forced branch of its own, and a Forced Poll
|
|
*request* — which only stamps the request — clears nothing.
|
|
- `poll_failures` is a poll-write table like `poll_lanes` and
|
|
`poll_passes`: the store's two methods live next to the Lane-pass
|
|
writers, and neither runs on a request path. A store failure is logged
|
|
and the poll continues — a failure row is best-effort, and no single bad
|
|
Series may stall a Lane.
|
|
- The due query does not join the table. A failing Series is polled at the
|
|
same pace as any other, because the moment the query started pacing by
|
|
failure state, the age would stop meaning what the worklist reads it as.
|
|
- Orphan removal stays a single statement: the composite foreign key's
|
|
`ON DELETE CASCADE` is what takes the failure row with the Series.
|
|
- The table starts empty at deploy, so nothing is findable for the first
|
|
twelve hours after deploy and a pre-existing breakage reads as new.
|
|
Accepted; the alternative is inventing history.
|
|
|
|
## Cost of reversing
|
|
|
|
The failure state is derived, not stored: a rollback drops the table and
|
|
every in-progress failure age with it, leaving the pass log's outcome
|
|
counts as the only trace — which is exactly the "log line nobody reads"
|
|
the ticket set out to replace. Recreating the table later starts the ages
|
|
over, so a reversal that is later reversed loses the evidence of the
|
|
intervening failures. The failure state itself, however, is the one thing
|
|
that is *not* lost by reversing: it is re-derived from the next pass, in
|
|
the same direction the original design derives it — the row is
|
|
recreated by the next failing read and deleted by the next good one.
|