Record poll failures as rows: the row is the failure state (#165)
This commit is contained in:
@@ -0,0 +1,147 @@
|
||||
# ADR-0016: The failure row is the state
|
||||
|
||||
Date: 2026-08-22
|
||||
Status: accepted
|
||||
|
||||
## Decision
|
||||
|
||||
A Series that used to work and has stopped is recorded in one table,
|
||||
`poll_failures`, keyed by the same `(site, series_id)` composite the rest of
|
||||
the system uses. One row per failing Series, carrying the outcome word and a
|
||||
`failing_since` stamp. The row's **existence** is the failure state: there is
|
||||
no success sentinel, no counter, no history, and nothing added to the Series
|
||||
row. A correct read deletes the row, and a repeated identical failure writes
|
||||
nothing — the stamp ages the run of failures, not the current word.
|
||||
|
||||
The write is one upsert, from the poll pass loop:
|
||||
|
||||
```sql
|
||||
INSERT INTO poll_failures (site, series_id, outcome, failing_since)
|
||||
VALUES ($1, $2, $3, $4)
|
||||
ON CONFLICT (site, series_id) DO UPDATE SET outcome = EXCLUDED.outcome
|
||||
WHERE poll_failures.outcome <> EXCLUDED.outcome
|
||||
```
|
||||
|
||||
`failing_since` is simply absent from the `SET` arm, so it is never
|
||||
overwritten, and the `WHERE` makes an unchanged word write nothing at all —
|
||||
one statement, no read-then-write race. The clear is a plain `DELETE`; a row
|
||||
that is not there is not an error, because the poller clears on every
|
||||
successful read and most clears find nothing. The composite foreign key
|
||||
cascades: deleting a Series takes its failure row, so orphan removal stays a
|
||||
single statement.
|
||||
|
||||
## Why a future reader will find this surprising
|
||||
|
||||
The failure state lives in a table of its own, not on the Series row and not
|
||||
in the pass log. A flag on `series` would be the familiar shape, and the
|
||||
pass log already records outcomes per pass — why a third place?
|
||||
|
||||
Because the failure must outlive the pass that observed it and mean
|
||||
something the pass row cannot say. The pass log is a per-Lane stream: it
|
||||
records that a Site returned N `errors` and M `not_found` on a given run,
|
||||
but "this exact Series has been failing for three months" is not a question
|
||||
any single pass row answers — it is a question across rows, and the Lanes
|
||||
page is a live view of the last 14 days, not a Series index. A Series-level
|
||||
fact needs Series-level storage, and a flag on the Series row is the wrong
|
||||
shape too: the failure is *transient by definition* (a correct read ends it)
|
||||
and *repeatable* (the same Series can fail again next year), so the honest
|
||||
record of "failing since" is a stamp that moves, not a column that flips.
|
||||
A row that is created and deleted by the read's outcome is that stamp, and
|
||||
nothing else: the moment a read succeeds the row is gone, so "is this Series
|
||||
failing?" is answered by one indexed existence check — no success sentinel
|
||||
to keep consistent with the failure, no counter to reset, no history to
|
||||
prune. The state cannot drift out of step with the reads that maintain it,
|
||||
because the reads *are* the maintenance.
|
||||
|
||||
The two no-write outcomes are the second surprise. `refused` (the Site is
|
||||
holding a challenge) and `unreachable` (the browser sidecar was lost
|
||||
mid-loop) issue **no statement at all** — neither an upsert nor a delete.
|
||||
The direction matters, and both directions are wrong to touch:
|
||||
|
||||
- **Writing would condemn a whole library for one Site's bad day.** A
|
||||
refusal is the Site's mood, not a fact about any particular Series: when
|
||||
Cloudflare turns a zone's JS detection on, every Series on that Site reads
|
||||
refused in the same pass. Writing those rows would stamp every Series in
|
||||
the library as failing on the evidence of one Site's configuration, and
|
||||
#166's worklist would present a Site-wide outage as thousands of broken
|
||||
Series.
|
||||
- **Deleting would claim a recovery nothing read.** A lost sidecar tells us
|
||||
nothing about the page — the read never happened. Deleting the row would
|
||||
report the Series healthy, and worse, it would reset `failing_since`: the
|
||||
age that makes a three-month failure findable would start over on a
|
||||
browser restart, erasing evidence no page read contradicted.
|
||||
|
||||
So a refusal or an unreachable pass leaves the table exactly as it found
|
||||
it — the pass neither adds rows that would condemn nor deletes rows that
|
||||
would claim recovery. The four outcomes that are real evidence about the
|
||||
page (`not_found`, `no_chapter`, `unfetchable`, `errors`) write; the two
|
||||
that are not write nothing; success deletes. There is no third case.
|
||||
|
||||
## Considered options
|
||||
|
||||
**A `failing_since` column on the Series row, alongside the failure word.**
|
||||
Rejected: it is two more columns on a table every due query and every
|
||||
bookmark join already touches, for a state that is transient and repeatable.
|
||||
The column would need its own "clear on success" writer anyway — the same
|
||||
maintenance the row has — while permanently widening the hottest table in
|
||||
the system for a value that is usually absent. And the worklist's join
|
||||
would have to distinguish "never failed" from "currently healthy", which
|
||||
is exactly the sentinel problem the row avoids: an empty column means both.
|
||||
|
||||
**A counter or last-outcome timestamp instead of (or beside) the row.**
|
||||
Rejected: nothing in the system consumes a failure *count* or a
|
||||
last-outcome time — the worklist needs "failing and since when". A counter
|
||||
invites "three failures = something" thresholds that the ticket explicitly
|
||||
keeps out of scope, and a last-outcome timestamp conflates "the word
|
||||
changed" with "the run restarted", which the age is deliberately defined
|
||||
against. The stamp is the only number that means anything, and the row
|
||||
carries exactly that.
|
||||
|
||||
**Write into `poll_passes` and derive the failure state from the pass
|
||||
stream.**
|
||||
Rejected: a pass row is per-Lane and per-run; deriving "this Series is
|
||||
failing since" means scanning 14 days of outcome counts per Series and
|
||||
guessing at continuity across retention boundaries. The pass log answers
|
||||
"what did this Site's lane do recently"; the failure table answers "which
|
||||
Series are broken right now". Two questions, two tables — the pass log is
|
||||
already pruned on a fixed cutoff, which would silently reset every
|
||||
failure's age the moment the evidence aged out.
|
||||
|
||||
**Refused/unreachable write nothing, but the delete happens anyway (or the
|
||||
upsert happens, with no delete).**
|
||||
Rejected in both directions, above: writing condemns a library for one
|
||||
Site's mood; deleting claims a recovery nothing read and resets the age
|
||||
that makes a long failure findable. The asymmetry is the point — the two
|
||||
outcomes are evidence about the *environment*, never about a page.
|
||||
|
||||
## Consequences
|
||||
|
||||
- The poller writes from the pass loop, one call after `checkOne`, and
|
||||
clears through the ordinary success path: a Forced Poll that reads the
|
||||
page deletes the row with no forced branch of its own, and a Forced Poll
|
||||
*request* — which only stamps the request — clears nothing.
|
||||
- `poll_failures` is a poll-write table like `poll_lanes` and
|
||||
`poll_passes`: the store's two methods live next to the Lane-pass
|
||||
writers, and neither runs on a request path. A store failure is logged
|
||||
and the poll continues — a failure row is best-effort, and no single bad
|
||||
Series may stall a Lane.
|
||||
- The due query does not join the table. A failing Series is polled at the
|
||||
same pace as any other, because the moment the query started pacing by
|
||||
failure state, the age would stop meaning what the worklist reads it as.
|
||||
- Orphan removal stays a single statement: the composite foreign key's
|
||||
`ON DELETE CASCADE` is what takes the failure row with the Series.
|
||||
- The table starts empty at deploy, so nothing is findable for the first
|
||||
twelve hours after deploy and a pre-existing breakage reads as new.
|
||||
Accepted; the alternative is inventing history.
|
||||
|
||||
## Cost of reversing
|
||||
|
||||
The failure state is derived, not stored: a rollback drops the table and
|
||||
every in-progress failure age with it, leaving the pass log's outcome
|
||||
counts as the only trace — which is exactly the "log line nobody reads"
|
||||
the ticket set out to replace. Recreating the table later starts the ages
|
||||
over, so a reversal that is later reversed loses the evidence of the
|
||||
intervening failures. The failure state itself, however, is the one thing
|
||||
that is *not* lost by reversing: it is re-derived from the next pass, in
|
||||
the same direction the original design derives it — the row is
|
||||
recreated by the next failing read and deleted by the next good one.
|
||||
Reference in New Issue
Block a user