Files
mangaBookmark/docs/adr/0016-the-failure-row-is-the-state.md
sulthan e93e79c1bb Spec #137: per-Series poll failure state, completion hint, and outbound owner notification (#174)
Implements spec #137 (spec 4 of 4 from wayfinder map #114).

Closes #137.

Tickets: #164, #165, #166, #167, #168, #169, #170, #171, #172, #173 — all closed, landed on this branch.

## Summary
- #164/#168: sixth outcome word `not_found`; per-Site completed marker predicate.
- #165/#169: `poll_failures` row is the failure state; the pass remembers a Site-reported completion.
- #166/#171: two new admin filters (`failing`, `unverified`); outbound owner notification + stall condition.
- #167/#170: a failure names itself on the Series page; the completion hint reaches the owner and decides nothing.
- #172: the other three fault conditions (no-browser-route, sidecar-down, adapter-broken) feeding the notifier.
- #173: the landing verdict line shares the same `latest.FaultsFrom` judgement the notifier uses, so the page and the push cannot disagree.

`cd backend && go test ./...` green on the merged branch (8 packages).

Reviewed-on: #174
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-23 11:31:10 +07:00

8.0 KiB

ADR-0016: The failure row is the state

Date: 2026-08-22 Status: accepted

Decision

A Series that used to work and has stopped is recorded in one table, poll_failures, keyed by the same (site, series_id) composite the rest of the system uses. One row per failing Series, carrying the outcome word and a failing_since stamp. The row's existence is the failure state: there is no success sentinel, no counter, no history, and nothing added to the Series row. A correct read deletes the row, and a repeated identical failure writes nothing — the stamp ages the run of failures, not the current word.

The write is one upsert, from the poll pass loop:

INSERT INTO poll_failures (site, series_id, outcome, failing_since)
VALUES ($1, $2, $3, $4)
ON CONFLICT (site, series_id) DO UPDATE SET outcome = EXCLUDED.outcome
WHERE poll_failures.outcome <> EXCLUDED.outcome

failing_since is simply absent from the SET arm, so it is never overwritten, and the WHERE makes an unchanged word write nothing at all — one statement, no read-then-write race. The clear is a plain DELETE; a row that is not there is not an error, because the poller clears on every successful read and most clears find nothing. The composite foreign key cascades: deleting a Series takes its failure row, so orphan removal stays a single statement.

Why a future reader will find this surprising

The failure state lives in a table of its own, not on the Series row and not in the pass log. A flag on series would be the familiar shape, and the pass log already records outcomes per pass — why a third place?

Because the failure must outlive the pass that observed it and mean something the pass row cannot say. The pass log is a per-Lane stream: it records that a Site returned N errors and M not_found on a given run, but "this exact Series has been failing for three months" is not a question any single pass row answers — it is a question across rows, and the Lanes page is a live view of the last 14 days, not a Series index. A Series-level fact needs Series-level storage, and a flag on the Series row is the wrong shape too: the failure is transient by definition (a correct read ends it) and repeatable (the same Series can fail again next year), so the honest record of "failing since" is a stamp that moves, not a column that flips. A row that is created and deleted by the read's outcome is that stamp, and nothing else: the moment a read succeeds the row is gone, so "is this Series failing?" is answered by one indexed existence check — no success sentinel to keep consistent with the failure, no counter to reset, no history to prune. The state cannot drift out of step with the reads that maintain it, because the reads are the maintenance.

The two no-write outcomes are the second surprise. refused (the Site is holding a challenge) and unreachable (the browser sidecar was lost mid-loop) issue no statement at all — neither an upsert nor a delete. The direction matters, and both directions are wrong to touch:

  • Writing would condemn a whole library for one Site's bad day. A refusal is the Site's mood, not a fact about any particular Series: when Cloudflare turns a zone's JS detection on, every Series on that Site reads refused in the same pass. Writing those rows would stamp every Series in the library as failing on the evidence of one Site's configuration, and #166's worklist would present a Site-wide outage as thousands of broken Series.
  • Deleting would claim a recovery nothing read. A lost sidecar tells us nothing about the page — the read never happened. Deleting the row would report the Series healthy, and worse, it would reset failing_since: the age that makes a three-month failure findable would start over on a browser restart, erasing evidence no page read contradicted.

So a refusal or an unreachable pass leaves the table exactly as it found it — the pass neither adds rows that would condemn nor deletes rows that would claim recovery. The four outcomes that are real evidence about the page (not_found, no_chapter, unfetchable, errors) write; the two that are not write nothing; success deletes. There is no third case.

Considered options

A failing_since column on the Series row, alongside the failure word. Rejected: it is two more columns on a table every due query and every bookmark join already touches, for a state that is transient and repeatable. The column would need its own "clear on success" writer anyway — the same maintenance the row has — while permanently widening the hottest table in the system for a value that is usually absent. And the worklist's join would have to distinguish "never failed" from "currently healthy", which is exactly the sentinel problem the row avoids: an empty column means both.

A counter or last-outcome timestamp instead of (or beside) the row. Rejected: nothing in the system consumes a failure count or a last-outcome time — the worklist needs "failing and since when". A counter invites "three failures = something" thresholds that the ticket explicitly keeps out of scope, and a last-outcome timestamp conflates "the word changed" with "the run restarted", which the age is deliberately defined against. The stamp is the only number that means anything, and the row carries exactly that.

Write into poll_passes and derive the failure state from the pass stream. Rejected: a pass row is per-Lane and per-run; deriving "this Series is failing since" means scanning 14 days of outcome counts per Series and guessing at continuity across retention boundaries. The pass log answers "what did this Site's lane do recently"; the failure table answers "which Series are broken right now". Two questions, two tables — the pass log is already pruned on a fixed cutoff, which would silently reset every failure's age the moment the evidence aged out.

Refused/unreachable write nothing, but the delete happens anyway (or the upsert happens, with no delete). Rejected in both directions, above: writing condemns a library for one Site's mood; deleting claims a recovery nothing read and resets the age that makes a long failure findable. The asymmetry is the point — the two outcomes are evidence about the environment, never about a page.

Consequences

  • The poller writes from the pass loop, one call after checkOne, and clears through the ordinary success path: a Forced Poll that reads the page deletes the row with no forced branch of its own, and a Forced Poll request — which only stamps the request — clears nothing.
  • poll_failures is a poll-write table like poll_lanes and poll_passes: the store's two methods live next to the Lane-pass writers, and neither runs on a request path. A store failure is logged and the poll continues — a failure row is best-effort, and no single bad Series may stall a Lane.
  • The due query does not join the table. A failing Series is polled at the same pace as any other, because the moment the query started pacing by failure state, the age would stop meaning what the worklist reads it as.
  • Orphan removal stays a single statement: the composite foreign key's ON DELETE CASCADE is what takes the failure row with the Series.
  • The table starts empty at deploy, so nothing is findable for the first twelve hours after deploy and a pre-existing breakage reads as new. Accepted; the alternative is inventing history.

Cost of reversing

The failure state is derived, not stored: a rollback drops the table and every in-progress failure age with it, leaving the pass log's outcome counts as the only trace — which is exactly the "log line nobody reads" the ticket set out to replace. Recreating the table later starts the ages over, so a reversal that is later reversed loses the evidence of the intervening failures. The failure state itself, however, is the one thing that is not lost by reversing: it is re-derived from the next pass, in the same direction the original design derives it — the row is recreated by the next failing read and deleted by the next good one.