Spec: per-Series poll failure state, completion hint, and outbound owner notification #137
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Spec derived from the wayfinder map #114, which is fully charted. This is spec 4 of 4; the map's
decisions were settled in #127, #129, #130 and #132.
Blocked by: #134, #135, #136 (specs 1, 2 and 3 of this series). Spec 1 creates the pass log, the filter vocabulary and
the pages; spec 2 fixes the rule that a human correction clears nothing and the cascade requirement
for any new per-Series table; spec 3 creates the Finished Series flag the completion hint sits beside.
Problem Statement
A Series can be broken for three months and every surface calls it healthy.
The check timestamp is stamped before the fetch, deliberately — otherwise a renamed or deleted
Series is retried on every pass for ever. So the column means attempted, never succeeded. A Series
whose page has quietly broken — a markup change, a wrong address, a permanent 404 — is stamped every
hour and reads as perfectly fine: it is not "never checked" (it has a timestamp), not "stale" (the
timestamp is minutes old), and not "unpollable" (it has an address). The only trace is log lines
nobody reads, and the per-Site outcome counts roll off in twelve hours without ever naming the Series.
And a number nothing can verify looks exactly like one a Poll confirmed. Where a Reader's report
raised a Latest Chapter on a Series the Poll cannot read, that number stands indefinitely with no
mark on it.
Separately: I have no idea which works have actually ended. I can mark a Series finished, but I
have to notice on my own — by remembering a series I have not seen a chapter of in a year. Every one
of the six Sites publishes its own completion status on the page the Poll already reads, and the
backend throws it away.
And the whole surface is pull-only. The one fault class that is invisible from every Reader
surface — a Lane that owed Polls, made none, and has nothing to say for it — is also the one there is
no reason to go looking for. The web UI dies visibly; the userscript degrades honestly to its cache
with a pending count in the panel and loses no Progress. Those faults find themselves within one
reading session. A stall does not: every Bookmark still opens, Progress still syncs, and Latest
Chapter is quietly wrong for as long as it lasts.
Solution
Three things a Series remembers about itself, and one message that leaves the building.
Per-Series failure state, as a row whose existence is the state. A row in a new table means "the
last Poll that learned anything about this Series failed", and a correct read deletes it. No success
sentinel, no counter, no history. Two new filters follow: Series failing for over twelve hours, and
the subset of those whose Latest Chapter came from a Reader and has never been confirmed.
Two outcomes deliberately say nothing. A Site refusing us and our own browser being unreachable
are not evidence about any particular Series. Writing a row would mark a whole library as failing when
one Site had a bad day; deleting one would claim recovery when nothing was read, resetting the age
that makes a three-month failure findable.
A completion hint, from the marker every Site already publishes. One column holding when the last
successful read saw the Site's own completed value. Decay is free: every successful read rewrites
the answer. It is a hint next to the Finish control and a filter to work through — never a shortcut,
never a writer of the Series' finished flag.
Four conditions push one Discord message each. The rule that selects them is the decision: silent
on every Reader surface, wrong, repairable only by the owner, and holding past twelve hours. A new
table remembers only that the owner was told, so a fault lasting a month sends one message rather
than thirty. No new timer — each Lane already wakes at least hourly and is its own clock.
User Stories
breakage does not survive for months behind a healthy-looking timestamp.
adapter cannot parse, and a database write that failed are three different words and three
different actions.
breakage from one that has stood since spring.
months of parse failures does not reset to "failing for a minute".
and broke on Tuesday, so that a never-implemented adapter and a regressed one are not the same
list.
recovery needs no sweep and no reconciliation.
the same as a healthy one every hour.
day does not mark its whole library as broken.
that a machine I turned off does not look like two hundred broken Series.
cannot erase the age of a three-month failure and hide it from the list.
never mistaken for the page becoming readable.
successful read does, so that there is no special case to remember.
confirmed in twelve hours to be its own list, so that I have a worklist of numbers nothing stands
behind.
is not a fault to correct.
identical on the page, so that I do not have to learn a visual rule the line cannot state in a
word — the age is right there.
opens an empty list.
from "this Site looks unwell" to "these rows are broken" in one click.
failure does not quietly slow a Lane.
per-Series state does not turn a delete into a multi-step dance.
have to remember which works ended.
dropped translation is never presented as a finished work.
broken selector degrades to a missing hint rather than to a suspicion.
stops being hinted without any expiry rule.
that a challenge is not treated as evidence about the work.
in one sitting.
without opening it.
row, so that a rare signal does not spend a field on every row.
so that the Site never decides the Lifecycle.
beside the Site's marker and my own decision.
work.
that an absent hint means "not completed" everywhere rather than "we never look here".
that the one fault no surface shows reaches me anyway.
that a Site turning hostile overnight is repaired by a deploy rather than discovered weeks later.
did not clear is not reported as a fault.
so that "asleep for an hour" and "gone for a week" stop looking the same to me.
a single fault is a single notification.
chapter for over twelve hours, so that a layout change is reported as a broken adapter rather than
as hundreds of unrelated bad rows.
two-hundred-Series Site both fire.
reading.
not swallowed by the first.
to me.
hundreds of messages.
not go hunting from my phone.
that I can decide whether to get up.
searchable.
without arithmetic.
stack runs without a webhook.
about a durable fact needs no machinery to survive.
message cannot become an availability concern for the login flow.
message never sends me to a page that looks healthy.
Implementation Decisions
Per-Series failure state: one table, the row's existence is the state
A row means "the last Poll that learned anything about this Series failed". No row means it succeeded,
or was never polled (a zero check stamp separates that), or every Poll so far ended in an outcome the
write rule ignores. So there is no
''success sentinel, noCHECKkeeping two columns consistent,and
failing_sinceis never zero.The cascade reaches this table only — never
bookmarks, whose refusal is the orphan guard, soremoving an orphan stays a single statement.
Recorded because it will be re-litigated: 3NF does not decide table-versus-columns here. The key is
(site, series_id)either way, the relation is 1:1, and the one rule between the attributes (a wordimplies an age) yields no value, so it is a constraint rather than a transitive dependency — both
designs are in BCNF. A 1:1 table with the same key is vertical partitioning, not normalisation, and a
lookup table for the outcome vocabulary would be a domain constraint a Go constant already enforces.
The table won on its own merits:
seriesstays narrow, the sentinel disappears, and the guarded writegets simpler.
Rejected: three columns on
series— a last outcome, a first-failure timestamp, and aconsecutive-failure counter. The counter goes because it is a proxy for duration measured in units that
differ per Site: each Lane carries its own rest, and a paused Lane stops the count without stopping the
failure.
failing since <age>is the fact the owner reads.The write
failing_sincesurvives a change of word because the conflict update never touches it — the age is theage of the run of failures, not of the current word. The same failure repeating writes nothing; a
healthy Series polling correctly writes nothing. The cost lands on transitions, not on Polls. This is
also why there is no
last_outcome_atcolumn: it would sit within seconds of the check stamp, which isalready stored.
Which outcomes write, and which are silent
no_chapterunfetchablenot_founderrorsrefusedunreachableNeither insert nor delete for the last two, and the distinction is load-bearing in both directions.
Writing a row would mark a whole library as failing when one Site refused for a day. Deleting one would
claim recovery when nothing was read — resetting the age for a Series broken three months, so a single
refusal from its Site would erase exactly the fact that makes it findable. A refused or unreachable
Poll therefore issues one fewer query, and the Series keeps the last thing it truly learned about
itself. Both facts are already recorded where they belong: per Site in the durable Lane refusal, and in
the poller's in-memory browser-down timestamp.
errorssplits:not_foundbecomes a sixth outcome word for a 4xx status other than 403. Today oneword covers a 404 (the page is gone — correct the address), a 503 (the Site is busy — wait) and our own
write failing (the database is unwell): three different owner actions behind the one word a page prints.
One branch in the read path, and the pass log gains a sixth count column.
Written by nothing else, ever. A Forced Poll request clears nothing, because a request is not
evidence; a forced Poll that reads the page clears the row exactly as any other correct read does, which
needs no special case in the code. A hand correction clears nothing, as #135 requires.
Two new filters
failing— the joined row existsAND s.latest_chapter_num IS NOT NULL AND f.failing_sinceolderthan
ownerWindow(12h; reuse the constant, add no new figure). Thelatest_chapter_num IS NOT NULLtest is what keeps it disjoint from
no-chapter(never once succeeded) rather than overlapping itfor no gain. Deliberately no
finished_at = 0guard: a Series already failing when the ownerfinished it keeps its evidence, and guarding it would erase evidence rather than prevent a false
figure. (Contrast the four predicates #136 does guard: those are computed from clocks that keep
ticking, so they lie about a finished row; this one is computed from stored outcomes, which simply
stop arriving.)
unverified—s.latest_raised_by IS NOT NULLplus thefailingtest. The Latest Chapter camefrom a Reader's Sighting and no Poll has confirmed it for over twelve hours. Spec 2's correction
control is the action it leads to, and the detail page adds one sentence beside that control for this
case.
Rejected: a stored column for the
unverifiedpair. It would hold the conjunction of a fact onseriesand a row in another table — derived data with five writers (a Sighting raising the Series, aPoll succeeding, a Poll failing, an owner correction, and clearing the attribution) for a value the join
computes free. Unlike the table-versus-columns question above, this one is a genuine 3NF violation, and
the admin read model acquires the join for other reasons anyway.
AdminSeriesgainsLEFT JOIN poll_failures f USING (site, series_id), andSeriesRowgains theoutcome word and the failing-since timestamp. Nothing is added to
series— the row's existence inthe new table is the failure state, so there is no success sentinel to project.
The per-Site outcome counts stay unlinked, permanently
Earlier charting promised that once a
failingfilter existed, all the per-Site outcome counts wouldlink to it. That promise is withdrawn, on the rule that a figure links to what it counts — the same
rule that already refused to link the
no_chaptercount to theno-chapterfilter.Two failures of the same test.
refusedandunreachablewrite no per-Series state at all, so two ofthe six counts would open an empty list beside a large number. And the other four count attempts inside
a twelve-hour window while the filter lists Series failing now for over twelve hours: a Series that
broke at 04:00 and healed at 05:00 is counted and not listed, while one broken for three months on a
paused Lane is listed and counted zero. Neither set contains the other, in either direction.
Instead each Site row on the Lanes page carries one separate link to
/admin/series?site=<site>&filter=failing, as navigation rather than attached to a number. Rejected: an?outcome=parameter to make the links exact — the surface is one filter at a time with only?site=and
?kind=stacking, and a second axis would not fix the window mismatch anyway.Completion hint: storage
The false-positive containment this was named for is not built, because the research removed the
threat. Measured 2026-08-19 against live pages (note:
docs/research/completion-markers.mdon branchresearch/completion-markers): all six Sites publish an explicit machine-readable completion marker, anda Poll that triggers only on each Site's completed value has no false-positive path on any Site. So
there is no N-consecutive-observation gate, no expiry, no confirmation counter and no suspicion damping.
The residual error is a false negative — a finished Series that gets no hint — and we accept it.
series.site_completed_at bigint NOT NULL DEFAULT 0, epoch-ms of the read that saw the Site's completedvalue; zero means the last successful read did not see it. Decay is therefore free: every successful
read rewrites the answer, so a stale hint cannot survive one Poll. The timestamp is the age the hint
renders with, not a history.
demonicscans and novelfull are binary and cannot express hiatus at all; asurascans'
droppedandkagane's upload status are scanlation-editorial rather than completion state; and kagane is the only
Site that separates the work's status from the translation's (proven divergent on one series). A
shared word would be a different fact per Site, and nothing here acts on any value but completed.
the admin surface never fetches — every intervention goes through the database.
Completion hint: which value counts
One predicate per Site, the completed value only, no mapping table:
completed(itsdroppedis editorial: one series stops at chapter 14 while the work continues)Completedfinished(on_hiatus,discontinued,not_yet_releasedare distinct)Completedonly — upload status is the release's state, and the two divergeCompletedrather than zeroing. Inherited free from the failure-write rule above: a challenge is not evidence
about the Series.
props, comix's initial-data detail entry, kagane's API body, lightnovelworld's head JSON-LD).
demonicscans (an info-block list-item pair) and novelfull (a status link) each need one new selector
against markup nobody controls. Paid anyway, on two grounds: a broken selector degrades to a missing
hint, which is the error class we already accept; and on exactly those two Sites
Completedis the onlysignal the Site can ever emit, so omitting them makes an absent hint unreadable — the owner could not
tell "not finished" from "we never look".
Completion hint: the write
Its own statement in the pass, after the read. It cannot ride the check stamp (written before the
fetch) and it cannot ride the chapter setter (skipped on the unchanged-number path). It cannot share the
failure write either: that one targets the failure table and fires on a failure, while this is a fact
learned from a successful read of
series. The two share only their position in the pass.site_completed_atjoins the due query's projection, so the statement fires only when the answer wouldchange. The ordinary Poll of an ongoing Series writes nothing, exactly as a healthy Poll writes nothing
to the failure table.
Completion hint: presentation
site-completed, ordered with the informational tail besidefinished, andcarrying
AND finished_at = 0— the same guard #136's hygiene predicates gained.lines, and this fact is true of few Series, so a per-row field spends space on every row for a rare
signal. The filter is the work list; the figure is how the owner sees the work without opening it.
it: it may not pre-fill, add a second button, or remove the confirm step. A hint that shortens the
path is the Site deciding the Lifecycle.
opinion on the row beside the Site's marker and the owner's own flag, and would need its own rule for
the day the Site's marker changes again. An ignored Series stays in a fifty-row filter; if that noise
turns out to be real, the column can be added then.
leaves the Poll query, so its value freezes at whatever the last Poll saw — old and unrefreshable,
therefore not work. Un-finishing puts it back in the query and the next Poll writes a fresh answer.
the two browser-gated Sites' hints age while the sidecar is asleep or unreachable — the same degradation
their Covers already have, and no new notification.
Outbound notification: the rule, then the conditions
The dashboard is not pull-only. Push exists for one class of fault: silent on every Reader surface,
wrong, repairable only by the owner, and persisting past twelve hours. That rule is the decision — the
four conditions below are what it currently selects, and a fifth candidate must pass all four tests.
A backend that stopped is not the case worth pushing, and the reason came from the Reader surfaces
rather than from the plan. The web UI dies visibly. The userscript degrades quietly but honestly: reads
come from the cache, a write that meets a dropped connection goes to the retry queue, the queue only
speaks up for a 400, a 401 or a full queue, and the sole visible marker is the pending count in the panel.
No Progress is lost — the only loss path is eviction at the queue cap, which toasts. So downtime costs a
delay and is found within one reading session, by every Reader. A stalled Lane is the opposite.
The four conditions, each at
ownerWindow(12h; no new constant):stall—due > 0 AND checked = 0 AND skip = '' AND refused = 0on the pass log.no-browser-route— a Site with no browser route refusing for more than 12h. The read pathmaps both a 403 with the mitigation header and an interstitial body to a held challenge on the
plain-TLS path too, so this is
refusedplus the durable Lane refusal. This is the comix event of2026-08-12, and the repair is a deploy. A browser Site refusing sends nothing — that is an ordinary
challenge that did not clear, and a Site setting its own pace is not a fault.
sidecar-down— no browser Lane has reached the sidecar for more than 12h. This is not a stalland cannot be found by the stall test: the pass returns early both when a sibling Lane lost Chrome
and when no fetcher exists, and both are recorded as skip values. The quiet degradation is deliberate,
but "asleep for an hour" and "gone for a week" are indistinguishable to the backend, and the second
freezes two Sites' marks.
adapter-broken— more than half of one Site's Series hold ano_chapterfailure row olderthan 12h. A layout change produces exactly that outcome; one such Series is a bad row, most of a Site
is a broken adapter. A share, not a count, so a four-Series Site and a two-hundred-Series Site both
fire. This is the one fault on the whole map that no other signal names.
Never pushed: a refusal by a browser Site (not a fault), a pause (the owner's own act), a single
failing Series (the filter is the tool, and one Site change makes hundreds of rows), a forced Poll ageing
(the same fault from the other end, and visible on the page the button lives on), a gated cover host (a
missing Cover is visible in the UI).
Two collisions found in the code, and their fixes
early return — the pass falls through to its normal return, writing
due > 0, checked = 0, skip = ''.A newly-gated Site would have sent both
stallandno-browser-route. So the stall test gainsAND refused = 0, amending the definition #134 lands. Rejected alternative: a tenth skip value,because the enum holds one value per return path and this is a break inside one — a tenth value would
describe a pass that did run.
suppression row's
siteholds the empty string forsidecar-down: the first Lane to notice writesthe row, the others find it present and stay quiet. The row's own presence is the lock — which is
exactly why the state is its own table and not columns on the per-Site Lane row, where an empty Site name
has no meaning. Rejected: naming one Lane the reporter, since that Lane's pass may be the skipped one.
Where it fires from, and suppression
Beside the pass-log insert, in the deferred call every return path passes through. Two of the four
conditions occur on early returns, so the success path can never see them.
No new timer. The Lane loop already does one pass, then waits the pace the pass reported, which is
never more than the default rest — each Lane is its own clock at least hourly. (The premise that the
backend has no clock was wrong; it has six.)
New table
owner_notices(kind, site, notified_at)— one row per episode, cleared when the conditionno longer holds, so a fault lasting a month sends one message rather than thirty at ~120 passes a day.
Every threshold is derived, not stored: a refusal's and a sidecar absence's age from the pass log, a
Series failure's age from its failing-since. The only new durable fact is that the owner was told,
which is why the table carries nothing else. Rows are bounded by four kinds times six Sites, so there is
no retention rule.
Transport
DISCORD_WEBHOOK_URL, and unset means the whole path is off, logged once at start, exactly as thebrowser URL behaves — a local stack must not need a webhook to run.
OAuth application. The existing Discord coupling gains nothing new.
irrelevant at this volume (roughly 30 messages a minute allowed per webhook; this sends a few a day).
condition holds. No queue, no backoff: the condition is durable, so the retry is free. (The userscript
needs a queue because a Reader's Progress is otherwise lost; a message about a durable fact is not.)
second identity to hold when the owner's Discord id already names the recipient; the public ntfy.sh
server — 250 messages a day and 60 burst, but every topic is public by design (no authentication;
anyone knowing the topic can read and publish) with limited retention; self-hosted ntfy or Gotify — a
container and a volume, which is the infrastructure this decision refuses.
Presentation
A Discord embed, colour
13589581— the dark branch of the danger token. One colour for all four,because all four are faults; ember stays forbidden, and no second colour is introduced because no
non-fault condition pushes.
title= the subject (the Site, orbrowser sidecar) carrying the deep link built fromPUBLIC_BASE_URL, so no body line is spent on an address.description= one sentence: condition, age, repair.timestamp= the pass time, which Discord renders in the reader's own zone — the one thing an embeddoes better than plain text.
footer= the machine word (stall,no-browser-route,sidecar-down,adapter-broken), making anold message searchable.
Only the colour of Cinder survives into Discord; the typography cannot.
The page and the push stay in step
The landing verdict line must read the same four conditions. A message naming a fault the page cannot
explain sends the owner to a page that looks healthy. This amends #134's verdict line, which reads Lane
attention generally; narrow it to these four.
Testing Decisions
What makes a good test here: it asserts a durable outcome — a row present or absent after a pass, a
filter's membership, a payload's shape — and it fails on a plausible bug. The two silent outcomes are the
highest-value tests in this spec, because their bug is invisible: writing where nothing should be written
produces a plausible-looking row.
The primary seam stays the router for everything the owner sees, and the poller's existing
round-at-a-time entry point for everything the Lane writes. One new seam, and only one:
Poller.Notifier func(...) errorwithnilmeaning off — the same field-injection pattern the fetchersalready use, so all four conditions are testable with no wire and no Discord.
Modules and coverage:
writing outcome inserts a row with the right word; a repeated identical failure writes nothing; a changed
word updates the word and keeps the failing-since (this is the test that protects the age); a
successful read deletes the row; a refused Poll neither inserts nor deletes, and an unreachable
Poll neither inserts nor deletes — asserted against a pre-seeded row so both directions are covered;
a forced Poll that succeeds deletes the row and a forced request alone does not; the completion write
fires only on a transition and not on an ongoing Series' ordinary Poll; a refused Poll leaves the
completion column unchanged; the six-word outcome taxonomy maps a 404 to
not_foundand a 503 toerrors.network-gated live smoke tests): each of the six Sites' completed predicate against a completed fixture
and an ongoing one; kagane's divergent case asserts that the upload status is ignored; a fixture with
the selector removed yields not-completed rather than an error. The live smoke tests stay
network-gated and are not part of the default run.
failingandunverifiedpredicates, including disjointness fromno-chapter(a Seriesthat never succeeded must be in
no-chapterand never infailing); a Series failing for under twelvehours is in neither filter; a finished Series still appears in
failing;site-completedexcludes afinished Series; the cascade removes failure rows when a Series is deleted and the orphan delete stays a
single statement; the admin join projects the outcome word and the age.
three-month failure; the two new filters and the completion figure and filter; the outcome counts render
unlinked while the per-Site navigation link is present; the
unverifiedsentence appears beside thecorrection control; the hint is hidden on a finished Series and the Finish control is unchanged when a
hint is present; the verdict line reflects each of the four conditions.
returns; a stall on a Site that also refused fires only
no-browser-route(the collision test — thisis the one a naive implementation gets wrong); one dead sidecar across three browser Lanes sends exactly
one message; a pause sends nothing; a browser Site's refusal sends nothing; a single failing Series sends
nothing; the adapter-broken share fires on a four-Series Site and does not fire at exactly half; a failed
send leaves the notified timestamp unset and the next pass retries. One separate test covers the embed
payload — colour, title link, footer word, timestamp, and the absence of a field grid.
Out of Scope
Series is polled at the same pace as any other. Pacing is not a dashboard question.
translation-status value.
Ongoing for ever; that is the accepted false negative.
removing its confirm step.
unverifiedpair.refusal, one failing Series, an ageing forced Poll, and a gated cover host.
service. A push cannot report its own process's death, so the honest coverage lives outside; but it is
another service to run, watch and update, and it watches a fault that is already visible from the
Reader surfaces: the web UI dies visibly and the userscript degrades to its cache with a pending count,
losing no Progress. Downtime therefore costs a delay and is found within one reading session.
Further Notes
both stop at 0011, and everything the map specified is unbuilt. This spec's schema work is
poll_failures;series.site_completed_at bigint NOT NULL DEFAULT 0;owner_notices(kind, site, notified_at); and one extra count column on the pass log. Take the next free numbers after the earlierspecs land.
two outcomes that deliberately write nothing. No ADR for the notification path — it is reversible by
deleting the path, so it fails the hardness test.
Stallis already defined inCONTEXT.md(it landed while charting).failing,unverifiedandsite-completedare filter names over facts the glossary already covers, so no newnouns. The completion column gets no glossary entry: the glossary describes the system that exists.
research/completion-markers(the note isdocs/research/completion-markers.md) andresearch/htmx-series-list. Neither was merged, deliberately — charting should not land build artifacts.Lanes-page per-Site navigation link, the detail-page
unverifiedsentence, and two more options on thefilter select.
any log line that prints configuration), the owner gate on any new route, and the store's query
construction. The notification path must not import anything from the session or OAuth packages — that
independence is the reason a webhook was chosen over a bot. Say which invariant you preserved in the PR
and run the full backend test suite before calling it done.