Spec: per-Series poll failure state, completion hint, and outbound owner notification #137

Closed
opened 2026-08-21 15:47:54 +07:00 by sulthan · 0 comments
Owner

Spec derived from the wayfinder map #114, which is fully charted. This is spec 4 of 4; the map's
decisions were settled in #127, #129, #130 and #132.

Blocked by: #134, #135, #136 (specs 1, 2 and 3 of this series). Spec 1 creates the pass log, the filter vocabulary and
the pages; spec 2 fixes the rule that a human correction clears nothing and the cascade requirement
for any new per-Series table; spec 3 creates the Finished Series flag the completion hint sits beside.

Problem Statement

A Series can be broken for three months and every surface calls it healthy.

The check timestamp is stamped before the fetch, deliberately — otherwise a renamed or deleted
Series is retried on every pass for ever. So the column means attempted, never succeeded. A Series
whose page has quietly broken — a markup change, a wrong address, a permanent 404 — is stamped every
hour and reads as perfectly fine: it is not "never checked" (it has a timestamp), not "stale" (the
timestamp is minutes old), and not "unpollable" (it has an address). The only trace is log lines
nobody reads, and the per-Site outcome counts roll off in twelve hours without ever naming the Series.

And a number nothing can verify looks exactly like one a Poll confirmed. Where a Reader's report
raised a Latest Chapter on a Series the Poll cannot read, that number stands indefinitely with no
mark on it.

Separately: I have no idea which works have actually ended. I can mark a Series finished, but I
have to notice on my own — by remembering a series I have not seen a chapter of in a year. Every one
of the six Sites publishes its own completion status on the page the Poll already reads, and the
backend throws it away.

And the whole surface is pull-only. The one fault class that is invisible from every Reader
surface — a Lane that owed Polls, made none, and has nothing to say for it — is also the one there is
no reason to go looking for. The web UI dies visibly; the userscript degrades honestly to its cache
with a pending count in the panel and loses no Progress. Those faults find themselves within one
reading session. A stall does not: every Bookmark still opens, Progress still syncs, and Latest
Chapter is quietly wrong for as long as it lasts.

Solution

Three things a Series remembers about itself, and one message that leaves the building.

Per-Series failure state, as a row whose existence is the state. A row in a new table means "the
last Poll that learned anything about this Series failed", and a correct read deletes it. No success
sentinel, no counter, no history. Two new filters follow: Series failing for over twelve hours, and
the subset of those whose Latest Chapter came from a Reader and has never been confirmed.

Two outcomes deliberately say nothing. A Site refusing us and our own browser being unreachable
are not evidence about any particular Series. Writing a row would mark a whole library as failing when
one Site had a bad day; deleting one would claim recovery when nothing was read, resetting the age
that makes a three-month failure findable.

A completion hint, from the marker every Site already publishes. One column holding when the last
successful read saw the Site's own completed value. Decay is free: every successful read rewrites
the answer. It is a hint next to the Finish control and a filter to work through — never a shortcut,
never a writer of the Series' finished flag.

Four conditions push one Discord message each. The rule that selects them is the decision: silent
on every Reader surface, wrong, repairable only by the owner, and holding past twelve hours. A new
table remembers only that the owner was told, so a fault lasting a month sends one message rather
than thirty. No new timer — each Lane already wakes at least hourly and is its own clock.

User Stories

  1. As the owner, I want a Series that used to work and has stopped to be findable, so that a quiet
    breakage does not survive for months behind a healthy-looking timestamp.
  2. As the owner, I want the failure to name what went wrong, so that a page that is gone, a page my
    adapter cannot parse, and a database write that failed are three different words and three
    different actions.
  3. As the owner, I want to know how long a Series has been failing, so that I can tell this morning's
    breakage from one that has stood since spring.
  4. As the owner, I want the age to survive the failure word changing, so that a page that 404s after
    months of parse failures does not reset to "failing for a minute".
  5. As the owner, I want a Series that never once worked kept separate from one that worked for a year
    and broke on Tuesday, so that a never-implemented adapter and a regressed one are not the same
    list.
  6. As the owner, I want a correct read to clear the failure state with no extra bookkeeping, so that
    recovery needs no sweep and no reconciliation.
  7. As the owner, I want a repeated identical failure to write nothing, so that a broken Series costs
    the same as a healthy one every hour.
  8. As the owner, I want a Site refusing us to write no per-Series state at all, so that one Site's bad
    day does not mark its whole library as broken.
  9. As the owner, I want my browser sidecar being unreachable to write no per-Series state either, so
    that a machine I turned off does not look like two hundred broken Series.
  10. As the owner, I want those two outcomes not to clear state either, so that a single refusal
    cannot erase the age of a three-month failure and hide it from the list.
  11. As the owner, I want my own hand-typed correction to clear nothing, so that typing a number is
    never mistaken for the page becoming readable.
  12. As the owner, I want a forced check that succeeds to clear the failure exactly as any other
    successful read does, so that there is no special case to remember.
  13. As the owner, I want a Series whose Latest Chapter came from a Reader and that no Poll has
    confirmed in twelve hours to be its own list, so that I have a worklist of numbers nothing stands
    behind.
  14. As the owner, I want a Series failing for one hour to be in no list at all, so that one bad fetch
    is not a fault to correct.
  15. As the owner, I want a Series failing for one hour and one failing for three months to look
    identical on the page, so that I do not have to learn a visual rule the line cannot state in a
    word — the age is right there.
  16. As the owner, I want the per-Site outcome counts to stay unlinked, so that a large number never
    opens an empty list.
  17. As the owner, I want one plain navigation link per Site to its failing Series, so that I still get
    from "this Site looks unwell" to "these rows are broken" in one click.
  18. As the owner, I want a failing Series polled at the same pace as any other, so that a persistent
    failure does not quietly slow a Lane.
  19. As the owner, I want removing an orphan Series to remain a single statement, so that adding
    per-Series state does not turn a delete into a multi-step dance.
  20. As the owner, I want to be told which Series their own Site marks as completed, so that I do not
    have to remember which works ended.
  21. As the owner, I want that hint to read only each Site's completed value, so that a hiatus or a
    dropped translation is never presented as a finished work.
  22. As the owner, I want a failed extraction to mean "not completed" rather than "unknown", so that a
    broken selector degrades to a missing hint rather than to a suspicion.
  23. As the owner, I want the hint refreshed by every successful read, so that a Series that resumes
    stops being hinted without any expiry rule.
  24. As the owner, I want a refused or unreachable Poll to leave the hint's previous value alone, so
    that a challenge is not treated as evidence about the work.
  25. As the owner, I want a filter listing every Site-completed Series, so that I can work through them
    in one sitting.
  26. As the owner, I want one figure on the landing page linking to that list, so that I see the work
    without opening it.
  27. As the owner, I want the hint on the detail page beside the Finish control and not on the list
    row, so that a rare signal does not spend a field on every row.
  28. As the owner, I want the Finish control unchanged and still confirm-gated when a hint is present,
    so that the Site never decides the Lifecycle.
  29. As the owner, I want no way to dismiss a hint, so that there is not a third opinion on the row
    beside the Site's marker and my own decision.
  30. As the owner, I want a finished Series to carry no hint, so that a retired row is not presented as
    work.
  31. As the owner, I want all six Sites read, including the two where a broken selector is likeliest, so
    that an absent hint means "not completed" everywhere rather than "we never look here".
  32. As the owner, I want a message when a Lane owed Polls, made none, and has nothing to say for it, so
    that the one fault no surface shows reaches me anyway.
  33. As the owner, I want a Site with no browser route refusing for over twelve hours to reach me, so
    that a Site turning hostile overnight is repaired by a deploy rather than discovered weeks later.
  34. As the owner, I want a browser Site's refusal to send nothing, so that an ordinary challenge that
    did not clear is not reported as a fault.
  35. As the owner, I want a message when no browser Lane has reached the sidecar for over twelve hours,
    so that "asleep for an hour" and "gone for a week" stop looking the same to me.
  36. As the owner, I want one dead sidecar to send one message rather than one per browser Lane, so that
    a single fault is a single notification.
  37. As the owner, I want a message when more than half of one Site's Series have failed to yield a
    chapter for over twelve hours, so that a layout change is reported as a broken adapter rather than
    as hundreds of unrelated bad rows.
  38. As the owner, I want that threshold expressed as a share, so that a four-Series Site and a
    two-hundred-Series Site both fire.
  39. As the owner, I want a fault lasting a month to send one message, so that the channel stays worth
    reading.
  40. As the owner, I want a new message when a condition lifts and returns, so that a second episode is
    not swallowed by the first.
  41. As the owner, I want a pause I set myself to send nothing, so that my own act is not reported back
    to me.
  42. As the owner, I want a single failing Series to send nothing, so that one Site change does not send
    hundreds of messages.
  43. As the owner, I want a message to carry a link straight to the page that explains it, so that I do
    not go hunting from my phone.
  44. As the owner, I want the message to state the condition, its age and the repair in one sentence, so
    that I can decide whether to get up.
  45. As the owner, I want the machine-readable condition word in the message, so that an old message is
    searchable.
  46. As the owner, I want the timestamp rendered in my own timezone, so that I can place the fault
    without arithmetic.
  47. As the owner, I want no notification path at all when I have not configured one, so that a local
    stack runs without a webhook.
  48. As the owner, I want a failed send retried on the next pass with nothing queued, so that a message
    about a durable fact needs no machinery to survive.
  49. As the owner, I want the notification path to touch nothing about signing in, so that an outbound
    message cannot become an availability concern for the login flow.
  50. As the owner, I want the landing page's verdict line to read the same four conditions, so that a
    message never sends me to a page that looks healthy.

Implementation Decisions

Per-Series failure state: one table, the row's existence is the state

CREATE TABLE poll_failures (
  site          text   NOT NULL,
  series_id     text   NOT NULL,
  outcome       text   NOT NULL,
  failing_since bigint NOT NULL,  -- unix ms, first failure of the current run
  PRIMARY KEY (site, series_id),
  FOREIGN KEY (site, series_id) REFERENCES series (site, series_id) ON DELETE CASCADE
);

A row means "the last Poll that learned anything about this Series failed". No row means it succeeded,
or was never polled (a zero check stamp separates that), or every Poll so far ended in an outcome the
write rule ignores. So there is no '' success sentinel, no CHECK keeping two columns consistent,
and failing_since is never zero.

The cascade reaches this table only — never bookmarks, whose refusal is the orphan guard, so
removing an orphan stays a single statement.

Recorded because it will be re-litigated: 3NF does not decide table-versus-columns here. The key is
(site, series_id) either way, the relation is 1:1, and the one rule between the attributes (a word
implies an age) yields no value, so it is a constraint rather than a transitive dependency — both
designs are in BCNF. A 1:1 table with the same key is vertical partitioning, not normalisation, and a
lookup table for the outcome vocabulary would be a domain constraint a Go constant already enforces.
The table won on its own merits: series stays narrow, the sentinel disappears, and the guarded write
gets simpler.

Rejected: three columns on series — a last outcome, a first-failure timestamp, and a
consecutive-failure counter. The counter goes because it is a proxy for duration measured in units that
differ per Site: each Lane carries its own rest, and a paused Lane stops the count without stopping the
failure. failing since <age> is the fact the owner reads.

The write

-- failure
INSERT INTO poll_failures (site, series_id, outcome, failing_since) VALUES ($1,$2,$3,$4)
ON CONFLICT (site, series_id) DO UPDATE SET outcome = excluded.outcome
  WHERE poll_failures.outcome <> excluded.outcome;
-- success
DELETE FROM poll_failures WHERE site = $1 AND series_id = $2;

failing_since survives a change of word because the conflict update never touches it — the age is the
age of the run of failures, not of the current word. The same failure repeating writes nothing; a
healthy Series polling correctly writes nothing.
The cost lands on transitions, not on Polls. This is
also why there is no last_outcome_at column: it would sit within seconds of the check stamp, which is
already stored.

Which outcomes write, and which are silent

outcome writes? why
no_chapter yes 200, real HTML, no chapter found — the Series' page or our adapter
unfetchable yes the stored address failed the fetch gate — the Series' own address
not_found yes 4xx other than 403 — the page is gone
errors yes transport failure, 5xx, or a failed check stamp
refused no statement at all a challenge held: the Site's mood, identical for every Series it hosts
unreachable no statement at all the browser was interrupted: our own sidecar, and nothing was read

Neither insert nor delete for the last two, and the distinction is load-bearing in both directions.
Writing a row would mark a whole library as failing when one Site refused for a day. Deleting one would
claim recovery when nothing was read — resetting the age for a Series broken three months, so a single
refusal from its Site would erase exactly the fact that makes it findable. A refused or unreachable
Poll therefore issues one fewer query, and the Series keeps the last thing it truly learned about
itself. Both facts are already recorded where they belong: per Site in the durable Lane refusal, and in
the poller's in-memory browser-down timestamp.

errors splits: not_found becomes a sixth outcome word for a 4xx status other than 403. Today one
word covers a 404 (the page is gone — correct the address), a 503 (the Site is busy — wait) and our own
write failing (the database is unwell): three different owner actions behind the one word a page prints.
One branch in the read path, and the pass log gains a sixth count column.

Written by nothing else, ever. A Forced Poll request clears nothing, because a request is not
evidence; a forced Poll that reads the page clears the row exactly as any other correct read does, which
needs no special case in the code. A hand correction clears nothing, as #135 requires.

Two new filters

  • failing — the joined row exists AND s.latest_chapter_num IS NOT NULL AND f.failing_since older
    than ownerWindow (12h; reuse the constant, add no new figure). The latest_chapter_num IS NOT NULL
    test is what keeps it disjoint from no-chapter (never once succeeded) rather than overlapping it
    for no gain. Deliberately no finished_at = 0 guard: a Series already failing when the owner
    finished it keeps its evidence, and guarding it would erase evidence rather than prevent a false
    figure. (Contrast the four predicates #136 does guard: those are computed from clocks that keep
    ticking, so they lie about a finished row; this one is computed from stored outcomes, which simply
    stop arriving.)
  • unverified — s.latest_raised_by IS NOT NULL plus the failing test. The Latest Chapter came
    from a Reader's Sighting and no Poll has confirmed it for over twelve hours. Spec 2's correction
    control is the action it leads to, and the detail page adds one sentence beside that control for this
    case.

Rejected: a stored column for the unverified pair. It would hold the conjunction of a fact on
series and a row in another table — derived data with five writers (a Sighting raising the Series, a
Poll succeeding, a Poll failing, an owner correction, and clearing the attribution) for a value the join
computes free. Unlike the table-versus-columns question above, this one is a genuine 3NF violation, and
the admin read model acquires the join for other reasons anyway.

AdminSeries gains LEFT JOIN poll_failures f USING (site, series_id), and SeriesRow gains the
outcome word and the failing-since timestamp. Nothing is added to series — the row's existence in
the new table is the failure state, so there is no success sentinel to project.

The per-Site outcome counts stay unlinked, permanently

Earlier charting promised that once a failing filter existed, all the per-Site outcome counts would
link to it. That promise is withdrawn, on the rule that a figure links to what it counts — the same
rule that already refused to link the no_chapter count to the no-chapter filter.

Two failures of the same test. refused and unreachable write no per-Series state at all, so two of
the six counts would open an empty list beside a large number. And the other four count attempts inside
a twelve-hour window
while the filter lists Series failing now for over twelve hours: a Series that
broke at 04:00 and healed at 05:00 is counted and not listed, while one broken for three months on a
paused Lane is listed and counted zero. Neither set contains the other, in either direction.

Instead each Site row on the Lanes page carries one separate link to
/admin/series?site=<site>&filter=failing, as navigation rather than attached to a number. Rejected: an
?outcome= parameter to make the links exact — the surface is one filter at a time with only ?site=
and ?kind= stacking, and a second axis would not fix the window mismatch anyway.

Completion hint: storage

The false-positive containment this was named for is not built, because the research removed the
threat.
Measured 2026-08-19 against live pages (note: docs/research/completion-markers.md on branch
research/completion-markers): all six Sites publish an explicit machine-readable completion marker, and
a Poll that triggers only on each Site's completed value has no false-positive path on any Site. So
there is no N-consecutive-observation gate, no expiry, no confirmation counter and no suspicion damping.
The residual error is a false negative — a finished Series that gets no hint — and we accept it.

series.site_completed_at bigint NOT NULL DEFAULT 0, epoch-ms of the read that saw the Site's completed
value; zero means the last successful read did not see it. Decay is therefore free: every successful
read rewrites the answer, so a stale hint cannot survive one Poll. The timestamp is the age the hint
renders with, not a history.

  • Rejected: a text column holding a normalised status word. The six vocabularies are not comparable:
    demonicscans and novelfull are binary and cannot express hiatus at all; asurascans' dropped and
    kagane's upload status are scanlation-editorial rather than completion state; and kagane is the only
    Site that separates the work's status from the translation's (proven divergent on one series). A
    shared word would be a different fact per Site, and nothing here acts on any value but completed.
  • Rejected: storing nothing and re-deriving on read. The marker exists only in a fetched page, and
    the admin surface never fetches — every intervention goes through the database.

Completion hint: which value counts

One predicate per Site, the completed value only, no mapping table:

Site Reads
asurascans status is completed (its dropped is editorial: one series stops at chapter 14 while the work continues)
demonicscans Completed
comix status is finished (on_hiatus, discontinued, not_yet_released are distinct)
kagane publication status is Completed only — upload status is the release's state, and the two diverge
novelfull Completed
lightnovelworld creative-work status is the completed action status
  • A failed extraction is not-completed, never unknown-as-suspicion.
  • A refused or unreachable Poll makes no statement at all — the column keeps its previous value
    rather than zeroing. Inherited free from the failure-write rule above: a challenge is not evidence
    about the Series.
  • All six adapters extract it. Four ride a payload the adapter already parses (asurascans' island
    props, comix's initial-data detail entry, kagane's API body, lightnovelworld's head JSON-LD).
    demonicscans (an info-block list-item pair) and novelfull (a status link) each need one new selector
    against markup nobody controls. Paid anyway, on two grounds: a broken selector degrades to a missing
    hint, which is the error class we already accept; and on exactly those two Sites Completed is the only
    signal the Site can ever emit, so omitting them makes an absent hint unreadable — the owner could not
    tell "not finished" from "we never look".

Completion hint: the write

Its own statement in the pass, after the read. It cannot ride the check stamp (written before the
fetch) and it cannot ride the chapter setter (skipped on the unchanged-number path). It cannot share the
failure write either
: that one targets the failure table and fires on a failure, while this is a fact
learned from a successful read of series. The two share only their position in the pass.

site_completed_at joins the due query's projection, so the statement fires only when the answer would
change
. The ordinary Poll of an ongoing Series writes nothing, exactly as a healthy Poll writes nothing
to the failure table.

Completion hint: presentation

  • One new filter name, site-completed, ordered with the informational tail beside finished, and
    carrying AND finished_at = 0 — the same guard #136's hygiene predicates gained.
  • One figure in the landing stats block linking to it, printing an unlinked digit at zero.
  • One line beside the Finish control on the detail page. No field on the list row — a row has two
    lines, and this fact is true of few Series, so a per-row field spends space on every row for a rare
    signal. The filter is the work list; the figure is how the owner sees the work without opening it.
  • The Finish control does not change. Still one control, still confirm-gated. The hint is text next to
    it: it may not pre-fill, add a second button, or remove the confirm step. A hint that shortens the
    path is the Site deciding the Lifecycle.
  • No dismissal column. The owner's answers are Finish or ignore. A stored dismissal would be a third
    opinion on the row beside the Site's marker and the owner's own flag, and would need its own rule for
    the day the Site's marker changes again. An ignored Series stays in a fifty-row filter; if that noise
    turns out to be real, the column can be added then.
  • After a finish: the filter's guard excludes it and the detail page hides the hint. A finished Series
    leaves the Poll query, so its value freezes at whatever the last Poll saw — old and unrefreshable,
    therefore not work. Un-finishing puts it back in the query and the next Poll writes a fresh answer.
  • Two consequences worth stating rather than rediscovering: the hint is only as fresh as the last Poll, so
    the two browser-gated Sites' hints age while the sidecar is asleep or unreachable — the same degradation
    their Covers already have, and no new notification.

Outbound notification: the rule, then the conditions

The dashboard is not pull-only. Push exists for one class of fault: silent on every Reader surface,
wrong, repairable only by the owner, and persisting past twelve hours.
That rule is the decision — the
four conditions below are what it currently selects, and a fifth candidate must pass all four tests.

A backend that stopped is not the case worth pushing, and the reason came from the Reader surfaces
rather than from the plan. The web UI dies visibly. The userscript degrades quietly but honestly: reads
come from the cache, a write that meets a dropped connection goes to the retry queue, the queue only
speaks up for a 400, a 401 or a full queue, and the sole visible marker is the pending count in the panel.
No Progress is lost — the only loss path is eviction at the queue cap, which toasts. So downtime costs a
delay and is found within one reading session, by every Reader. A stalled Lane is the opposite.

The four conditions, each at ownerWindow (12h; no new constant):

  1. stall — due > 0 AND checked = 0 AND skip = '' AND refused = 0 on the pass log.
  2. no-browser-route — a Site with no browser route refusing for more than 12h. The read path
    maps both a 403 with the mitigation header and an interstitial body to a held challenge on the
    plain-TLS path too, so this is refused plus the durable Lane refusal. This is the comix event of
    2026-08-12, and the repair is a deploy. A browser Site refusing sends nothing — that is an ordinary
    challenge that did not clear, and a Site setting its own pace is not a fault.
  3. sidecar-down — no browser Lane has reached the sidecar for more than 12h. This is not a stall
    and cannot be found by the stall test
    : the pass returns early both when a sibling Lane lost Chrome
    and when no fetcher exists, and both are recorded as skip values. The quiet degradation is deliberate,
    but "asleep for an hour" and "gone for a week" are indistinguishable to the backend, and the second
    freezes two Sites' marks.
  4. adapter-broken — more than half of one Site's Series hold a no_chapter failure row older
    than 12h. A layout change produces exactly that outcome; one such Series is a bad row, most of a Site
    is a broken adapter. A share, not a count, so a four-Series Site and a two-hundred-Series Site both
    fire. This is the one fault on the whole map that no other signal names.

Never pushed: a refusal by a browser Site (not a fault), a pause (the owner's own act), a single
failing Series (the filter is the tool, and one Site change makes hundreds of rows), a forced Poll ageing
(the same fault from the other end, and visible on the page the button lives on), a gated cover host (a
missing Cover is visible in the UI).

Two collisions found in the code, and their fixes

  • The stall test catches a refusing Site. The twice-refused break inside the pass loop is not an
    early return — the pass falls through to its normal return, writing due > 0, checked = 0, skip = ''.
    A newly-gated Site would have sent both stall and no-browser-route. So the stall test gains
    AND refused = 0
    , amending the definition #134 lands. Rejected alternative: a tenth skip value,
    because the enum holds one value per return path and this is a break inside one — a tenth value would
    describe a pass that did run.
  • One dead sidecar is three Lanes. Each browser Lane notices the loss on its own pass. So the
    suppression row's site holds the empty string for sidecar-down: the first Lane to notice writes
    the row, the others find it present and stay quiet. The row's own presence is the lock — which is
    exactly why the state is its own table and not columns on the per-Site Lane row, where an empty Site name
    has no meaning. Rejected: naming one Lane the reporter, since that Lane's pass may be the skipped one.

Where it fires from, and suppression

Beside the pass-log insert, in the deferred call every return path passes through. Two of the four
conditions occur on early returns, so the success path can never see them.

No new timer. The Lane loop already does one pass, then waits the pace the pass reported, which is
never more than the default rest — each Lane is its own clock at least hourly. (The premise that the
backend has no clock was wrong; it has six.)

New table owner_notices(kind, site, notified_at) — one row per episode, cleared when the condition
no longer holds, so a fault lasting a month sends one message rather than thirty at ~120 passes a day.
Every threshold is derived, not stored: a refusal's and a sidecar absence's age from the pass log, a
Series failure's age from its failing-since. The only new durable fact is that the owner was told,
which is why the table carries nothing else. Rows are bounded by four kinds times six Sites, so there is
no retention rule.

Transport

  • DISCORD_WEBHOOK_URL, and unset means the whole path is off, logged once at start, exactly as the
    browser URL behaves — a local stack must not need a webhook to run.
  • A webhook is a plain POST and does not touch the login path: no bot, no gateway, no scope, not the
    OAuth application. The existing Discord coupling gains nothing new.
  • The address is a secret in the class of the token key and is never logged. Rate limits are
    irrelevant at this volume (roughly 30 messages a minute allowed per webhook; this sends a few a day).
  • A failed POST is logged and the notified timestamp is left unset, so the next pass retries while the
    condition holds. No queue, no backoff: the condition is durable, so the retry is free. (The userscript
    needs a queue because a Reader's Progress is otherwise lost; a message about a durable fact is not.)
  • Rejected transports, measured 2026-08-21: a Telegram bot — free and unlimited at this volume, but a
    second identity to hold when the owner's Discord id already names the recipient; the public ntfy.sh
    server — 250 messages a day and 60 burst, but every topic is public by design (no authentication;
    anyone knowing the topic can read and publish) with limited retention; self-hosted ntfy or Gotify — a
    container and a volume, which is the infrastructure this decision refuses.

Presentation

A Discord embed, colour 13589581 — the dark branch of the danger token. One colour for all four,
because all four are faults; ember stays forbidden, and no second colour is introduced because no
non-fault condition pushes.

  • title = the subject (the Site, or browser sidecar) carrying the deep link built from
    PUBLIC_BASE_URL, so no body line is spent on an address.
  • description = one sentence: condition, age, repair.
  • timestamp = the pass time, which Discord renders in the reader's own zone — the one thing an embed
    does better than plain text.
  • footer = the machine word (stall, no-browser-route, sidecar-down, adapter-broken), making an
    old message searchable.
  • No field grid, no thumbnail, no author block — the figures belong on the page the title links to.
    Only the colour of Cinder survives into Discord; the typography cannot.

The page and the push stay in step

The landing verdict line must read the same four conditions. A message naming a fault the page cannot
explain sends the owner to a page that looks healthy. This amends #134's verdict line, which reads Lane
attention generally; narrow it to these four.

Testing Decisions

What makes a good test here: it asserts a durable outcome — a row present or absent after a pass, a
filter's membership, a payload's shape — and it fails on a plausible bug. The two silent outcomes are the
highest-value tests in this spec, because their bug is invisible: writing where nothing should be written
produces a plausible-looking row.

The primary seam stays the router for everything the owner sees, and the poller's existing
round-at-a-time entry point for everything the Lane writes. One new seam, and only one:
Poller.Notifier func(...) error with nil meaning off — the same field-injection pattern the fetchers
already use, so all four conditions are testable with no wire and no Discord.

Modules and coverage:

  • Poller (prior art: the round-at-a-time pass tests with a fake fetcher and an injected clock): each
    writing outcome inserts a row with the right word; a repeated identical failure writes nothing; a changed
    word updates the word and keeps the failing-since (this is the test that protects the age); a
    successful read deletes the row; a refused Poll neither inserts nor deletes, and an unreachable
    Poll neither inserts nor deletes
    — asserted against a pre-seeded row so both directions are covered;
    a forced Poll that succeeds deletes the row and a forced request alone does not; the completion write
    fires only on a transition and not on an ongoing Series' ordinary Poll; a refused Poll leaves the
    completion column unchanged; the six-word outcome taxonomy maps a 404 to not_found and a 503 to
    errors.
  • Adapters (prior art: the existing per-Site parsing tests over stored page fixtures, plus the
    network-gated live smoke tests): each of the six Sites' completed predicate against a completed fixture
    and an ongoing one; kagane's divergent case asserts that the upload status is ignored; a fixture with
    the selector removed yields not-completed rather than an error. The live smoke tests stay
    network-gated and are not part of the default run.
  • Store: the failing and unverified predicates, including disjointness from no-chapter (a Series
    that never succeeded must be in no-chapter and never in failing); a Series failing for under twelve
    hours is in neither filter; a finished Series still appears in failing; site-completed excludes a
    finished Series; the cascade removes failure rows when a Series is deleted and the orphan delete stays a
    single statement; the admin join projects the outcome word and the age.
  • Router / web: the failure marker renders on the fact line and identically for a one-hour and a
    three-month failure; the two new filters and the completion figure and filter; the outcome counts render
    unlinked while the per-Site navigation link is present; the unverified sentence appears beside the
    correction control; the hint is hidden on a finished Series and the Finish control is unchanged when a
    hint is present; the verdict line reflects each of the four conditions.
  • Notifier: each of the four conditions fires exactly once while it holds and again after it lifts and
    returns; a stall on a Site that also refused fires only no-browser-route (the collision test — this
    is the one a naive implementation gets wrong); one dead sidecar across three browser Lanes sends exactly
    one message; a pause sends nothing; a browser Site's refusal sends nothing; a single failing Series sends
    nothing; the adapter-broken share fires on a four-Series Site and does not fire at exactly half; a failed
    send leaves the notified timestamp unset and the next pass retries. One separate test covers the embed
    payload
    — colour, title link, footer word, timestamp, and the absence of a field grid.

Out of Scope

  • A consecutive-failure counter, a last-outcome timestamp, or any per-Series failure history.
  • Slowing a Lane for a persistent failure. The due query does not join the failure table; a failing
    Series is polled at the same pace as any other. Pacing is not a dashboard question.
  • A normalised status word for the completion hint, or any use of a hiatus, dropped or
    translation-status value.
  • Surfacing a stalled-but-not-completed serialisation. On the two binary Sites a stalled series reads
    Ongoing for ever; that is the accepted false negative.
  • A machine write of the Series' finished flag, a pre-filled Finish control, a second Finish button, or
    removing its confirm step.
  • A dismissal column for the hint.
  • Linking the per-Site outcome counts to any filter. Withdrawn permanently, with reasons above.
  • A stored column for the unverified pair.
  • Notification for anything outside the four conditions, including a paused Lane, a browser Site's
    refusal, one failing Series, an ageing forced Poll, and a gated cover host.
  • A queue, a backoff, or any delivery machinery beyond retry-on-next-pass.
  • Telegram, ntfy, Gotify, email, or a second transport of any kind.
  • A monitor outside the backend — an external prober against the health endpoint, or a hosted uptime
    service. A push cannot report its own process's death, so the honest coverage lives outside; but it is
    another service to run, watch and update, and it watches a fault that is already visible from the
    Reader surfaces
    : the web UI dies visibly and the userscript degrades to its cache with a pending count,
    losing no Progress. Downtime therefore costs a delay and is found within one reading session.

Further Notes

  • Migration numbers are claimed in landing order. Measured 2026-08-21: the tree's migrations and ADRs
    both stop at 0011, and everything the map specified is unbuilt. This spec's schema work is
    poll_failures; series.site_completed_at bigint NOT NULL DEFAULT 0; owner_notices(kind, site, notified_at); and one extra count column on the pass log. Take the next free numbers after the earlier
    specs land.
  • One ADR is worth writing with the implementation: a table whose row is the failure state, plus the
    two outcomes that deliberately write nothing. No ADR for the notification path — it is reversible by
    deleting the path, so it fails the hardness test.
  • Glossary: Stall is already defined in CONTEXT.md (it landed while charting). failing,
    unverified and site-completed are filter names over facts the glossary already covers, so no new
    nouns. The completion column gets no glossary entry: the glossary describes the system that exists.
  • The three research branches this consumes should be deleted once implemented:
    research/completion-markers (the note is docs/research/completion-markers.md) and
    research/htmx-series-list. Neither was merged, deliberately — charting should not land build artifacts.
  • The prototype amendments this spec carries into the design project: the fact-line failure marker, the
    Lanes-page per-Site navigation link, the detail-page unverified sentence, and two more options on the
    filter select.
  • Security-critical surfaces touched: a new secret in the environment (never logged, and it must not reach
    any log line that prints configuration), the owner gate on any new route, and the store's query
    construction. The notification path must not import anything from the session or OAuth packages — that
    independence is the reason a webhook was chosen over a bot. Say which invariant you preserved in the PR
    and run the full backend test suite before calling it done.
Spec derived from the wayfinder map #114, which is fully charted. This is spec 4 of 4; the map's decisions were settled in #127, #129, #130 and #132. Blocked by: #134, #135, #136 (specs 1, 2 and 3 of this series). Spec 1 creates the pass log, the filter vocabulary and the pages; spec 2 fixes the rule that a human correction clears nothing and the cascade requirement for any new per-Series table; spec 3 creates the Finished Series flag the completion hint sits beside. ## Problem Statement **A Series can be broken for three months and every surface calls it healthy.** The check timestamp is stamped *before* the fetch, deliberately — otherwise a renamed or deleted Series is retried on every pass for ever. So the column means *attempted*, never *succeeded*. A Series whose page has quietly broken — a markup change, a wrong address, a permanent 404 — is stamped every hour and reads as perfectly fine: it is not "never checked" (it has a timestamp), not "stale" (the timestamp is minutes old), and not "unpollable" (it has an address). The only trace is log lines nobody reads, and the per-Site outcome counts roll off in twelve hours without ever naming the Series. **And a number nothing can verify looks exactly like one a Poll confirmed.** Where a Reader's report raised a Latest Chapter on a Series the Poll cannot read, that number stands indefinitely with no mark on it. **Separately: I have no idea which works have actually ended.** I can mark a Series finished, but I have to notice on my own — by remembering a series I have not seen a chapter of in a year. Every one of the six Sites publishes its own completion status on the page the Poll already reads, and the backend throws it away. **And the whole surface is pull-only.** The one fault class that is invisible from every Reader surface — a Lane that owed Polls, made none, and has nothing to say for it — is also the one there is no reason to go looking for. The web UI dies visibly; the userscript degrades honestly to its cache with a pending count in the panel and loses no Progress. Those faults find themselves within one reading session. A stall does not: every Bookmark still opens, Progress still syncs, and Latest Chapter is quietly wrong for as long as it lasts. ## Solution Three things a Series remembers about itself, and one message that leaves the building. **Per-Series failure state, as a row whose existence is the state.** A row in a new table means "the last Poll that learned anything about this Series failed", and a correct read *deletes* it. No success sentinel, no counter, no history. Two new filters follow: Series failing for over twelve hours, and the subset of those whose Latest Chapter came from a Reader and has never been confirmed. **Two outcomes deliberately say nothing.** A Site refusing us and our own browser being unreachable are not evidence about any particular Series. Writing a row would mark a whole library as failing when one Site had a bad day; deleting one would claim recovery when nothing was read, resetting the age that makes a three-month failure findable. **A completion hint, from the marker every Site already publishes.** One column holding when the last successful read saw the Site's own *completed* value. Decay is free: every successful read rewrites the answer. It is a hint next to the Finish control and a filter to work through — never a shortcut, never a writer of the Series' finished flag. **Four conditions push one Discord message each.** The rule that selects them is the decision: silent on every Reader surface, wrong, repairable only by the owner, and holding past twelve hours. A new table remembers only *that the owner was told*, so a fault lasting a month sends one message rather than thirty. No new timer — each Lane already wakes at least hourly and is its own clock. ## User Stories 1. As the owner, I want a Series that used to work and has stopped to be findable, so that a quiet breakage does not survive for months behind a healthy-looking timestamp. 2. As the owner, I want the failure to name what went wrong, so that a page that is gone, a page my adapter cannot parse, and a database write that failed are three different words and three different actions. 3. As the owner, I want to know how long a Series has been failing, so that I can tell this morning's breakage from one that has stood since spring. 4. As the owner, I want the age to survive the failure *word* changing, so that a page that 404s after months of parse failures does not reset to "failing for a minute". 5. As the owner, I want a Series that never once worked kept separate from one that worked for a year and broke on Tuesday, so that a never-implemented adapter and a regressed one are not the same list. 6. As the owner, I want a correct read to clear the failure state with no extra bookkeeping, so that recovery needs no sweep and no reconciliation. 7. As the owner, I want a repeated identical failure to write nothing, so that a broken Series costs the same as a healthy one every hour. 8. As the owner, I want a Site refusing us to write no per-Series state at all, so that one Site's bad day does not mark its whole library as broken. 9. As the owner, I want my browser sidecar being unreachable to write no per-Series state either, so that a machine I turned off does not look like two hundred broken Series. 10. As the owner, I want those two outcomes not to *clear* state either, so that a single refusal cannot erase the age of a three-month failure and hide it from the list. 11. As the owner, I want my own hand-typed correction to clear nothing, so that typing a number is never mistaken for the page becoming readable. 12. As the owner, I want a forced check that succeeds to clear the failure exactly as any other successful read does, so that there is no special case to remember. 13. As the owner, I want a Series whose Latest Chapter came from a Reader and that no Poll has confirmed in twelve hours to be its own list, so that I have a worklist of numbers nothing stands behind. 14. As the owner, I want a Series failing for one hour to be in no list at all, so that one bad fetch is not a fault to correct. 15. As the owner, I want a Series failing for one hour and one failing for three months to look identical on the page, so that I do not have to learn a visual rule the line cannot state in a word — the age is right there. 16. As the owner, I want the per-Site outcome counts to stay unlinked, so that a large number never opens an empty list. 17. As the owner, I want one plain navigation link per Site to its failing Series, so that I still get from "this Site looks unwell" to "these rows are broken" in one click. 18. As the owner, I want a failing Series polled at the same pace as any other, so that a persistent failure does not quietly slow a Lane. 19. As the owner, I want removing an orphan Series to remain a single statement, so that adding per-Series state does not turn a delete into a multi-step dance. 20. As the owner, I want to be told which Series their own Site marks as completed, so that I do not have to remember which works ended. 21. As the owner, I want that hint to read only each Site's *completed* value, so that a hiatus or a dropped translation is never presented as a finished work. 22. As the owner, I want a failed extraction to mean "not completed" rather than "unknown", so that a broken selector degrades to a missing hint rather than to a suspicion. 23. As the owner, I want the hint refreshed by every successful read, so that a Series that resumes stops being hinted without any expiry rule. 24. As the owner, I want a refused or unreachable Poll to leave the hint's previous value alone, so that a challenge is not treated as evidence about the work. 25. As the owner, I want a filter listing every Site-completed Series, so that I can work through them in one sitting. 26. As the owner, I want one figure on the landing page linking to that list, so that I see the work without opening it. 27. As the owner, I want the hint on the detail page beside the Finish control and *not* on the list row, so that a rare signal does not spend a field on every row. 28. As the owner, I want the Finish control unchanged and still confirm-gated when a hint is present, so that the Site never decides the Lifecycle. 29. As the owner, I want no way to dismiss a hint, so that there is not a third opinion on the row beside the Site's marker and my own decision. 30. As the owner, I want a finished Series to carry no hint, so that a retired row is not presented as work. 31. As the owner, I want all six Sites read, including the two where a broken selector is likeliest, so that an absent hint means "not completed" everywhere rather than "we never look here". 32. As the owner, I want a message when a Lane owed Polls, made none, and has nothing to say for it, so that the one fault no surface shows reaches me anyway. 33. As the owner, I want a Site with no browser route refusing for over twelve hours to reach me, so that a Site turning hostile overnight is repaired by a deploy rather than discovered weeks later. 34. As the owner, I want a browser Site's refusal to send nothing, so that an ordinary challenge that did not clear is not reported as a fault. 35. As the owner, I want a message when no browser Lane has reached the sidecar for over twelve hours, so that "asleep for an hour" and "gone for a week" stop looking the same to me. 36. As the owner, I want one dead sidecar to send one message rather than one per browser Lane, so that a single fault is a single notification. 37. As the owner, I want a message when more than half of one Site's Series have failed to yield a chapter for over twelve hours, so that a layout change is reported as a broken adapter rather than as hundreds of unrelated bad rows. 38. As the owner, I want that threshold expressed as a share, so that a four-Series Site and a two-hundred-Series Site both fire. 39. As the owner, I want a fault lasting a month to send one message, so that the channel stays worth reading. 40. As the owner, I want a new message when a condition lifts and returns, so that a second episode is not swallowed by the first. 41. As the owner, I want a pause I set myself to send nothing, so that my own act is not reported back to me. 42. As the owner, I want a single failing Series to send nothing, so that one Site change does not send hundreds of messages. 43. As the owner, I want a message to carry a link straight to the page that explains it, so that I do not go hunting from my phone. 44. As the owner, I want the message to state the condition, its age and the repair in one sentence, so that I can decide whether to get up. 45. As the owner, I want the machine-readable condition word in the message, so that an old message is searchable. 46. As the owner, I want the timestamp rendered in my own timezone, so that I can place the fault without arithmetic. 47. As the owner, I want no notification path at all when I have not configured one, so that a local stack runs without a webhook. 48. As the owner, I want a failed send retried on the next pass with nothing queued, so that a message about a durable fact needs no machinery to survive. 49. As the owner, I want the notification path to touch nothing about signing in, so that an outbound message cannot become an availability concern for the login flow. 50. As the owner, I want the landing page's verdict line to read the same four conditions, so that a message never sends me to a page that looks healthy. ## Implementation Decisions ### Per-Series failure state: one table, the row's existence is the state ```sql CREATE TABLE poll_failures ( site text NOT NULL, series_id text NOT NULL, outcome text NOT NULL, failing_since bigint NOT NULL, -- unix ms, first failure of the current run PRIMARY KEY (site, series_id), FOREIGN KEY (site, series_id) REFERENCES series (site, series_id) ON DELETE CASCADE ); ``` A row means "the last Poll that learned anything about this Series failed". No row means it succeeded, or was never polled (a zero check stamp separates that), or every Poll so far ended in an outcome the write rule ignores. **So there is no `''` success sentinel, no `CHECK` keeping two columns consistent, and `failing_since` is never zero.** **The cascade reaches this table only — never `bookmarks`, whose refusal *is* the orphan guard**, so removing an orphan stays a single statement. Recorded because it will be re-litigated: **3NF does not decide table-versus-columns here.** The key is `(site, series_id)` either way, the relation is 1:1, and the one rule between the attributes (a word implies an age) yields no value, so it is a constraint rather than a transitive dependency — both designs are in BCNF. A 1:1 table with the same key is vertical partitioning, not normalisation, and a lookup table for the outcome vocabulary would be a domain constraint a Go constant already enforces. The table won on its own merits: `series` stays narrow, the sentinel disappears, and the guarded write gets *simpler*. **Rejected: three columns on `series`** — a last outcome, a first-failure timestamp, and a consecutive-failure counter. The counter goes because it is a proxy for duration measured in units that differ per Site: each Lane carries its own rest, and a paused Lane stops the count without stopping the failure. `failing since <age>` is the fact the owner reads. ### The write ```sql -- failure INSERT INTO poll_failures (site, series_id, outcome, failing_since) VALUES ($1,$2,$3,$4) ON CONFLICT (site, series_id) DO UPDATE SET outcome = excluded.outcome WHERE poll_failures.outcome <> excluded.outcome; -- success DELETE FROM poll_failures WHERE site = $1 AND series_id = $2; ``` `failing_since` survives a change of word because the conflict update never touches it — the age is the age of the *run* of failures, not of the current word. **The same failure repeating writes nothing; a healthy Series polling correctly writes nothing.** The cost lands on transitions, not on Polls. This is also why there is no `last_outcome_at` column: it would sit within seconds of the check stamp, which is already stored. ### Which outcomes write, and which are silent | outcome | writes? | why | |---|---|---| | `no_chapter` | yes | 200, real HTML, no chapter found — the Series' page or our adapter | | `unfetchable` | yes | the stored address failed the fetch gate — the Series' own address | | `not_found` | yes | 4xx other than 403 — the page is gone | | `errors` | yes | transport failure, 5xx, or a failed check stamp | | `refused` | **no statement at all** | a challenge held: the *Site's* mood, identical for every Series it hosts | | `unreachable` | **no statement at all** | the browser was interrupted: *our own* sidecar, and nothing was read | Neither insert nor delete for the last two, **and the distinction is load-bearing in both directions.** Writing a row would mark a whole library as failing when one Site refused for a day. Deleting one would claim recovery when nothing was read — resetting the age for a Series broken three months, so a single refusal from its Site would erase exactly the fact that makes it findable. A refused or unreachable Poll therefore issues **one fewer query**, and the Series keeps the last thing it truly learned about itself. Both facts are already recorded where they belong: per Site in the durable Lane refusal, and in the poller's in-memory browser-down timestamp. **`errors` splits: `not_found` becomes a sixth outcome word** for a 4xx status other than 403. Today one word covers a 404 (the page is gone — correct the address), a 503 (the Site is busy — wait) and our own write failing (the database is unwell): three different owner actions behind the one word a page prints. One branch in the read path, and the pass log gains a sixth count column. **Written by nothing else, ever.** A Forced Poll *request* clears nothing, because a request is not evidence; a forced Poll that reads the page clears the row exactly as any other correct read does, which needs no special case in the code. A hand correction clears nothing, as #135 requires. ### Two new filters - **`failing`** — the joined row exists `AND s.latest_chapter_num IS NOT NULL AND f.failing_since` older than `ownerWindow` (12h; reuse the constant, add no new figure). The `latest_chapter_num IS NOT NULL` test is what keeps it **disjoint** from `no-chapter` (never once succeeded) rather than overlapping it for no gain. Deliberately **no `finished_at = 0` guard**: a Series already failing when the owner finished it keeps its evidence, and guarding it would erase evidence rather than prevent a false figure. (Contrast the four predicates #136 does guard: those are computed from clocks that keep ticking, so they lie about a finished row; this one is computed from stored outcomes, which simply stop arriving.) - **`unverified`** — `s.latest_raised_by IS NOT NULL` plus the `failing` test. The Latest Chapter came from a Reader's Sighting and no Poll has confirmed it for over twelve hours. Spec 2's correction control is the action it leads to, and the detail page adds one sentence beside that control for this case. **Rejected: a stored column for the `unverified` pair.** It would hold the conjunction of a fact on `series` and a row in another table — derived data with five writers (a Sighting raising the Series, a Poll succeeding, a Poll failing, an owner correction, and clearing the attribution) for a value the join computes free. Unlike the table-versus-columns question above, *this* one is a genuine 3NF violation, and the admin read model acquires the join for other reasons anyway. `AdminSeries` gains `LEFT JOIN poll_failures f USING (site, series_id)`, and `SeriesRow` gains the outcome word and the failing-since timestamp. **Nothing is added to `series`** — the row's existence in the new table is the failure state, so there is no success sentinel to project. ### The per-Site outcome counts stay unlinked, permanently Earlier charting promised that once a `failing` filter existed, all the per-Site outcome counts would link to it. **That promise is withdrawn**, on the rule that a figure links to what it counts — the same rule that already refused to link the `no_chapter` count to the `no-chapter` filter. Two failures of the same test. `refused` and `unreachable` write no per-Series state at all, so two of the six counts would open an empty list beside a large number. And the other four count *attempts inside a twelve-hour window* while the filter lists *Series failing now for over twelve hours*: a Series that broke at 04:00 and healed at 05:00 is counted and not listed, while one broken for three months on a paused Lane is listed and counted zero. **Neither set contains the other, in either direction.** Instead each Site row on the Lanes page carries **one separate link** to `/admin/series?site=<site>&filter=failing`, as navigation rather than attached to a number. Rejected: an `?outcome=` parameter to make the links exact — the surface is one filter at a time with only `?site=` and `?kind=` stacking, and a second axis would not fix the window mismatch anyway. ### Completion hint: storage **The false-positive containment this was named for is not built, because the research removed the threat.** Measured 2026-08-19 against live pages (note: `docs/research/completion-markers.md` on branch `research/completion-markers`): all six Sites publish an explicit machine-readable completion marker, and **a Poll that triggers only on each Site's *completed* value has no false-positive path on any Site.** So there is no N-consecutive-observation gate, no expiry, no confirmation counter and no suspicion damping. The residual error is a **false negative** — a finished Series that gets no hint — and we accept it. `series.site_completed_at bigint NOT NULL DEFAULT 0`, epoch-ms of the read that saw the Site's completed value; zero means the last successful read did not see it. **Decay is therefore free**: every successful read rewrites the answer, so a stale hint cannot survive one Poll. The timestamp is the age the hint renders with, not a history. - **Rejected: a text column holding a normalised status word.** The six vocabularies are not comparable: demonicscans and novelfull are binary and cannot express hiatus at all; asurascans' `dropped` and kagane's upload status are scanlation-editorial rather than completion state; and kagane is the only Site that separates the *work's* status from the *translation's* (proven divergent on one series). A shared word would be a different fact per Site, and nothing here acts on any value but completed. - **Rejected: storing nothing and re-deriving on read.** The marker exists only in a fetched page, and the admin surface never fetches — every intervention goes through the database. ### Completion hint: which value counts One predicate per Site, the **completed value only**, no mapping table: | Site | Reads | |---|---| | asurascans | status is `completed` (its `dropped` is editorial: one series stops at chapter 14 while the work continues) | | demonicscans | `Completed` | | comix | status is `finished` (`on_hiatus`, `discontinued`, `not_yet_released` are distinct) | | kagane | publication status is `Completed` **only** — upload status is the release's state, and the two diverge | | novelfull | `Completed` | | lightnovelworld | creative-work status is the completed action status | - **A failed extraction is not-completed, never unknown-as-suspicion.** - **A refused or unreachable Poll makes no statement at all** — the column keeps its previous value rather than zeroing. Inherited free from the failure-write rule above: a challenge is not evidence about the Series. - **All six adapters extract it.** Four ride a payload the adapter already parses (asurascans' island props, comix's initial-data detail entry, kagane's API body, lightnovelworld's head JSON-LD). demonicscans (an info-block list-item pair) and novelfull (a status link) each need **one new selector** against markup nobody controls. Paid anyway, on two grounds: a broken selector degrades to a missing hint, which is the error class we already accept; and on exactly those two Sites `Completed` is the only signal the Site can ever emit, so omitting them makes an absent hint unreadable — the owner could not tell "not finished" from "we never look". ### Completion hint: the write **Its own statement in the pass, after the read.** It cannot ride the check stamp (written *before* the fetch) and it cannot ride the chapter setter (skipped on the unchanged-number path). **It cannot share the failure write either**: that one targets the failure table and fires on a *failure*, while this is a fact learned from a *successful* read of `series`. The two share only their position in the pass. `site_completed_at` joins the due query's projection, so the statement fires **only when the answer would change**. The ordinary Poll of an ongoing Series writes nothing, exactly as a healthy Poll writes nothing to the failure table. ### Completion hint: presentation - **One new filter name, `site-completed`**, ordered with the informational tail beside `finished`, and carrying `AND finished_at = 0` — the same guard #136's hygiene predicates gained. - **One figure in the landing stats block** linking to it, printing an unlinked digit at zero. - **One line beside the Finish control** on the detail page. **No field on the list row** — a row has two lines, and this fact is true of few Series, so a per-row field spends space on every row for a rare signal. The filter is the work list; the figure is how the owner sees the work without opening it. - **The Finish control does not change.** Still one control, still confirm-gated. The hint is text next to it: it may not pre-fill, add a second button, or remove the confirm step. **A hint that shortens the path is the Site deciding the Lifecycle.** - **No dismissal column.** The owner's answers are Finish or ignore. A stored dismissal would be a third opinion on the row beside the Site's marker and the owner's own flag, and would need its own rule for the day the Site's marker changes again. An ignored Series stays in a fifty-row filter; if that noise turns out to be real, the column can be added then. - **After a finish**: the filter's guard excludes it and the detail page hides the hint. A finished Series leaves the Poll query, so its value freezes at whatever the last Poll saw — old and unrefreshable, therefore not work. Un-finishing puts it back in the query and the next Poll writes a fresh answer. - Two consequences worth stating rather than rediscovering: the hint is only as fresh as the last Poll, so the two browser-gated Sites' hints age while the sidecar is asleep or unreachable — the same degradation their Covers already have, and **no new notification**. ### Outbound notification: the rule, then the conditions **The dashboard is not pull-only.** Push exists for one class of fault: **silent on every Reader surface, wrong, repairable only by the owner, and persisting past twelve hours.** *That rule is the decision* — the four conditions below are what it currently selects, and a fifth candidate must pass all four tests. **A backend that stopped is not the case worth pushing**, and the reason came from the Reader surfaces rather than from the plan. The web UI dies visibly. The userscript degrades quietly but honestly: reads come from the cache, a write that meets a dropped connection goes to the retry queue, the queue only speaks up for a 400, a 401 or a full queue, and the sole visible marker is the pending count in the panel. No Progress is lost — the only loss path is eviction at the queue cap, which toasts. So downtime costs a delay and is found within one reading session, by every Reader. A stalled Lane is the opposite. The four conditions, each at `ownerWindow` (12h; **no new constant**): 1. **`stall`** — `due > 0 AND checked = 0 AND skip = '' AND refused = 0` on the pass log. 2. **`no-browser-route`** — a Site with **no browser route** refusing for more than 12h. The read path maps both a 403 with the mitigation header and an interstitial body to a held challenge on the plain-TLS path too, so this is `refused` plus the durable Lane refusal. This is the comix event of 2026-08-12, and the repair is a deploy. **A browser Site refusing sends nothing** — that is an ordinary challenge that did not clear, and a Site setting its own pace is not a fault. 3. **`sidecar-down`** — no browser Lane has reached the sidecar for more than 12h. **This is not a stall and cannot be found by the stall test**: the pass returns early both when a sibling Lane lost Chrome and when no fetcher exists, and both are recorded as skip values. The quiet degradation is deliberate, but "asleep for an hour" and "gone for a week" are indistinguishable to the backend, and the second freezes two Sites' marks. 4. **`adapter-broken`** — more than **half** of one Site's Series hold a `no_chapter` failure row older than 12h. A layout change produces exactly that outcome; one such Series is a bad row, most of a Site is a broken adapter. **A share, not a count**, so a four-Series Site and a two-hundred-Series Site both fire. This is the one fault on the whole map that no other signal names. **Never pushed**: a refusal by a browser Site (not a fault), a pause (the owner's own act), a single failing Series (the filter is the tool, and one Site change makes hundreds of rows), a forced Poll ageing (the same fault from the other end, and visible on the page the button lives on), a gated cover host (a missing Cover is visible in the UI). ### Two collisions found in the code, and their fixes - **The stall test catches a refusing Site.** The twice-refused break inside the pass loop is *not* an early return — the pass falls through to its normal return, writing `due > 0, checked = 0, skip = ''`. A newly-gated Site would have sent both `stall` and `no-browser-route`. **So the stall test gains `AND refused = 0`**, amending the definition #134 lands. Rejected alternative: a tenth skip value, because the enum holds one value per *return* path and this is a break inside one — a tenth value would describe a pass that did run. - **One dead sidecar is three Lanes.** Each browser Lane notices the loss on its own pass. So the suppression row's `site` holds **the empty string** for `sidecar-down`: the first Lane to notice writes the row, the others find it present and stay quiet. **The row's own presence is the lock** — which is exactly why the state is its own table and not columns on the per-Site Lane row, where an empty Site name has no meaning. Rejected: naming one Lane the reporter, since that Lane's pass may be the skipped one. ### Where it fires from, and suppression Beside the pass-log insert, **in the deferred call every return path passes through**. Two of the four conditions occur on early returns, so the success path can never see them. **No new timer.** The Lane loop already does one pass, then waits the pace the pass reported, which is never more than the default rest — each Lane is its own clock at least hourly. (The premise that the backend has no clock was wrong; it has six.) New table `owner_notices(kind, site, notified_at)` — **one row per episode**, cleared when the condition no longer holds, so a fault lasting a month sends one message rather than thirty at ~120 passes a day. **Every threshold is derived, not stored**: a refusal's and a sidecar absence's age from the pass log, a Series failure's age from its failing-since. **The only new durable fact is that the owner was told**, which is why the table carries nothing else. Rows are bounded by four kinds times six Sites, so there is no retention rule. ### Transport - **`DISCORD_WEBHOOK_URL`, and unset means the whole path is off**, logged once at start, exactly as the browser URL behaves — a local stack must not need a webhook to run. - A webhook is a plain POST and **does not touch the login path**: no bot, no gateway, no scope, not the OAuth application. The existing Discord coupling gains nothing new. - **The address is a secret in the class of the token key and is never logged.** Rate limits are irrelevant at this volume (roughly 30 messages a minute allowed per webhook; this sends a few a day). - **A failed POST is logged and the notified timestamp is left unset**, so the next pass retries while the condition holds. No queue, no backoff: the condition is durable, so the retry is free. (The userscript needs a queue because a Reader's Progress is otherwise lost; a message about a durable fact is not.) - **Rejected transports**, measured 2026-08-21: a Telegram bot — free and unlimited at this volume, but a second identity to hold when the owner's Discord id already names the recipient; the public ntfy.sh server — 250 messages a day and 60 burst, but **every topic is public by design** (no authentication; anyone knowing the topic can read and publish) with limited retention; self-hosted ntfy or Gotify — a container and a volume, which is the infrastructure this decision refuses. ### Presentation **A Discord embed**, colour `13589581` — the dark branch of the danger token. One colour for all four, because all four are faults; **ember stays forbidden**, and no second colour is introduced because no non-fault condition pushes. - `title` = the subject (the Site, or `browser sidecar`) carrying the deep link built from `PUBLIC_BASE_URL`, so no body line is spent on an address. - `description` = one sentence: condition, age, repair. - `timestamp` = the pass time, which Discord renders in the reader's own zone — the one thing an embed does better than plain text. - `footer` = the machine word (`stall`, `no-browser-route`, `sidecar-down`, `adapter-broken`), making an old message searchable. - **No field grid, no thumbnail, no author block** — the figures belong on the page the title links to. Only the colour of Cinder survives into Discord; the typography cannot. ### The page and the push stay in step **The landing verdict line must read the same four conditions.** A message naming a fault the page cannot explain sends the owner to a page that looks healthy. This amends #134's verdict line, which reads Lane attention generally; narrow it to these four. ## Testing Decisions **What makes a good test here**: it asserts a durable outcome — a row present or absent after a pass, a filter's membership, a payload's shape — and it fails on a plausible bug. The two silent outcomes are the highest-value tests in this spec, because their bug is invisible: writing where nothing should be written produces a plausible-looking row. **The primary seam stays the router** for everything the owner sees, and the poller's existing round-at-a-time entry point for everything the Lane writes. **One new seam, and only one:** `Poller.Notifier func(...) error` with `nil` meaning off — the same field-injection pattern the fetchers already use, so all four conditions are testable with no wire and no Discord. Modules and coverage: - **Poller** (prior art: the round-at-a-time pass tests with a fake fetcher and an injected clock): each writing outcome inserts a row with the right word; a repeated identical failure writes nothing; a changed word updates the word and **keeps** the failing-since (this is the test that protects the age); a successful read deletes the row; **a refused Poll neither inserts nor deletes**, and **an unreachable Poll neither inserts nor deletes** — asserted against a pre-seeded row so both directions are covered; a forced Poll that succeeds deletes the row and a forced *request* alone does not; the completion write fires only on a transition and not on an ongoing Series' ordinary Poll; a refused Poll leaves the completion column unchanged; the six-word outcome taxonomy maps a 404 to `not_found` and a 503 to `errors`. - **Adapters** (prior art: the existing per-Site parsing tests over stored page fixtures, plus the network-gated live smoke tests): each of the six Sites' completed predicate against a completed fixture and an ongoing one; kagane's divergent case asserts that the *upload* status is ignored; a fixture with the selector removed yields not-completed rather than an error. The live smoke tests stay network-gated and are not part of the default run. - **Store**: the `failing` and `unverified` predicates, including disjointness from `no-chapter` (a Series that never succeeded must be in `no-chapter` and never in `failing`); a Series failing for under twelve hours is in neither filter; a finished Series still appears in `failing`; `site-completed` excludes a finished Series; the cascade removes failure rows when a Series is deleted and the orphan delete stays a single statement; the admin join projects the outcome word and the age. - **Router / web**: the failure marker renders on the fact line and identically for a one-hour and a three-month failure; the two new filters and the completion figure and filter; the outcome counts render unlinked while the per-Site navigation link is present; the `unverified` sentence appears beside the correction control; the hint is hidden on a finished Series and the Finish control is unchanged when a hint is present; the verdict line reflects each of the four conditions. - **Notifier**: each of the four conditions fires exactly once while it holds and again after it lifts and returns; a stall on a Site that also refused fires **only** `no-browser-route` (the collision test — this is the one a naive implementation gets wrong); one dead sidecar across three browser Lanes sends exactly one message; a pause sends nothing; a browser Site's refusal sends nothing; a single failing Series sends nothing; the adapter-broken share fires on a four-Series Site and does not fire at exactly half; a failed send leaves the notified timestamp unset and the next pass retries. **One separate test covers the embed payload** — colour, title link, footer word, timestamp, and the absence of a field grid. ## Out of Scope - **A consecutive-failure counter, a last-outcome timestamp, or any per-Series failure history.** - **Slowing a Lane for a persistent failure.** The due query does **not** join the failure table; a failing Series is polled at the same pace as any other. Pacing is not a dashboard question. - **A normalised status word for the completion hint**, or any use of a hiatus, dropped or translation-status value. - **Surfacing a stalled-but-not-completed serialisation.** On the two binary Sites a stalled series reads Ongoing for ever; that is the accepted false negative. - **A machine write of the Series' finished flag**, a pre-filled Finish control, a second Finish button, or removing its confirm step. - **A dismissal column for the hint.** - **Linking the per-Site outcome counts to any filter.** Withdrawn permanently, with reasons above. - **A stored column for the `unverified` pair.** - **Notification for anything outside the four conditions**, including a paused Lane, a browser Site's refusal, one failing Series, an ageing forced Poll, and a gated cover host. - **A queue, a backoff, or any delivery machinery** beyond retry-on-next-pass. - **Telegram, ntfy, Gotify, email, or a second transport of any kind.** - **A monitor outside the backend** — an external prober against the health endpoint, or a hosted uptime service. A push cannot report its own process's death, so the honest coverage lives outside; but it is another service to run, watch and update, and it watches a fault that is **already visible from the Reader surfaces**: the web UI dies visibly and the userscript degrades to its cache with a pending count, losing no Progress. Downtime therefore costs a delay and is found within one reading session. ## Further Notes - **Migration numbers are claimed in landing order.** Measured 2026-08-21: the tree's migrations and ADRs both stop at 0011, and everything the map specified is unbuilt. This spec's schema work is `poll_failures`; `series.site_completed_at bigint NOT NULL DEFAULT 0`; `owner_notices(kind, site, notified_at)`; and one extra count column on the pass log. Take the next free numbers after the earlier specs land. - **One ADR is worth writing with the implementation**: a table whose row *is* the failure state, plus the two outcomes that deliberately write nothing. **No ADR for the notification path** — it is reversible by deleting the path, so it fails the hardness test. - **Glossary**: `Stall` is already defined in `CONTEXT.md` (it landed while charting). `failing`, `unverified` and `site-completed` are filter names over facts the glossary already covers, so no new nouns. The completion column gets no glossary entry: the glossary describes the system that exists. - **The three research branches this consumes should be deleted once implemented**: `research/completion-markers` (the note is `docs/research/completion-markers.md`) and `research/htmx-series-list`. Neither was merged, deliberately — charting should not land build artifacts. - The prototype amendments this spec carries into the design project: the fact-line failure marker, the Lanes-page per-Site navigation link, the detail-page `unverified` sentence, and two more options on the filter select. - Security-critical surfaces touched: a new secret in the environment (never logged, and it must not reach any log line that prints configuration), the owner gate on any new route, and the store's query construction. The notification path must not import anything from the session or OAuth packages — that independence is the reason a webhook was chosen over a bot. Say which invariant you preserved in the PR and run the full backend test suite before calling it done.
sulthan added the ready-for-agent label 2026-08-21 15:47:54 +07:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: sulthan/mangaBookmark#137