Sightings: a Reader report defers a Poll where being wrong hurts only them #103
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Blocked by: #102
Problem Statement
As the owner of this backend, most of its work is wasted. My userscript already reads the Latest Chapter off the Series page every time I visit it, and it already sends that value to the backend. Then, minutes later, the Poll Lane fetches the same page and reads the same number. The Series was freshly known and got checked anyway.
That waste has a cost that scales exactly where I do not want it to. Every redundant Poll is a request against a Site, and the Poll Lane's pace is computed from how many Series need checking. As my Library grows past a few hundred Series, the automatic shortening starts squeezing the gap toward the floor — and a meaningful share of what it is squeezing to make room for is work the Reader already did.
I cannot simply trust those reports. Registration is open to any member of my Discord guild, which is not the same as a hand-picked handful of friends. Series are shared: one row, read by everyone who bookmarks it. So a Reader who sends a wrong Latest Chapter — through malice or a broken adapter — changes what everyone sees, and if that report were also allowed to postpone the Poll, nothing would come along to correct it.
Solution
A Reader's report becomes a named thing — a Sighting — and it earns the right to postpone a Poll only where being wrong can hurt nobody but the Reader who sent it.
On a Series only one Reader bookmarks, a Sighting defers that Series' next Poll. On a Series several Readers share, the Sighting still updates the Latest Chapter instantly, exactly as today, but the Poll happens on schedule anyway — so a wrong value on a shared Series is corrected within the hour by a check that was never postponed.
A ceiling caps the whole mechanism: however many Sightings arrive, every Series is Polled at least every six hours. And detection is free, because the Poll already compares. A Poll that finds a lower chapter number than the one stored means whoever last raised it was wrong; a Poll finding a higher number is just the Site publishing, and means nothing. Three contradictions and that Reader's Sightings stop deferring anything until they earn it back.
User Stories
Implementation Decisions
Sighting. Now a glossary term: what a Reader's browser happened to see of a Series's Latest Chapter while that Reader was present. It is a client report, and the Latest Chapter definition is amended to say the value is established by a Poll and, between Polls, by a Sighting. A Sighting that a later Poll contradicts downwards is a false Sighting, and enough of those cost the Reader the right to defer at all.
What a Sighting always does. Updates the shared Latest Chapter, immediately, on every Series, exactly as the current write path does. Never touches Progress, never touches the Bookmark's ordering timestamp. This is unchanged behaviour and must stay unchanged.
What a Sighting sometimes does. Defers that Series' next Poll, but only when the Series has exactly one Bookmark — that is, only the reporting Reader bookmarks it. On a Series with more than one Bookmark the Poll is not deferred. The reasoning is the decisive one in the whole design: on a shared Series, a Poll that was never postponed corrects a lie within the hour, and the correction is visible in the log; on a solitary Series, the only person a lie can reach is the Reader who told it. Restricting deferral to solitary Series therefore closes the "one Reader can poison what others see, undetected" hole by construction rather than by rule, while keeping the instant update everyone benefits from.
The ceiling. Six hours. Regardless of how many Sightings arrive, a Series that has not been Polled in six hours is Polled. This bounds how long any deferral can last and is what makes the whole mechanism safe by default: poison dies in at most six hours, deterministically, rather than in expectation. It is a deliberate, harder version of the randomised-auditing technique the literature offers.
Detection. No new comparison and no new request. The Poll already reads the true value and already writes it. At that moment it knows whether the stored value was higher than what the Site actually publishes. A Poll finding a lower number than stored means the last Sighting to raise it was false. A Poll finding a higher number is the Site publishing and means nothing. This asymmetry is what keeps detection from producing false alarms.
Attribution. The Series remembers which Reader last raised its Latest Chapter — one identifier on the Series row. Without it a contradiction can only flag the Series, which patches one row and lets the same Reader do it again to another. With it, the guard is about the Reader.
Reputation. Two counters on the Reader. A contradiction increments the disagreement count; a Poll that confirms a Reader's Sighting increments the agreement count. At three disagreements that Reader's Sightings stop deferring anything — they still write the Latest Chapter. Twenty consecutive agreements clear the disagreements back to zero. The published form of this technique uses a ratio; a threshold was chosen instead so that an attacker cannot bank free lies by first building credit, and so that recovery does not require the owner to watch a log.
Note the real cost of that recovery rule: an agreement is only recorded when a Poll later confirms a Sighting, so twenty agreements take twenty Polls of Series that Reader bookmarks — hours to days of real time, not twenty page views. Combined with the six-hour ceiling, a determined attacker buys three deferral windows and then waits days, and every step is named in the log. That is the intended price.
Clearing marks. The owner clears a Reader's counters from the administrative page. This exists because the guard has one known false-positive mode: a Site changing its page shape can make a correct adapter read a wrong high number, marking an honest Reader. That failure is visible in the log — it will mark every Reader of that Site at once — but it needs a remedy that is not SQL.
New Readers start trusted. Zero counters means trusted. The alternative — earning trust — makes the feature broken for a new guild member until they have done nothing wrong for a while, and the exposure it avoids is one hour of a wrong value on a solitary Series.
Why not consensus. Truth discovery, weighted majority and reputation-by-comparison all need multiple independent sources per object to work at all; the standard survey states plainly that an object provided by very few sources cannot have its confidence evaluated. With two Readers, disagreement is a coin flip. The authoritative Poll is used as the oracle instead, which is the same role the gold-question audit plays in the crowdsourcing literature — except deterministic, because the ceiling guarantees the audit rather than sampling it.
Threat model, on record. Registration is open to any member of the Discord guild; there is no invite table, by an earlier decision. The question this guards is not "would my friend lie" but "would any guild member lie". That is why the guard exists at all.
Testing Decisions
A good test here asserts two observable things: whether a Poll happened, and what the stored Latest Chapter is afterwards. It must not assert counter arithmetic through internal calls, nor how the deferral decision is represented in a row.
Everything lands at the existing Poller seam, with the store as the way in. A Sighting is seeded through the store — the same write path the request handler uses — then a round is run against the injected fetcher and a frozen clock, and the assertion is whether that Series was fetched. This keeps the whole spec at the seam that already exists and avoids a combined harness that drives the router and the poller together. No new seam.
The behaviours to cover:
Prior art: the existing cooldown-across-passes test is the closest shape — seed state, freeze the clock, run a round, assert what was and was not fetched. The tests that assert a shared Series is fetched exactly once already establish how to set up multi-Reader Series. The dashboard's roster tests cover the clear-marks control at the router seam and are not repeated here.
Out of Scope
Further Notes
The mechanism's saving is real but bounded by how the deferral is restricted: only solitary Series defer, so a Library whose Series are widely shared saves less. That is the correct trade — shared Series are exactly the ones where a Poll is amortised across several Readers anyway, so they are the cheapest Polls in the system per Reader served.
This spec needs its own decision record, separate from the Poll Lanes one, covering the trust model, the six-hour ceiling, the three-and-twenty thresholds, the solitary-Series restriction, and the rejected alternatives with their citations. A future reader finding that a client report can postpone a server check will reasonably think it is madness; the record is where the answer lives.
The administrative page must land before this work, so that a mark can be cleared as soon as one can be created.
sulthan referenced this issue2026-08-16 16:36:53 +07:00