Sightings: a Reader report defers a Poll where being wrong hurts only them #103

Closed
opened 2026-08-16 11:28:27 +07:00 by sulthan · 0 comments
Owner

Blocked by: #102

Problem Statement

As the owner of this backend, most of its work is wasted. My userscript already reads the Latest Chapter off the Series page every time I visit it, and it already sends that value to the backend. Then, minutes later, the Poll Lane fetches the same page and reads the same number. The Series was freshly known and got checked anyway.

That waste has a cost that scales exactly where I do not want it to. Every redundant Poll is a request against a Site, and the Poll Lane's pace is computed from how many Series need checking. As my Library grows past a few hundred Series, the automatic shortening starts squeezing the gap toward the floor — and a meaningful share of what it is squeezing to make room for is work the Reader already did.

I cannot simply trust those reports. Registration is open to any member of my Discord guild, which is not the same as a hand-picked handful of friends. Series are shared: one row, read by everyone who bookmarks it. So a Reader who sends a wrong Latest Chapter — through malice or a broken adapter — changes what everyone sees, and if that report were also allowed to postpone the Poll, nothing would come along to correct it.

Solution

A Reader's report becomes a named thing — a Sighting — and it earns the right to postpone a Poll only where being wrong can hurt nobody but the Reader who sent it.

On a Series only one Reader bookmarks, a Sighting defers that Series' next Poll. On a Series several Readers share, the Sighting still updates the Latest Chapter instantly, exactly as today, but the Poll happens on schedule anyway — so a wrong value on a shared Series is corrected within the hour by a check that was never postponed.

A ceiling caps the whole mechanism: however many Sightings arrive, every Series is Polled at least every six hours. And detection is free, because the Poll already compares. A Poll that finds a lower chapter number than the one stored means whoever last raised it was wrong; a Poll finding a higher number is just the Site publishing, and means nothing. Three contradictions and that Reader's Sightings stop deferring anything until they earn it back.

User Stories

  1. As the owner, I want the backend to skip checking a Series my own browser just checked, so that its limited request budget goes to Series nobody has looked at.
  2. As the owner, I want that saving to grow as my Library grows, so that the hourly promise survives a larger collection.
  3. As a Reader, I want the Latest Chapter I just saw to appear for everyone immediately, so that reporting it is still instant.
  4. As a Reader, I want my report to be believed about my own Series, so that the backend is not duplicating work I just did.
  5. As a Reader, I want a report about a Series others also read to change what we all see but never postpone the check, so that my mistake is corrected quickly rather than persisting.
  6. As a Reader, I want a Series to be Polled eventually no matter how often anyone reports it, so that no Series can be talked into never being checked.
  7. As a Reader, I want a Series I stop visiting to return to the normal schedule on its own, so that abandoning a Series does not freeze its Latest Chapter.
  8. As a Reader, I want a Sighting never to reorder my list, so that reporting a chapter is not mistaken for reading one.
  9. As a Reader, I want a Sighting never to change my Progress, so that where I am up to stays mine to set.
  10. As a Reader, I want a Sighting that reports the number already stored to still count as a check, so that visiting a Series with nothing new is not wasted.
  11. As a Reader, I want the New Chapter signal to behave exactly as it does today, so that nothing about how I notice updates changes.
  12. As the owner, I want a Reader who reports chapters that do not exist to be detected automatically, so that I do not have to audit reports by hand.
  13. As the owner, I want that detection to cost no extra requests, so that a guard against waste is not itself wasteful.
  14. As the owner, I want a Poll finding a higher number treated as normal, so that a Site publishing new chapters never looks like an attack.
  15. As the owner, I want the Reader who caused a contradicted value to be named rather than the Series merely flagged, so that a repeat offender can be stopped rather than one Series patched.
  16. As the owner, I want a small number of contradictions to remove only the ability to defer, so that a Reader with a broken adapter is not locked out of their own Library.
  17. As the owner, I want an honest Reader with a long good record to survive a single bad report, so that one adapter bug does not punish someone permanently.
  18. As the owner, I want a marked Reader to be able to earn their standing back through subsequent correct reports, so that no manual step is required for the common case.
  19. As the owner, I want earning it back to require many confirmed reports rather than a wait, so that a patient attacker cannot simply outlast the guard.
  20. As the owner, I want each contradiction logged with the Reader, the Series and both numbers, so that I can tell a broken adapter from a deliberate lie.
  21. As the owner, I want to clear a Reader's marks from the administrative page, so that a false mark from a broken adapter is fixed without touching the database.
  22. As the owner, I want a new Reader to be trusted from their first day, so that the feature is not broken for everyone until they earn it.
  23. As the owner, I want a marked Reader's reports to keep updating the Latest Chapter, so that the penalty removes a privilege rather than silencing them.
  24. As the owner, I want the guard sized for an open guild rather than for hand-picked friends, so that the trust model matches how registration actually works.
  25. As a maintainer, I want the deferral decision to depend on facts the scheduler already has, so that the Poll Lanes do not gain a new query per round.
  26. As a maintainer, I want the added state to be a small fixed amount per Reader and per Series, so that this does not become a schema of its own.
  27. As a maintainer, I want the rejected alternatives recorded with citations, so that nobody re-proposes consensus scoring for a system with a handful of Readers.

Implementation Decisions

Sighting. Now a glossary term: what a Reader's browser happened to see of a Series's Latest Chapter while that Reader was present. It is a client report, and the Latest Chapter definition is amended to say the value is established by a Poll and, between Polls, by a Sighting. A Sighting that a later Poll contradicts downwards is a false Sighting, and enough of those cost the Reader the right to defer at all.

What a Sighting always does. Updates the shared Latest Chapter, immediately, on every Series, exactly as the current write path does. Never touches Progress, never touches the Bookmark's ordering timestamp. This is unchanged behaviour and must stay unchanged.

What a Sighting sometimes does. Defers that Series' next Poll, but only when the Series has exactly one Bookmark — that is, only the reporting Reader bookmarks it. On a Series with more than one Bookmark the Poll is not deferred. The reasoning is the decisive one in the whole design: on a shared Series, a Poll that was never postponed corrects a lie within the hour, and the correction is visible in the log; on a solitary Series, the only person a lie can reach is the Reader who told it. Restricting deferral to solitary Series therefore closes the "one Reader can poison what others see, undetected" hole by construction rather than by rule, while keeping the instant update everyone benefits from.

The ceiling. Six hours. Regardless of how many Sightings arrive, a Series that has not been Polled in six hours is Polled. This bounds how long any deferral can last and is what makes the whole mechanism safe by default: poison dies in at most six hours, deterministically, rather than in expectation. It is a deliberate, harder version of the randomised-auditing technique the literature offers.

Detection. No new comparison and no new request. The Poll already reads the true value and already writes it. At that moment it knows whether the stored value was higher than what the Site actually publishes. A Poll finding a lower number than stored means the last Sighting to raise it was false. A Poll finding a higher number is the Site publishing and means nothing. This asymmetry is what keeps detection from producing false alarms.

Attribution. The Series remembers which Reader last raised its Latest Chapter — one identifier on the Series row. Without it a contradiction can only flag the Series, which patches one row and lets the same Reader do it again to another. With it, the guard is about the Reader.

Reputation. Two counters on the Reader. A contradiction increments the disagreement count; a Poll that confirms a Reader's Sighting increments the agreement count. At three disagreements that Reader's Sightings stop deferring anything — they still write the Latest Chapter. Twenty consecutive agreements clear the disagreements back to zero. The published form of this technique uses a ratio; a threshold was chosen instead so that an attacker cannot bank free lies by first building credit, and so that recovery does not require the owner to watch a log.

Note the real cost of that recovery rule: an agreement is only recorded when a Poll later confirms a Sighting, so twenty agreements take twenty Polls of Series that Reader bookmarks — hours to days of real time, not twenty page views. Combined with the six-hour ceiling, a determined attacker buys three deferral windows and then waits days, and every step is named in the log. That is the intended price.

Clearing marks. The owner clears a Reader's counters from the administrative page. This exists because the guard has one known false-positive mode: a Site changing its page shape can make a correct adapter read a wrong high number, marking an honest Reader. That failure is visible in the log — it will mark every Reader of that Site at once — but it needs a remedy that is not SQL.

New Readers start trusted. Zero counters means trusted. The alternative — earning trust — makes the feature broken for a new guild member until they have done nothing wrong for a while, and the exposure it avoids is one hour of a wrong value on a solitary Series.

Why not consensus. Truth discovery, weighted majority and reputation-by-comparison all need multiple independent sources per object to work at all; the standard survey states plainly that an object provided by very few sources cannot have its confidence evaluated. With two Readers, disagreement is a coin flip. The authoritative Poll is used as the oracle instead, which is the same role the gold-question audit plays in the crowdsourcing literature — except deterministic, because the ceiling guarantees the audit rather than sampling it.

Threat model, on record. Registration is open to any member of the Discord guild; there is no invite table, by an earlier decision. The question this guards is not "would my friend lie" but "would any guild member lie". That is why the guard exists at all.

Testing Decisions

A good test here asserts two observable things: whether a Poll happened, and what the stored Latest Chapter is afterwards. It must not assert counter arithmetic through internal calls, nor how the deferral decision is represented in a row.

Everything lands at the existing Poller seam, with the store as the way in. A Sighting is seeded through the store — the same write path the request handler uses — then a round is run against the injected fetcher and a frozen clock, and the assertion is whether that Series was fetched. This keeps the whole spec at the seam that already exists and avoids a combined harness that drives the router and the poller together. No new seam.

The behaviours to cover:

  • A Sighting on a Series with one Bookmark defers the next Poll: the round does not fetch it.
  • A Sighting on a Series with two Bookmarks does not defer: the round fetches it, and the stored Latest Chapter was updated by the Sighting beforehand.
  • A Sighting updates the Latest Chapter, and leaves Progress and the Bookmark's ordering timestamp untouched, in both the shared and solitary cases.
  • A Series last Polled longer ago than the ceiling is Polled despite a recent Sighting.
  • A Poll that finds a lower chapter number than stored records a disagreement against the Reader who last raised it, and logs both numbers.
  • A Poll that finds a higher number records no disagreement.
  • A Poll that confirms the stored value records an agreement for that Reader.
  • A Reader at the disagreement threshold no longer defers: their Sighting on a solitary Series is followed by a Poll in the next round.
  • A marked Reader's Sighting still updates the Latest Chapter.
  • Twenty confirmations clear the marks and deferral resumes.
  • Clearing marks through the store restores deferral immediately.
  • A Series whose sole Bookmark is deleted, or joined by a second Bookmark, stops deferring from the next round.

Prior art: the existing cooldown-across-passes test is the closest shape — seed state, freeze the clock, run a round, assert what was and was not fetched. The tests that assert a shared Series is fetched exactly once already establish how to set up multi-Reader Series. The dashboard's roster tests cover the clear-marks control at the router seam and are not repeated here.

Out of Scope

  • Any cross-Reader agreement, voting, weighting or consensus scoring. Rejected on evidence; the decision record carries the citations.
  • Rate limiting the Sighting write path itself. Separate concern, and the existing body-size and validation limits already apply.
  • Letting a Sighting change Progress, favourite state, or Lifecycle bucket.
  • Any Reader-visible indication of their own reputation. The counters are the owner's operational view, not a score shown to Readers.
  • Automatic expiry of marks by elapsed time. Recovery is by confirmed agreements only, deliberately, so waiting is not a strategy.
  • Any change to the userscripts. They already send what a Sighting needs; this spec changes only what the backend does with it.
  • Extending deferral to Acquisition, which is a single read at a Series' creation and never queues behind a Lane.

Further Notes

The mechanism's saving is real but bounded by how the deferral is restricted: only solitary Series defer, so a Library whose Series are widely shared saves less. That is the correct trade — shared Series are exactly the ones where a Poll is amortised across several Readers anyway, so they are the cheapest Polls in the system per Reader served.

This spec needs its own decision record, separate from the Poll Lanes one, covering the trust model, the six-hour ceiling, the three-and-twenty thresholds, the solitary-Series restriction, and the rejected alternatives with their citations. A future reader finding that a client report can postpone a server check will reasonably think it is madness; the record is where the answer lives.

The administrative page must land before this work, so that a mark can be cleared as soon as one can be created.

Blocked by: #102 ## Problem Statement As the owner of this backend, most of its work is wasted. My userscript already reads the Latest Chapter off the Series page every time I visit it, and it already sends that value to the backend. Then, minutes later, the Poll Lane fetches the same page and reads the same number. The Series was freshly known and got checked anyway. That waste has a cost that scales exactly where I do not want it to. Every redundant Poll is a request against a Site, and the Poll Lane's pace is computed from how many Series need checking. As my Library grows past a few hundred Series, the automatic shortening starts squeezing the gap toward the floor — and a meaningful share of what it is squeezing to make room for is work the Reader already did. I cannot simply trust those reports. Registration is open to any member of my Discord guild, which is not the same as a hand-picked handful of friends. Series are shared: one row, read by everyone who bookmarks it. So a Reader who sends a wrong Latest Chapter — through malice or a broken adapter — changes what everyone sees, and if that report were also allowed to postpone the Poll, nothing would come along to correct it. ## Solution A Reader's report becomes a named thing — a **Sighting** — and it earns the right to postpone a Poll only where being wrong can hurt nobody but the Reader who sent it. On a Series only one Reader bookmarks, a Sighting defers that Series' next Poll. On a Series several Readers share, the Sighting still updates the Latest Chapter instantly, exactly as today, but the Poll happens on schedule anyway — so a wrong value on a shared Series is corrected within the hour by a check that was never postponed. A ceiling caps the whole mechanism: however many Sightings arrive, every Series is Polled at least every six hours. And detection is free, because the Poll already compares. A Poll that finds a **lower** chapter number than the one stored means whoever last raised it was wrong; a Poll finding a higher number is just the Site publishing, and means nothing. Three contradictions and that Reader's Sightings stop deferring anything until they earn it back. ## User Stories 1. As the owner, I want the backend to skip checking a Series my own browser just checked, so that its limited request budget goes to Series nobody has looked at. 2. As the owner, I want that saving to grow as my Library grows, so that the hourly promise survives a larger collection. 3. As a Reader, I want the Latest Chapter I just saw to appear for everyone immediately, so that reporting it is still instant. 4. As a Reader, I want my report to be believed about my own Series, so that the backend is not duplicating work I just did. 5. As a Reader, I want a report about a Series others also read to change what we all see but never postpone the check, so that my mistake is corrected quickly rather than persisting. 6. As a Reader, I want a Series to be Polled eventually no matter how often anyone reports it, so that no Series can be talked into never being checked. 7. As a Reader, I want a Series I stop visiting to return to the normal schedule on its own, so that abandoning a Series does not freeze its Latest Chapter. 8. As a Reader, I want a Sighting never to reorder my list, so that reporting a chapter is not mistaken for reading one. 9. As a Reader, I want a Sighting never to change my Progress, so that where I am up to stays mine to set. 10. As a Reader, I want a Sighting that reports the number already stored to still count as a check, so that visiting a Series with nothing new is not wasted. 11. As a Reader, I want the New Chapter signal to behave exactly as it does today, so that nothing about how I notice updates changes. 12. As the owner, I want a Reader who reports chapters that do not exist to be detected automatically, so that I do not have to audit reports by hand. 13. As the owner, I want that detection to cost no extra requests, so that a guard against waste is not itself wasteful. 14. As the owner, I want a Poll finding a higher number treated as normal, so that a Site publishing new chapters never looks like an attack. 15. As the owner, I want the Reader who caused a contradicted value to be named rather than the Series merely flagged, so that a repeat offender can be stopped rather than one Series patched. 16. As the owner, I want a small number of contradictions to remove only the ability to defer, so that a Reader with a broken adapter is not locked out of their own Library. 17. As the owner, I want an honest Reader with a long good record to survive a single bad report, so that one adapter bug does not punish someone permanently. 18. As the owner, I want a marked Reader to be able to earn their standing back through subsequent correct reports, so that no manual step is required for the common case. 19. As the owner, I want earning it back to require many confirmed reports rather than a wait, so that a patient attacker cannot simply outlast the guard. 20. As the owner, I want each contradiction logged with the Reader, the Series and both numbers, so that I can tell a broken adapter from a deliberate lie. 21. As the owner, I want to clear a Reader's marks from the administrative page, so that a false mark from a broken adapter is fixed without touching the database. 22. As the owner, I want a new Reader to be trusted from their first day, so that the feature is not broken for everyone until they earn it. 23. As the owner, I want a marked Reader's reports to keep updating the Latest Chapter, so that the penalty removes a privilege rather than silencing them. 24. As the owner, I want the guard sized for an open guild rather than for hand-picked friends, so that the trust model matches how registration actually works. 25. As a maintainer, I want the deferral decision to depend on facts the scheduler already has, so that the Poll Lanes do not gain a new query per round. 26. As a maintainer, I want the added state to be a small fixed amount per Reader and per Series, so that this does not become a schema of its own. 27. As a maintainer, I want the rejected alternatives recorded with citations, so that nobody re-proposes consensus scoring for a system with a handful of Readers. ## Implementation Decisions **Sighting.** Now a glossary term: what a Reader's browser happened to see of a Series's Latest Chapter while that Reader was present. It is a client report, and the Latest Chapter definition is amended to say the value is established by a Poll and, between Polls, by a Sighting. A Sighting that a later Poll contradicts downwards is a false Sighting, and enough of those cost the Reader the right to defer at all. **What a Sighting always does.** Updates the shared Latest Chapter, immediately, on every Series, exactly as the current write path does. Never touches Progress, never touches the Bookmark's ordering timestamp. This is unchanged behaviour and must stay unchanged. **What a Sighting sometimes does.** Defers that Series' next Poll, but only when the Series has exactly one Bookmark — that is, only the reporting Reader bookmarks it. On a Series with more than one Bookmark the Poll is not deferred. The reasoning is the decisive one in the whole design: on a shared Series, a Poll that was never postponed corrects a lie within the hour, and the correction is visible in the log; on a solitary Series, the only person a lie can reach is the Reader who told it. Restricting deferral to solitary Series therefore closes the "one Reader can poison what others see, undetected" hole by construction rather than by rule, while keeping the instant update everyone benefits from. **The ceiling.** Six hours. Regardless of how many Sightings arrive, a Series that has not been Polled in six hours is Polled. This bounds how long any deferral can last and is what makes the whole mechanism safe by default: poison dies in at most six hours, deterministically, rather than in expectation. It is a deliberate, harder version of the randomised-auditing technique the literature offers. **Detection.** No new comparison and no new request. The Poll already reads the true value and already writes it. At that moment it knows whether the stored value was higher than what the Site actually publishes. A Poll finding a lower number than stored means the last Sighting to raise it was false. A Poll finding a higher number is the Site publishing and means nothing. This asymmetry is what keeps detection from producing false alarms. **Attribution.** The Series remembers which Reader last raised its Latest Chapter — one identifier on the Series row. Without it a contradiction can only flag the Series, which patches one row and lets the same Reader do it again to another. With it, the guard is about the Reader. **Reputation.** Two counters on the Reader. A contradiction increments the disagreement count; a Poll that confirms a Reader's Sighting increments the agreement count. At three disagreements that Reader's Sightings stop deferring anything — they still write the Latest Chapter. Twenty consecutive agreements clear the disagreements back to zero. The published form of this technique uses a ratio; a threshold was chosen instead so that an attacker cannot bank free lies by first building credit, and so that recovery does not require the owner to watch a log. Note the real cost of that recovery rule: an agreement is only recorded when a Poll later confirms a Sighting, so twenty agreements take twenty Polls of Series that Reader bookmarks — hours to days of real time, not twenty page views. Combined with the six-hour ceiling, a determined attacker buys three deferral windows and then waits days, and every step is named in the log. That is the intended price. **Clearing marks.** The owner clears a Reader's counters from the administrative page. This exists because the guard has one known false-positive mode: a Site changing its page shape can make a correct adapter read a wrong high number, marking an honest Reader. That failure is visible in the log — it will mark every Reader of that Site at once — but it needs a remedy that is not SQL. **New Readers start trusted.** Zero counters means trusted. The alternative — earning trust — makes the feature broken for a new guild member until they have done nothing wrong for a while, and the exposure it avoids is one hour of a wrong value on a solitary Series. **Why not consensus.** Truth discovery, weighted majority and reputation-by-comparison all need multiple independent sources per object to work at all; the standard survey states plainly that an object provided by very few sources cannot have its confidence evaluated. With two Readers, disagreement is a coin flip. The authoritative Poll is used as the oracle instead, which is the same role the gold-question audit plays in the crowdsourcing literature — except deterministic, because the ceiling guarantees the audit rather than sampling it. **Threat model, on record.** Registration is open to any member of the Discord guild; there is no invite table, by an earlier decision. The question this guards is not "would my friend lie" but "would any guild member lie". That is why the guard exists at all. ## Testing Decisions A good test here asserts two observable things: whether a Poll happened, and what the stored Latest Chapter is afterwards. It must not assert counter arithmetic through internal calls, nor how the deferral decision is represented in a row. **Everything lands at the existing Poller seam, with the store as the way in.** A Sighting is seeded through the store — the same write path the request handler uses — then a round is run against the injected fetcher and a frozen clock, and the assertion is whether that Series was fetched. This keeps the whole spec at the seam that already exists and avoids a combined harness that drives the router and the poller together. No new seam. The behaviours to cover: - A Sighting on a Series with one Bookmark defers the next Poll: the round does not fetch it. - A Sighting on a Series with two Bookmarks does not defer: the round fetches it, and the stored Latest Chapter was updated by the Sighting beforehand. - A Sighting updates the Latest Chapter, and leaves Progress and the Bookmark's ordering timestamp untouched, in both the shared and solitary cases. - A Series last Polled longer ago than the ceiling is Polled despite a recent Sighting. - A Poll that finds a lower chapter number than stored records a disagreement against the Reader who last raised it, and logs both numbers. - A Poll that finds a higher number records no disagreement. - A Poll that confirms the stored value records an agreement for that Reader. - A Reader at the disagreement threshold no longer defers: their Sighting on a solitary Series is followed by a Poll in the next round. - A marked Reader's Sighting still updates the Latest Chapter. - Twenty confirmations clear the marks and deferral resumes. - Clearing marks through the store restores deferral immediately. - A Series whose sole Bookmark is deleted, or joined by a second Bookmark, stops deferring from the next round. Prior art: the existing cooldown-across-passes test is the closest shape — seed state, freeze the clock, run a round, assert what was and was not fetched. The tests that assert a shared Series is fetched exactly once already establish how to set up multi-Reader Series. The dashboard's roster tests cover the clear-marks control at the router seam and are not repeated here. ## Out of Scope - Any cross-Reader agreement, voting, weighting or consensus scoring. Rejected on evidence; the decision record carries the citations. - Rate limiting the Sighting write path itself. Separate concern, and the existing body-size and validation limits already apply. - Letting a Sighting change Progress, favourite state, or Lifecycle bucket. - Any Reader-visible indication of their own reputation. The counters are the owner's operational view, not a score shown to Readers. - Automatic expiry of marks by elapsed time. Recovery is by confirmed agreements only, deliberately, so waiting is not a strategy. - Any change to the userscripts. They already send what a Sighting needs; this spec changes only what the backend does with it. - Extending deferral to Acquisition, which is a single read at a Series' creation and never queues behind a Lane. ## Further Notes The mechanism's saving is real but bounded by how the deferral is restricted: only solitary Series defer, so a Library whose Series are widely shared saves less. That is the correct trade — shared Series are exactly the ones where a Poll is amortised across several Readers anyway, so they are the cheapest Polls in the system per Reader served. This spec needs its own decision record, separate from the Poll Lanes one, covering the trust model, the six-hour ceiling, the three-and-twenty thresholds, the solitary-Series restriction, and the rejected alternatives with their citations. A future reader finding that a client report can postpone a server check will reasonably think it is madness; the record is where the answer lives. The administrative page must land before this work, so that a mark can be cleared as soon as one can be created.
sulthan added the enhancementready-for-agent labels 2026-08-16 11:28:27 +07:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: sulthan/mangaBookmark#103