58014eb8dd5c598279864eee68baebebdf99ce67
2 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
17ee0bd3f8 |
docs: correct the bot-score claims behind the browser poller (#99)
Docs only. No code changes - `git diff origin/main --stat` touches five Markdown files and adds one research note. ## What was wrong Several docs explained Cloudflare challenges as a "bot score" that our request rate could worsen. That mechanism does not exist on these sites. Researched live on 2026-08-12 against Cloudflare's own documentation and blog plus RFC 9309 - 22 primary pages, every claim carrying a source URL and read date, seven areas explicitly marked `Not publicly documented`. The note is `docs/research/cloudflare-bot-scoring-and-poll-cadence.md`. - The 1-99 bot score is **Enterprise Bot Management only**. A free-plan zone has no score at all; it gets Bot Fight Mode, which matches *signatures* (headless browsers, cloud-hosting IPs). - **No per-IP request rate is documented as an input to challenge issuance.** Volume is policed by Rate Limiting Rules, a separate opt-in product: one rule, IP-only counting, 10-second windows on Free. Published DDoS thresholds are ~1,000 errors/sec. - **`cf_clearance` defaults to 30 minutes**, so every cadence at or above 1 hour re-solves the challenge anyway. Cadence changes how many ~4s solves happen per day and nothing else. - The documented risk is **fingerprint quality**, which this repo already solved (real Chrome, stock UA, non-UTC clock). ## What changed | File | Correction | |---|---| | `AGENTS.md` | The block is per-zone configuration plus request fingerprint, not IP reputation. comix.to turning its gate on 2026-08-12 is the worked example. Residential egress avoids the cloud-hosting-IP *signature* rather than earning a better score. The UTC measurement stands; its mechanism is now marked undocumented. | | `backend/AGENTS.md` | Says why `_BROWSER_COOLDOWN` is longer: cost, not safety. | | `docs/adr/0003` | Dated correction - the sites do not "bot-score" the VPS IP. Decision stands on its sweep-depth argument. | | `docs/adr/0006` | Dated correction - no score to be better at. Decision stands on VPS memory. | | `DEPLOY.md` | A red kagane smoke run means the Site's settings or this Chrome's fingerprint moved, not "Cloudflare's scoring". | ADRs got dated `Corrected 2026-08-12:` paragraphs rather than silent rewrites - the record of what was decided stays intact, only the wrong mechanism is retracted. ## Deliberately not in this PR - **The 6h browser cooldown is unchanged.** I had lowered it to 1h and reverted that; cadence is a behaviour change and belongs with the comix work in #98, not in a docs correction. - **Two code comments still carry the myth**: `backend/main.go:83-84` ("a hammer against sites that are already bot-scoring us"). Left alone to keep this diff docs-only. Related: #98. Reviewed-on: #99 Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com> Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com> |
||
|
|
f1eb7d514c |
Record the lightnovelworld series-identity decision (#77) (#81)
Docs only. No code, no tests, nothing to run. Implementation is specified in #80. Outcome of a grilling session on 2026-08-11 against #77, backed by live measurement of lightnovelworld over 2026-08-10/11. ## What changed **`docs/adr/0008-series-identity-is-discovered-not-derived.md`** (new) A Series identity is discovered from the Site's own links, never derived from an address. On lightnovelworld the userscript reads the chapter page's `All Chapter` anchor instead of building a `/novel/<slug>/` address by string manipulation. A Chapter Slug is not an identity and is not stored. The backend's chapter scan drops its per-Series scoping and runs against the body truncated before the visitor comment thread. Evidence in the ADR: 3 of 41 sampled novels serve chapters under a slug that differs from their series slug, divergence runs in both directions, one novel serves chapters under two slugs, and neither slug is computable from the other. The pointer was checked on 8 chapter pages and agreed every time. Three narrower selectors are recorded as rejected, each with the measurement that killed it. Three rejected options are recorded with reasons: correcting the stored address only, which keeps an identity the Site does not guarantee; scoping the scan to a container, which the probe refuted; and a SQL migration, which is impossible because the database holds no source for the correct slug. **`CONTEXT.md`** - **Series** - identity is the canonical slug the Site publishes, never the title and never a Chapter Slug. - **Chapter Slug** - new term. A slug a Site builds its chapter addresses from. Not an identity: one Series may have several, and none is computable from another. - **Latest Chapter** - now the highest-numbered chapter, explicitly not a date and not the Site's own newest-chapter banner. Settles #79. **`docs/research/lightnovelworld-chapter-vs-series-slug.md`** (new, committed with its corrections) The 41-novel survey behind the ADR. Two claims are struck through and corrected in place, with the date and sample size of the probe that refuted each: the `ul.clstyle` container it named is the hidden, empty "Latest Reading" template rather than the chapter list, and its caveat about the comment region understated the risk, because that region is writable by any visitor while the scan takes an unbounded maximum into a Series row shared by every Reader (ADR-0003). ## Review notes Nothing here constrains code that exists today - the ADR describes work not yet written. The part worth disagreeing with, if any of it is wrong, is the fail-closed rule: a missing truncation marker means skip the Series and log, never scan the whole page. Related: #77 (the defect), #80 (the spec), #79 (the numbering anomaly, closed by decision), #71 (the same size cap seen from the cover side). Reviewed-on: #81 Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com> Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com> |