Commit Graph

3 Commits

Author SHA1 Message Date
sulthan ac50c428a9 fix: give covers their own 10 MiB byte cap (#71)
The cover fetch reused maxBodyBytes, the 4 MiB ceiling sized for series
pages, so any cover above it was rejected, logged, and retried forever
while the Series kept a monogram. Measured against asurascans on
2026-08-17 that is not an edge case: p90 is 4.52 MB and 3 of 25 covers
exceed 4 MiB, two of them plain JPEGs rather than the 8.57 MB animated
GIF the issue names.

Covers now have maxCoverBytes = 10 MiB, separate from the page cap: a
cover is one bounded binary asset, the page cap still has 3.5x headroom
over measured pages and should not be loosened along with it. 10 MiB is
~18% over the largest cover observed and matches the GitHub and Discord
image limits. The format offers no help in picking the number — GIF has
no maximum size at all — so docs/research/gif-maximum-byte-size.md
records the spec reading, the decoder behaviour, and the live
distribution the cap is derived from.

Transcoding was rejected: decoding is the OOM path, since Go's
image/gif allocates width x height per frame with no dimension guard
and a legal 65535^2 GIF would ask for ~4.29 GB on a 1974 MiB swapless
host.
2026-08-17 12:05:42 +07:00
sulthan 17ee0bd3f8 docs: correct the bot-score claims behind the browser poller (#99)
Docs only. No code changes - `git diff origin/main --stat` touches five Markdown files and adds one research note.

## What was wrong

Several docs explained Cloudflare challenges as a "bot score" that our request rate could worsen. That mechanism does not exist on these sites.

Researched live on 2026-08-12 against Cloudflare's own documentation and blog plus RFC 9309 - 22 primary pages, every claim carrying a source URL and read date, seven areas explicitly marked `Not publicly documented`. The note is `docs/research/cloudflare-bot-scoring-and-poll-cadence.md`.

- The 1-99 bot score is **Enterprise Bot Management only**. A free-plan zone has no score at all; it gets Bot Fight Mode, which matches *signatures* (headless browsers, cloud-hosting IPs).
- **No per-IP request rate is documented as an input to challenge issuance.** Volume is policed by Rate Limiting Rules, a separate opt-in product: one rule, IP-only counting, 10-second windows on Free. Published DDoS thresholds are ~1,000 errors/sec.
- **`cf_clearance` defaults to 30 minutes**, so every cadence at or above 1 hour re-solves the challenge anyway. Cadence changes how many ~4s solves happen per day and nothing else.
- The documented risk is **fingerprint quality**, which this repo already solved (real Chrome, stock UA, non-UTC clock).

## What changed

| File | Correction |
|---|---|
| `AGENTS.md` | The block is per-zone configuration plus request fingerprint, not IP reputation. comix.to turning its gate on 2026-08-12 is the worked example. Residential egress avoids the cloud-hosting-IP *signature* rather than earning a better score. The UTC measurement stands; its mechanism is now marked undocumented. |
| `backend/AGENTS.md` | Says why `_BROWSER_COOLDOWN` is longer: cost, not safety. |
| `docs/adr/0003` | Dated correction - the sites do not "bot-score" the VPS IP. Decision stands on its sweep-depth argument. |
| `docs/adr/0006` | Dated correction - no score to be better at. Decision stands on VPS memory. |
| `DEPLOY.md` | A red kagane smoke run means the Site's settings or this Chrome's fingerprint moved, not "Cloudflare's scoring". |

ADRs got dated `Corrected 2026-08-12:` paragraphs rather than silent rewrites - the record of what was decided stays intact, only the wrong mechanism is retracted.

## Deliberately not in this PR

- **The 6h browser cooldown is unchanged.** I had lowered it to 1h and reverted that; cadence is a behaviour change and belongs with the comix work in #98, not in a docs correction.
- **Two code comments still carry the myth**: `backend/main.go:83-84` ("a hammer against sites that are already bot-scoring us"). Left alone to keep this diff docs-only.

Related: #98.
Reviewed-on: #99
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-12 09:32:07 +07:00
sulthan f1eb7d514c Record the lightnovelworld series-identity decision (#77) (#81)
Docs only. No code, no tests, nothing to run. Implementation is specified in #80.

Outcome of a grilling session on 2026-08-11 against #77, backed by live measurement of lightnovelworld over 2026-08-10/11.

## What changed

**`docs/adr/0008-series-identity-is-discovered-not-derived.md`** (new)

A Series identity is discovered from the Site's own links, never derived from an address.
On lightnovelworld the userscript reads the chapter page's `All Chapter` anchor instead of
building a `/novel/<slug>/` address by string manipulation. A Chapter Slug is not an
identity and is not stored. The backend's chapter scan drops its per-Series scoping and
runs against the body truncated before the visitor comment thread.

Evidence in the ADR: 3 of 41 sampled novels serve chapters under a slug that differs from
their series slug, divergence runs in both directions, one novel serves chapters under two
slugs, and neither slug is computable from the other. The pointer was checked on 8 chapter
pages and agreed every time. Three narrower selectors are recorded as rejected, each with
the measurement that killed it.

Three rejected options are recorded with reasons: correcting the stored address only, which
keeps an identity the Site does not guarantee; scoping the scan to a container, which the
probe refuted; and a SQL migration, which is impossible because the database holds no
source for the correct slug.

**`CONTEXT.md`**

- **Series** - identity is the canonical slug the Site publishes, never the title and never a Chapter Slug.
- **Chapter Slug** - new term. A slug a Site builds its chapter addresses from. Not an identity: one Series may have several, and none is computable from another.
- **Latest Chapter** - now the highest-numbered chapter, explicitly not a date and not the Site's own newest-chapter banner. Settles #79.

**`docs/research/lightnovelworld-chapter-vs-series-slug.md`** (new, committed with its corrections)

The 41-novel survey behind the ADR. Two claims are struck through and corrected in place,
with the date and sample size of the probe that refuted each: the `ul.clstyle` container it
named is the hidden, empty "Latest Reading" template rather than the chapter list, and its
caveat about the comment region understated the risk, because that region is writable by
any visitor while the scan takes an unbounded maximum into a Series row shared by every
Reader (ADR-0003).

## Review notes

Nothing here constrains code that exists today - the ADR describes work not yet written.
The part worth disagreeing with, if any of it is wrong, is the fail-closed rule: a missing
truncation marker means skip the Series and log, never scan the whole page.

Related: #77 (the defect), #80 (the spec), #79 (the numbering anomaly, closed by decision),
#71 (the same size cap seen from the cover side).

Reviewed-on: #81
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-11 09:34:26 +07:00