Compare commits

...

1 Commits

Author SHA1 Message Date
sulthan 1f5d0695ad docs: correct the bot-score claims behind the browser poller
Every doc statement that explained a Cloudflare challenge as a "score"
was wrong. Researched against Cloudflare's own docs on 2026-08-12
(docs/research/cloudflare-bot-scoring-and-poll-cadence.md, 22 primary
pages plus RFC 9309): the 1-99 bot score is Enterprise Bot Management
only, free-plan zones get Bot Fight Mode signature matching and no score
at all, and no per-IP request rate is documented as an input to
challenge issuance. cf_clearance also expires in 30 minutes, so every
cadence at or above 1h re-solves the challenge regardless.

Docs only - no behaviour change. The 6h browser cooldown stays; its
justification is now cost (a serialized single-tab solve costs seconds,
a plain read costs one request), not a risk reduction nothing documents.

- AGENTS.md: the block is per-zone configuration plus request
  fingerprint, not IP reputation; comix.to turning its gate on
  2026-08-12 is the example. Residential egress avoids the
  cloud-hosting-IP signature rather than earning a better score. The UTC
  measurement stands but its mechanism is marked undocumented.
- backend/AGENTS.md: states why the browser cooldown is longer.
- ADR-0003, ADR-0006: dated corrections rather than rewrites. Both
  decisions stand on their other arguments (sweep depth, VPS memory).
- DEPLOY.md: a red smoke run means the Site's settings or this Chrome's
  fingerprint moved, not that "Cloudflare's scoring" did.
2026-08-12 09:29:27 +07:00
6 changed files with 275 additions and 10 deletions
+3 -3
View File
@@ -18,11 +18,11 @@ Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free
- Cross-origin `fetch()` work **only** against CORS-enabled backend. Manga sites `https://`, so backend **must be HTTPS** (else mixed-content block).
- Every site is its **own origin with its own `localStorage`** — a shared remote store is the only way to unify bookmarks. Cloud sync required, not optional.
- Userscript run in **isolated world**, so embedded API token safe from site's JS.
- Cloudflare's block on manga sites **IP-reputation-based, not universal — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — Cloudflare's bot scoring can flip previously-clean IP without notice. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test.
- Cloudflare's block on manga sites is **per-zone configuration plus request fingerprint, not IP reputation — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — a Site can turn its protection on overnight, which is exactly what comix.to did on 2026-08-12. An earlier version of this line blamed "Cloudflare's bot scoring"; that was wrong. The 1-99 bot score is Enterprise Bot Management only and does not exist for a free-plan zone, and no per-IP request rate is documented as an input to challenge issuance — `docs/research/cloudflare-bot-scoring-and-poll-cadence.md`. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test.
- **kagane.to and novelfull.com are the exception to the above** — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`). When that's unset, kagane is skipped entirely (a plain fetch would only retrieve a challenge page) while novelfull pages are still attempted over plain TLS — its challenge is a live time-varying fact and its cover bytes never need the browser. The four other sites poll fine over plain TLS.
- **The CDP browser must look like a real browser, and stock headless images don't.** Measured 2026-08-08 against kagane.to, all from the same IP: `chromedp/headless-shell:stable` never cleared the challenge in 90s (`navigator.webdriver` true, empty plugin list, Chromium-branded client hints — suppressing `webdriver` alone changed nothing); `zenika/alpine-chrome` ships Chrome 124, refused outright; real Chrome with the default `--headless=new` UA never cleared, because the UA says `HeadlessChrome`; real Chrome with a stock UA **and** a non-UTC clock zone cleared in ~4s. Hence `chrome/` — a Debian image with `google-chrome-stable`, a version-derived UA, and `TZ`/`BROWSER_TZ`. Chrome reads the zone *name* through ICU from `/etc/localtime`'s symlink target, ignoring the file's contents, so mounting the host's `/etc/localtime` does **not** work; `/etc/timezone` is mounted instead.
- **The browser is not in the API stack and must not be put back.** It's its own compose unit (`chrome/docker-compose.yml`) on a second machine, reached over the tailnet — it held 471 MiB on a 1974 MiB swapless VPS, and a residential egress scores better with Cloudflare anyway (ADR-0006). Consequences that constrain code: `BROWSER_WS_URL` must be a tailnet **IP** (a MagicDNS name 500s at `/json/version`, same trap as the old Docker service name); the CDP port binds to the tailnet address only, since CDP authenticates nothing and that host has a real LAN; and the browser is on-demand (ADR-0005), so an unreachable or asleep one must degrade exactly as an unset `BROWSER_WS_URL` — plain-TLS libraries unaffected, kagane/novelfull logged and skipped, stored covers still served. Never add `chromedp.NoModifyURL`: discovery per fetch is what makes a restarted Chrome invisible.
- **UTC is the tell, not a country mismatch.** A UTC clock is the datacenter default, so Cloudflare scores it as one; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while `Asia/Jakarta` **and** `America/New_York` both cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (`Asia/Bangkok`) rather than the measured egress. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one.
- **The browser is not in the API stack and must not be put back.** It's its own compose unit (`chrome/docker-compose.yml`) on a second machine, reached over the tailnet — it held 471 MiB on a 1974 MiB swapless VPS, and a residential egress avoids the cloud-hosting-IP signature Bot Fight Mode documentedly challenges (ADR-0006; not a better "score" — free-plan zones have no score). Consequences that constrain code: `BROWSER_WS_URL` must be a tailnet **IP** (a MagicDNS name 500s at `/json/version`, same trap as the old Docker service name); the CDP port binds to the tailnet address only, since CDP authenticates nothing and that host has a real LAN; and the browser is on-demand (ADR-0005), so an unreachable or asleep one must degrade exactly as an unset `BROWSER_WS_URL` — plain-TLS libraries unaffected, kagane/novelfull logged and skipped, stored covers still served. Never add `chromedp.NoModifyURL`: discovery per fetch is what makes a restarted Chrome invisible.
- **UTC is the tell, not a country mismatch.** A UTC clock is the datacenter default, and the challenge refuses it; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while `Asia/Jakarta` **and** `America/New_York` both cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (`Asia/Bangkok`) rather than the measured egress. A second earlier claim, that Cloudflare "scores" a UTC clock, was also wrong: the measurement is real but the mechanism is not documented anywhere — Cloudflare publishes no timezone signal, and free-plan zones carry no score at all. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one.
- **A challenged page needs the tab kept open.** The interstitial takes seconds to solve and only then writes clearance into the browser's shared cookie jar. Navigate-read-close never clears anything; `BrowserFetcher.run` holds one tab and re-reads until the payload arrives.
## Architecture
+4 -2
View File
@@ -275,7 +275,8 @@ copy immediately — reinstall on all devices, or they silently stop syncing.
Kagane and novelfull sit behind a Cloudflare JavaScript challenge no TLS
fingerprint clears, so the poller reaches them through a real Chrome over CDP.
That browser does **not** run on the VPS: it held 471 MiB of a 1974 MiB box
with no swap, and it scores better from a residential IP anyway (ADR-0006). It
with no swap, and a residential IP avoids the cloud-hosting-IP signature Bot
Fight Mode challenges anyway (ADR-0006). It
is its own compose unit, deployed and updated independently of everything
above.
@@ -454,7 +455,8 @@ SMOKE_BROWSER_WS_URL=ws://100.x.y.z:9222 go test -run TestSmokeKagane ./internal
```
A red run means "not clearing from this address right now", which is a live
fact to re-check before it is a defect — Cloudflare's scoring moves. Then, from
fact to re-check before it is a defect — a Site's Cloudflare settings, and the
fingerprint this Chrome presents after an update, both move. Then, from
the web UI, open a bookmarked kagane series and confirm the cover renders. Once
a cover is stored it is served from Postgres forever after, so the browser being
asleep, unreachable, or mid-power-outage costs chapter freshness and nothing
+7 -1
View File
@@ -149,7 +149,13 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
`LATEST_CHAPTER_POLL_ENABLED`/`_COOLDOWN`/`_BROWSER_COOLDOWN`/`_INTERVAL`/
`_BATCH`/`_STAGGER` (background latest-chapter poller; defaults on,
`1h` plain-TLS cooldown, `6h` browser cooldown, `10m`/`14`/`20s`; both
cooldowns have a `15m` floor).
cooldowns have a `15m` floor). The browser cooldown is longer for cost, not
for safety: a challenged page costs seconds of a serialized single-tab
browser, while a plain read costs one request. It buys no documented
reduction in challenge risk — free-plan zones have no bot score and no
published per-IP rate input, and `cf_clearance` expires in 30 minutes so
every cadence at or above 1h re-solves anyway —
`docs/research/cloudflare-bot-scoring-and-poll-cadence.md`.
`USERSCRIPT_PATH` and `NOVEL_USERSCRIPT_PATH` (files served at
`/u/{token}/manga-bookmark.user.js` and `/u/{token}/novel-bookmark.user.js`,
defaults `/userscript/manga-bookmark.user.js` and
@@ -19,8 +19,15 @@ and because the Series row now knows how many Readers hold it, the poll queue is
absorbs the shortfall. That ordering is only expressible because the split happened.
Raising throughput instead was rejected: sweeping 400 Series hourly needs the stagger
cut from 20s to ~9s, doubling request rate against sites that already bot-score the
single VPS IP.
cut from 20s to ~9s, doubling request rate against sites already fronted by Cloudflare
from the single VPS IP.
Corrected 2026-08-12: the original wording said those sites "bot-score" the VPS IP.
They do not — the 1-99 bot score is Enterprise Bot Management only, and no per-IP
request rate is documented as an input to challenge issuance
(`docs/research/cloudflare-bot-scoring-and-poll-cadence.md`). The decision stands on
its first argument, sweep depth versus the 1-hour cooldown; the rate-limit fear was
never evidenced.
## Only the Poll writes Series fields
+9 -2
View File
@@ -23,8 +23,15 @@ requests: the poller's due query joins bookmarks, production held four kagane
series and no bookmarks on any of them, and with no kagane bookmark the web UI
never rendered a kagane cover either.
The home machine has 5.9 GiB of swap and a residential egress, which Cloudflare
scores better than a datacenter IP. Both machines were already on the tailnet.
The home machine has 5.9 GiB of swap and a residential egress, which avoids the
cloud-hosting-IP signature Cloudflare's Bot Fight Mode documentedly challenges. Both
machines were already on the tailnet.
Corrected 2026-08-12: the original wording said Cloudflare "scores" a residential
egress better than a datacenter IP. There is no score on a free-plan zone; what is
documented is signature matching, and hosting-provider IP space is one of the
signatures (`docs/research/cloudflare-bot-scoring-and-poll-cadence.md`). Memory was
the load-bearing reason regardless.
This move is only safe because covers are persisted (ADR-0005's sibling work,
issue #43/#45) and the browser is on-demand (ADR-0005). Without stored covers a
@@ -0,0 +1,243 @@
# Cloudflare bot scoring and poll cadence — what is actually documented
Research note for the browser-backed poller cadence decision (kagane.to, novelfull.com, comix.to). All pages were fetched live from **developers.cloudflare.com / blog.cloudflare.com on 2026-08-12**. Primary sources only: Cloudflare's own documentation, Cloudflare blog posts, and RFCs/standards where noted. Where Cloudflare does not publicly document something, this note says **`Not publicly documented`** instead of guessing. Repo-measured facts from the existing poller work are reused without re-derivation and marked as such.
The site configuration of the three challenged sites (which plan, which bot product, Challenge Passage setting, whether Precursor is enabled) is **not observable from outside** — Cloudflare does not expose a zone's security configuration to anonymous clients. Anything in this note that depends on those unknowns is flagged `[INFERENCE]`.
---
## Short answer
**No — polling once per hour per series, from one residential IP through one real Chrome holding a valid `cf_clearance`, carries no documented challenge risk beyond polling every six hours.** Challenge issuance on Free/Pro-grade protection (Bot Fight Mode, WAF rules) is signature- and fingerprint-driven (headless browsers, cloud-hosted IPs, browser signals); the only rate-aware detector — the per-request bot score — exists solely on Enterprise Bot Management, and free-plan Rate Limiting Rules count per-IP over 10-second windows, which 20–60 requests/hour cannot trip. Both cadences re-solve the challenge every visit anyway: `cf_clearance` expires after **30 minutes by default** (site-configurable), so a 1-hour gap always finds it expired. The documented lever that matters — already verified in this repo — is **fingerprint quality**: real Chrome + real timezone clears in ~4 s; headless variants never do. Residual, undocumented risk is site-specific: Challenge Passage, Precursor (behavior-bound re-challenge), and custom WAF rules are zone settings not observable from outside.
---
## Summary answer table
| Question | Answer | Section |
|---|---|---|
| What is a bot score? | Integer 1–99 = Cloudflare's certainty a request is automated; **Enterprise Bot Management only**; everyone else gets coarse "bot groupings" (Pro+ analytics) or nothing. | §1 |
| Is request frequency documented as a bot-score input? | Partially: ML inputs are "headers, session characteristics, and browser signals"; the `__cf_bm` cookie "measures a single user's request pattern". **No numeric rate threshold is documented.** Volume policing is Rate Limiting, a separate product. | §1, §5 |
| What can a free-plan site deploy? | Bot Fight Mode only: challenges *signatures* (headless browsers, cloud-hosting IPs) with a computational challenge; JavaScript Detections forced on; no scores, no analytics, cannot be skipped/customized. | §2 |
| Does the free tier score continuously? | **No.** No score exists on Free at all — granular scores need Enterprise Bot Management, groupings need Pro+. BFM just challenges signature matches. | §2 |
| What does `cf-mitigated: challenge` mean? | The response was a Cloudflare Challenge Page (any type); `challenge` is the only value; body is always `text/html`. | §3 |
| How is a successful solve remembered? | `cf_clearance` cookie, issued with `SameSite=None; Secure; Partitioned`; suppresses challenges while valid. | §3, §4 |
| `cf_clearance` lifetime? | **30 minutes by default**, configurable by the site via Challenge Passage (15–45 min recommended); +skew minutes; +1 h for XHR. | §4 |
| Is `cf_clearance` bound to IP / device? | Documented: "securely tied to the specific visitor and device it was issued to"; the *solve request* must come from the same IP that received the challenge (different IP → invalid solve → challenge loop). Replay from another machine/IP is therefore **not** valid. | §4 |
| What invalidates clearance early? | Precursor (if enabled): suspicious session → clearance reduced/invalidated, re-challenge even before expiry. Zone-level toggle; unobservable from outside. | §4 |
| Does polling more often raise challenge risk? | **No documented mechanism at 20–60 req/hour.** Free-plan rate limiting is 10 s/IP-only; DDoS thresholds are ~1,000 errors/sec. Scores (the only rate-aware thing) are Enterprise-only. | §5 |
| Is there a documented "legitimate poller" path? | Yes, but it requires **self-identification** (Web Bot Auth signature or published IP list + stable UA) via the verified-bots application — not anonymity. robots.txt is voluntary; nothing exempts anonymous scrapers. | §6 |
---
## 1. What a bot score is and what feeds it
### 1.1 The score itself
Cloudflare documents the bot score as "a score from _1_ to _99_ that indicates how likely that request came from a bot" — 1 = quite certain automated, 99 = quite certain human. Source: [Bot scores — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot-score/), read 2026-08-12.
Two access tiers, both gated:
- **Granular 1–99 scores are only available to Enterprise customers who purchased Bot Management.** "All other customers can only access this information through bot groupings in Bot Analytics" (categories: `Not computed` = 0, `Automated` = 1, `Likely automated` = 2–29, `Likely human` = 30–99, `Verified bot`). Bot groupings themselves require "a Pro plan or higher". Source: [Bot scores — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot-score/), read 2026-08-12.
- A score of 0 means "Bot Management did not evaluate the request" (redirected, handled by another feature) — "does not indicate the request is safe or human". Same source.
So: **on a Free-plan site there is no bot score at all, for anyone.** `[INFERENCE]` the three challenged sites are almost certainly not Enterprise Bot Management customers, but this is not externally verifiable.
### 1.2 The detection engines (Enterprise Bot Management)
Cloudflare documents four engines, all stated to apply to Enterprise Bot Management (the bot-score page: "The following detection engines only apply to Enterprise Bot Management"). Sources: [Bot scores — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot-score/) and [Bot detection engines — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot-detection-engines/), both read 2026-08-12.
| Engine | Documented behavior | Score it produces |
|---|---|---|
| **Heuristics** | "Processes all requests"; pattern matching against "a growing database of malicious fingerprints". | 1 for high-confidence deterministic detections; occasionally 29 "where Cloudflare has identified automated traffic and is still assessing traffic overlap" |
| **Machine learning** | Supervised model, trained on "billions of daily requests". Input variables: "headers, session characteristics, and browser signals". Output: "predicted probability that a client is human (such as the probability of successfully solving a Challenge)". | Most scores 2–99 |
| **Anomaly detection** | Unsupervised; learns a per-domain baseline, flags outlier requests; **deprecated, not onboarding new customers**. | 1 |
| **JavaScript detections** | "Identifies headless browsers and other automation tools" via "a lightweight, invisible JavaScript injection"; runs client-side; "blocks, challenges, or passes requests to other engines". Enabled by default (but optional) in Bot Management. | Pass/fail (`cf.bot_management.js_detection.passed`), not a score |
Crucially, the ML engine's documented inputs are *headers, session characteristics, and browser signals* — **no rate or per-IP volume parameter is listed.** The only place request patterns appear is the `__cf_bm` cookie note: "Cloudflare uses the `__cf_bm` cookie to smooth out the bot score and reduce false positives… The Bot Management cookie measures a single user's request pattern and applies it to the machine learning data to generate a reliable bot score for all of that user's requests." ([Bot scores — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot-score/), read 2026-08-12). So frequency is *a* signal inside Enterprise Bot Management via the session cookie — **but no numeric threshold, window, or per-IP rate is published anywhere.** `Not publicly documented`: any specific requests-per-hour / requests-per-IP value that raises or lowers a bot score.
### 1.3 Rate limiting is a separate product
Volume enforcement is not part of bot scoring at all. Rate Limiting Rules are a distinct WAF product with their own evaluation phase (`http_ratelimit`, running after custom rules and before SBFM). Sources: [Rate limiting rules — Cloudflare docs](https://developers.cloudflare.com/waf/rate-limiting-rules/) and [Security features interoperability — Cloudflare docs](https://developers.cloudflare.com/waf/feature-interoperability/), read 2026-08-12. See §5 for what the Free plan's version of that product can actually do.
---
## 2. The free-plan reality
### 2.1 What each plan gets
Cloudflare's plan table ([Plans — Cloudflare docs](https://developers.cloudflare.com/bots/plans/), read 2026-08-12):
| Plan | Bot product | Documented detections | Action | Control |
|---|---|---|---|---|
| **Free** | **Bot Fight Mode** (BFM) | "Simple bots from cloud hosting providers and headless browsers" | "Cloudflare issues a computationally expensive challenge" | Applied to all traffic across the domain; no exceptions possible |
| **Pro / Business / Enterprise (no BM)** | **Super Bot Fight Mode** (SBFM) | Configurable actions per bot category (Definitely automated / Likely automated / Verified bots) | Challenge or block | Runs on Ruleset Engine; **can** be skipped via custom rules |
| **Enterprise + Bot Management** | Bot Management | "Simple and sophisticated bots, headless browsers, and domain-specific anomalies" | Customer-chosen (block, challenges) | Per-path / per-IP rules; access to bot score, JA3/JA4, bot tags, detection IDs |
Sources: [Plans — Free](https://developers.cloudflare.com/bots/plans/free/), [Plans — Bot Management for Enterprise](https://developers.cloudflare.com/bots/plans/bm-subscription/), [Bot Fight Mode — Cloudflare docs](https://developers.cloudflare.com/bots/get-started/bot-fight-mode/), [Super Bot Fight Mode — Cloudflare docs](https://developers.cloudflare.com/bots/get-started/super-bot-fight-mode/) — all read 2026-08-12.
### 2.2 Bot Fight Mode specifics (the Free-plan product)
- Identifies "traffic matching patterns of known bots" and "issues computationally expensive challenges that force the requesting client to perform CPU-intensive calculations". ([Bot Fight Mode — Cloudflare docs](https://developers.cloudflare.com/bots/get-started/bot-fight-mode/), read 2026-08-12.)
- It "does not run on the Ruleset Engine — it operates in a separate evaluation pipeline where _Skip_, _Bypass_, and _Allow_ actions have no effect"; **you cannot bypass or skip BFM** with custom rules or Page Rules. ([Bot Fight Mode — Cloudflare docs](https://developers.cloudflare.com/bots/get-started/bot-fight-mode/) and [Security features interoperability — Cloudflare docs](https://developers.cloudflare.com/waf/feature-interoperability/), read 2026-08-12.)
- **JavaScript Detections is automatically enabled for BFM customers and cannot be disabled.** ([Bot Fight Mode — Cloudflare docs](https://developers.cloudflare.com/bots/get-started/bot-fight-mode/), read 2026-08-12.) This is the documented hook that explains the repo's `headless-shell` / `HeadlessChrome` failures: JSD "identifies headless browsers" ([JavaScript detections — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/javascript-detections/), read 2026-08-12).
- False positives on *legitimate automated* traffic are acknowledged as expected behavior: "false positives can occur where legitimate human or automated traffic is incorrectly challenged or blocked", and the only remedies are disabling BFM or upgrading to Bot Management. ([Handle False Positives — Cloudflare docs](https://developers.cloudflare.com/bots/troubleshooting/false-positives/), read 2026-08-12.)
### 2.3 Does the free tier "score" continuously?
**No.** The Free plan exposes no score, no bot analytics (groupings need Pro+), and no per-request decision data. BFM is a static on/off toggle that challenges signature matches; there is no continuous per-request score on Free. ([Plans — Free](https://developers.cloudflare.com/bots/plans/free/) and [Bot scores — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot-score/), read 2026-08-12.) JSD *runs* on every HTML request even on Free (forced by BFM), but the only documented way to act on its result — the `cf.bot_management.js_detection.passed` field — is gated behind an Enterprise Bot Management subscription ("Prerequisites: You must have an Enterprise Bot Management subscription"). ([JavaScript detections — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/javascript-detections/), read 2026-08-12.) On Free, JSD output feeds BFM's internal challenge decision; `Not publicly documented` exactly how.
`[INFERENCE]` The observed behavior on the three sites (real Chrome + real timezone passes in ~4 s; headless variants never do) is consistent with either BFM or a WAF custom rule using a challenge action — Cloudflare does not expose which product a site runs, and the failure signature (HeadlessChrome UA / headless-shell never passing) matches JSD's documented headless-browser detection either way.
---
## 3. Managed Challenge / JS challenge mechanics and `cf-mitigated`
### 3.1 What the observed response is
A `403` with `cf-mitigated: challenge`, `server: cloudflare`, and a "Just a moment…" body is a **Cloudflare Challenge Page**. Cloudflare documents: "the Challenge Page response (regardless of the Challenge Page type) will have the `cf-mitigated` header present and set to `challenge`… `challenge` is the only valid value. The header is set for all Challenge Page types", and "the content-type of a challenge will be `text/html`". ([Detect a Challenge Page response — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/detect-response/), read 2026-08-12.)
### 3.2 What a Challenge Page does
"An interstitial Challenge Page… acts as a gate between the visitor and your website… The Challenge Page intercepts the visitor… by holding the request and evaluating the browser environment for automated signals, and serving a challenge. The visitor cannot reach their destination without passing the challenge." ([Interstitial Challenge Pages — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/), read 2026-08-12.)
Three variants, in increasing severity ([Interstitial Challenge Pages — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/), read 2026-08-12):
- **Non-Interactive**: Cloudflare judges automation from browser signals gathered by injected JS; the page needs no human interaction, typically < 5 s of JS processing.
- **Managed Challenge**: "Cloudflare dynamically chooses the appropriate type of challenge… based on the characteristics of a request from the signals indicated by their browser. Most human visitors are automatically verified and the Challenge Page will display **Successful**. However, if Cloudflare detects non-human attributes… they may be required to interact." Cloudflare's stated recommendation for WAF rules.
- **Interactive**: requires explicit human interaction (CAPTCHA-style). Cloudflare's "End the CAPTCHA era" position is that Managed Challenges should make this rare.
Cloudflare's own framing matches the repo's ~4 s real-Chrome solve: a normal browser passes with no interaction (Managed Challenge auto-verify or Non-Interactive JS processing).
### 3.3 Which product issues which challenge
Documented mapping ([How Challenges work — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/concepts/how-challenges-work/), read 2026-08-12):
| Trigger | Challenge type |
|---|---|
| WAF custom rules, **rate limiting rules**, IP Access rules | Interstitial Challenge Page |
| Bot Management | JavaScript Detections (invisible, per-request) |
| **Bot Fight Mode / Super Bot Fight Mode** | Interstitial Challenge Page |
| Under Attack Mode | Managed Challenge |
### 3.4 How a successful solve is remembered
Solving issues the **`cf_clearance`** cookie: "Clearance Cookie stores the proof of challenge passed. It is used to no longer issue a challenge if present. It is required to reach an origin server." ([Cloudflare Cookies — Cloudflare docs](https://developers.cloudflare.com/fundamentals/reference/policies-compliances/cloudflare-cookies/), read 2026-08-12.) "When that visitor tries to access other parts of your website, Cloudflare evaluates the cookie before presenting another challenge. If the cookie is still valid, no challenges will be shown." ([Challenge Passage — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/challenge-passage/), read 2026-08-12.)
`cf_clearance` is set with `SameSite=None; Secure; Partitioned`; because of the `Partitioned` (CHIPS) attribute, "a clearance obtained in one top-level context is not reused in a different top-level context" — so a clearance from kagane.to does not carry to comix.to even on the same browser. ([SameSite cookie interaction — Cloudflare docs](https://developers.cloudflare.com/waf/troubleshooting/samesite-cookie-interaction/), read 2026-08-12.)
---
## 4. `cf_clearance` — lifetime, binding, invalidation (the key question)
### 4.1 Lifetime
- **Default: 30 minutes.** "By default, the `cf_clearance` cookie has a lifetime of 30 minutes. Cloudflare recommends a setting between 15 and 45 minutes." The site owner can change it via the **Challenge Passage** setting. ([Challenge Passage — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/challenge-passage/), read 2026-08-12; also [SameSite cookie interaction — Cloudflare docs](https://developers.cloudflare.com/waf/troubleshooting/samesite-cookie-interaction/), read 2026-08-12.)
- Validation grace: "a few extra minutes are included to account for clock skew. For XmlHTTP requests, an extra hour is added to the validation time." ([Challenge Passage — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/challenge-passage/), read 2026-08-12.)
- "The Challenge Passage does not apply to rate limiting rules." Same source.
- `Not publicly documented`: whether the three target sites have changed the Challenge Passage from the 30-minute default, and whether there is any maximum value Cloudflare enforces.
**Consequence for cadence:** with the default 30-minute TTL, *any* poll cadence ≥ 1 hour finds the cookie expired and re-solves the challenge on every visit. A 1-hour and a 6-hour cadence therefore differ only in *how many times per day* the browser re-solves (~4× for the same series), not in whether a re-solve happens. This repo already measured the re-solve cost: ~4 s with real Chrome + real timezone. `[INFERENCE]` a site could raise the Challenge Passage to hours/days, which would make a 1-hour cadence *cheaper* (cookie still valid, no re-solve) — but that setting is unobservable and unlikely to be long on free manga sites.
### 4.2 What it is bound to — can clearance be replayed from another IP?
Documented statements, both from Cloudflare's own docs:
1. **Device/visitor binding:** "The cookie is securely tied to the specific visitor and device it was issued to, preventing reuse across machines." ([Clearance — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/concepts/clearance/), read 2026-08-12.)
2. **IP binding of the solve:** under Challenge limitations, Cloudflare lists "Client software where the solve request of a Managed Challenge comes from a different IP than the original IP a Challenge request was issued to. For example, if you receive the Challenge from one IP and solve it using another IP, the solve is not valid and you may encounter a Challenge loop." ([How Challenges work — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/concepts/how-challenges-work/), read 2026-08-12.)
3. **Top-level-site binding** via CHIPS partitioning: clearance is not reused across embedding contexts. ([SameSite cookie interaction — Cloudflare docs](https://developers.cloudflare.com/waf/troubleshooting/samesite-cookie-interaction/), read 2026-08-12.)
So the repo's belief is **verified by primary sources**: a `cf_clearance` obtained on one machine cannot be replayed from a different IP — the solve is IP-bound and the cookie is device-bound. `Not publicly documented`: whether the cookie value is also cryptographically bound to the User-Agent or TLS/JA3 fingerprint. The only UA-adjacent documented statement is the reverse direction: challenge *solving* breaks when a browser extension modifies the User-Agent or Canvas/WebGL APIs ("Cloudflare Challenges cannot support… Browser extensions that modify the browser's User-Agent value or Web APIs such as Canvas and WebGL") — i.e., tampering with browser signals is documented to *fail* challenges ([How Challenges work — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/concepts/how-challenges-work/), read 2026-08-12).
### 4.3 Two-tier clearance, and the behavior-bound invalidation (Precursor)
`cf_clearance` now carries two kinds of clearance ([Clearance — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/concepts/clearance/), read 2026-08-12):
- **Challenge clearance** — granted by solving a challenge; level-gated (Interactive > Managed > Non-Interactive; higher clears bypass lower challenges); "remains valid for the duration configured by the customer (Challenge Passage), **unless Precursor determines the session is suspicious**".
- **Precursor clearance** — "continuously updated based on session behavior"; an ongoing client-side process that periodically reassesses behavior. "If Precursor determines that a session is suspicious: the visitor's effective Challenge clearance may be **reduced or invalidated**; the visitor may be **re-challenged, even if the cookie has not expired**."
Precursor is documented as "client-side, session-based verification that continuously evaluates visitor behavior to identify automation… to detect automation that appears legitimate in individual requests but exhibits non-human patterns across a session", writing session state back into `cf_clearance`. It is a **zone-level toggle** (Security → Settings → Precursor; modes: Minimize Friction default, Maximize Security recommended), and "Precursor supersedes JavaScript Detections (JSD)". ([Precursor — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/precursor/), read 2026-08-12.)
**Implication:** the only documented mechanism by which *behavior over time* (as opposed to a single request's fingerprint) can revoke a valid clearance is Precursor — and it is opt-in per zone. `[INFERENCE]` It is unlikely to be enabled on free manga sites, but this is not observable from outside. With Precursor off (the default posture `[INFERENCE]`), a valid `cf_clearance` is honored purely on TTL + device/IP binding.
---
## 5. Does polling more often raise challenge risk?
### 5.1 Is per-IP request rate an input to challenge issuance? (documented answer: no such lever on non-Enterprise protection)
- **Bot scores** (the only per-request automated-detection output) are Enterprise-Bot-Management-only (§1.1); the ML engine's documented inputs are headers/session/browser signals, with request *pattern* entering only via `__cf_bm` — and no numeric rate is published (§1.2). `Not publicly documented`: any requests-per-hour value that changes a bot score or challenge probability.
- **BFM/SBFM** match "patterns of known bots" — signatures, not volumes ([Bot Fight Mode — Cloudflare docs](https://developers.cloudflare.com/bots/get-started/bot-fight-mode/), read 2026-08-12).
- **Rate limiting** is the product that polices volume, and it is opt-in per zone with explicit per-plan constraints (§5.2). Cloudflare's only documented link between "one valid clearance + high volume" is a *recommendation to site owners*: "Cloudflare recommends that customers add a rate limiting rule based on the `cf_clearance` cookie value. This helps ensure that a single, valid cookie cannot be abused by one machine to send an excessive volume of requests." ([Clearance — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/concepts/clearance/), read 2026-08-12.) Note: counting by cookie value is only available on Enterprise (see table below); a Free/Pro site cannot even build that rule.
### 5.2 What Rate Limiting Rules would do to 20–60 requests/hour
Rate limiting rules are opt-in; nothing runs them unless the site creates a rule. The documented per-plan capabilities ([Rate limiting rules — Cloudflare docs](https://developers.cloudflare.com/waf/rate-limiting-rules/), read 2026-08-12):
| Capability | Free | Pro | Business |
|---|---|---|---|
| Number of rules | **1** | 2 | 5 |
| Counting characteristics | **IP only** | IP only | IP, IP w/ NAT |
| Counting period | **10 s only** | ≤ 1 min | ≤ 10 min |
| Mitigation timeout | **10 s** | ≤ 1 h | ≤ 1 day |
| Fields in expression | **Path, Verified Bot** | + Host, URI, Full URI, Query | + Method, Source IP, User Agent |
Even the most aggressive free-plan rule (1 request per 10 s = 360/hour) is 6–18× above our 20–60/hour volume; the counting window is 10 s, so a per-hour burst is invisible to it. At 1 request per minute worst-case, **20–60 requests/hour from one residential IP cannot trip any rate limiting rule the Free plan can express.** Also documented: rate limiting is approximate, not precise — "there may be a delay of up to a few seconds between detecting a request and updating rate counters… excess requests could still reach the origin", and counters are per-data-center. ([Rate limiting rules — Cloudflare docs](https://developers.cloudflare.com/waf/rate-limiting-rules/), read 2026-08-12.) `[INFERENCE]` a site could also challenge on rate via WAF custom rules, but that requires Pro+ (custom rules are not on Free) and its own configuration.
### 5.3 DDoS protection (always on, all plans)
HTTP DDoS Attack Protection "is always enabled" and can only be tuned, not disabled. The only *published numeric* thresholds are error-rate-based: origin-error floods mitigate at the default "High" sensitivity of **1,000 errors per second** (Pro+ also requires 5× normal origin traffic). Per-IP volumetric thresholds are adaptive and `Not publicly documented` in the managed ruleset docs. 20–60 requests/hour is ~9 orders of magnitude below the published figure. ([HTTP DDoS Attack Protection — Cloudflare docs](https://developers.cloudflare.com/ddos-protection/managed-rulesets/http/), read 2026-08-12.)
### 5.4 Execution order (which product fires first)
Documented phase order: `ddos_l7` → custom rules → `http_ratelimit` (rate limiting) → managed rules → `http_request_sbfm` (SBFM); BFM runs outside this pipeline and cannot be skipped; a terminating action (block/challenge) stops later phases. ([Security features interoperability — Cloudflare docs](https://developers.cloudflare.com/waf/feature-interoperability/), read 2026-08-12.) Practical reading: on the three sites, the challenge we see could come from any of these stages; none of them documents a volume input at our scale (§5.1–5.3).
---
## 6. The documented legitimate side
### 6.1 Verified bots — the only "treated well" path, and it requires self-identification
Cloudflare documents a Verified bot as one meeting two bars ([Verified bots — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot/verified-bots/), read 2026-08-12):
1. **Honest self-identification** — "through a cryptographic Web Bot Auth signature, a published IP list with a stable user-agent, or reverse DNS".
2. **Non-abusive behavior** — "it obeys `robots.txt` and crawl directives, **maintains reasonable request rates**, and has not been observed evading website owner preferences or attacking sites".
Relevant verified-bot *behavior classes* exist for exactly this kind of client: "**Feed Fetching** — RSS readers, podcast aggregators, and news feed bots" and "**Monitoring & Operations** — Uptime monitoring, webhooks, and health checks". Becoming verified requires an application via the dashboard and validation via Web Bot Auth or IP validation; breach of the policy (e.g. "An AI Crawler that does not respect the crawl-delay directive") removes the bot from the allowlist. ([Verified bots — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot/verified-bots/), read 2026-08-12.)
"Historically, Verified bots have been excluded in default bot configurations across all plans" (same source) — i.e., verified bots are *default-allowed* under SBFM/Bot Management. **But** this path is the opposite of what a scraper wants: it requires the poller to publicly identify itself (stable, published IPs or cryptographic signatures) and to have its identity vetted by Cloudflare — and the *site* still decides via verified-bot policy whether to allow the category. There is **no documented mechanism for an anonymous low-volume automated client to be treated well.** `[INFERENCE]` a manga-site scraper would never qualify (it would be classified as Data Collection / scraping behavior, which is not a default-allowed class).
### 6.2 robots.txt and crawl control
- `robots.txt` **compliance is voluntary** — "The file expresses your preferences, but it does not prevent crawlers from accessing your content at a technical level." Enforcement requires Cloudflare's AI Crawl Control. ([robots.txt setting — Cloudflare docs](https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/), read 2026-08-12.)
- The managed `robots.txt` feature (all plans) is aimed at AI crawlers; it prepends `Disallow` rules for AI bots and a Content Signals Policy. It does not create any allowance for generic scrapers. (Same source.)
- RFC-side: the robots exclusion standard is an unauthenticated convention; nothing in it grants access rights. ([RFC 9309 "Robots Exclusion Protocol"](https://www.rfc-editor.org/rfc/rfc9309.html) — read 2026-08-12.) The standard defines crawl-delay etc. as voluntary directives; Cloudflare's docs are the operative statement for CF-protected sites.
### 6.3 The documented takeaway for "slowing down vs. fingerprint quality"
Cloudflare's own documentation repeatedly points at **browser/device signals** as the decision input on non-Enterprise protection (JSD detecting headless browsers; Managed Challenge choosing based on "signals indicated by their browser"; challenges failing when UA/Canvas/WebGL are modified — §3.2, §4.2), and at **identity** (verified bots) as the only legitimacy signal for automation (§6.1). Request *rate* appears only as: (a) an unnamed component of Enterprise-ML "session characteristics", (b) a voluntary verified-bot behavioral bar, and (c) the separate, opt-in, Free-plan-impotent Rate Limiting product (§5). Nothing documented says "slow down and you'll be challenged less" for a free-plan site — **fingerprint quality is the lever that the documentation actually describes**, which matches this repo's measurements (real Chrome + real timezone passes; every headless variant fails regardless of rate).
---
## 7. Implications for cadence design
### Documented facts (with sources above)
1. **1-hour vs 6-hour cadence is not a documented risk lever.** Challenge issuance on Free/Pro-grade protection is signature-based; the rate-aware scoring only exists on Enterprise Bot Management; free-plan rate limiting cannot express a limit our volume could trip (§1, §2, §5).
2. **Every poll ≥ 1 hour re-solves the challenge anyway.** `cf_clearance` defaults to 30 minutes; Challenge Passage is site-configurable and unobservable. The re-solve cost is what this repo measured (~4 s, real Chrome + real timezone) (§4.1, repo measurements).
3. **The documented failure modes are fingerprint, not rate:** headless browsers (JSD), cloud-hosting IPs (BFM heuristics), modified UA/Canvas/WebGL (challenge solve failure) (§2.2, §3, §4.2).
4. **Clearance is not portable:** device-bound + solve-IP-bound + CHIPS-partitioned; replaying a cookie from another IP is documented invalid (§4.2).
5. **Anonymity has no documented "good citizen" path:** the only legitimate-automation route (verified bots) requires self-identification and site-side allowance (§6).
6. **The one behavior-bound revocation mechanism (Precursor) is opt-in per zone**, not a default documented behavior (§4.3).
### Inferences (not documented)
- `[INFERENCE]` The three sites run Free/Pro-grade protection (BFM, SBFM, or WAF challenge rules), not Enterprise Bot Management; therefore no continuous per-request bot score exists for our traffic.
- `[INFERENCE]` The sites have not changed Challenge Passage to hours/days (free manga sites default to the 30-minute default); if they had, hourly polling would get *cheaper* (valid cookie, no re-solve).
- `[INFERENCE]` Precursor is not enabled on these sites; if it were, hourly re-visits from an automated Chrome could accumulate session-behavior signals and trigger re-challenge even with a valid cookie — the only documented scenario in which polling *frequency* (via session behavior) could matter.
- `[INFERENCE]` 6-hour cooldowns buy nothing documented beyond raw request-count reduction (fewer challenge solves per day, less origin load); the risk profile at 1 request/hour/series is not documented to differ from 6 request/hour/series.
- `[INFERENCE]` If the owner wants belt-and-braces, the engineering levers that match the documentation are: keep the real-Chrome fingerprint (no UA spoofing, no headless-shell, real timezone — already done), keep a persistent user-data profile so `cf_clearance`/`__cf_bm` persist across visits, and treat any change of exit IP (e.g. home connection rebooting to a new IP) as a guaranteed re-solve, since clearance does not travel with the IP.
### Bottom line
Moving browser-backed sites from 6-hour to 1-hour per-series cooldown is **not contradicted by any documented Cloudflare mechanism** at 20–60 requests/hour from one residential IP through one real Chrome. The documented risk is carried by fingerprint quality (already solved in this repo) and by unobservable site configuration (Challenge Passage, Precursor, possible custom WAF rules). The residual, non-documented risk is that these sites sit behind Cloudflare's *proprietary* detection, and Cloudflare publishes neither its per-IP thresholds nor the ML feature set — so "no documented lever" is not "no lever".