Every doc statement that explained a Cloudflare challenge as a "score" was wrong. Researched against Cloudflare's own docs on 2026-08-12 (docs/research/cloudflare-bot-scoring-and-poll-cadence.md, 22 primary pages plus RFC 9309): the 1-99 bot score is Enterprise Bot Management only, free-plan zones get Bot Fight Mode signature matching and no score at all, and no per-IP request rate is documented as an input to challenge issuance. cf_clearance also expires in 30 minutes, so every cadence at or above 1h re-solves the challenge regardless. Docs only - no behaviour change. The 6h browser cooldown stays; its justification is now cost (a serialized single-tab solve costs seconds, a plain read costs one request), not a risk reduction nothing documents. - AGENTS.md: the block is per-zone configuration plus request fingerprint, not IP reputation; comix.to turning its gate on 2026-08-12 is the example. Residential egress avoids the cloud-hosting-IP signature rather than earning a better score. The UTC measurement stands but its mechanism is marked undocumented. - backend/AGENTS.md: states why the browser cooldown is longer. - ADR-0003, ADR-0006: dated corrections rather than rewrites. Both decisions stand on their other arguments (sweep depth, VPS memory). - DEPLOY.md: a red smoke run means the Site's settings or this Chrome's fingerprint moved, not that "Cloudflare's scoring" did.
4.2 KiB
ADR-0006: The browser runs on the home machine, over the tailnet
Date: 2026-08-09 Status: accepted
Decision
The headless browser is no longer part of the API stack. It is its own compose
unit (chrome/docker-compose.yml), deployed on the home machine, and the API on
the VPS reaches it over the existing tailnet through BROWSER_WS_URL. No
fallback sidecar remains on the VPS.
The backend needs no code change for this. The CDP endpoint was already a configuration seam and the fetcher only ever holds the endpoint URL, so relocation — and reversal — is one environment variable.
Why
The sidecar held 471 MiB working set (645 MiB peak) on a 1974 MiB VPS with no swap, which also hosts Traefik, Gitea and its Postgres. That is 24% of the host and 86% of this project's memory, for a service that at the time answered zero requests: the poller's due query joins bookmarks, production held four kagane series and no bookmarks on any of them, and with no kagane bookmark the web UI never rendered a kagane cover either.
The home machine has 5.9 GiB of swap and a residential egress, which avoids the cloud-hosting-IP signature Cloudflare's Bot Fight Mode documentedly challenges. Both machines were already on the tailnet.
Corrected 2026-08-12: the original wording said Cloudflare "scores" a residential
egress better than a datacenter IP. There is no score on a free-plan zone; what is
documented is signature matching, and hosting-provider IP space is one of the
signatures (docs/research/cloudflare-bot-scoring-and-poll-cadence.md). Memory was
the load-bearing reason regardless.
This move is only safe because covers are persisted (ADR-0005's sibling work, issue #43/#45) and the browser is on-demand (ADR-0005). Without stored covers a sleeping home machine would blank the library; without on-demand start the CI runner that already holds ~1.2 GiB of that box's 1.8 GiB would be squeezed around the clock.
Constraints
BROWSER_WS_URL must be the tailnet IP, never a MagicDNS hostname. Chrome's
DevTools HTTP handler answers /json/version with a 500 for any Host header
that is not an IP or localhost. This is the same trap that previously forced a
pinned Docker IP; the pinned subnet is gone, the constraint is not.
The CDP port binds to the tailnet address only, never 0.0.0.0. CDP
authenticates nothing: whatever reaches the port drives the browser and, through
it, the host. On the VPS the safety came from Docker network membership; the
home machine has a real LAN, so a 0.0.0.0 bind is a hole punched into it. The
bind address is the enforcement and Tailscale device identity plus a per-device
ACL is the policy. BROWSER_BIND_ADDR deliberately has no default, so an unset
value fails the deploy instead of publishing CDP to the LAN.
No bearer-token proxy is added in front of CDP. It would only defend against a device already inside the tailnet, and it would be one more thing between the poller and a browser that is already hard enough to keep clearing challenges.
Resource limits are load-bearing, not decorative. The browser is the
newcomer on that box, not the incumbent. A hard 512 MiB cap with 1 GiB
memory+swap makes Chrome reclaim its own cold pages onto the machine's SATA swap
instead of taking resident memory from the runner; untuned Chrome peaked at
645 MiB cgroup, which is more than is free there. oom_score_adj biases the
kernel to kill the browser first and never CI. Reduced CPU weight makes a
challenge solve yield to a running build — cold start degrades to about 3 s at
half a CPU, immaterial against a 45-second challenge budget. The shared-memory
reservation drops from 1 GiB to 128 MiB against a measured 19 MiB peak.
Consequences
An unreachable browser degrades exactly as an unset BROWSER_WS_URL already
does: plain-TLS libraries are unaffected, kagane and novelfull log and skip, the
series waits out its cooldown, and stored covers keep serving. A power outage at
home costs chapter freshness on two sites, never the appearance of the library.
The two units are deployed and updated independently. REDEPLOY.md §8 covers
the browser; everything before it covers the API stack. A local docker compose up now brings up two services, not three, and polls kagane only if
BROWSER_WS_URL is pointed somewhere.