Relocate the headless browser to the home machine, make it on-demand, and persist kagane covers #38
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Problem Statement
The headless Chrome sidecar idles at 471 MiB working set (595 MiB cgroup, 645 MiB peak) on a
1974 MiB VPS with no swap, which also hosts Traefik, Gitea and its Postgres, and one unrelated
app. Chrome is 24% of the entire host and 86% of this project's memory. The box sits at ~899 MiB
available with zero slack for a spike.
Worse, it is currently earning none of that. The sidecar exists only for kagane and novelfull —
the two sites behind a Cloudflare JavaScript challenge that no TLS fingerprint clears. Production
holds 4 kagane series, 0 novelfull series, and 0 bookmarks on any of them. The poller's due
query joins
bookmarks, so a series nobody has bookmarked is never fetched; and with no kaganebookmark, the web UI never renders a kagane cover, so the cover proxy is never called either. The
browser is a 471 MiB tenant serving zero requests, kept because a kagane/novelfull library is
coming.
The owner has a second machine ("home machine"): 1.8 GiB RAM with ~1.2 GiB already used by a Gitea
runner, ~616 MiB available, and 5.9 GiB of SATA-SSD swap. Both machines are already on the same
tailnet. The home machine is behind CGNAT with no public IP and is almost always on, barring power
outages.
Solution
Move the browser off the VPS entirely and make it cost nothing while idle.
The sidecar relocates to the home machine as its own deployable unit, reached over the existing
tailnet. The backend needs no code change for this: the CDP endpoint is already a configuration
seam, and the fetcher only ever holds the endpoint URL.
On the home machine the browser becomes on demand. Chrome is no longer the container's
foreground process; a supervisor owns the CDP port for the container's lifetime, spawns Chrome on
the first connection, and reaps it once the endpoint has been idle past a threshold. Idle cost
falls from 186 MiB anonymous to 2.9 MiB. Chrome exists for roughly ten minutes a day, so it barely
contends with the Gitea runner that owns that box.
Cover images stop depending on the browser being awake. The poll cycle prefetches each kagane
cover once and persists the bytes, so the web UI serves them from local storage forever after —
instant, and unaffected by the browser being asleep, unreachable, or mid-power-outage.
Browser-backed sites get their own, longer cooldown, because they are the only ones that cost a
browser wake. Plain-TLS sites keep the cadence they have.
User Stories
471 MiB it holds is returned to a 2 GiB box that has no swap.
project's memory appetite never takes down unrelated services.
so that the Gitea runner already using 1.2 GiB of that box is not squeezed.
runaway or leaking Chrome cannot drive the box into swap thrash and disrupt CI jobs.
victim, so that if memory does run out it is the browser that dies and never the CI runner.
so that a challenge solve never starves a running build.
home machine without touching the API stack on the VPS.
endpoint that is unauthenticated by design is never exposed to my LAN or the internet.
that relocation is a configuration change I can reverse in one line.
so that Cloudflare clearance is reused instead of re-solved after every deploy.
commit window, so that reaping the browser never discards the clearance it just earned.
that my reading list stays fresh whether or not I am browsing those sites.
a window I actually care about.
freshness, so that my main libraries do not get slower to satisfy a constraint that only
applies to two sites.
library never waits on a browser waking up.
not blank my library until every cover is refetched.
power outage costs me chapter freshness but never the appearance of my library.
do not wait up to six hours for the next poll cycle to fetch it.
for it exactly once ever, not once per restart.
plain-TLS libraries are unaffected by anything happening on the home machine.
resumes updating when the home machine comes back.
the log from a fetch that timed out, so that I am not misdiagnosing a healthy reap as a site
problem.
series behind a broken browser waits out a full cooldown instead of retrying every tick.
stalls or fails the chapter-list poll it rode along with.
so that relocating and reaping does not reintroduce the challenge failures already solved — a
non-UTC timezone, a User-Agent carrying no HeadlessChrome token and derived from the installed
binary, and no automation flag.
a real cover and a real chapter list come back through the tunnel before trusting it.
that Cloudflare scores the traffic better than it scores a datacenter IP.
socat front-end recorded, so that a later simplification does not silently break challenge
clearing.
that the diagram and config notes are not describing a same-host sidecar that no longer exists.
metadata, so that no future query accidentally loads image bytes into a hot path.
so that serving a cover is a primary-key lookup rather than a scan.
so that this remains a self-hosted tool with no external account to maintain.
reclaims is not quietly given back.
Implementation Decisions
Deployment topology
the home machine. The API stack drops the service, its
depends_on, and the dedicated browsernetwork along with it.
BROWSER_WS_URLmust name the tailnet IPaddress, not a MagicDNS hostname: Chrome's DevTools HTTP handler answers
/json/versionwith a500 for any Host header that is not an IP or
localhost. This is the same trap alreadydocumented for the Docker service name.
0.0.0.0. CDP has no authentication of its own — anything that reaches it has full control ofthat browser and, by extension, a foothold on that host. Tailscale identity plus a per-device ACL
is the access control; no additional bearer-token proxy is added, on the grounds that it would
only defend against a device already inside the tailnet.
BROWSER_WS_URLalready does today.Browser lifecycle
holds the CDP port for the container's lifetime, spawns Chrome on the first inbound connection
under a lock, and a periodic reaper terminates Chrome once there are no live connections and the
last use is older than the idle threshold.
cookies to disk on a roughly 31-second batch timer, so a shorter threshold would discard the
Cloudflare clearance just earned; and the effective reap happens at threshold + 90 seconds
because Go's default HTTP transport parks the discovery connection for
IdleConnTimeout.chromedp's remote allocator re-runs/json/versiondiscovery on every allocation, and the fetcher builds a fresh context per fetch,so the process-lifetime allocator holds only the bare endpoint URL and picks up a restarted
Chrome's new debugger UUID automatically. This was verified against a real restart with a changed
UUID. Correspondingly,
chromedp.NoModifyURLmust never be added — it is the single changethat would break restart survival, and the existing comment forbidding it becomes load-bearing
for a second reason.
legacy: Chromium forcibly remaps a non-loopback remote-debugging address to loopback since M113,
and Chrome ignores the remote-debugging flags entirely on a default profile since Chrome 136.
recreation. Without it, every recreate re-solves the challenge.
non-UTC timezone (a UTC clock is itself the bot signal), a User-Agent with the major version read
out of the installed binary and no HeadlessChrome token, and no automation flag.
roughly 30%, but with an on-demand browser that memory exists for seconds per wake, and every
flag is an untested change to the fingerprint of the one thing that is hard to get right.
Resource limits on the home machine
pressure rather than taking memory from the CI runner, and cold pages go to that machine's swap.
Sizing evidence: a tuned Chrome survived a 350 MiB hard cap with 114 MiB of headroom and zero
kills; untuned Chrome peaked at 645 MiB cgroup, which exceeds what is free on the home machine —
so the cap is load-bearing here, not decorative.
never takes the CI runner.
around 3 seconds at half a CPU, which is immaterial against a 45-second challenge budget.
Poller: per-class cooldown
LatestPollgains a browser cooldown alongside the existing one, read from a new environmentvariable and defaulting to 6 hours; the existing cooldown keeps its 1 hour default and continues
to govern plain-TLS sites. It is clamped like its siblings.
Store.DueForLatestCheckapplies both cutoffs in a single query, choosing per row with aconditional on the site column. The browser-backed site list is passed in by the poller as a
parameter — the store stays ignorant of which sites are behind a challenge, which is the
poller's existing responsibility (it already decides which fetcher a site gets). All values are
bound parameters; nothing is concatenated into query text.
finished series skipped, archived still polled, orphan series excluded by the bookmark join.
Cover persistence
coverstable keyed by the kagane image id, holding the image bytes, the content type, anda fetch timestamp. Keyed by image id rather than by series key because that is the identifier the
request actually carries; serving becomes a primary-key lookup instead of a scan for a series row
whose stored cover URL happens to contain that UUID.
series: it makes it structurallyimpossible for a future query or a
SELECT *to pull image bytes into the poller's hot path,rather than relying on the convention that the column list stays explicit. Orphan rows are
theoretical — nothing in the store deletes a series row, and a test already pins that.
column list used by the due query.
image is fetched through the browser and persisted. This is strictly best-effort — a cover
failure is logged and never affects the chapter-list result or the series' check stamp.
through the browser, as today, and writes through to the store so the browser is asked at
most once ever for a given image. The in-process cover cache and its accompanying note are
deleted: the store is the cache now, and unlike the map it survives a restart.
the long immutable cache headers all stay exactly as they are. The stored value is
client-supplied, so the id is validated before it reaches either the database or the browser.
a backend-served path needs no modification.
Error reporting
the caller's own deadline expiring. It gets a distinguishing wrap so the log tells the two apart.
This matters operationally: with an on-demand browser the interrupted-by-reap case becomes a
normal event, and misreading it as a site timeout would send someone hunting a Cloudflare problem
that does not exist.
Documentation
sidecar and must be updated: the topology diagram, the local-stack instructions, the new
environment variable, and the removal of the browser service from the API stack.
invisible from the code and expensive to rediscover: that the remote allocator survives Chrome
restarting behind a stable endpoint and therefore why the discovery-skipping option is forbidden;
that Chromium commits cookies on a ~31-second timer, so any stop/start lifecycle must outlive it;
and that the non-loopback debugging address is a no-op since M113 with the flags ignored on a
default profile since Chrome 136, which is why the socat front-end and explicit profile directory
are load-bearing.
Testing Decisions
A good test here asserts observable behaviour at the boundary a caller actually uses, and would
fail if the behaviour regressed. It does not assert on call sequences, internal field values, or
the shape of a private helper. The existing suite is the model: fakes are injected at the
already-existing struct fields, the store is a real Postgres from the throwaway-container helper,
and assertions are on what ends up stored or served.
Seams, fewest possible — three of the four already exist:
and batch. The browser cooldown and the cover fetcher become fields on it, and both behaviours
are driven through one poll cycle with fakes. Prior art: the existing tests that prove a kagane
row goes to the browser fetcher and never to the TLS one, and that a kagane row is skipped
entirely when no browser fetcher is wired.
art: the existing due-query tests covering cutoff behaviour, oldest-first ordering, the limit,
reader-count precedence, finished-versus-archived, and orphan exclusion.
handler tests and their counting fake keep working; they gain assertions that a second request
does not reach the browser because the first persisted, and that a cover already in the store is
served without any browser at all.
What gets tested:
after the shorter one, in the same cycle, from the same query.
already has.
and a subsequent request is served from storage.
trailing segment, a challenged fetch, a content type outside the allowlist, a path-traversal
attempt, and an unauthenticated request.
matching how the sibling poller settings are already covered.
a real browser.
What is not unit-tested, deliberately: the supervisor shell, the compose split, the tailnet
binding, and the resource limits. Inventing a Go seam for them would test the seam rather than the
behaviour. They are verified by the existing live smoke checks — which fetch a real cover and a
real chapter list through the configured CDP endpoint and are skipped unless that endpoint is set —
run against the home machine after deploy, plus observation that the container records no OOM kills
and no restarts over several days.
The full suite requires Docker, as it already does: each test package starts its own throwaway
Postgres.
Out of Scope
APIs cannot return kagane covers at all, because those images are served with an origin-restricted
resource policy and are only reachable via a same-origin fetch from inside the cleared page; and
hosted CDP providers would put a third-party account in the critical path of a self-hosted tool
to save memory that relocation already saves.
rejected: the cookie carries a continuously re-evaluated behavioural component in addition to a
roughly 30-minute lifetime, and a replay from a different client is re-challenged on fingerprint
mismatch. The browser must make the request itself, which is what the code already does.
meaningfully lighter one does not implement the graphics APIs the challenge probes, and the
alternative engines that do clear it are heavier than the current Chrome.
immutable per image id.
parallel signal it already is.
order.
Further Notes
Measured evidence behind the numbers in this spec, all gathered 2026-08-09:
anonymous and roughly 407 MiB is reclaimable page cache and shared memory. Zero OOM kills, zero
restarts, drift of 1.9 MiB over 62 seconds — this is a floor reached within five minutes of boot,
not a leak.
reclaimable by shrinking the cap to 32 MiB with the container staying up.
forceful termination; fresher cookies are lost under either. The disk write was first observed at
30.8 seconds, matching Chromium's persistent cookie store commit timer.
UUID; an in-flight fetch during a forced kill failed with a bare cancellation while the allocator
itself stayed healthy and subsequent fetches succeeded.
One live fact worth re-checking rather than assuming: whether the challenge clears from the home
machine's residential address. It clears from the VPS today in about 4 seconds, and a residential
address should score better than a datacenter one, but Cloudflare's scoring is time-varying and
this has not been measured from that egress. The live smoke check is exactly the instrument for
that, and a red run there means "not clearing from this address right now", which is a fact to
re-check rather than necessarily a defect.
The two sites this whole apparatus serves currently hold four series and zero bookmarks. The work
is justified by the library that is coming, not by present traffic — and the on-demand lifecycle
means that if that library never arrives, the cost of having built this is 2.9 MiB.
Last implementation child is done. Status of the epic:
Nothing in this epic needs further agent work. What is left is operator work: provision the home machine per
DEPLOY.md§7, add the Tailscale ACL, pointBROWSER_WS_URLat its tailnet IP, and then observe the three criteria that only production can show — covers still rendering with the machine off, several days without an OOM kill or restart, and the VPS memory improvement (DEPLOY.md§7 now takes afree -mreading before and after, which was previously unmeasurable).Verified before handover, on this branch: the browser unit builds and clears kagane's challenge live (56710 bytes of
image/webp, plus a 200 chapter list) at a 321 MiB peak against the 512 MiB cap, with the shm reduced to 128 MiB and zero OOM kills. Bind isolation holds — refused on the host's other address, accepted on the configured one.One thing the epic did not foresee: removing the
browsernetwork leftbookmark-apiondbalone, which isinternal: true, so the poller lost all egress. Fixed by putting the API back ondefault; only surfaced by running the stack, not by reading the config.Topology and the reasoning behind it are recorded in ADR-0006; ADR-0005 is cross-referenced as superseded in part.
Moving to
ready-for-human.