Relocate the headless browser to the home machine, make it on-demand, and persist kagane covers #38

Closed
opened 2026-08-09 03:30:36 +07:00 by sulthan · 1 comment
Owner

Problem Statement

The headless Chrome sidecar idles at 471 MiB working set (595 MiB cgroup, 645 MiB peak) on a
1974 MiB VPS with no swap, which also hosts Traefik, Gitea and its Postgres, and one unrelated
app. Chrome is 24% of the entire host and 86% of this project's memory. The box sits at ~899 MiB
available with zero slack for a spike.

Worse, it is currently earning none of that. The sidecar exists only for kagane and novelfull —
the two sites behind a Cloudflare JavaScript challenge that no TLS fingerprint clears. Production
holds 4 kagane series, 0 novelfull series, and 0 bookmarks on any of them. The poller's due
query joins bookmarks, so a series nobody has bookmarked is never fetched; and with no kagane
bookmark, the web UI never renders a kagane cover, so the cover proxy is never called either. The
browser is a 471 MiB tenant serving zero requests, kept because a kagane/novelfull library is
coming.

The owner has a second machine ("home machine"): 1.8 GiB RAM with ~1.2 GiB already used by a Gitea
runner, ~616 MiB available, and 5.9 GiB of SATA-SSD swap. Both machines are already on the same
tailnet. The home machine is behind CGNAT with no public IP and is almost always on, barring power
outages.

Solution

Move the browser off the VPS entirely and make it cost nothing while idle.

The sidecar relocates to the home machine as its own deployable unit, reached over the existing
tailnet. The backend needs no code change for this: the CDP endpoint is already a configuration
seam, and the fetcher only ever holds the endpoint URL.

On the home machine the browser becomes on demand. Chrome is no longer the container's
foreground process; a supervisor owns the CDP port for the container's lifetime, spawns Chrome on
the first connection, and reaps it once the endpoint has been idle past a threshold. Idle cost
falls from 186 MiB anonymous to 2.9 MiB. Chrome exists for roughly ten minutes a day, so it barely
contends with the Gitea runner that owns that box.

Cover images stop depending on the browser being awake. The poll cycle prefetches each kagane
cover once and persists the bytes, so the web UI serves them from local storage forever after —
instant, and unaffected by the browser being asleep, unreachable, or mid-power-outage.

Browser-backed sites get their own, longer cooldown, because they are the only ones that cost a
browser wake. Plain-TLS sites keep the cadence they have.

User Stories

  1. As the operator of the VPS, I want the headless browser to run somewhere else, so that the
    471 MiB it holds is returned to a 2 GiB box that has no swap.
  2. As the operator of the VPS, I want Traefik and Gitea to keep running normally, so that this
    project's memory appetite never takes down unrelated services.
  3. As the operator of the home machine, I want the browser to consume almost nothing while idle,
    so that the Gitea runner already using 1.2 GiB of that box is not squeezed.
  4. As the operator of the home machine, I want a hard memory ceiling on the browser, so that a
    runaway or leaking Chrome cannot drive the box into swap thrash and disrupt CI jobs.
  5. As the operator of the home machine, I want the kernel to pick the browser as its first OOM
    victim, so that if memory does run out it is the browser that dies and never the CI runner.
  6. As the operator of the home machine, I want the browser to give way to CI under CPU contention,
    so that a challenge solve never starves a running build.
  7. As the operator, I want the browser deployable as its own unit, so that I can update it on the
    home machine without touching the API stack on the VPS.
  8. As the operator, I want the browser's CDP port reachable only over the tailnet, so that an
    endpoint that is unauthenticated by design is never exposed to my LAN or the internet.
  9. As the operator, I want the backend to require no code change to talk to a remote browser, so
    that relocation is a configuration change I can reverse in one line.
  10. As the operator, I want the browser's profile to survive container restarts and recreations,
    so that Cloudflare clearance is reused instead of re-solved after every deploy.
  11. As the operator, I want the browser's idle timeout to sit well clear of Chromium's cookie
    commit window, so that reaping the browser never discards the clearance it just earned.
  12. As a Reader, I want kagane and novelfull chapter lists to keep updating in the background, so
    that my reading list stays fresh whether or not I am browsing those sites.
  13. As a Reader, I want a series to be at most six hours stale, so that new chapters surface within
    a window I actually care about.
  14. As a Reader, I want asura, demonic, comix and lightnovelworld to keep their current one-hour
    freshness, so that my main libraries do not get slower to satisfy a constraint that only
    applies to two sites.
  15. As a Reader, I want kagane cover images to appear instantly in the web UI, so that browsing my
    library never waits on a browser waking up.
  16. As a Reader, I want cover images to keep working after a backend restart, so that a deploy does
    not blank my library until every cover is refetched.
  17. As a Reader, I want cover images to keep working while the home machine is offline, so that a
    power outage costs me chapter freshness but never the appearance of my library.
  18. As a Reader, I want a newly bookmarked kagane series to show its cover on first view, so that I
    do not wait up to six hours for the next poll cycle to fetch it.
  19. As a Reader, I want a cover fetched on first view to be persisted, so that the browser is asked
    for it exactly once ever, not once per restart.
  20. As a Reader, I want the poller to keep working when the browser is unreachable, so that my
    plain-TLS libraries are unaffected by anything happening on the home machine.
  21. As a Reader, I want an unreachable browser to cost nothing but freshness, so that kagane simply
    resumes updating when the home machine comes back.
  22. As the operator, I want a browser fetch that was interrupted by a reap to be distinguishable in
    the log from a fetch that timed out, so that I am not misdiagnosing a healthy reap as a site
    problem.
  23. As the operator, I want the poller to keep stamping its check time before fetching, so that a
    series behind a broken browser waits out a full cooldown instead of retrying every tick.
  24. As the operator, I want cover prefetch to be best-effort, so that a failed cover fetch never
    stalls or fails the chapter-list poll it rode along with.
  25. As the operator, I want the browser's anti-bot properties preserved through the new lifecycle,
    so that relocating and reaping does not reintroduce the challenge failures already solved — a
    non-UTC timezone, a User-Agent carrying no HeadlessChrome token and derived from the installed
    binary, and no automation flag.
  26. As the operator, I want a live end-to-end check I can run after deploying, so that I can prove
    a real cover and a real chapter list come back through the tunnel before trusting it.
  27. As the operator, I want the residential egress of the home machine used for these fetches, so
    that Cloudflare scores the traffic better than it scores a datacenter IP.
  28. As a future maintainer, I want the reasons behind the supervisor, the profile volume and the
    socat front-end recorded, so that a later simplification does not silently break challenge
    clearing.
  29. As a future maintainer, I want the architecture documentation to describe a remote browser, so
    that the diagram and config notes are not describing a same-host sidecar that no longer exists.
  30. As a future maintainer, I want the cover storage to be structurally separate from series
    metadata, so that no future query accidentally loads image bytes into a hot path.
  31. As a future maintainer, I want cover bytes keyed by the identifier the request actually carries,
    so that serving a cover is a primary-key lookup rather than a scan.
  32. As the operator, I want no third-party scraping or hosted-browser service in the critical path,
    so that this remains a self-hosted tool with no external account to maintain.
  33. As the operator, I want no fallback browser left on the VPS, so that the memory this work
    reclaims is not quietly given back.

Implementation Decisions

Deployment topology

  • The browser sidecar leaves the API's compose stack and becomes its own compose unit, deployed on
    the home machine. The API stack drops the service, its depends_on, and the dedicated browser
    network along with it.
  • The API reaches the browser over the tailnet. BROWSER_WS_URL must name the tailnet IP
    address
    , not a MagicDNS hostname: Chrome's DevTools HTTP handler answers /json/version with a
    500 for any Host header that is not an IP or localhost. This is the same trap already
    documented for the Docker service name.
  • The browser's CDP port is published only on the home machine's tailnet address, never
    0.0.0.0. CDP has no authentication of its own — anything that reaches it has full control of
    that browser and, by extension, a foothold on that host. Tailscale identity plus a per-device ACL
    is the access control; no additional bearer-token proxy is added, on the grounds that it would
    only defend against a device already inside the tailnet.
  • No fallback sidecar remains on the VPS. An unreachable browser degrades exactly as an unset
    BROWSER_WS_URL already does today.

Browser lifecycle

  • Chrome stops being the container's foreground process. A supervisor in the image's entrypoint
    holds the CDP port for the container's lifetime, spawns Chrome on the first inbound connection
    under a lock, and a periodic reaper terminates Chrome once there are no live connections and the
    last use is older than the idle threshold.
  • The idle threshold is 300 seconds. Two measured constraints set the floor: Chromium commits
    cookies to disk on a roughly 31-second batch timer, so a shorter threshold would discard the
    Cloudflare clearance just earned; and the effective reap happens at threshold + 90 seconds
    because Go's default HTTP transport parks the discovery connection for IdleConnTimeout.
  • No change is required in the browser fetcher. chromedp's remote allocator re-runs
    /json/version discovery on every allocation, and the fetcher builds a fresh context per fetch,
    so the process-lifetime allocator holds only the bare endpoint URL and picks up a restarted
    Chrome's new debugger UUID automatically. This was verified against a real restart with a changed
    UUID. Correspondingly, chromedp.NoModifyURL must never be added — it is the single change
    that would break restart survival, and the existing comment forbidding it becomes load-bearing
    for a second reason.
  • The socat front-end and the explicit non-default user-data directory both stay. They are not
    legacy: Chromium forcibly remaps a non-loopback remote-debugging address to loopback since M113,
    and Chrome ignores the remote-debugging flags entirely on a default profile since Chrome 136.
  • The profile directory gets a named volume, so clearance survives both a reap and a container
    recreation. Without it, every recreate re-solves the challenge.
  • The anti-bot properties already established are preserved verbatim through the new spawn path: a
    non-UTC timezone (a UTC clock is itself the bot signal), a User-Agent with the major version read
    out of the installed binary and no HeadlessChrome token, and no automation flag.
  • Chrome flag tuning is deliberately not adopted. A measured set of flags cut idle memory by
    roughly 30%, but with an on-demand browser that memory exists for seconds per wake, and every
    flag is an untested change to the fingerprint of the one thing that is hard to get right.

Resource limits on the home machine

  • Hard memory cap of 512 MiB with 1 GiB of memory+swap, so Chrome reclaims its own page cache under
    pressure rather than taking memory from the CI runner, and cold pages go to that machine's swap.
    Sizing evidence: a tuned Chrome survived a 350 MiB hard cap with 114 MiB of headroom and zero
    kills; untuned Chrome peaked at 645 MiB cgroup, which exceeds what is free on the home machine —
    so the cap is load-bearing here, not decorative.
  • OOM score adjustment biases the kernel to kill the browser first, so a memory crisis on that box
    never takes the CI runner.
  • CPU weight is reduced so that a challenge solve yields to a running build. Cold start degrades to
    around 3 seconds at half a CPU, which is immaterial against a 45-second challenge budget.
  • The shared-memory reservation is reduced. Measured usage is 19 MiB against a 1 GiB reservation.

Poller: per-class cooldown

  • LatestPoll gains a browser cooldown alongside the existing one, read from a new environment
    variable and defaulting to 6 hours; the existing cooldown keeps its 1 hour default and continues
    to govern plain-TLS sites. It is clamped like its siblings.
  • Store.DueForLatestCheck applies both cutoffs in a single query, choosing per row with a
    conditional on the site column. The browser-backed site list is passed in by the poller as a
    parameter — the store stays ignorant of which sites are behind a challenge, which is the
    poller's existing responsibility (it already decides which fetcher a site gets). All values are
    bound parameters; nothing is concatenated into query text.
  • Ordering and exclusions are unchanged: reader count descending then least-recently-checked,
    finished series skipped, archived still polled, orphan series excluded by the bookmark join.

Cover persistence

  • A new covers table keyed by the kagane image id, holding the image bytes, the content type, and
    a fetch timestamp. Keyed by image id rather than by series key because that is the identifier the
    request actually carries; serving becomes a primary-key lookup instead of a scan for a series row
    whose stored cover URL happens to contain that UUID.
  • Deliberately a separate table rather than columns on series: it makes it structurally
    impossible for a future query or a SELECT * to pull image bytes into the poller's hot path,
    rather than relying on the convention that the column list stays explicit. Orphan rows are
    theoretical — nothing in the store deletes a series row, and a test already pins that.
  • The store gains a read and a write for cover bytes. The cover blob is not added to the series
    column list used by the due query.
  • The poll cycle prefetches: when a kagane series is checked and no cover bytes are stored yet, the
    image is fetched through the browser and persisted. This is strictly best-effort — a cover
    failure is logged and never affects the chapter-list result or the series' check stamp.
  • The web handler reads cover bytes from the store first. On a miss it falls back to fetching
    through the browser, as today, and writes through to the store so the browser is asked at
    most once ever for a given image. The in-process cover cache and its accompanying note are
    deleted: the store is the cache now, and unlike the map it survives a restart.
  • The content-type allowlist, the UUID validation at the request boundary, the session gate, and
    the long immutable cache headers all stay exactly as they are. The stored value is
    client-supplied, so the id is validated before it reaches either the database or the browser.
  • The URL shape the templates render is unchanged, so the rewrite from a stored kagane image URL to
    a backend-served path needs no modification.

Error reporting

  • A fetch interrupted by a reap currently surfaces as a bare cancellation, indistinguishable from
    the caller's own deadline expiring. It gets a distinguishing wrap so the log tells the two apart.
    This matters operationally: with an on-demand browser the interrupted-by-reap case becomes a
    normal event, and misreading it as a site timeout would send someone hunting a Cloudflare problem
    that does not exist.

Documentation

  • The project architecture notes and the backend configuration notes both describe a same-host
    sidecar and must be updated: the topology diagram, the local-stack instructions, the new
    environment variable, and the removal of the browser service from the API stack.
  • Three measured facts are recorded alongside the existing dated Cloudflare notes, because they are
    invisible from the code and expensive to rediscover: that the remote allocator survives Chrome
    restarting behind a stable endpoint and therefore why the discovery-skipping option is forbidden;
    that Chromium commits cookies on a ~31-second timer, so any stop/start lifecycle must outlive it;
    and that the non-loopback debugging address is a no-op since M113 with the flags ignored on a
    default profile since Chrome 136, which is why the socat front-end and explicit profile directory
    are load-bearing.

Testing Decisions

A good test here asserts observable behaviour at the boundary a caller actually uses, and would
fail if the behaviour regressed. It does not assert on call sequences, internal field values, or
the shape of a private helper. The existing suite is the model: fakes are injected at the
already-existing struct fields, the store is a real Postgres from the throwaway-container helper,
and assertions are on what ends up stored or served.

Seams, fewest possible — three of the four already exist:

  • The poller struct. Already the injection point for store, fetchers, clock, cooldown, interval
    and batch. The browser cooldown and the cover fetcher become fields on it, and both behaviours
    are driven through one poll cycle with fakes. Prior art: the existing tests that prove a kagane
    row goes to the browser fetcher and never to the TLS one, and that a kagane row is skipped
    entirely when no browser fetcher is wired.
  • The due-check store query. Signature grows a second cutoff and the browser-site list. Prior
    art: the existing due-query tests covering cutoff behaviour, oldest-first ordering, the limit,
    reader-count precedence, finished-versus-archived, and orphan exclusion.
  • The cover fetcher interface used by the web handler. Unchanged in shape, so the existing
    handler tests and their counting fake keep working; they gain assertions that a second request
    does not reach the browser because the first persisted, and that a cover already in the store is
    served without any browser at all.
  • Cover storage on the store is the one new surface: a read and a write, tested directly.

What gets tested:

  • Browser-backed sites are only due after the longer cooldown while plain-TLS sites remain due
    after the shorter one, in the same cycle, from the same query.
  • A poll cycle persists a kagane cover it did not previously have, and does not refetch one it
    already has.
  • A cover fetch that fails leaves the chapter-list result and the series check stamp untouched.
  • The handler serves a stored cover without consulting the browser; on a miss it fetches, persists,
    and a subsequent request is served from storage.
  • Every existing cover rejection path still yields the same result: a non-UUID id, a UUID with a
    trailing segment, a challenged fetch, a content type outside the allowlist, a path-traversal
    attempt, and an unauthenticated request.
  • Configuration loading for the new cooldown: default, override, clamp, and unparseable fallback,
    matching how the sibling poller settings are already covered.
  • The distinguishing wrap on an interrupted fetch is asserted on the classification, not by killing
    a real browser.

What is not unit-tested, deliberately: the supervisor shell, the compose split, the tailnet
binding, and the resource limits. Inventing a Go seam for them would test the seam rather than the
behaviour. They are verified by the existing live smoke checks — which fetch a real cover and a
real chapter list through the configured CDP endpoint and are skipped unless that endpoint is set —
run against the home machine after deploy, plus observation that the container records no OOM kills
and no restarts over several days.

The full suite requires Docker, as it already does: each test package starts its own throwaway
Postgres.

Out of Scope

  • Any hosted browser endpoint or scraping API. Researched and rejected: request/response scraping
    APIs cannot return kagane covers at all, because those images are served with an origin-restricted
    resource policy and are only reachable via a same-origin fetch from inside the cleared page; and
    hosted CDP providers would put a third-party account in the critical path of a self-hosted tool
    to save memory that relocation already saves.
  • Extracting and replaying the Cloudflare clearance cookie through the TLS client. Researched and
    rejected: the cookie carries a continuously re-evaluated behavioural component in addition to a
    roughly 30-minute lifetime, and a replay from a different client is re-challenged on fingerprint
    mismatch. The browser must make the request itself, which is what the code already does.
  • Swapping the browser engine. No lighter engine clears this class of challenge; the only
    meaningfully lighter one does not implement the graphics APIs the challenge probes, and the
    alternative engines that do clear it are heavier than the current Chrome.
  • Chrome flag tuning, for the reasons in the implementation decisions.
  • Host-level changes on the VPS, including adding swap. Out of bounds by the owner's decision.
  • An LRU or size-eviction policy for cover storage. A library holds tens of series and covers are
    immutable per image id.
  • Prefetching covers for any site other than kagane. No other site's images are challenge-gated.
  • Changing the userscripts. Their independent latest-chapter capture is unaffected and stays as the
    parallel signal it already is.
  • Changing the wire format, the bookmark ownership rules, or the timestamp rule that governs list
    order.

Further Notes

Measured evidence behind the numbers in this spec, all gathered 2026-08-09:

  • VPS sidecar at idle: 471 MiB working set, 595 MiB cgroup, 645 MiB peak, of which 186 MiB is
    anonymous and roughly 407 MiB is reclaimable page cache and shared memory. Zero OOM kills, zero
    restarts, drift of 1.9 MiB over 62 seconds — this is a floor reached within five minutes of boot,
    not a leak.
  • Supervisor prototype at idle with Chrome reaped: 2.9 MiB anonymous, 4.5 MiB cgroup. Proven
    reclaimable by shrinking the cap to 32 MiB with the container staying up.
  • Cold start to a completed page read: 0.4 s unconstrained, 0.9 s at one CPU, 3.0 s at half a CPU.
  • Clearance survives a reap for cookies older than about 31 seconds, under both graceful and
    forceful termination; fresher cookies are lost under either. The disk write was first observed at
    30.8 seconds, matching Chromium's persistent cookie store commit timer.
  • The remote allocator survived a Chrome restart behind a stable endpoint with a changed debugger
    UUID; an in-flight fetch during a forced kill failed with a bare cancellation while the allocator
    itself stayed healthy and subsequent fetches succeeded.

One live fact worth re-checking rather than assuming: whether the challenge clears from the home
machine's residential address. It clears from the VPS today in about 4 seconds, and a residential
address should score better than a datacenter one, but Cloudflare's scoring is time-varying and
this has not been measured from that egress. The live smoke check is exactly the instrument for
that, and a red run there means "not clearing from this address right now", which is a fact to
re-check rather than necessarily a defect.

The two sites this whole apparatus serves currently hold four series and zero bookmarks. The work
is justified by the library that is coming, not by present traffic — and the on-demand lifecycle
means that if that library never arrives, the cost of having built this is 2.9 MiB.

## Problem Statement The headless Chrome sidecar idles at 471 MiB working set (595 MiB cgroup, 645 MiB peak) on a 1974 MiB VPS with **no swap**, which also hosts Traefik, Gitea and its Postgres, and one unrelated app. Chrome is 24% of the entire host and 86% of this project's memory. The box sits at ~899 MiB available with zero slack for a spike. Worse, it is currently earning none of that. The sidecar exists only for kagane and novelfull — the two sites behind a Cloudflare JavaScript challenge that no TLS fingerprint clears. Production holds 4 kagane series, 0 novelfull series, and **0 bookmarks on any of them**. The poller's due query joins `bookmarks`, so a series nobody has bookmarked is never fetched; and with no kagane bookmark, the web UI never renders a kagane cover, so the cover proxy is never called either. The browser is a 471 MiB tenant serving zero requests, kept because a kagane/novelfull library is coming. The owner has a second machine ("home machine"): 1.8 GiB RAM with ~1.2 GiB already used by a Gitea runner, ~616 MiB available, and 5.9 GiB of SATA-SSD swap. Both machines are already on the same tailnet. The home machine is behind CGNAT with no public IP and is almost always on, barring power outages. ## Solution Move the browser off the VPS entirely and make it cost nothing while idle. The sidecar relocates to the home machine as its own deployable unit, reached over the existing tailnet. The backend needs no code change for this: the CDP endpoint is already a configuration seam, and the fetcher only ever holds the endpoint URL. On the home machine the browser becomes **on demand**. Chrome is no longer the container's foreground process; a supervisor owns the CDP port for the container's lifetime, spawns Chrome on the first connection, and reaps it once the endpoint has been idle past a threshold. Idle cost falls from 186 MiB anonymous to 2.9 MiB. Chrome exists for roughly ten minutes a day, so it barely contends with the Gitea runner that owns that box. Cover images stop depending on the browser being awake. The poll cycle prefetches each kagane cover once and persists the bytes, so the web UI serves them from local storage forever after — instant, and unaffected by the browser being asleep, unreachable, or mid-power-outage. Browser-backed sites get their own, longer cooldown, because they are the only ones that cost a browser wake. Plain-TLS sites keep the cadence they have. ## User Stories 1. As the operator of the VPS, I want the headless browser to run somewhere else, so that the 471 MiB it holds is returned to a 2 GiB box that has no swap. 2. As the operator of the VPS, I want Traefik and Gitea to keep running normally, so that this project's memory appetite never takes down unrelated services. 3. As the operator of the home machine, I want the browser to consume almost nothing while idle, so that the Gitea runner already using 1.2 GiB of that box is not squeezed. 4. As the operator of the home machine, I want a hard memory ceiling on the browser, so that a runaway or leaking Chrome cannot drive the box into swap thrash and disrupt CI jobs. 5. As the operator of the home machine, I want the kernel to pick the browser as its first OOM victim, so that if memory does run out it is the browser that dies and never the CI runner. 6. As the operator of the home machine, I want the browser to give way to CI under CPU contention, so that a challenge solve never starves a running build. 7. As the operator, I want the browser deployable as its own unit, so that I can update it on the home machine without touching the API stack on the VPS. 8. As the operator, I want the browser's CDP port reachable only over the tailnet, so that an endpoint that is unauthenticated by design is never exposed to my LAN or the internet. 9. As the operator, I want the backend to require no code change to talk to a remote browser, so that relocation is a configuration change I can reverse in one line. 10. As the operator, I want the browser's profile to survive container restarts and recreations, so that Cloudflare clearance is reused instead of re-solved after every deploy. 11. As the operator, I want the browser's idle timeout to sit well clear of Chromium's cookie commit window, so that reaping the browser never discards the clearance it just earned. 12. As a Reader, I want kagane and novelfull chapter lists to keep updating in the background, so that my reading list stays fresh whether or not I am browsing those sites. 13. As a Reader, I want a series to be at most six hours stale, so that new chapters surface within a window I actually care about. 14. As a Reader, I want asura, demonic, comix and lightnovelworld to keep their current one-hour freshness, so that my main libraries do not get slower to satisfy a constraint that only applies to two sites. 15. As a Reader, I want kagane cover images to appear instantly in the web UI, so that browsing my library never waits on a browser waking up. 16. As a Reader, I want cover images to keep working after a backend restart, so that a deploy does not blank my library until every cover is refetched. 17. As a Reader, I want cover images to keep working while the home machine is offline, so that a power outage costs me chapter freshness but never the appearance of my library. 18. As a Reader, I want a newly bookmarked kagane series to show its cover on first view, so that I do not wait up to six hours for the next poll cycle to fetch it. 19. As a Reader, I want a cover fetched on first view to be persisted, so that the browser is asked for it exactly once ever, not once per restart. 20. As a Reader, I want the poller to keep working when the browser is unreachable, so that my plain-TLS libraries are unaffected by anything happening on the home machine. 21. As a Reader, I want an unreachable browser to cost nothing but freshness, so that kagane simply resumes updating when the home machine comes back. 22. As the operator, I want a browser fetch that was interrupted by a reap to be distinguishable in the log from a fetch that timed out, so that I am not misdiagnosing a healthy reap as a site problem. 23. As the operator, I want the poller to keep stamping its check time before fetching, so that a series behind a broken browser waits out a full cooldown instead of retrying every tick. 24. As the operator, I want cover prefetch to be best-effort, so that a failed cover fetch never stalls or fails the chapter-list poll it rode along with. 25. As the operator, I want the browser's anti-bot properties preserved through the new lifecycle, so that relocating and reaping does not reintroduce the challenge failures already solved — a non-UTC timezone, a User-Agent carrying no HeadlessChrome token and derived from the installed binary, and no automation flag. 26. As the operator, I want a live end-to-end check I can run after deploying, so that I can prove a real cover and a real chapter list come back through the tunnel before trusting it. 27. As the operator, I want the residential egress of the home machine used for these fetches, so that Cloudflare scores the traffic better than it scores a datacenter IP. 28. As a future maintainer, I want the reasons behind the supervisor, the profile volume and the socat front-end recorded, so that a later simplification does not silently break challenge clearing. 29. As a future maintainer, I want the architecture documentation to describe a remote browser, so that the diagram and config notes are not describing a same-host sidecar that no longer exists. 30. As a future maintainer, I want the cover storage to be structurally separate from series metadata, so that no future query accidentally loads image bytes into a hot path. 31. As a future maintainer, I want cover bytes keyed by the identifier the request actually carries, so that serving a cover is a primary-key lookup rather than a scan. 32. As the operator, I want no third-party scraping or hosted-browser service in the critical path, so that this remains a self-hosted tool with no external account to maintain. 33. As the operator, I want no fallback browser left on the VPS, so that the memory this work reclaims is not quietly given back. ## Implementation Decisions ### Deployment topology - The browser sidecar leaves the API's compose stack and becomes its own compose unit, deployed on the home machine. The API stack drops the service, its `depends_on`, and the dedicated browser network along with it. - The API reaches the browser over the tailnet. `BROWSER_WS_URL` must name the **tailnet IP address**, not a MagicDNS hostname: Chrome's DevTools HTTP handler answers `/json/version` with a 500 for any Host header that is not an IP or `localhost`. This is the same trap already documented for the Docker service name. - The browser's CDP port is published **only** on the home machine's tailnet address, never `0.0.0.0`. CDP has no authentication of its own — anything that reaches it has full control of that browser and, by extension, a foothold on that host. Tailscale identity plus a per-device ACL is the access control; no additional bearer-token proxy is added, on the grounds that it would only defend against a device already inside the tailnet. - No fallback sidecar remains on the VPS. An unreachable browser degrades exactly as an unset `BROWSER_WS_URL` already does today. ### Browser lifecycle - Chrome stops being the container's foreground process. A supervisor in the image's entrypoint holds the CDP port for the container's lifetime, spawns Chrome on the first inbound connection under a lock, and a periodic reaper terminates Chrome once there are no live connections and the last use is older than the idle threshold. - The idle threshold is **300 seconds**. Two measured constraints set the floor: Chromium commits cookies to disk on a roughly 31-second batch timer, so a shorter threshold would discard the Cloudflare clearance just earned; and the effective reap happens at threshold + 90 seconds because Go's default HTTP transport parks the discovery connection for `IdleConnTimeout`. - **No change is required in the browser fetcher.** `chromedp`'s remote allocator re-runs `/json/version` discovery on every allocation, and the fetcher builds a fresh context per fetch, so the process-lifetime allocator holds only the bare endpoint URL and picks up a restarted Chrome's new debugger UUID automatically. This was verified against a real restart with a changed UUID. Correspondingly, `chromedp.NoModifyURL` must **never** be added — it is the single change that would break restart survival, and the existing comment forbidding it becomes load-bearing for a second reason. - The socat front-end and the explicit non-default user-data directory both stay. They are not legacy: Chromium forcibly remaps a non-loopback remote-debugging address to loopback since M113, and Chrome ignores the remote-debugging flags entirely on a default profile since Chrome 136. - The profile directory gets a named volume, so clearance survives both a reap and a container recreation. Without it, every recreate re-solves the challenge. - The anti-bot properties already established are preserved verbatim through the new spawn path: a non-UTC timezone (a UTC clock is itself the bot signal), a User-Agent with the major version read out of the installed binary and no HeadlessChrome token, and no automation flag. - Chrome flag tuning is deliberately **not** adopted. A measured set of flags cut idle memory by roughly 30%, but with an on-demand browser that memory exists for seconds per wake, and every flag is an untested change to the fingerprint of the one thing that is hard to get right. ### Resource limits on the home machine - Hard memory cap of 512 MiB with 1 GiB of memory+swap, so Chrome reclaims its own page cache under pressure rather than taking memory from the CI runner, and cold pages go to that machine's swap. Sizing evidence: a tuned Chrome survived a 350 MiB hard cap with 114 MiB of headroom and zero kills; untuned Chrome peaked at 645 MiB cgroup, which exceeds what is free on the home machine — so the cap is load-bearing here, not decorative. - OOM score adjustment biases the kernel to kill the browser first, so a memory crisis on that box never takes the CI runner. - CPU weight is reduced so that a challenge solve yields to a running build. Cold start degrades to around 3 seconds at half a CPU, which is immaterial against a 45-second challenge budget. - The shared-memory reservation is reduced. Measured usage is 19 MiB against a 1 GiB reservation. ### Poller: per-class cooldown - `LatestPoll` gains a browser cooldown alongside the existing one, read from a new environment variable and defaulting to 6 hours; the existing cooldown keeps its 1 hour default and continues to govern plain-TLS sites. It is clamped like its siblings. - `Store.DueForLatestCheck` applies both cutoffs in a single query, choosing per row with a conditional on the site column. The browser-backed site list is passed in by the poller as a parameter — the store stays ignorant of which sites are behind a challenge, which is the poller's existing responsibility (it already decides which fetcher a site gets). All values are bound parameters; nothing is concatenated into query text. - Ordering and exclusions are unchanged: reader count descending then least-recently-checked, finished series skipped, archived still polled, orphan series excluded by the bookmark join. ### Cover persistence - A new `covers` table keyed by the kagane image id, holding the image bytes, the content type, and a fetch timestamp. Keyed by image id rather than by series key because that is the identifier the request actually carries; serving becomes a primary-key lookup instead of a scan for a series row whose stored cover URL happens to contain that UUID. - Deliberately a separate table rather than columns on `series`: it makes it structurally impossible for a future query or a `SELECT *` to pull image bytes into the poller's hot path, rather than relying on the convention that the column list stays explicit. Orphan rows are theoretical — nothing in the store deletes a series row, and a test already pins that. - The store gains a read and a write for cover bytes. The cover blob is **not** added to the series column list used by the due query. - The poll cycle prefetches: when a kagane series is checked and no cover bytes are stored yet, the image is fetched through the browser and persisted. This is strictly best-effort — a cover failure is logged and never affects the chapter-list result or the series' check stamp. - The web handler reads cover bytes from the store first. On a miss it falls back to fetching through the browser, as today, and **writes through** to the store so the browser is asked at most once ever for a given image. The in-process cover cache and its accompanying note are deleted: the store is the cache now, and unlike the map it survives a restart. - The content-type allowlist, the UUID validation at the request boundary, the session gate, and the long immutable cache headers all stay exactly as they are. The stored value is client-supplied, so the id is validated before it reaches either the database or the browser. - The URL shape the templates render is unchanged, so the rewrite from a stored kagane image URL to a backend-served path needs no modification. ### Error reporting - A fetch interrupted by a reap currently surfaces as a bare cancellation, indistinguishable from the caller's own deadline expiring. It gets a distinguishing wrap so the log tells the two apart. This matters operationally: with an on-demand browser the interrupted-by-reap case becomes a normal event, and misreading it as a site timeout would send someone hunting a Cloudflare problem that does not exist. ### Documentation - The project architecture notes and the backend configuration notes both describe a same-host sidecar and must be updated: the topology diagram, the local-stack instructions, the new environment variable, and the removal of the browser service from the API stack. - Three measured facts are recorded alongside the existing dated Cloudflare notes, because they are invisible from the code and expensive to rediscover: that the remote allocator survives Chrome restarting behind a stable endpoint and therefore why the discovery-skipping option is forbidden; that Chromium commits cookies on a ~31-second timer, so any stop/start lifecycle must outlive it; and that the non-loopback debugging address is a no-op since M113 with the flags ignored on a default profile since Chrome 136, which is why the socat front-end and explicit profile directory are load-bearing. ## Testing Decisions A good test here asserts observable behaviour at the boundary a caller actually uses, and would fail if the behaviour regressed. It does not assert on call sequences, internal field values, or the shape of a private helper. The existing suite is the model: fakes are injected at the already-existing struct fields, the store is a real Postgres from the throwaway-container helper, and assertions are on what ends up stored or served. Seams, fewest possible — three of the four already exist: - **The poller struct.** Already the injection point for store, fetchers, clock, cooldown, interval and batch. The browser cooldown and the cover fetcher become fields on it, and both behaviours are driven through one poll cycle with fakes. Prior art: the existing tests that prove a kagane row goes to the browser fetcher and never to the TLS one, and that a kagane row is skipped entirely when no browser fetcher is wired. - **The due-check store query.** Signature grows a second cutoff and the browser-site list. Prior art: the existing due-query tests covering cutoff behaviour, oldest-first ordering, the limit, reader-count precedence, finished-versus-archived, and orphan exclusion. - **The cover fetcher interface used by the web handler.** Unchanged in shape, so the existing handler tests and their counting fake keep working; they gain assertions that a second request does not reach the browser because the first persisted, and that a cover already in the store is served without any browser at all. - **Cover storage on the store** is the one new surface: a read and a write, tested directly. What gets tested: - Browser-backed sites are only due after the longer cooldown while plain-TLS sites remain due after the shorter one, in the same cycle, from the same query. - A poll cycle persists a kagane cover it did not previously have, and does not refetch one it already has. - A cover fetch that fails leaves the chapter-list result and the series check stamp untouched. - The handler serves a stored cover without consulting the browser; on a miss it fetches, persists, and a subsequent request is served from storage. - Every existing cover rejection path still yields the same result: a non-UUID id, a UUID with a trailing segment, a challenged fetch, a content type outside the allowlist, a path-traversal attempt, and an unauthenticated request. - Configuration loading for the new cooldown: default, override, clamp, and unparseable fallback, matching how the sibling poller settings are already covered. - The distinguishing wrap on an interrupted fetch is asserted on the classification, not by killing a real browser. What is **not** unit-tested, deliberately: the supervisor shell, the compose split, the tailnet binding, and the resource limits. Inventing a Go seam for them would test the seam rather than the behaviour. They are verified by the existing live smoke checks — which fetch a real cover and a real chapter list through the configured CDP endpoint and are skipped unless that endpoint is set — run against the home machine after deploy, plus observation that the container records no OOM kills and no restarts over several days. The full suite requires Docker, as it already does: each test package starts its own throwaway Postgres. ## Out of Scope - Any hosted browser endpoint or scraping API. Researched and rejected: request/response scraping APIs cannot return kagane covers at all, because those images are served with an origin-restricted resource policy and are only reachable via a same-origin fetch from inside the cleared page; and hosted CDP providers would put a third-party account in the critical path of a self-hosted tool to save memory that relocation already saves. - Extracting and replaying the Cloudflare clearance cookie through the TLS client. Researched and rejected: the cookie carries a continuously re-evaluated behavioural component in addition to a roughly 30-minute lifetime, and a replay from a different client is re-challenged on fingerprint mismatch. The browser must make the request itself, which is what the code already does. - Swapping the browser engine. No lighter engine clears this class of challenge; the only meaningfully lighter one does not implement the graphics APIs the challenge probes, and the alternative engines that do clear it are heavier than the current Chrome. - Chrome flag tuning, for the reasons in the implementation decisions. - Host-level changes on the VPS, including adding swap. Out of bounds by the owner's decision. - An LRU or size-eviction policy for cover storage. A library holds tens of series and covers are immutable per image id. - Prefetching covers for any site other than kagane. No other site's images are challenge-gated. - Changing the userscripts. Their independent latest-chapter capture is unaffected and stays as the parallel signal it already is. - Changing the wire format, the bookmark ownership rules, or the timestamp rule that governs list order. ## Further Notes Measured evidence behind the numbers in this spec, all gathered 2026-08-09: - VPS sidecar at idle: 471 MiB working set, 595 MiB cgroup, 645 MiB peak, of which 186 MiB is anonymous and roughly 407 MiB is reclaimable page cache and shared memory. Zero OOM kills, zero restarts, drift of 1.9 MiB over 62 seconds — this is a floor reached within five minutes of boot, not a leak. - Supervisor prototype at idle with Chrome reaped: 2.9 MiB anonymous, 4.5 MiB cgroup. Proven reclaimable by shrinking the cap to 32 MiB with the container staying up. - Cold start to a completed page read: 0.4 s unconstrained, 0.9 s at one CPU, 3.0 s at half a CPU. - Clearance survives a reap for cookies older than about 31 seconds, under both graceful and forceful termination; fresher cookies are lost under either. The disk write was first observed at 30.8 seconds, matching Chromium's persistent cookie store commit timer. - The remote allocator survived a Chrome restart behind a stable endpoint with a changed debugger UUID; an in-flight fetch during a forced kill failed with a bare cancellation while the allocator itself stayed healthy and subsequent fetches succeeded. One live fact worth re-checking rather than assuming: whether the challenge clears from the home machine's residential address. It clears from the VPS today in about 4 seconds, and a residential address should score better than a datacenter one, but Cloudflare's scoring is time-varying and this has not been measured from that egress. The live smoke check is exactly the instrument for that, and a red run there means "not clearing from this address right now", which is a fact to re-check rather than necessarily a defect. The two sites this whole apparatus serves currently hold four series and zero bookmarks. The work is justified by the library that is coming, not by present traffic — and the on-demand lifecycle means that if that library never arrives, the cost of having built this is 2.9 MiB.
sulthan added the ready-for-agent label 2026-08-09 03:30:36 +07:00
Author
Owner

Last implementation child is done. Status of the epic:

  • #43 persist kagane covers — closed (PR #49)
  • #44 on-demand browser sidecar — closed (PR #50)
  • #45 prefetch covers during the poll cycle — closed (PR #51)
  • #46 move the browser to the home machine — implemented in PR #52, awaiting deploy

Nothing in this epic needs further agent work. What is left is operator work: provision the home machine per DEPLOY.md §7, add the Tailscale ACL, point BROWSER_WS_URL at its tailnet IP, and then observe the three criteria that only production can show — covers still rendering with the machine off, several days without an OOM kill or restart, and the VPS memory improvement (DEPLOY.md §7 now takes a free -m reading before and after, which was previously unmeasurable).

Verified before handover, on this branch: the browser unit builds and clears kagane's challenge live (56710 bytes of image/webp, plus a 200 chapter list) at a 321 MiB peak against the 512 MiB cap, with the shm reduced to 128 MiB and zero OOM kills. Bind isolation holds — refused on the host's other address, accepted on the configured one.

One thing the epic did not foresee: removing the browser network left bookmark-api on db alone, which is internal: true, so the poller lost all egress. Fixed by putting the API back on default; only surfaced by running the stack, not by reading the config.

Topology and the reasoning behind it are recorded in ADR-0006; ADR-0005 is cross-referenced as superseded in part.

Moving to ready-for-human.

Last implementation child is done. Status of the epic: - #43 persist kagane covers — closed (PR #49) - #44 on-demand browser sidecar — closed (PR #50) - #45 prefetch covers during the poll cycle — closed (PR #51) - #46 move the browser to the home machine — implemented in PR #52, awaiting deploy Nothing in this epic needs further agent work. What is left is operator work: provision the home machine per `DEPLOY.md` §7, add the Tailscale ACL, point `BROWSER_WS_URL` at its tailnet IP, and then observe the three criteria that only production can show — covers still rendering with the machine off, several days without an OOM kill or restart, and the VPS memory improvement (`DEPLOY.md` §7 now takes a `free -m` reading before and after, which was previously unmeasurable). Verified before handover, on this branch: the browser unit builds and clears kagane's challenge live (56710 bytes of `image/webp`, plus a 200 chapter list) at a 321 MiB peak against the 512 MiB cap, with the shm reduced to 128 MiB and zero OOM kills. Bind isolation holds — refused on the host's other address, accepted on the configured one. One thing the epic did not foresee: removing the `browser` network left `bookmark-api` on `db` alone, which is `internal: true`, so the poller lost all egress. Fixed by putting the API back on `default`; only surfaced by running the stack, not by reading the config. Topology and the reasoning behind it are recorded in ADR-0006; ADR-0005 is cross-referenced as superseded in part. Moving to `ready-for-human`.
sulthan added ready-for-human and removed ready-for-agent labels 2026-08-09 15:25:51 +07:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: sulthan/mangaBookmark#38