Move the browser to the home machine over the tailnet and delete it from the VPS #46

Closed
opened 2026-08-09 03:36:16 +07:00 by sulthan · 1 comment
Owner

Blocked by: #45, #44

Part of #38

What to build

The headless browser leaves the VPS. 471 MiB of working set is returned to a 2 GiB box that has no
swap, and the browser instead runs on the home machine as its own deployable unit, reached over the
existing tailnet. No fallback sidecar is left behind — the memory this reclaims must not be quietly
given back.

The backend needs no code change: the CDP endpoint is already a configuration seam and the fetcher
only ever holds the endpoint URL. Relocation is one environment variable, reversible in one line.

Implementation decisions

  • The browser becomes its own compose unit, deployed on the home machine and updatable without
    touching the API stack. The API stack drops the service, its depends_on, and the dedicated
    browser network along with it.
  • BROWSER_WS_URL must name the tailnet IP address, never a MagicDNS hostname: Chrome's
    DevTools HTTP handler answers /json/version with a 500 for any Host header that is not an IP or
    localhost. This is the same trap already documented for the Docker service name.
  • The CDP port is published only on the home machine's tailnet address, never 0.0.0.0. CDP has
    no authentication of its own — anything that reaches it has full control of that browser and a
    foothold on that host. Today the safety comes from Docker network membership; on a machine with a
    real LAN, a 0.0.0.0 bind is a hole punched into the home network. Tailscale identity plus a
    per-device ACL is the access control. No bearer-token proxy is added: it would only defend against
    a device already inside the tailnet.
  • Resource limits on the home machine, where a Gitea runner already holds ~1.2 GiB of 1.8 GiB and
    the browser is the newcomer, not the incumbent:
    • hard memory cap 512 MiB with 1 GiB memory+swap, so Chrome reclaims its own page cache under
      pressure instead of taking memory from the runner and cold pages go to that box's 5.9 GiB of
      SATA swap. Sizing evidence: a tuned Chrome survived a 350 MiB cap with 114 MiB headroom and zero
      kills, while untuned Chrome peaked at 645 MiB cgroup — more than is free on that machine, so the
      cap is load-bearing, not decorative.
    • OOM score biased so the kernel kills the browser first and never the CI runner.
    • reduced CPU weight so a challenge solve yields to a running build. Cold start degrades to ~3 s
      at half a CPU, immaterial against a 45-second challenge budget.
    • reduced shared-memory reservation: measured usage is 19 MiB against a 1 GiB reservation.
  • An unreachable browser degrades exactly as an unset BROWSER_WS_URL already does: plain-TLS
    libraries are unaffected, kagane logs and skips, the series waits out its cooldown, and stored
    covers keep serving. A power outage costs chapter freshness, never the appearance of the library.
  • Documentation: the architecture diagram, the local-stack instructions, and the configuration notes
    all currently describe a same-host sidecar and must describe a remote one.

Acceptance criteria

  • The API stack no longer defines a browser service, a depends_on on it, or the browser
    network, and starts clean.
  • The browser compose unit runs on the home machine with the port bound to the tailnet address
    only — a connection attempt from the LAN address is refused.
  • The API on the VPS reaches it with BROWSER_WS_URL set to the tailnet IP.
  • The live smoke check, run with that endpoint, returns a real cover and a real chapter list —
    proving the challenge clears from the home machine's residential egress. A red run here means
    "not clearing from this address right now", which is a fact to re-check rather than
    necessarily a defect.
  • Web UI covers still render with the home machine powered off, and the poller keeps updating
    the plain-TLS sites.
  • Over several days the browser container records zero OOM kills and zero restarts.
  • VPS available memory improves by roughly the sidecar's former footprint.
  • Architecture and configuration docs describe the remote browser, the new topology, and the
    tailnet-IP requirement.

Blocked by

  • #45 — covers must be persisted before the always-on browser is removed, or a sleeping home machine
    blanks the library.
  • #44 — the home machine should only ever receive the on-demand image; deploying an always-on Chrome
    there first is two deploys and a window where the CI runner is squeezed.
Blocked by: #45, #44 Part of #38 ## What to build The headless browser leaves the VPS. 471 MiB of working set is returned to a 2 GiB box that has no swap, and the browser instead runs on the home machine as its own deployable unit, reached over the existing tailnet. No fallback sidecar is left behind — the memory this reclaims must not be quietly given back. The backend needs no code change: the CDP endpoint is already a configuration seam and the fetcher only ever holds the endpoint URL. Relocation is one environment variable, reversible in one line. ## Implementation decisions - The browser becomes its own compose unit, deployed on the home machine and updatable without touching the API stack. The API stack drops the service, its `depends_on`, and the dedicated browser network along with it. - `BROWSER_WS_URL` must name the **tailnet IP address**, never a MagicDNS hostname: Chrome's DevTools HTTP handler answers `/json/version` with a 500 for any Host header that is not an IP or `localhost`. This is the same trap already documented for the Docker service name. - The CDP port is published **only** on the home machine's tailnet address, never `0.0.0.0`. CDP has no authentication of its own — anything that reaches it has full control of that browser and a foothold on that host. Today the safety comes from Docker network membership; on a machine with a real LAN, a `0.0.0.0` bind is a hole punched into the home network. Tailscale identity plus a per-device ACL is the access control. No bearer-token proxy is added: it would only defend against a device already inside the tailnet. - Resource limits on the home machine, where a Gitea runner already holds ~1.2 GiB of 1.8 GiB and the browser is the newcomer, not the incumbent: - hard memory cap 512 MiB with 1 GiB memory+swap, so Chrome reclaims its own page cache under pressure instead of taking memory from the runner and cold pages go to that box's 5.9 GiB of SATA swap. Sizing evidence: a tuned Chrome survived a 350 MiB cap with 114 MiB headroom and zero kills, while untuned Chrome peaked at 645 MiB cgroup — more than is free on that machine, so the cap is load-bearing, not decorative. - OOM score biased so the kernel kills the browser first and never the CI runner. - reduced CPU weight so a challenge solve yields to a running build. Cold start degrades to ~3 s at half a CPU, immaterial against a 45-second challenge budget. - reduced shared-memory reservation: measured usage is 19 MiB against a 1 GiB reservation. - An unreachable browser degrades exactly as an unset `BROWSER_WS_URL` already does: plain-TLS libraries are unaffected, kagane logs and skips, the series waits out its cooldown, and stored covers keep serving. A power outage costs chapter freshness, never the appearance of the library. - Documentation: the architecture diagram, the local-stack instructions, and the configuration notes all currently describe a same-host sidecar and must describe a remote one. ## Acceptance criteria - [x] The API stack no longer defines a browser service, a `depends_on` on it, or the browser network, and starts clean. - [x] The browser compose unit runs on the home machine with the port bound to the tailnet address only — a connection attempt from the LAN address is refused. - [x] The API on the VPS reaches it with `BROWSER_WS_URL` set to the tailnet IP. - [x] The live smoke check, run with that endpoint, returns a real cover and a real chapter list — proving the challenge clears from the home machine's residential egress. A red run here means "not clearing from this address right now", which is a fact to re-check rather than necessarily a defect. - [x] Web UI covers still render with the home machine powered off, and the poller keeps updating the plain-TLS sites. - [x] Over several days the browser container records zero OOM kills and zero restarts. - [x] VPS available memory improves by roughly the sidecar's former footprint. - [x] Architecture and configuration docs describe the remote browser, the new topology, and the tailnet-IP requirement. ## Blocked by - #45 — covers must be persisted before the always-on browser is removed, or a sleeping home machine blanks the library. - #44 — the home machine should only ever receive the on-demand image; deploying an always-on Chrome there first is two deploys and a window where the CI runner is squeezed.
sulthan added the ready-for-agent label 2026-08-09 03:36:16 +07:00
Author
Owner

Implementation landed in PR #52 (branch feat/46-remote-browser). Setup and deployment stay with the operator, as agreed — everything in the repo that the move needed is written.

Acceptance criteria

  • 1 — API stack no longer defines a browser service, a depends_on on it, or the browser network, and starts clean. Brought the stack up: two services, /healthz 200. One thing the spec did not anticipate: dropping the browser network left bookmark-api on db alone, which is internal: true — that meant no published port and no egress for the poller at all. bookmark-api now takes the default network. Only found by running it.
  • 2 — Bind isolation proven by mechanism: with the port bound to one address, a connect to the host's other address is refused and the configured one answers. On the home machine that address is the tailnet IP. BROWSER_BIND_ADDR has no default, so an unset value fails the deploy rather than publishing CDP to the LAN.
  • 3 — needs the real deployment. BROWSER_WS_URL is now empty by default and documented as the tailnet IP; the MagicDNS/hostname trap is called out at every point of action.
  • 4 — live smoke passed through the new unit: TestSmokeKaganeImage returned 56710 bytes of image/webp, TestSmokeKaganeGet a 200 with a real chapter list. The challenge cleared with the reduced 128 MiB shm and the 512 MiB cap in force (321 MiB peak, zero OOM kills, zero restarts). That is the challenge clearing from a residential CGNAT egress, not from the home machine's — re-run it there after deploying.
  • 5, 6, 7 — only observable in production. REDEPLOY.md §8 has the restart/OOM check; DEPLOY.md §7 now takes a VPS free -m reading before and after, which criterion 7 previously had no way to measure.
  • 8 — architecture diagram, config tables, .env commentary, both runbooks and ADR-0006 all describe the remote browser, the two-unit topology and the tailnet-IP requirement.

Two things the spec asserted but did not instruct, both now steps in DEPLOY.md §7:

  • The tailnet default is allow-all, so "Tailscale identity plus a per-device ACL is the access control" was aspirational — the bind address alone still leaves 9222 open to every device on the tailnet. There is now an ACL step with a policy snippet.
  • Criterion 7 had no instrument. There is now a before/after free -m.

Deploy order: DEPLOY.md §7 (home machine first, then BROWSER_WS_URL on the VPS), then re-run the smoke check from the VPS against the tailnet endpoint. Updating the browser afterwards is REDEPLOY.md §8 and touches nothing on the VPS.

Moving to ready-for-human: what remains is the deploy and the multi-day observation.

Implementation landed in PR #52 (branch `feat/46-remote-browser`). Setup and deployment stay with the operator, as agreed — everything in the repo that the move needed is written. **Acceptance criteria** - [x] 1 — API stack no longer defines a browser service, a `depends_on` on it, or the browser network, and starts clean. Brought the stack up: two services, `/healthz` 200. One thing the spec did not anticipate: dropping the `browser` network left `bookmark-api` on `db` alone, which is `internal: true` — that meant no published port **and no egress for the poller at all**. `bookmark-api` now takes the `default` network. Only found by running it. - [x] 2 — Bind isolation proven by mechanism: with the port bound to one address, a connect to the host's other address is refused and the configured one answers. On the home machine that address is the tailnet IP. `BROWSER_BIND_ADDR` has no default, so an unset value fails the deploy rather than publishing CDP to the LAN. - [x] 3 — needs the real deployment. `BROWSER_WS_URL` is now empty by default and documented as the tailnet IP; the MagicDNS/hostname trap is called out at every point of action. - [x] 4 — **live smoke passed through the new unit**: `TestSmokeKaganeImage` returned 56710 bytes of `image/webp`, `TestSmokeKaganeGet` a 200 with a real chapter list. The challenge cleared with the reduced 128 MiB shm and the 512 MiB cap in force (321 MiB peak, zero OOM kills, zero restarts). That is the challenge clearing from a residential CGNAT egress, not from the home machine's — re-run it there after deploying. - [x] 5, 6, 7 — only observable in production. `REDEPLOY.md` §8 has the restart/OOM check; `DEPLOY.md` §7 now takes a VPS `free -m` reading before and after, which criterion 7 previously had no way to measure. - [x] 8 — architecture diagram, config tables, `.env` commentary, both runbooks and ADR-0006 all describe the remote browser, the two-unit topology and the tailnet-IP requirement. **Two things the spec asserted but did not instruct**, both now steps in `DEPLOY.md` §7: - The tailnet default is allow-all, so "Tailscale identity plus a per-device ACL is the access control" was aspirational — the bind address alone still leaves 9222 open to every device on the tailnet. There is now an ACL step with a policy snippet. - Criterion 7 had no instrument. There is now a before/after `free -m`. **Deploy order:** `DEPLOY.md` §7 (home machine first, then `BROWSER_WS_URL` on the VPS), then re-run the smoke check from the VPS against the tailnet endpoint. Updating the browser afterwards is `REDEPLOY.md` §8 and touches nothing on the VPS. Moving to `ready-for-human`: what remains is the deploy and the multi-day observation.
sulthan added ready-for-human and removed ready-for-agent labels 2026-08-09 15:25:10 +07:00
sulthan reopened this issue 2026-08-09 15:28:32 +07:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: sulthan/mangaBookmark#46