Clear the stale Chrome singleton lock at browser boot #76

Merged
sulthan merged 1 commits from fix/stale-chrome-singleton-lock into main 2026-08-10 19:11:24 +07:00
Owner

Symptom

The VPS could not reach the browser over the tailnet. Ping to the tailnet IP worked, docker-proxy was listening on 100.x:9222, and curl connected — then died with Recv failure: Connection reset by peer. docker logs bookmark-browser held one line, from the previous run's shutdown.

Cause

Chrome's SingletonLock in the chrome-profile volume names the hostname and pid that took it. Both change on docker compose up --build, so a Chrome that died uncleanly leaves a lock the next container reads as "The profile appears to be in use by another Google Chrome process (28349) on another computer (dd6e8d098e59)" and exits immediately. The volume outlives every recreate, so this never self-heals.

The reset follows from entrypoint.sh: socat accepts, start_browser fires, Chrome exits, wait_for_browser bails on browser_alive (:104), the helper returns 1 and closes the socket.

Change

  • rm -f "$profile"/Singleton* in the boot-time state reset, beside the connection markers and pid file. Safe because container_name pins this volume to one container — nothing can hold the profile when that line runs. Clearance cookies sit beside it and are untouched.
  • Chrome's stderr now reaches the container log instead of /dev/null, and the give-up path says why it gave up. Without these the failure is indistinguishable from a network fault, which is what made it expensive to find.

Scope: a lock left mid-container-lifetime is self-healing (same hostname, dead pid), so boot is the entire exposure.

Verification

sh -n chrome/entrypoint.sh clean. Deployed with docker compose up -d --build; /json/version answers.

Upstream, still open: whatever killed Chrome uncleanly — mem_limit: 512m against the Gitea runner is the suspect (ADR-0006). This makes that failure recoverable rather than terminal.

## Symptom The VPS could not reach the browser over the tailnet. Ping to the tailnet IP worked, `docker-proxy` was listening on `100.x:9222`, and `curl` connected — then died with `Recv failure: Connection reset by peer`. `docker logs bookmark-browser` held one line, from the previous run's shutdown. ## Cause Chrome's `SingletonLock` in the `chrome-profile` volume names the hostname and pid that took it. Both change on `docker compose up --build`, so a Chrome that died uncleanly leaves a lock the next container reads as *"The profile appears to be in use by another Google Chrome process (28349) on another computer (dd6e8d098e59)"* and exits immediately. The volume outlives every recreate, so this never self-heals. The reset follows from `entrypoint.sh`: socat accepts, `start_browser` fires, Chrome exits, `wait_for_browser` bails on `browser_alive` (`:104`), the helper returns 1 and closes the socket. ## Change - `rm -f "$profile"/Singleton*` in the boot-time state reset, beside the connection markers and pid file. Safe because `container_name` pins this volume to one container — nothing can hold the profile when that line runs. Clearance cookies sit beside it and are untouched. - Chrome's stderr now reaches the container log instead of `/dev/null`, and the give-up path says why it gave up. Without these the failure is indistinguishable from a network fault, which is what made it expensive to find. Scope: a lock left mid-container-lifetime is self-healing (same hostname, dead pid), so boot is the entire exposure. ## Verification `sh -n chrome/entrypoint.sh` clean. Deployed with `docker compose up -d --build`; `/json/version` answers. Upstream, still open: whatever killed Chrome uncleanly — `mem_limit: 512m` against the Gitea runner is the suspect (ADR-0006). This makes that failure recoverable rather than terminal.
sulthan added 1 commit 2026-08-10 19:11:13 +07:00
A Chrome killed uncleanly leaves SingletonLock in the chrome-profile
volume, naming the hostname and pid that took it. Both change when the
container is rebuilt, so the next Chrome reads it as "another computer
holds this profile" and exits at startup — permanently, since the volume
outlives every recreate.

The only symptom was a bare connection reset at 9222: socat accepts,
start_browser fires, Chrome dies, wait_for_browser bails on browser_alive
and the helper closes the socket. Chrome's stderr went to /dev/null and
the give-up path logged nothing, so neither docker logs nor the client
could tell a dead browser from a dead network — which is most of why this
cost an afternoon. Both are now on the container's stderr.

Clearing the lock at boot is safe because container_name pins the volume
to one container: nothing can hold the profile when that line runs. A
lock left mid-lifetime is self-healing (same hostname, dead pid), so boot
is the whole exposure.
sulthan merged commit 1ee5eb67ea into main 2026-08-10 19:11:24 +07:00
Sign in to join this conversation.