The browser sidecar needs a non-UTC zone to clear Cloudflare's challenge,
which left the two services logging on different clocks - the API in UTC,
the sidecar in local time. Correlating a poll failure against anything
else then means doing arithmetic at the worst possible moment.
This is cosmetic and nothing else in the service has a zone. Bookmark
timestamps are unix milliseconds (0001_bookmarks.sql pins that: the
userscripts send Date.now() and the ordering rule compares the integers
directly), the only two real time columns are timestamptz, the poller
works in durations, and nothing anywhere formats a wall clock - there is
no .Format( call in non-test backend code. So the zone can only reach
Go's log package, which stamps lines in local time.
No code change and no image change: distroless/static already carries the
full tzdata (1247 entries, Asia/Jakarta among them), so Go resolves the
name straight out of /usr/share/zoneinfo.
Named API_TZ rather than TZ because compose also reads the invoking
shell's environment, and a bare TZ would let an operator's exported zone
silently become the container's.
Verified against the built image, same instant:
no TZ 2026/08/08 16:24:30 open store: ...
TZ=Asia/Jakarta 2026/08/08 23:24:30 open store: ...
matching `date -u` 16:24:33 and Jakarta 23:24:33.
The previous commit documented the clock-zone requirement as "the zone
must match the egress IP's country". Re-measuring against the deployment
case shows that is wrong.
The original inference came from reading the host's /etc/timezone
(Asia/Bangkok) and assuming the egress IP was Thai. It is not: this host
egresses from an Indonesian IP. Asia/Bangkok cleared the challenge not
because it matched a country but because it is simply not UTC, and the
two share +07, which hid the distinction.
Measured 2026-08-08, identical container, one Indonesian egress IP:
TZ=UTC never cleared (60s, twice)
TZ=Asia/Jakarta cleared in 4s
TZ=America/New_York cleared in 4s
America/New_York matches neither the country nor the offset nor the
hemisphere, and clears just as fast. So a UTC clock is itself the bot
signal - Cloudflare scores it as the datacenter default - and any real
zone satisfies the check.
This makes the knob considerably less fragile than documented: BROWSER_TZ
needs a plausible zone, not a geolocated one, and a deployment that moves
region does not have to keep it in sync. Comments in chrome/entrypoint.sh
and docker-compose.yml, the hard constraint in AGENTS.md, and the PR
description are corrected accordingly.
BROWSER_TZ is also documented in .env.example for the first time, which
is the file an operator actually copies - the setting decides whether
kagane works at all, and a UTC server (the common case) is exactly the
one that fails with it unset.
Re-verified against the shipped image with TZ=Asia/Jakarta:
TestSmokeKaganeImage PASS (5.00s, 56710 bytes of image/webp),
TestSmokeKaganeGet PASS (1.24s, status 200).
kagane serves cover images from behind the same Cloudflare challenge as
its pages and with cross-origin-resource-policy: same-origin. The second
header is the decisive one: no <img> on the web UI's origin can load a
kagane cover even from a browser that already holds the clearance cookie,
verified 2026-08-08 by loading one from a foreign origin with and without
a referrer. Hot-linking cannot be made to work, so every kagane series
rendered the monogram placeholder.
Bookmark.CoverURL rewrites a stored kagane og:image to /img/kagane/{id}
and returns every other cover untouched; the templates render .CoverURL
in place of .Cover. The endpoint is session-gated like every other UI
route, and hands the id to the shared headless browser, whose fetch is
same-origin with kagane and therefore satisfies both the challenge and
the CORP header. Results are memoised in-process, so a cover costs one
navigation per deployment lifetime.
The id is matched against a UUID regex before it reaches the browser.
That gate is load-bearing rather than tidiness: the cover is a stored
client-supplied string, so an unvalidated one turns the endpoint into an
SSRF primitive aimed at the deployment's own network. ServeMux
path-cleans a traversal into a redirect before the handler runs, but the
handler does not rely on that, and a test pins it.
With BROWSER_WS_URL unset there is no browser and the endpoint answers
404 rather than reaching for a nil fetcher - the same degrade-to-
userscript behaviour the poller already has for these sites.
chromedp/headless-shell cannot clear kagane.to's managed challenge. It is
a stripped Chrome build, and the tells are structural rather than a
header: navigator.webdriver is true, the plugin list is empty, and the
client hints are Chromium- rather than Chrome-branded. Overriding
webdriver through CDP was tried on its own and changed nothing.
Everything below was measured on 2026-08-08 from a single IP, against the
same kagane cover, so the comparisons are like for like:
chromedp/headless-shell:stable never cleared (90s)
zenika/alpine-chrome never cleared - ships Chrome 124, old
enough that Cloudflare refuses it and
old enough to break chromedp's CDP structs
google-chrome, default UA never cleared (60s) - --headless=new
advertises "HeadlessChrome"
google-chrome, stock UA, UTC never cleared (90s)
google-chrome, stock UA, TZ set cleared in ~4s
So both remaining tells are load-bearing, and each was tested in
isolation. chrome/ is a Debian image with google-chrome-stable, a UA
whose version is read back out of the binary at startup (a hardcoded one
would drift out of step with the Sec-CH-UA hints on the next Chrome
update and become a fresh tell), and no --enable-automation.
The timezone matters because Cloudflare scores a browser whose clock zone
disagrees with its egress IP's country as a proxy. Note that the usual
`-v /etc/localtime:/etc/localtime:ro` does not work here: Chrome resolves
the zone through ICU, which takes the name from that path's symlink
target and ignores the file's contents, so glibc reports the host zone
while Chrome still reports UTC. /etc/timezone carries the name and is
mounted instead; BROWSER_TZ overrides it for a host whose clock is UTC in
a country that is not.
Chrome also binds its DevTools port to loopback and silently ignores
--remote-debugging-address, which is why headless-shell fronted it with
socat. This image does the same, so it stays a drop-in: the compose
service keeps the headless-shell name and its pinned address, and
BROWSER_WS_URL is unchanged.
Deploying needs `docker compose build headless-shell`.
BrowserFetcher.run navigated, waited for "body", read once, and closed
the tab - about half a second end to end. The Cloudflare interstitial has
a body too, so WaitReady was satisfied by the challenge page itself, and
the read that followed was of the interstitial rather than the site.
That made the challenge unclearable rather than merely slow. An
interstitial needs several seconds of a live page to solve itself and
write clearance into the browser's shared cookie jar; tearing the tab
down first means the clearance that would have unblocked every later
fetch is never obtained, so each call is challenged exactly like the one
before it.
run now holds one tab and re-reads until the caller's predicate reports
an answer, bounded by challengeTimeout and by the caller's own deadline.
Each caller supplies the predicate that fits its payload: kagane's
in-page fetch simply returns nothing while challenged, whereas
novelfull's payload is the DOM, and the interstitial has a DOM as well,
so that one excludes the challenge markup explicitly.
Exhausting the budget is now reported as errChallengeHeld and mapped back
to the 403 the poller already expects, keeping a challenged site distinct
from a broken transport.
Image loses its own retry loop, which run now subsumes.
Measured against a real kagane cover from a cold browser profile: no
image at all before, 4.9s to a 56710-byte image/webp after. The live
proof is TestSmokeKagane* in smoke_image_test.go, which skips unless
SMOKE_BROWSER_WS_URL names a sidecar, so `go test ./...` stays hermetic.
comix.to is an SPA whose client router rewrites document.title but never
touches the server-rendered og:title. The adapter read og:title, so a
bookmark taken after a cold load got the homepage's title ("Comix - Read
Comics online for free") and one taken after an in-page hop got the
previous series' title. Titles now come from document.title, with the
chapter page's " - Ch.<n>" tail stripped.
comix also serves no og:image at all, which is why every comix bookmark
fell back to the monogram placeholder. The cover is now the img whose alt
matches the cleaned title.
Both fixes need the page to have finished its client-side route change,
and comix fills document.title a beat after the URL changes - later than
the nav watcher's 300ms snapshot. The watcher therefore also re-detects
when the detect() signature changes, not only when the URL does.
kagane reader URLs carry no chapter number, so it comes out of og:title.
Volume-numbered series render "<Series> - Volume <v> Chapter <n>" with no
episode name, a shape the suffix regex did not match. One unmatched title
caused both reported symptoms: the volume tail stayed in the stored title
("SP Baby - Volume 1 Chapter 1"), and chapterNum came back null so no
chapter was ever recorded for the series. The regex now takes an optional
"Volume <v> " segment.
All three page shapes were captured live on 2026-08-08 and are pinned as
regression tests in userscript/test/logic.test.js.