0a79e5f3d7af624caf7f8e081532878561114121
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
cce3d61799 |
Delete the kagane-specific cover path (#63)
The web proxy for kagane covers is dead: since #62 every Site's cover bytes land in the content-addressed store at creation or on the poll, and the one public route serves them all. Remove the second way to reach a Cover: - Bookmark.CoverURL() and the templates' use of it: templates render the wire value (.Cover) and nothing else. - GET /img/kagane/{id}, web.CoverFetcher, coverIDRe: the route and its identifier validation are gone, and with them web/cover.go. - store.KaganeImageID, GetKaganeCover, PutKaganeCover, kaganeCoverSourceURL: the proxy's persistence. - Bookmark.CoverSource: dead once CoverURL is gone. Acquisition keeps the browser where kagane genuinely needs it, but the Site name leaves the routing: kaganeImageURLRe lives in browser.go with the rest of the per-Site knowledge, browserCoverURL claims the URLs the sidecar alone can fetch, and fetchCoverBytes routes by URL shape with no Site argument. No plain-TLS fallback for a claimed URL — that would only retrieve a challenge page. Cover tests: kagane route tests removed, the general-route guarantees they pinned kept and re-pinned — unstored and traversal-shaped addresses serve nothing (TestPublicCoverRejectsUnknownAddress), non-image content types are never echoed back (TestPublicCoverNeverEchoesNonImage + TestCoverStoreAcceptsAnySourceURL). Store content-addressing and reopen-persistence tests rewritten against PutCover/GetCover. No Site name remains in a cover code path outside the acquisition module; go test ./... green. |
||
|
|
78234f3c19 |
Browser-backed Sites join the Cover pipeline (#62) (#72)
Fixes #62
Browser-backed Sites join the Cover pipeline: kagane and novelfull Series now get their Covers at creation, through the same acquisition path as every other Site, instead of waiting for a poll pass.
## What changed
`latest.Acquirer` (creation-time acquisition, fired by the first Bookmark of a Series) previously skipped kagane and novelfull entirely — their pages only yield a Cloudflare challenge to the TLS client, so the request was spent for nothing. It now routes them like the poller does, with the two Sites split exactly as the issue demands:
- **kagane** — page fetched through the browser sidecar, cover URL extracted from the API JSON, bytes fetched through the browser sidecar (the only path that clears the challenge) into the content-addressed store. With no `BROWSER_WS_URL` configured, acquisition is skipped entirely and nothing falls back to a plain fetch.
- **novelfull** — page fetched through the browser sidecar, cover URL extracted from the HTML, bytes fetched over plain TLS through the ordinary gated fetcher (its image paths answer 200 with `access-control-allow-origin: *`, measured 2026-08-09). With no browser configured, the page fetch falls back to the TLS client — novelfull's challenge is a live time-varying fact (AGENTS.md), so when the page body answers, the Cover still lands; when it is challenged, nothing happens.
The byte-routing rule (kagane → browser, every other Site → TLS) is now one shared function (`latest.fetchCoverBytes`) used by both the Poller and the Acquirer, so the two cannot drift apart.
## Acceptance criteria
- [x] kagane cover bytes are fetched through the browser sidecar and stored in the content-addressed store — `TestAcquireKaganeCoverThroughBrowser`
- [x] novelfull cover URLs are extracted from the browser-fetched HTML, and its bytes are fetched over plain TLS — `TestAcquireNovelfullCoverOverPlainTLS`
- [x] With no browser sidecar configured, kagane Covers are absent and nothing falls back to a plain fetch — `TestAcquireKaganeSkippedWithoutBrowser`
- [x] With no browser sidecar configured, novelfull Covers still work if its page body is available — `TestAcquireNovelfullCoverWithoutBrowser`
- [x] Manually verified on-device: a kagane Series shows its Cover in the panel, not a broken-image glyph — being run by a separate manual-verification agent against a mocked scenario (no prod data); not part of this PR
- [x] `go test ./...` is green, with live-network checks gated behind `SMOKE_BROWSER_WS_URL` like the existing kagane image smoke test — new `TestSmokeAcquireKaganeCover` proves the end-to-end acquire path against the real browser when the env var is set
## Verification
- `go test ./...` green across all packages
- New unit tests exercise every routing decision with fakes — no network in the default suite
- Smoke test gated behind `SMOKE_BROWSER_WS_URL`, skipped by default
## Post-review changes (
|
||
|
|
2a3bb6922d |
Move the browser off the VPS to its own unit (#46) (#52)
Closes #46 once deployed. The headless browser leaves the API stack and becomes its own compose unit (`chrome/docker-compose.yml`) intended for the home machine, reached over the tailnet. No fallback sidecar is left on the VPS. The backend needs no code change — `BROWSER_WS_URL` was already the only coupling. Its default is now empty rather than a pinned Docker IP, so an unconfigured or unreachable browser degrades exactly as it always has: plain-TLS libraries unaffected, kagane/novelfull logged and skipped, stored covers still served. ### What shipped - `chrome/docker-compose.yml` + `chrome/.env.example` — the browser unit, with the CDP port bound to `${BROWSER_BIND_ADDR}` (no default) and the resource limits from the epic: 512 MiB / 1 GiB memory+swap, `oom_score_adj 800`, halved CPU weight, shm 1 GiB -> 128 MiB. - API stack drops the service, its `depends_on` and the `browser` network. - `bookmark-api` gains the `default` network. Dropping `browser` had left it on `db` alone, which is `internal: true` — no published port and, worse, no egress for the poller at all. Caught by actually bringing the stack up. - ADR-0006 for the topology; `DEPLOY.md` §7 for first-time setup of the browser machine; `REDEPLOY.md` §8 for its independent update cadence; architecture diagrams, config tables and troubleshooting rows across README/AGENTS/env. ### Verified locally - Browser unit builds and runs: Chrome 151, UA carries no `HeadlessChrome`, all limits applied as declared. - **Live smoke passes through the new unit**: `TestSmokeKaganeImage` fetched 56710 bytes of `image/webp`, `TestSmokeKaganeGet` got a 200 with a real chapter list. The challenge cleared under the reduced 128 MiB shm. - Bind isolation proven: refused on the host's non-loopback address, accepted on the configured one. - 321 MiB peak of the 512 MiB cap after a full solve; 0 restarts, no OOM kill. - API stack comes up clean, `/healthz` 200; egress confirmed present on `default` and absent on `db`. - `go test ./...`, `go vet`, `gofmt` clean. ### Left to the operator Provisioning the home machine, the Tailscale ACL, setting `BROWSER_WS_URL` in production, and observing acceptance criteria 5-7 (covers with the machine off, several days of zero OOM/restarts, VPS memory improvement). `DEPLOY.md` §7 now carries the before/after `free -m` reading those need. Reviewed-on: #52 Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com> Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com> |
||
|
|
741b23322b |
Fix comix titles and covers, kagane volume chapters, and kagane cover rendering (#37)
Fixes five reported symptoms across comix.to and kagane.to. Diagnosing them turned up two latent bugs underneath, both of which had to be fixed for the kagane cover work to function at all.
## Reported symptoms and their causes
| # | Symptom | Cause |
|---|---------|-------|
| 1 | comix bookmark titled `Comix - Read Comics online for free` | comix is an SPA that rewrites `document.title` on client routing but never touches the server-rendered `og:title`. The adapter read `og:title`, so a cold load stored the homepage's title. |
| 2 | next comix bookmark gets the *previous* series' title | Same cause. After an in-page hop, `og:title` still holds whatever page loaded first. |
| 3 | comix cover shows the placeholder | comix serves no `og:image` at all, so `coverFromPage()` had nothing to read. |
| 4 | kagane chapter never appears in the bookmark list | Reader URLs carry no chapter number, so it is parsed out of `og:title`. Volume-numbered series render `"<Series> - Volume <v> Chapter <n>"`, which the suffix regex did not match, so `chapterNum` came back null and nothing was recorded. |
| 5 | kagane title includes the chapter, e.g. `SP Baby - Volume 1 Chapter 1` | Same unmatched regex — the tail was never stripped. One fix covers 4 and 5. |
| 6 | kagane cover blocked in the web UI | kagane serves covers behind its Cloudflare challenge **and** with `cross-origin-resource-policy: same-origin`. No `<img>` on the UI's origin can load one even from a browser holding the clearance cookie. Hot-linking cannot be made to work. |
## What changed
**Userscript.** comix titles now come from `document.title` with the chapter page's `" - Ch.<n>"` tail stripped, and the cover is the `img` whose `alt` matches the cleaned title. comix fills `document.title` a beat *after* the URL changes — later than the nav watcher's 300 ms snapshot — so the watcher also re-detects when the `detect()` signature changes, not only when the URL does. The kagane suffix regex takes an optional `Volume <v> ` segment. All three page shapes were captured live on 2026-08-08 and pinned as regression tests.
**Cover proxy.** `Bookmark.CoverURL()` rewrites a stored kagane `og:image` to `/img/kagane/{id}`; templates render `.CoverURL` instead of `.Cover`. The endpoint is session-gated like every other UI route and fetches through the shared headless browser, which is same-origin with kagane and so satisfies both the challenge and the CORP header. Results are memoised in-process, so a cover costs one navigation per deployment lifetime. With `BROWSER_WS_URL` unset the endpoint answers 404 rather than reaching for a nil fetcher — the same degrade-to-userscript behaviour the poller already has.
The image id is matched against a UUID regex before it reaches the browser. That gate is load-bearing rather than tidiness: the cover is a stored client-supplied string, so an unvalidated one turns this endpoint into an SSRF primitive aimed at the deployment's own network. `ServeMux` path-cleans a traversal into a redirect before the handler runs, but the handler does not depend on that, and a test pins it.
## Two latent bugs found underneath
**`BrowserFetcher.run` never let a challenge solve.** It navigated, waited for `body`, read once, and closed the tab — roughly half a second end to end. The Cloudflare interstitial has a `body` too, so `WaitReady` was satisfied by the challenge page itself. This made the challenge *unclearable* rather than merely slow: an interstitial needs several seconds of a live page to solve itself and write clearance into the browser's shared cookie jar, so tearing the tab down first means every subsequent call is challenged exactly like the one before it. `run` now holds one tab and re-reads until the caller's predicate reports an answer, bounded by `challengeTimeout` and the caller's own deadline. Exhausting the budget maps back to the 403 the poller already expects, keeping a challenged site distinct from a broken transport.
**`chromedp/headless-shell` cannot clear kagane's challenge at all.** It is a stripped Chrome build and the tells are structural rather than a header: `navigator.webdriver` is true, the plugin list is empty, and the client hints are Chromium- rather than Chrome-branded. Overriding `webdriver` through CDP was tried on its own and changed nothing.
All measured 2026-08-08 from one IP against the same cover, so the comparisons are like for like:
| Browser | Result |
|---------|--------|
| `chromedp/headless-shell:stable` | never cleared (90 s) |
| `zenika/alpine-chrome` | never cleared — ships Chrome 124, old enough that Cloudflare refuses it and old enough to break chromedp's CDP structs |
| `google-chrome`, default UA | never cleared (60 s) — `--headless=new` advertises `HeadlessChrome` |
| `google-chrome`, stock UA, `TZ=UTC` | never cleared (90 s) |
| `google-chrome`, stock UA, any non-UTC `TZ` | **cleared in ~4 s** |
Both remaining tells are load-bearing, and each was tested in isolation. `chrome/` is a Debian image with `google-chrome-stable`, a UA whose version is read back out of the binary at startup (a hardcoded one would drift out of step with the `Sec-CH-UA` hints on the next Chrome update and become a fresh tell), and no `--enable-automation`.
### The timezone tell: UTC, not a country mismatch
The first pass concluded the zone had to match the egress IP's country. Re-measuring against the actual deployment case shows that was wrong, and the correction is in `1552dd1`.
The original inference read the host's `/etc/timezone` (`Asia/Bangkok`) and assumed a Thai egress. It isn't — this host egresses from an Indonesian IP. `Asia/Bangkok` cleared not because it matched a country but because it simply isn't UTC, and the two share +07, which hid the distinction. Same container, same Indonesian IP:
| `TZ` | Result |
|------|--------|
| `UTC` | never cleared (60 s, **twice**) |
| `Asia/Jakarta` | cleared in 4 s |
| `America/New_York` | cleared in 4 s |
`America/New_York` matches neither the country nor the offset nor the hemisphere and clears just as fast. A UTC clock is itself the bot signal — Cloudflare scores it as the datacenter default — and any real zone satisfies the check. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one, and a deployment that changes region need not keep it in sync.
One sharp edge remains: the usual `-v /etc/localtime:/etc/localtime:ro` does **not** work. Chrome resolves the zone through ICU, which takes the name from that path's symlink target and ignores the file's contents, so glibc reports the host zone while Chrome still reports UTC. `/etc/timezone` carries the name and is mounted instead.
Chrome also binds its DevTools port to loopback and silently ignores `--remote-debugging-address`, which is why headless-shell fronted it with socat. This image does the same, so it stays a drop-in: the compose service keeps the `headless-shell` name and its pinned address, and `BROWSER_WS_URL` is unchanged.
## Verification
```
go test ./... all packages ok
node --test 37 + 12 pass, 0 fail
SMOKE_BROWSER_WS_URL=... go test -run TestSmokeKagane ./internal/latest
TestSmokeKaganeImage PASS (5.29s) fetched 56710 bytes of image/webp
TestSmokeKaganeGet PASS (1.17s) status=200, real chapter-list JSON
```
The smoke test ran against the exact compose configuration — built image, empty `BROWSER_TZ`, `/etc/timezone` mounted, cold profile — hitting real kagane.to. It skips unless `SMOKE_BROWSER_WS_URL` names a sidecar, so `go test ./...` stays hermetic and Docker-only.
A red smoke run means the challenge is not clearing from that IP, which is a live, time-varying fact to re-check rather than necessarily a defect.
## Security invariants
- Auth unchanged. `/img/kagane/{id}` is session-gated by `requireSession`, the same guard as every other UI route.
- Outbound fetch gated: the id is UUID-validated before it reaches the browser, keeping the existing rule that a client-supplied string never selects a fetch target unchecked.
- No new secrets, no new logging of credentials, no change to CORS, sessions, or crypto.
- Templates still escape everything; `.CoverURL` returns a plain string and is not wrapped in `template.HTML`/`URL`.
- One new dependency-free image (`chrome/`) built from Debian plus Google's own apt repo; no new Go modules.
## Deploying
Needs `docker compose build headless-shell`.
**A UTC host must set `BROWSER_TZ`, or kagane silently stops working.** With it unset the sidecar falls back to the host's `/etc/timezone`; on a UTC server that yields UTC, which is the one value that never clears. Any real zone works — `BROWSER_TZ=Asia/Jakarta` for the current deployment. `.env.example` now documents this; it previously did not mention the knob at all.
Only the browser sidecar reads `BROWSER_TZ`. The backend keeps its UTC clock, and stored timestamps are unix ms, so nothing else shifts.
## Deliberately not done
Retry/backoff around the cover proxy, and a panel-side cover fix. The panel renders no covers, and covers cache in-process after the first fetch. Worth adding if kagane starts rate-limiting.
## Correction after review of the deployment case
`1552dd1` was added after the branch was first pushed: the deployment host runs UTC with an Indonesian egress IP, which prompted re-measuring the timezone claim and falsifying it. The earlier commits' reasoning is left intact rather than rebased away, so the diagnostic trail — including the wrong turn and what disproved it — stays readable.
Reviewed-on: #37
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
|