v1.4.1
5 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
672c16ffbf |
Remove the client latest-chapter scan for lightnovelworld (#91) (#93)
Follows the spec published on #91: remove the client-side latest-chapter scan for lightnovelworld rather than porting #87's truncation into a second codebase. - lightnovelworld.latestChapterFromAnchors deleted, not stubbed: absence is what the background-fetch guard keys off. - computeLatestChapter tolerates an adapter with no scanner (yields null) and is exported as the test seam. - backgroundRefreshLatest skips a scanner-less Site before the due filter: no Series page fetched, no freshness timestamp recorded, no batch slot consumed. The on-page path (maybeCaptureLatestOnSeriesPage) routes through the same null-tolerant computation. - novelfull's scanner, the shared max-chapter helper and all four manga Sites untouched. - userscript/AGENTS.md records the Poll-only contract for this Site and why. All seven acceptance criteria from the spec met. node --check clean; novel suite 30/30 (the regression pin fails if a lnw scan is reintroduced, scoped or not); manga suite 35/35, manga userscript byte-for-byte unchanged. Two-axis code review: no hard standard violations, spec-clean; one follow-up commit matching the sibling adapter guard from the manga script. Reviewed-on: #93 Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com> Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com> |
||
|
|
e7e22a12a5 |
lightnovelworld Series identity is read from the chapter page (#80) (#92)
Implements spec #80 / ADR-0008 — Gitea issues #86, #87, #88, #89, #90, all closed. A Reader bookmarks a novel on lightnovelworld and it never shows a New Chapter, because the Series identity was derived from the chapter address instead of read from the page. One Series can publish under several Chapter Slugs, so the derived key points at a slug that 404s. - **#89** — the userscript's lnw adapter stops deriving `seriesUrl`/`seriesId` from the path. It reads the page's own pointer (`a[aria-label='All Chapter']`), falling back to the microdata breadcrumb's second crumb, and carries `chapterSlug` on the page object, stored nowhere. - **#87** — the Poll's lnw chapter scan is unscoped (no stored-slug pattern can cover a Series' whole list) and truncated at the `wpd-threads` comment thread, the one region a visitor can write to. Marker absent means skip and log with the body length, never scan whole. Corrects the `maxBodyBytes` headroom comment to the measured 3.5x. - **#86** — the scan fixture is now text trimmed from a real, wholly-fetched Series page instead of a hand-written cross-series anchor that no live page carries. - **#90** — stale stored rows repair themselves on the next chapter visit: a pure transform over cache, queue and last-checked map, silent to the Reader, with progress, favourite and lifecycle bucket preserved when two rows merge. - **#88** — an env-gated live canary (`SMOKE_LNW_SERIES_URL`) proving the marker still occurs exactly once and still follows the last chapter anchor, asserted against the production symbols themselves. Verified on the merged branch: `go test ./...` green, `gofmt -l internal/latest/` silent, both userscripts `node --check` clean, 35/35 + 29/29 logic tests. Live canary green (marker once at byte 612,182 of 651,795). #90 verified on device with Playwright. Open follow-up: **#91** — the userscript's client-side latest-chapter scan is still scoped to the derived slug. Reviewed-on: #92 Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com> Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com> |
||
|
|
e2c054e7ce |
Covers render in the userscript panel, from a public route (#60) (#69)
Closes #60. Spec: #55. Originating bug: #47. Architecture: `docs/adr/0007-backend-hosts-cover-bytes.md`. Neither #47 nor #55 is closed from here. ## What this branch does The panel now renders Covers from the deployment's own origin, and both userscripts stop having an opinion about where a Cover lives. **The public route was already in place.** `GET /covers/{address}` landed with #59 (`92eba07`) and is registered on the bare mux, outside `httpmw.Auth` and outside the web UI's Discord session — `backend/main.go:210-214`, handler `backend/internal/api/handlers.go:142-158`. It reads no cookie and no header, answers `404` for an address that was never stored (and for a row whose file has gone missing — recorded-but-gone is not-found, never a fabricated body), refuses anything that is not `^[0-9a-f]{64}$` *before* the value becomes a path, and sets `Cache-Control: public, max-age=604800, immutable`. Those four properties are asserted by `backend/cover_test.go:231-278`. This branch re-verified them rather than re-implementing them; the only backend line it touches is a comment. **Both userscripts lose cover scraping entirely.** Every adapter's `cover:` field is gone, along with the two helpers that fed them: the manga script's `coverFromPage()` (the `img[alt]` DOM scan comix needed, because comix publishes no `og:image`) and the novel script's `metaName()` plus the now-callerless module-level `meta()`. Nothing under `userscript/` reads `og:image`, `meta[name=image]`, or `img[alt]` any more. **Nothing sends a cover either.** `delete body.cover` sits in `apiPut` — `manga-bookmark.user.js:486`, `novel-bookmark.user.js:275` — which is the single chokepoint every write passes through (`pushBookmark`, the retry-queue flush, `toggleFavorite`, `toggleArchive`). It operates on the `Object.assign` copy, so the in-memory row keeps the cover it renders with. This matters beyond tidiness: a Reader upgrading from an older copy has `localStorage` rows carrying third-party scraped URLs, and without the strip those would ride back up on the next write. The handler discards the field regardless (`handlers.go:53-59`) — it is permanently inert, not pending removal. **Failed loads get the designed empty state, not the broken-image glyph.** `onerror: (e) => e.target.replaceWith(el("div", { class: "cover ph" }))` on the cover `<img>` in both card renderers (`manga:1380-1390`, `novel:1134-1144`). The replacement is byte-identical to the existing no-cover branch on the very next line, so it picks up the `.cover.ph` styling already in the panel CSS — no new tokens, no new rule. `el()` routes any `on*` prop through `addEventListener`, so this is a listener, not an inline attribute string, and the swap is a `createElement` + DOM call with no markup parsing anywhere near it. This is the half of #47 that was visible on kagane. **The deleted scraping's tests went with it**: the two comix cover cases, the `pageImages` and `namedMetas` fixtures, the `img[alt]` and `meta[name=...]` stub branches, the now-dead `querySelectorAll` stub member, and every stale `og:image` fixture and `p.cover` assertion across both suites. The export lists needed no change and that was checked, not assumed — `coverFromPage` and `metaName` were module-private on `origin/main` and no cover symbol ever appeared in `module.exports`. Docs that described the deleted behaviour were corrected in the same breath, because leaving them would instruct the next agent to put the scraping back: `userscript/AGENTS.md` (adapter contract + the per-site notes for comix, kagane and novelfull), the README's adapter reference, and the userscript testing skill's stub table. ## Verification - `go test -count=1 ./...` — green across all nine packages (`backend` 29.8s, `latest`, `store`, `session`, `token`, `userscript`, `web`). - `node --check` clean on both userscripts; `node --test` on both logic suites — 46 tests, 46 pass. - `gofmt -l` clean; `go build ./...` clean. - The `onerror` swap is DOM behaviour and deliberately has no coverage in the Node harness — that harness stubs a browser precisely so it never needs a DOM, and #60 says not to invent coverage for it. It was instead exercised for real: the `el()` helper and the exact render expression were loaded into a headless Chromium with a deliberately unloadable `src`, and the resulting DOM was `<div class="cover ph"></div>`. Ad hoc, not committed. - **Not done, needs you:** the on-device criterion — a comix Series bookmarked mid-chapter showing its Cover in the panel. That needs a real install against the deployment and is the one box left unticked on #60. ## Reviewed Both `/code-review` axes ran against `cc0fa92`. Spec found no missed requirement and no scope creep; standards found the diff clean on the four areas it scrutinised (the `delete body.cover` placement, the `onerror` handler's DOM safety, comment quality, dead-code removal). Their combined findings — the dead `querySelectorAll` stub, the stale README and skill text, and the handler comment whose premise this change invalidates — are fixed in `8b58019`. ## Out of scope, deliberately The kagane-specific cover proxy still exists and still carries its session gate (#63 deletes it). The poll's blank-Cover fill (#61) and browser-backed Sites joining the pipeline (#62) are untouched. Reviewed-on: #69 Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com> Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com> |
||
|
|
741b23322b |
Fix comix titles and covers, kagane volume chapters, and kagane cover rendering (#37)
Fixes five reported symptoms across comix.to and kagane.to. Diagnosing them turned up two latent bugs underneath, both of which had to be fixed for the kagane cover work to function at all.
## Reported symptoms and their causes
| # | Symptom | Cause |
|---|---------|-------|
| 1 | comix bookmark titled `Comix - Read Comics online for free` | comix is an SPA that rewrites `document.title` on client routing but never touches the server-rendered `og:title`. The adapter read `og:title`, so a cold load stored the homepage's title. |
| 2 | next comix bookmark gets the *previous* series' title | Same cause. After an in-page hop, `og:title` still holds whatever page loaded first. |
| 3 | comix cover shows the placeholder | comix serves no `og:image` at all, so `coverFromPage()` had nothing to read. |
| 4 | kagane chapter never appears in the bookmark list | Reader URLs carry no chapter number, so it is parsed out of `og:title`. Volume-numbered series render `"<Series> - Volume <v> Chapter <n>"`, which the suffix regex did not match, so `chapterNum` came back null and nothing was recorded. |
| 5 | kagane title includes the chapter, e.g. `SP Baby - Volume 1 Chapter 1` | Same unmatched regex — the tail was never stripped. One fix covers 4 and 5. |
| 6 | kagane cover blocked in the web UI | kagane serves covers behind its Cloudflare challenge **and** with `cross-origin-resource-policy: same-origin`. No `<img>` on the UI's origin can load one even from a browser holding the clearance cookie. Hot-linking cannot be made to work. |
## What changed
**Userscript.** comix titles now come from `document.title` with the chapter page's `" - Ch.<n>"` tail stripped, and the cover is the `img` whose `alt` matches the cleaned title. comix fills `document.title` a beat *after* the URL changes — later than the nav watcher's 300 ms snapshot — so the watcher also re-detects when the `detect()` signature changes, not only when the URL does. The kagane suffix regex takes an optional `Volume <v> ` segment. All three page shapes were captured live on 2026-08-08 and pinned as regression tests.
**Cover proxy.** `Bookmark.CoverURL()` rewrites a stored kagane `og:image` to `/img/kagane/{id}`; templates render `.CoverURL` instead of `.Cover`. The endpoint is session-gated like every other UI route and fetches through the shared headless browser, which is same-origin with kagane and so satisfies both the challenge and the CORP header. Results are memoised in-process, so a cover costs one navigation per deployment lifetime. With `BROWSER_WS_URL` unset the endpoint answers 404 rather than reaching for a nil fetcher — the same degrade-to-userscript behaviour the poller already has.
The image id is matched against a UUID regex before it reaches the browser. That gate is load-bearing rather than tidiness: the cover is a stored client-supplied string, so an unvalidated one turns this endpoint into an SSRF primitive aimed at the deployment's own network. `ServeMux` path-cleans a traversal into a redirect before the handler runs, but the handler does not depend on that, and a test pins it.
## Two latent bugs found underneath
**`BrowserFetcher.run` never let a challenge solve.** It navigated, waited for `body`, read once, and closed the tab — roughly half a second end to end. The Cloudflare interstitial has a `body` too, so `WaitReady` was satisfied by the challenge page itself. This made the challenge *unclearable* rather than merely slow: an interstitial needs several seconds of a live page to solve itself and write clearance into the browser's shared cookie jar, so tearing the tab down first means every subsequent call is challenged exactly like the one before it. `run` now holds one tab and re-reads until the caller's predicate reports an answer, bounded by `challengeTimeout` and the caller's own deadline. Exhausting the budget maps back to the 403 the poller already expects, keeping a challenged site distinct from a broken transport.
**`chromedp/headless-shell` cannot clear kagane's challenge at all.** It is a stripped Chrome build and the tells are structural rather than a header: `navigator.webdriver` is true, the plugin list is empty, and the client hints are Chromium- rather than Chrome-branded. Overriding `webdriver` through CDP was tried on its own and changed nothing.
All measured 2026-08-08 from one IP against the same cover, so the comparisons are like for like:
| Browser | Result |
|---------|--------|
| `chromedp/headless-shell:stable` | never cleared (90 s) |
| `zenika/alpine-chrome` | never cleared — ships Chrome 124, old enough that Cloudflare refuses it and old enough to break chromedp's CDP structs |
| `google-chrome`, default UA | never cleared (60 s) — `--headless=new` advertises `HeadlessChrome` |
| `google-chrome`, stock UA, `TZ=UTC` | never cleared (90 s) |
| `google-chrome`, stock UA, any non-UTC `TZ` | **cleared in ~4 s** |
Both remaining tells are load-bearing, and each was tested in isolation. `chrome/` is a Debian image with `google-chrome-stable`, a UA whose version is read back out of the binary at startup (a hardcoded one would drift out of step with the `Sec-CH-UA` hints on the next Chrome update and become a fresh tell), and no `--enable-automation`.
### The timezone tell: UTC, not a country mismatch
The first pass concluded the zone had to match the egress IP's country. Re-measuring against the actual deployment case shows that was wrong, and the correction is in `1552dd1`.
The original inference read the host's `/etc/timezone` (`Asia/Bangkok`) and assumed a Thai egress. It isn't — this host egresses from an Indonesian IP. `Asia/Bangkok` cleared not because it matched a country but because it simply isn't UTC, and the two share +07, which hid the distinction. Same container, same Indonesian IP:
| `TZ` | Result |
|------|--------|
| `UTC` | never cleared (60 s, **twice**) |
| `Asia/Jakarta` | cleared in 4 s |
| `America/New_York` | cleared in 4 s |
`America/New_York` matches neither the country nor the offset nor the hemisphere and clears just as fast. A UTC clock is itself the bot signal — Cloudflare scores it as the datacenter default — and any real zone satisfies the check. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one, and a deployment that changes region need not keep it in sync.
One sharp edge remains: the usual `-v /etc/localtime:/etc/localtime:ro` does **not** work. Chrome resolves the zone through ICU, which takes the name from that path's symlink target and ignores the file's contents, so glibc reports the host zone while Chrome still reports UTC. `/etc/timezone` carries the name and is mounted instead.
Chrome also binds its DevTools port to loopback and silently ignores `--remote-debugging-address`, which is why headless-shell fronted it with socat. This image does the same, so it stays a drop-in: the compose service keeps the `headless-shell` name and its pinned address, and `BROWSER_WS_URL` is unchanged.
## Verification
```
go test ./... all packages ok
node --test 37 + 12 pass, 0 fail
SMOKE_BROWSER_WS_URL=... go test -run TestSmokeKagane ./internal/latest
TestSmokeKaganeImage PASS (5.29s) fetched 56710 bytes of image/webp
TestSmokeKaganeGet PASS (1.17s) status=200, real chapter-list JSON
```
The smoke test ran against the exact compose configuration — built image, empty `BROWSER_TZ`, `/etc/timezone` mounted, cold profile — hitting real kagane.to. It skips unless `SMOKE_BROWSER_WS_URL` names a sidecar, so `go test ./...` stays hermetic and Docker-only.
A red smoke run means the challenge is not clearing from that IP, which is a live, time-varying fact to re-check rather than necessarily a defect.
## Security invariants
- Auth unchanged. `/img/kagane/{id}` is session-gated by `requireSession`, the same guard as every other UI route.
- Outbound fetch gated: the id is UUID-validated before it reaches the browser, keeping the existing rule that a client-supplied string never selects a fetch target unchecked.
- No new secrets, no new logging of credentials, no change to CORS, sessions, or crypto.
- Templates still escape everything; `.CoverURL` returns a plain string and is not wrapped in `template.HTML`/`URL`.
- One new dependency-free image (`chrome/`) built from Debian plus Google's own apt repo; no new Go modules.
## Deploying
Needs `docker compose build headless-shell`.
**A UTC host must set `BROWSER_TZ`, or kagane silently stops working.** With it unset the sidecar falls back to the host's `/etc/timezone`; on a UTC server that yields UTC, which is the one value that never clears. Any real zone works — `BROWSER_TZ=Asia/Jakarta` for the current deployment. `.env.example` now documents this; it previously did not mention the knob at all.
Only the browser sidecar reads `BROWSER_TZ`. The backend keeps its UTC clock, and stored timestamps are unix ms, so nothing else shifts.
## Deliberately not done
Retry/backoff around the cover proxy, and a panel-side cover fix. The panel renders no covers, and covers cache in-process after the first fetch. Worth adding if kagane starts rate-limiting.
## Correction after review of the deployment case
`1552dd1` was added after the branch was first pushed: the deployment host runs UTC with an Indonesian egress IP, which prompted re-measuring the timezone claim and falsifying it. The earlier commits' reasoning is left intact rather than rebased away, so the diagnostic trail — including the wrong turn and what disproved it — stays readable.
Reviewed-on: #37
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
|
||
|
|
4229c179b0 |
rebrand: MangaBM → BookmarkManager, add novel library support (#15)
Two intertwined changes — the rebrand and the novel library were developed on
the same branch because the novel UI plumbing is part of the new "Bookmark
Manager" wordmark in the web shell.
## What it does
- **Rebrand**: MangaBM → BookmarkManager across the Go module, compose stack,
env vars, Traefik hostnames, container/image names, userscript storage
prefixes (`mangabm:cache` → `bmgr:manga:cache`, `mangabm:queue` → `bmgr:manga:queue`),
and docs.
- **Novel library**: same backend, two libraries. New `kind` column splits
bookmarks into `manga` / `novel`; PUT validates it. Two userscripts:
- `manga-bookmark.user.js` — unchanged behaviour, just stamps its own `kind`.
- `novel-bookmark.user.js` — separate Violentmonkey install with adapters
for **novelfull.com** (polled via headless browser — Cloudflare JS
challenge) and **lightnovelworld.net** (polled via plain TLS).
- **Web UI**: library switch on the app shell. Login art, libswitch, and
novel-site colours from the Cinder design snapshot.
## Plumbing
- `addedColumns` ALTER for `kind` runs on first start after upgrade; every
pre-existing row is backfilled to `'manga'`. No manual SQL, no down-time.
- `ALLOWED_ORIGINS` gains the two novel sites.
- New `NOVEL_USERSCRIPT_PATH` env (default `/userscript/novel-bookmark.user.js`),
bindmounted alongside the manga script.
- Traefik router names `mangabm*` → `bmapi*` / `bmweb*`.
## Test status
- `go test ./...` — green
- `node --test userscript/test/logic.test.js` — 34 pass
- `node --test userscript/test/novel-logic.test.js` — 11 pass
- `node --check` on both userscripts — clean
## Notes for the redeploy
.env keys were renamed (`MANGA_API_HOST` → `BOOKMARK_API_HOST`,
`MANGA_WEB_HOST` → `BOOKMARK_WEB_HOST`). Update DNS / Traefik labels on the
prod override before pulling, otherwise the public hostnames go dark.
See the redeploy instructions I'll post next to this PR.
Reviewed-on: #15
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
|