ad09569f2e30dacc30480c568642a737ceb734de
5 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ddbd57070d |
Poll comix.to through the browser sidecar (#98) (#105)
Closes #98.
comix.to began answering plain-TLS fetches with a Cloudflare JavaScript
challenge on 2026-08-12, so every poll got a 403 interstitial. Its cover host
`static.comix.to` is gated the same way. comix therefore joins kagane and
novelfull as a browser-backed Site.
## What changed
- **Registry** (`internal/latest/sites.go`): comix gains a `Browser` entry —
`comixRead`, `Done: body != "" && !isInterstitial(body)`, `Fallback: false`.
Skip-when-no-browser falls out of the existing routing; no site-string compare
was added anywhere.
- **Read shape** (`internal/latest/browser.go`): an in-tab `fetch()` of the
Series URL, not a DOM render. comix is an SPA — rendering it costs ~65
requests for the same server-rendered HTML one fetch returns (24.5 KB,
~480 ms measured). `comixSeriesPageURL` pins scheme + host + `/title/<slug>`
and rebuilds the address, so a client-supplied `series_url` cannot aim the
browser anywhere else.
- **Cover bytes**: `comixImageURLRe` pins `https://static.comix.to/<path>.<ext>`;
`BrowserFetcher.Image` now gates on `browserOnlyCoverURL` rather than a
kagane-only regex, so both Sites' image URLs route through the one path.
Bytes come from direct navigation, not a page-context fetch — comix's Series
page sets `cross-origin-embedder-policy: require-corp`, which fails one.
- **Parsers and stored Series identity: untouched.** The in-tab body is the same
server-rendered HTML the existing fixtures were cut from.
## Verification
- `go test ./...` green (needs Docker).
- New seam tests: comix routes to the browser when one is configured, and is
not fetched at all when none is (`TestComixUsesBrowserFetcher`,
`TestComixSkippedWhenNoBrowserFetcher`); URL-pin and cover-gate table tests.
- Live proof against the real browser unit, `TestSmokeComix` (env-gated):
page 24793 bytes in one in-tab fetch, chapter 53, cover accepted by the pin,
26862 bytes of `image/jpg` retrieved.
- Two-axis review run; findings were stale comments on `BrowserFetcher`, `Get`
and the `Fallback` field, fixed in
|
||
|
|
e7e22a12a5 |
lightnovelworld Series identity is read from the chapter page (#80) (#92)
Implements spec #80 / ADR-0008 — Gitea issues #86, #87, #88, #89, #90, all closed. A Reader bookmarks a novel on lightnovelworld and it never shows a New Chapter, because the Series identity was derived from the chapter address instead of read from the page. One Series can publish under several Chapter Slugs, so the derived key points at a slug that 404s. - **#89** — the userscript's lnw adapter stops deriving `seriesUrl`/`seriesId` from the path. It reads the page's own pointer (`a[aria-label='All Chapter']`), falling back to the microdata breadcrumb's second crumb, and carries `chapterSlug` on the page object, stored nowhere. - **#87** — the Poll's lnw chapter scan is unscoped (no stored-slug pattern can cover a Series' whole list) and truncated at the `wpd-threads` comment thread, the one region a visitor can write to. Marker absent means skip and log with the body length, never scan whole. Corrects the `maxBodyBytes` headroom comment to the measured 3.5x. - **#86** — the scan fixture is now text trimmed from a real, wholly-fetched Series page instead of a hand-written cross-series anchor that no live page carries. - **#90** — stale stored rows repair themselves on the next chapter visit: a pure transform over cache, queue and last-checked map, silent to the Reader, with progress, favourite and lifecycle bucket preserved when two rows merge. - **#88** — an env-gated live canary (`SMOKE_LNW_SERIES_URL`) proving the marker still occurs exactly once and still follows the last chapter anchor, asserted against the production symbols themselves. Verified on the merged branch: `go test ./...` green, `gofmt -l internal/latest/` silent, both userscripts `node --check` clean, 35/35 + 29/29 logic tests. Live canary green (marker once at byte 612,182 of 651,795). #90 verified on device with Playwright. Open follow-up: **#91** — the userscript's client-side latest-chapter scan is still scoped to the derived slug. Reviewed-on: #92 Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com> Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com> |
||
|
|
b6b88bde8a |
feat(latest): extract per-site covers (#58) (#67)
Closes #58 ## Summary - Add pure per-Site cover extraction beside latest-chapter parsing for all six Sites. - Read Asura, Demonic, LightNovelWorld, and NovelFull metadata; read the Comix target detail state; read Kagane's browser-fetched `series_covers[].image_id` JSON. - Preserve published cover URLs, percent-encode Demonic raw spaces, select Comix's smaller published `medium`, and avoid thumbnail rendition URL synthesis. - Add live-source fixtures plus no-cover and Cloudflare challenge coverage for every Site. ## Correctness - Scope Comix extraction to the requested series detail key, avoiding recommended posters. - Parse Kagane's current live API shape and emit its canonical compressed image route from the published image ID; unrelated JSON fields are ignored. - Validate Kagane image IDs against the existing UUID-shaped route constraint. - Keep extraction pure; storage, polling, and wire integration remain outside issue #58. ## Acceptance criteria - [x] Cover extraction exists for all six Sites in the existing latest parser module. - [x] Each Site has a live-source fixture with source URL and date. - [x] Comix reads the state blob, not metadata. - [x] Demonic raw spaces are percent-encoded. - [x] Comix returns the smaller published rendition. - [x] No-cover pages return empty. - [x] Cloudflare challenge pages return empty. - [x] No thumbnail URL is synthesized by editing a published URL. - [x] `go test ./...` passes. ## Verification - `go test ./...` - `go vet ./...` - `git diff --check` Parent issues #47 and #55 remain open as requested. Reviewed-on: #67 Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com> Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com> |
||
|
|
4229c179b0 |
rebrand: MangaBM → BookmarkManager, add novel library support (#15)
Two intertwined changes — the rebrand and the novel library were developed on
the same branch because the novel UI plumbing is part of the new "Bookmark
Manager" wordmark in the web shell.
## What it does
- **Rebrand**: MangaBM → BookmarkManager across the Go module, compose stack,
env vars, Traefik hostnames, container/image names, userscript storage
prefixes (`mangabm:cache` → `bmgr:manga:cache`, `mangabm:queue` → `bmgr:manga:queue`),
and docs.
- **Novel library**: same backend, two libraries. New `kind` column splits
bookmarks into `manga` / `novel`; PUT validates it. Two userscripts:
- `manga-bookmark.user.js` — unchanged behaviour, just stamps its own `kind`.
- `novel-bookmark.user.js` — separate Violentmonkey install with adapters
for **novelfull.com** (polled via headless browser — Cloudflare JS
challenge) and **lightnovelworld.net** (polled via plain TLS).
- **Web UI**: library switch on the app shell. Login art, libswitch, and
novel-site colours from the Cinder design snapshot.
## Plumbing
- `addedColumns` ALTER for `kind` runs on first start after upgrade; every
pre-existing row is backfilled to `'manga'`. No manual SQL, no down-time.
- `ALLOWED_ORIGINS` gains the two novel sites.
- New `NOVEL_USERSCRIPT_PATH` env (default `/userscript/novel-bookmark.user.js`),
bindmounted alongside the manga script.
- Traefik router names `mangabm*` → `bmapi*` / `bmweb*`.
## Test status
- `go test ./...` — green
- `node --test userscript/test/logic.test.js` — 34 pass
- `node --test userscript/test/novel-logic.test.js` — 11 pass
- `node --check` on both userscripts — clean
## Notes for the redeploy
.env keys were renamed (`MANGA_API_HOST` → `BOOKMARK_API_HOST`,
`MANGA_WEB_HOST` → `BOOKMARK_WEB_HOST`). Update DNS / Traefik labels on the
prod override before pulling, otherwise the public hostnames go dark.
See the redeploy instructions I'll post next to this PR.
Reviewed-on: #15
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
|
||
|
|
180ee78b1f |
Add comix.to and kagane.to support (#13)
Tracks read progress on comix.to and kagane.to alongside asura and demonic, in both the userscript and the backend. Implements `docs/superpowers/plans/2026-08-03-comix-kagane-support.md`. ## Userscript - `comix` adapter — `/title/<id>-<slug>`; only the id prefix is identity (the slug follows the title). No `og:image`, so the cover is matched by `alt`. - `kagane` adapter — reader URLs are uuids with no chapter number, so it comes out of `og:title`; anchor scanning is structurally impossible, replaced by `latestChapterFromApi` against kagane's same-origin JSON API. - `seriesId` threaded through `latestChapterFromAnchors` so comix can scope its scan to its own series and a recommendation strip cannot win the maximum. - `@match` for both hosts, panel chips, v1.6.0. ## Backend - `latestChapterFrom` cases: comix parses the SSR JSON state blob (`latestChapterUrl`, scoped to the series id); kagane parses API JSON (`chapter_no`). - Poller allowlist extended; `Poller.BrowserFetch` with `fetcherFor(site)` routes kagane to a browser fetcher. Nil means kagane is not polled at all — never a fallback to the TLS fetcher, which would only ever retrieve a challenge page. - `BrowserFetcher`: chromedp against a `headless-shell` sidecar. kagane sits behind a Cloudflare JS challenge that no TLS fingerprint clears, and the request is made inside the page rather than by replaying `cf_clearance`. - `BROWSER_WS_URL` wiring, sidecar in both compose files (no `ports:`, dedicated non-external network), Dockerfile on `golang:1.26-alpine` — chromedp requires go 1.26. - Web UI `--comix` / `--kagane` tokens in both colour branches. ## Notes for review - `series_url` is client-supplied and a headless browser is a strong SSRF primitive, so kagane's host is pinned twice: in `fetchableSeriesURL` and again in `kaganeAPIURL`. - Three chained defects found during verification made the browser path dead under Compose (sidecar flag collision, Chrome's Host-header DNS-rebinding check, the wrong chromedp option). Fixed; the compose comments record the wrong configurations too, so they don't get "simplified" back. - `ALLOWED_ORIGINS` now includes both new origins. Without it every write from comix/kagane silently fails CORS preflight, parks in the retry queue, and drops at the cap. ## Verification 221 backend tests, 32 userscript tests, static `CGO_ENABLED=0` build, both compose configs. Two gaps, both real: 1. The userscript on live pages via Violentmonkey needs a human browser profile — not run. Check: comix series page (title/cover, no chapter), comix chapter page (records the number; an *older* chapter must not regress it), comix SPA navigation without reload, kagane series page (og:image cover), kagane reader (number from `og:title`), both chips opening the right sites. 2. The kagane browser path has not completed end-to-end anywhere. Dial/navigate/fetch is confirmed, but Cloudflare 403'd headless-shell's Chrome on every attempt from the dev sandbox, and comix's poll-through-Docker was blocked by that environment's TLS interception. Both environment-dependent rather than branch defects — the first real deploy is the actual verification. Reviewed-on: #13 Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com> Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com> |