2d134fb05c9b4fa7d14255a477ba2f15117bb678
7 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2d134fb05c |
feat: one registry entry per Site, derived browser route and gate (#94)
Collapse the six per-site comparison points into a sites map in sites.go: Latest Chapter parse, Cover parse, browser-backed list, fetcher route, host pins, and the browser payload read all become lookups into it. All six hosts are now pinned in fetchableSeriesURL; asura/demonic/comix were previously accepted on any https host. |
||
|
|
e7e22a12a5 |
lightnovelworld Series identity is read from the chapter page (#80) (#92)
Implements spec #80 / ADR-0008 — Gitea issues #86, #87, #88, #89, #90, all closed. A Reader bookmarks a novel on lightnovelworld and it never shows a New Chapter, because the Series identity was derived from the chapter address instead of read from the page. One Series can publish under several Chapter Slugs, so the derived key points at a slug that 404s. - **#89** — the userscript's lnw adapter stops deriving `seriesUrl`/`seriesId` from the path. It reads the page's own pointer (`a[aria-label='All Chapter']`), falling back to the microdata breadcrumb's second crumb, and carries `chapterSlug` on the page object, stored nowhere. - **#87** — the Poll's lnw chapter scan is unscoped (no stored-slug pattern can cover a Series' whole list) and truncated at the `wpd-threads` comment thread, the one region a visitor can write to. Marker absent means skip and log with the body length, never scan whole. Corrects the `maxBodyBytes` headroom comment to the measured 3.5x. - **#86** — the scan fixture is now text trimmed from a real, wholly-fetched Series page instead of a hand-written cross-series anchor that no live page carries. - **#90** — stale stored rows repair themselves on the next chapter visit: a pure transform over cache, queue and last-checked map, silent to the Reader, with progress, favourite and lifecycle bucket preserved when two rows merge. - **#88** — an env-gated live canary (`SMOKE_LNW_SERIES_URL`) proving the marker still occurs exactly once and still follows the last chapter anchor, asserted against the production symbols themselves. Verified on the merged branch: `go test ./...` green, `gofmt -l internal/latest/` silent, both userscripts `node --check` clean, 35/35 + 29/29 logic tests. Live canary green (marker once at byte 612,182 of 651,795). #90 verified on device with Playwright. Open follow-up: **#91** — the userscript's client-side latest-chapter scan is still scoped to the derived slug. Reviewed-on: #92 Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com> Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com> |
||
|
|
7c7d597019 |
Delete the kagane-specific cover path (#63) (#73)
Closes #63 Deletes the second way to reach a Cover. Since #62, every Site's cover bytes land in the content-addressed store at creation or on the poll, and the one public route serves them all — nothing needs the kagane proxy anymore. ## What went - **Template-level rewrite:** `Bookmark.CoverURL()` and both templates' use of it. Cards and chrome now render `.Cover` — the wire value — and nothing else. `Bookmark.CoverSource` was dead once `CoverURL` went, so it and its `bookmarkColumns` entry are gone too. - **Kagane-only cover route and its identifier validation:** `GET /img/kagane/{id}`, `web.CoverFetcher`, `coverIDRe`, and the whole `internal/web/cover.go`. - **The proxy's persistence:** `store.KaganeImageID`, `GetKaganeCover`, `PutKaganeCover`, `kaganeCoverSourceURL`, `kaganeCoverRe`. - **The kagane-shaped branch in the byte-fetch routing:** `fetchCoverBytes` no longer takes a `site` argument and no longer names a Site. The URL shape kagane's API publishes is claimed by the browser module itself — `kaganeImageURLRe` + `browserCoverURL` live in `latest/browser.go` with the rest of the per-Site knowledge — and `BrowserFetcher.Image` is now URL-driven (it validates the URL it will navigate to, same SSRF discipline as before). The no-plain-TLS-fallback rule for a claimed URL is preserved: a claimed address with no browser is an error, never a challenge-page fetch. ## What stayed (deliberately) - `BrowserFetcher.Image` and the browser-backed acquisition path: kagane genuinely serves cover bytes behind the challenge + `cross-origin-resource-policy: same-origin`, so the sidecar remains the only fetcher for them — it just routes by URL claim now instead of by Site name. - `fetcherFor`'s per-Site page routing (kagane/novelfull page fetches) — that is the page path, not a cover path. ## Acceptance criteria - [x] Template-level kagane cover rewrite gone - [x] Kagane-only cover route and its identifier validation gone - [x] Tests removed/rewritten against the general route, guarantees kept: unstored + traversal-shaped addresses serve nothing (`TestPublicCoverRejectsUnknownAddress`), non-image content types never echoed (`TestPublicCoverNeverEchoesNonImage` — new; the store-side gate was already pinned by `TestCoverStoreAcceptsAnySourceURL`). Store reopen-persistence and filesystem content-addressing tests rewritten against `PutCover`/`GetCover`, no guarantee lost. - [x] No Site name in a cover code path outside the acquisition module (`grep kagane backend`: store/web/templates/api are clean; remaining hits are `latest/browser.go` + `latest/sites.go`, tests, docs) - [x] Web UI and panel render Covers for all six Sites (templates render the wire address; panel renders `b.cover` — untouched, it never had a kagane path) - [x] `go test ./...` green ## Verification - `go vet ./...` clean - `go test ./...` — all packages pass (root 16.9s, latest 12.7s, store 12.7s, web 0.004s) - `CGO_ENABLED=0 go build` produces the static binary - Cover-path tests run verbosely: `TestPublicCoverServesStoredBytesUnauthenticated`, `TestPublicCoverRejectsUnknownAddress` (unknown/malformed/traversal/empty), `TestPublicCoverNeverEchoesNonImage`, `TestListRendersAcquiredCover`, `TestAcquireKaganeCoverThroughBrowser`, `TestRunOncePrefetchesKaganeCover`, `TestRunOnceRoutesNonKaganeCoverToPublicFetcher` all pass; the three `SMOKE_*` tests skip without the browser sidecar, as designed Live browser verification of the "web UI and panel render Covers for all six Sites" criterion is being run separately with Playwright against real Site pages and a locally mocked backend. Reviewed-on: #73 Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com> Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com> |
||
|
|
b6b88bde8a |
feat(latest): extract per-site covers (#58) (#67)
Closes #58 ## Summary - Add pure per-Site cover extraction beside latest-chapter parsing for all six Sites. - Read Asura, Demonic, LightNovelWorld, and NovelFull metadata; read the Comix target detail state; read Kagane's browser-fetched `series_covers[].image_id` JSON. - Preserve published cover URLs, percent-encode Demonic raw spaces, select Comix's smaller published `medium`, and avoid thumbnail rendition URL synthesis. - Add live-source fixtures plus no-cover and Cloudflare challenge coverage for every Site. ## Correctness - Scope Comix extraction to the requested series detail key, avoiding recommended posters. - Parse Kagane's current live API shape and emit its canonical compressed image route from the published image ID; unrelated JSON fields are ignored. - Validate Kagane image IDs against the existing UUID-shaped route constraint. - Keep extraction pure; storage, polling, and wire integration remain outside issue #58. ## Acceptance criteria - [x] Cover extraction exists for all six Sites in the existing latest parser module. - [x] Each Site has a live-source fixture with source URL and date. - [x] Comix reads the state blob, not metadata. - [x] Demonic raw spaces are percent-encoded. - [x] Comix returns the smaller published rendition. - [x] No-cover pages return empty. - [x] Cloudflare challenge pages return empty. - [x] No thumbnail URL is synthesized by editing a published URL. - [x] `go test ./...` passes. ## Verification - `go test ./...` - `go vet ./...` - `git diff --check` Parent issues #47 and #55 remain open as requested. Reviewed-on: #67 Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com> Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com> |
||
|
|
08749df050 |
feat(backend)!: run on Postgres with a migration-owned schema (#28)
Swap modernc.org/sqlite for jackc/pgx/v5 with no observable change: same endpoints, same wire format, same updated_at ordering rule. The schema now comes from numbered SQL embedded in the binary and applied on startup, one transaction each, recorded in schema_migrations. That replaces two pieces of SQLite-era machinery, both deleted rather than ported: the column probing (Postgres has ADD COLUMN IF NOT EXISTS, and there is no legacy database left to probe) and the Asura key rewrite, which has run clean on every start for months now that the userscripts strip build hashes before writing. Its regexp survives as latest.asuraBuildHash, where the poller still needs it to scope chapter links to a series whose slug carries a rotating hash. Types get real: favorite is a boolean, chapter numbers double precision, timestamps stay unix-ms bigint. SQLite's null-safe IS NOT becomes IS DISTINCT FROM, which is what implements the rule that only reading progress reorders a list. Inside COALESCE/NULLIF the status and kind parameters need an explicit ::text -- there is no target column to infer from and Postgres refuses to guess. Tests lose their free t.TempDir() database, so Docker is now a hard prerequisite for `go test ./...`: internal/pgtest starts one postgres:17-alpine per test binary and hands each test a database of its own. Also lands CONTEXT.md and the four ADRs written while scoping #18. BREAKING CHANGE: DB_PATH is retired for DATABASE_URL, which is required and has no default. Compose gains a postgres service on an internal network with its own volume; POSTGRES_PASSWORD joins .env. The old bookmarks-data volume is deliberately left undeclared so `docker compose down -v` cannot take the pre-migration database with it. main is not deployable until #25 and #26 land. Closes #20 Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com> Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com> |
||
|
|
4229c179b0 |
rebrand: MangaBM → BookmarkManager, add novel library support (#15)
Two intertwined changes — the rebrand and the novel library were developed on
the same branch because the novel UI plumbing is part of the new "Bookmark
Manager" wordmark in the web shell.
## What it does
- **Rebrand**: MangaBM → BookmarkManager across the Go module, compose stack,
env vars, Traefik hostnames, container/image names, userscript storage
prefixes (`mangabm:cache` → `bmgr:manga:cache`, `mangabm:queue` → `bmgr:manga:queue`),
and docs.
- **Novel library**: same backend, two libraries. New `kind` column splits
bookmarks into `manga` / `novel`; PUT validates it. Two userscripts:
- `manga-bookmark.user.js` — unchanged behaviour, just stamps its own `kind`.
- `novel-bookmark.user.js` — separate Violentmonkey install with adapters
for **novelfull.com** (polled via headless browser — Cloudflare JS
challenge) and **lightnovelworld.net** (polled via plain TLS).
- **Web UI**: library switch on the app shell. Login art, libswitch, and
novel-site colours from the Cinder design snapshot.
## Plumbing
- `addedColumns` ALTER for `kind` runs on first start after upgrade; every
pre-existing row is backfilled to `'manga'`. No manual SQL, no down-time.
- `ALLOWED_ORIGINS` gains the two novel sites.
- New `NOVEL_USERSCRIPT_PATH` env (default `/userscript/novel-bookmark.user.js`),
bindmounted alongside the manga script.
- Traefik router names `mangabm*` → `bmapi*` / `bmweb*`.
## Test status
- `go test ./...` — green
- `node --test userscript/test/logic.test.js` — 34 pass
- `node --test userscript/test/novel-logic.test.js` — 11 pass
- `node --check` on both userscripts — clean
## Notes for the redeploy
.env keys were renamed (`MANGA_API_HOST` → `BOOKMARK_API_HOST`,
`MANGA_WEB_HOST` → `BOOKMARK_WEB_HOST`). Update DNS / Traefik labels on the
prod override before pulling, otherwise the public hostnames go dark.
See the redeploy instructions I'll post next to this PR.
Reviewed-on: #15
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
|
||
|
|
180ee78b1f |
Add comix.to and kagane.to support (#13)
Tracks read progress on comix.to and kagane.to alongside asura and demonic, in both the userscript and the backend. Implements `docs/superpowers/plans/2026-08-03-comix-kagane-support.md`. ## Userscript - `comix` adapter — `/title/<id>-<slug>`; only the id prefix is identity (the slug follows the title). No `og:image`, so the cover is matched by `alt`. - `kagane` adapter — reader URLs are uuids with no chapter number, so it comes out of `og:title`; anchor scanning is structurally impossible, replaced by `latestChapterFromApi` against kagane's same-origin JSON API. - `seriesId` threaded through `latestChapterFromAnchors` so comix can scope its scan to its own series and a recommendation strip cannot win the maximum. - `@match` for both hosts, panel chips, v1.6.0. ## Backend - `latestChapterFrom` cases: comix parses the SSR JSON state blob (`latestChapterUrl`, scoped to the series id); kagane parses API JSON (`chapter_no`). - Poller allowlist extended; `Poller.BrowserFetch` with `fetcherFor(site)` routes kagane to a browser fetcher. Nil means kagane is not polled at all — never a fallback to the TLS fetcher, which would only ever retrieve a challenge page. - `BrowserFetcher`: chromedp against a `headless-shell` sidecar. kagane sits behind a Cloudflare JS challenge that no TLS fingerprint clears, and the request is made inside the page rather than by replaying `cf_clearance`. - `BROWSER_WS_URL` wiring, sidecar in both compose files (no `ports:`, dedicated non-external network), Dockerfile on `golang:1.26-alpine` — chromedp requires go 1.26. - Web UI `--comix` / `--kagane` tokens in both colour branches. ## Notes for review - `series_url` is client-supplied and a headless browser is a strong SSRF primitive, so kagane's host is pinned twice: in `fetchableSeriesURL` and again in `kaganeAPIURL`. - Three chained defects found during verification made the browser path dead under Compose (sidecar flag collision, Chrome's Host-header DNS-rebinding check, the wrong chromedp option). Fixed; the compose comments record the wrong configurations too, so they don't get "simplified" back. - `ALLOWED_ORIGINS` now includes both new origins. Without it every write from comix/kagane silently fails CORS preflight, parks in the retry queue, and drops at the cap. ## Verification 221 backend tests, 32 userscript tests, static `CGO_ENABLED=0` build, both compose configs. Two gaps, both real: 1. The userscript on live pages via Violentmonkey needs a human browser profile — not run. Check: comix series page (title/cover, no chapter), comix chapter page (records the number; an *older* chapter must not regress it), comix SPA navigation without reload, kagane series page (og:image cover), kagane reader (number from `og:title`), both chips opening the right sites. 2. The kagane browser path has not completed end-to-end anywhere. Dial/navigate/fetch is confirmed, but Cloudflare 403'd headless-shell's Chrome on every attempt from the dev sandbox, and comix's poll-through-Docker was blocked by that environment's TLS interception. Both environment-dependent rather than branch defects — the first real deploy is the actual verification. Reviewed-on: #13 Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com> Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com> |