Per-Site Cover extraction, fixture-backed #58

Closed
opened 2026-08-09 23:12:11 +07:00 by sulthan · 3 comments
Owner

Parent

Spec: #55. Originating bug: #47. Architecture and rejected alternatives: docs/adr/0007-backend-hosts-cover-bytes.md. Domain vocabulary: CONTEXT.md.

Do not close #47 or #55 from this ticket.

What to build

Given the body the poller already fetches for a Series, produce that Series' cover URL — for all six Sites, with each Site's rules pinned by a fixture trimmed from a live page. This sits beside the existing per-Site latest-chapter extraction, in the same module, driven by the same response, because a Site is one thing to learn and splitting its rules across two modules is how the second convention starts.

Extraction is a pure transformation and is independent of storage, fetching and the wire, so it can proceed in parallel with the other foundation work.

The rules, all verified live on 2026-08-09:

  • asura — og:image in the server-rendered HTML. Note its .webp cover URL answers Content-Type: image/jpeg; anything downstream must trust the header, never the extension.
  • demonic — og:image. The published value contains a raw unencoded space and must be percent-encoded before it can be fetched.
  • comix — there is no og:image anywhere on the page. The cover lives in the same server-rendered state blob the latest-chapter URL is already read from, under a poster object carrying both a medium (roughly 36 KB) and a large (full size) absolute URL. Take medium.
  • lightnovelworld — og:image.
  • novelfull — meta[name=image], absolute. Its HTML is challenge-gated, so this branch is fed by the browser-backed body the poller already retrieves for this Site.
  • kagane — the cover URL is carried in the series API JSON the browser-backed fetch already retrieves.

Where a Site publishes several renditions, prefer the smaller deterministically-named one — that is comix's medium. The largest surface a Cover is ever rendered into is a card, so storing full-resolution images spends disk and phone bandwidth on pixels nobody sees. Do not synthesise a thumbnail URL by editing a published one. asura's page uses a -400 variant that its metadata tag does not publish; deriving it by string surgery is precisely the guess that breaks silently when the Site changes.

Every fixture carries a comment naming the URL it was trimmed from and the date it was taken, matching the existing fixtures' convention. A page with no cover, and the Cloudflare challenge body the fixtures already include, must both yield nothing rather than a wrong value.

Acceptance criteria

  • Cover extraction exists for all six Sites, in the module that already holds latest-chapter extraction
  • Each Site has at least one fixture trimmed from a live page, commented with its source URL and date
  • comix extracts from the state blob, not from a metadata tag
  • demonic's raw space is percent-encoded in the produced URL
  • comix yields the smaller published rendition
  • A page carrying no cover yields empty, not a wrong value
  • The Cloudflare challenge fixture yields empty
  • No thumbnail URL is synthesised by editing a published URL
  • go test ./... is green

Blocked by

  • None — can start immediately.
## Parent Spec: #55. Originating bug: #47. Architecture and rejected alternatives: `docs/adr/0007-backend-hosts-cover-bytes.md`. Domain vocabulary: `CONTEXT.md`. Do not close #47 or #55 from this ticket. ## What to build Given the body the poller already fetches for a Series, produce that Series' cover URL — for all six Sites, with each Site's rules pinned by a fixture trimmed from a live page. This sits beside the existing per-Site latest-chapter extraction, in the same module, driven by the same response, because a Site is one thing to learn and splitting its rules across two modules is how the second convention starts. Extraction is a pure transformation and is independent of storage, fetching and the wire, so it can proceed in parallel with the other foundation work. The rules, all verified live on 2026-08-09: - **asura** — `og:image` in the server-rendered HTML. Note its `.webp` cover URL answers `Content-Type: image/jpeg`; anything downstream must trust the header, never the extension. - **demonic** — `og:image`. The published value contains a **raw unencoded space** and must be percent-encoded before it can be fetched. - **comix** — there is **no `og:image` anywhere on the page**. The cover lives in the same server-rendered state blob the latest-chapter URL is already read from, under a `poster` object carrying both a `medium` (roughly 36 KB) and a `large` (full size) absolute URL. Take `medium`. - **lightnovelworld** — `og:image`. - **novelfull** — `meta[name=image]`, absolute. Its HTML is challenge-gated, so this branch is fed by the browser-backed body the poller already retrieves for this Site. - **kagane** — the cover URL is carried in the series API JSON the browser-backed fetch already retrieves. Where a Site publishes several renditions, prefer the smaller deterministically-named one — that is comix's `medium`. The largest surface a Cover is ever rendered into is a card, so storing full-resolution images spends disk and phone bandwidth on pixels nobody sees. **Do not synthesise a thumbnail URL by editing a published one.** asura's page uses a `-400` variant that its metadata tag does not publish; deriving it by string surgery is precisely the guess that breaks silently when the Site changes. Every fixture carries a comment naming the URL it was trimmed from and the date it was taken, matching the existing fixtures' convention. A page with no cover, and the Cloudflare challenge body the fixtures already include, must both yield nothing rather than a wrong value. ## Acceptance criteria - [x] Cover extraction exists for all six Sites, in the module that already holds latest-chapter extraction - [x] Each Site has at least one fixture trimmed from a live page, commented with its source URL and date - [x] comix extracts from the state blob, not from a metadata tag - [x] demonic's raw space is percent-encoded in the produced URL - [x] comix yields the smaller published rendition - [x] A page carrying no cover yields empty, not a wrong value - [x] The Cloudflare challenge fixture yields empty - [x] No thumbnail URL is synthesised by editing a published URL - [x] `go test ./...` is green ## Blocked by - None — can start immediately.
sulthan added the ready-for-agent label 2026-08-09 23:12:11 +07:00
Author
Owner

Implemented in ff84eec. All nine acceptance criteria are checked: six site extractors, live-source fixtures, Comix medium state extraction, Demonic space encoding, no-cover and challenge empty results, and no URL rewriting. Added target-series scoping for Comix and field-based Kagane JSON extraction. Verification: go test ./... passes.

Implemented in ff84eec. All nine acceptance criteria are checked: six site extractors, live-source fixtures, Comix medium state extraction, Demonic space encoding, no-cover and challenge empty results, and no URL rewriting. Added target-series scoping for Comix and field-based Kagane JSON extraction. Verification: go test ./... passes.
sulthan added ready-for-human and removed ready-for-agent labels 2026-08-10 01:00:20 +07:00
Author
Owner

PR #67 opened: #67. Final commit d1801f4 includes Comix JSON detail parsing, Kagane top-level JSON decoding, shared Comix ID parsing, deterministic challenge tests, and empty-metadata fallback. go test ./... and go vet ./... pass.

PR #67 opened: https://gitea.violetcrown.my.id/sulthan/mangaBookmark/pulls/67. Final commit d1801f4 includes Comix JSON detail parsing, Kagane top-level JSON decoding, shared Comix ID parsing, deterministic challenge tests, and empty-metadata fallback. `go test ./...` and `go vet ./...` pass.
Author
Owner

Updated implementation and PR #67 (Closes #58).

Acceptance checklist remains fully complete:

  • all six Site extractors live beside latest-chapter parsing
  • six live-source fixtures with source URL/date comments
  • Comix reads target detail state and returns published medium rendition
  • Demonic raw space is percent-encoded
  • no-cover and Cloudflare challenge bodies return empty
  • no thumbnail rendition URL is synthesized
  • go test ./... is green

Follow-up correction in 65edb92: the live Kagane series response publishes series_covers[].image_id, not a top-level cover URL. The extractor now reads that field, validates the ID with the existing constraint, and emits Kagane's canonical compressed image route. go vet ./... and git diff --check also pass. Parent issues #47 and #55 remain open.

Updated implementation and PR #67 (Closes #58). Acceptance checklist remains fully complete: - all six Site extractors live beside latest-chapter parsing - six live-source fixtures with source URL/date comments - Comix reads target detail state and returns published medium rendition - Demonic raw space is percent-encoded - no-cover and Cloudflare challenge bodies return empty - no thumbnail rendition URL is synthesized - `go test ./...` is green Follow-up correction in 65edb92: the live Kagane series response publishes `series_covers[].image_id`, not a top-level `cover` URL. The extractor now reads that field, validates the ID with the existing constraint, and emits Kagane's canonical compressed image route. `go vet ./...` and `git diff --check` also pass. Parent issues #47 and #55 remain open.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: sulthan/mangaBookmark#58