Per-Site Cover extraction, fixture-backed #58
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Parent
Spec: #55. Originating bug: #47. Architecture and rejected alternatives:
docs/adr/0007-backend-hosts-cover-bytes.md. Domain vocabulary:CONTEXT.md.Do not close #47 or #55 from this ticket.
What to build
Given the body the poller already fetches for a Series, produce that Series' cover URL — for all six Sites, with each Site's rules pinned by a fixture trimmed from a live page. This sits beside the existing per-Site latest-chapter extraction, in the same module, driven by the same response, because a Site is one thing to learn and splitting its rules across two modules is how the second convention starts.
Extraction is a pure transformation and is independent of storage, fetching and the wire, so it can proceed in parallel with the other foundation work.
The rules, all verified live on 2026-08-09:
og:imagein the server-rendered HTML. Note its.webpcover URL answersContent-Type: image/jpeg; anything downstream must trust the header, never the extension.og:image. The published value contains a raw unencoded space and must be percent-encoded before it can be fetched.og:imageanywhere on the page. The cover lives in the same server-rendered state blob the latest-chapter URL is already read from, under aposterobject carrying both amedium(roughly 36 KB) and alarge(full size) absolute URL. Takemedium.og:image.meta[name=image], absolute. Its HTML is challenge-gated, so this branch is fed by the browser-backed body the poller already retrieves for this Site.Where a Site publishes several renditions, prefer the smaller deterministically-named one — that is comix's
medium. The largest surface a Cover is ever rendered into is a card, so storing full-resolution images spends disk and phone bandwidth on pixels nobody sees. Do not synthesise a thumbnail URL by editing a published one. asura's page uses a-400variant that its metadata tag does not publish; deriving it by string surgery is precisely the guess that breaks silently when the Site changes.Every fixture carries a comment naming the URL it was trimmed from and the date it was taken, matching the existing fixtures' convention. A page with no cover, and the Cloudflare challenge body the fixtures already include, must both yield nothing rather than a wrong value.
Acceptance criteria
go test ./...is greenBlocked by
Implemented in
ff84eec. All nine acceptance criteria are checked: six site extractors, live-source fixtures, Comix medium state extraction, Demonic space encoding, no-cover and challenge empty results, and no URL rewriting. Added target-series scoping for Comix and field-based Kagane JSON extraction. Verification: go test ./... passes.sulthan referenced this issue2026-08-10 01:05:38 +07:00
sulthan referenced this issue2026-08-10 01:05:56 +07:00
PR #67 opened: #67. Final commit
d1801f4includes Comix JSON detail parsing, Kagane top-level JSON decoding, shared Comix ID parsing, deterministic challenge tests, and empty-metadata fallback.go test ./...andgo vet ./...pass.sulthan referenced this issue2026-08-10 01:14:46 +07:00
Updated implementation and PR #67 (Closes #58).
Acceptance checklist remains fully complete:
go test ./...is greenFollow-up correction in
65edb92: the live Kagane series response publishesseries_covers[].image_id, not a top-levelcoverURL. The extractor now reads that field, validates the ID with the existing constraint, and emits Kagane's canonical compressed image route.go vet ./...andgit diff --checkalso pass. Parent issues #47 and #55 remain open.