Files
mangaBookmark/docs/adr/0007-backend-hosts-cover-bytes.md
sulthan 30c57bd39c Define Cover and record hosting its bytes (#47) (#64)
Defines **Cover** in the glossary and records ADR-0007, the decision behind #47's fix.

## Why these two files, and why now

`CONTEXT.md` named Cover inside the **Series** entry — "facts true regardless of who is reading — title, cover, Latest Chapter" — but never said *what* one is. That gap is the bug. Nothing in the model distinguished "an address on a Site" from "an image a Reader's browser can display", so both clients were left to work it out independently, and one of them got it wrong. kagane serves covers with `cross-origin-resource-policy: same-origin`, the web UI rewrote them to a proxy in its templates, the JSON API did not, and the panel rendered a broken-image glyph. The new entry closes the ambiguity: *an address no client can load is not a Cover, it is a missing one.*

ADR-0007 records what follows from that — the backend fetches, stores and serves every Site's cover bytes — plus the alternatives that were rejected and, more importantly, the two places this deliberately departs from existing precedent:

- **Destination-class control instead of a host allowlist.** `fetchableSeriesURL` sets the allowlist precedent for `series_url`, and covers do not follow it. Cover hosts are CDNs that move independently of their Site — demonicscans serves its covers from `readermc.org` — so an allowlist would stop producing Covers the day a Site switched CDN, and that failure would look exactly like #47. The resolve-then-classify step is what actually stops the SSRF.
- **A public cover route where the kagane proxy is session-gated.** An `<img>` cannot send a bearer token, and it cannot be given one either: the panel's shadow root is `mode: "open"`, so the host page's JavaScript can read any `src` the script sets.

Both are security-adjacent departures, which is precisely why they are written down rather than left in a commit message.

## Scope

Documentation only — no code, no schema, no behaviour. The implementation is #56–#63.

## Why this should merge promptly rather than sit

All eight implementation tickets cite `docs/adr/0007-backend-hosts-cover-bytes.md` as the authority for decisions they must not relitigate, and they are written in the vocabulary this glossary entry defines. An agent picking up #56 reads both from `main`. Until this lands they get a 404 and either invent a rationale or stall — so this PR gates the tickets, not the other way round.

Related: #47 (bug), #55 (spec), #54 (deferred admin refetch).
Reviewed-on: #64
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-09 23:19:37 +07:00

5.7 KiB

ADR-0007: The backend hosts every Site's Cover bytes

Date: 2026-08-09 Status: accepted

Decision

A Cover is the image a Reader's browser can display for a Series. Where the Site keeps the picture is the backend's problem, not the client's: the backend fetches the bytes, stores them, and serves them from its own origin. No client ever renders a third-party URL, and no client ever supplies one.

Concretely:

  • Acquisition is server-side. The Cover is extracted from the same series-page fetch that already yields Latest Chapter. It runs once at Series creation rather than waiting for the poll queue, so a newly bookmarked Series has both facts in seconds instead of up to a queue's depth. The poll fills a blank Cover and never overwrites a non-blank one.
  • Bytes live on a filesystem volume, content-addressed by the SHA-256 of the source URL, sharded ${COVER_DIR}/ab/cd/<sha256>. The database holds the path and content type, not the bytes.
  • One public route serves them. No session, no credential.
  • The wire carries an absolute URL built from a configured public base, and carries "" until the bytes exist.

Why a future reader will find this surprising

Four of the six Sites let anyone hot-link their covers — static.comix.to even answers access-control-allow-origin: *. Hosting copies looks like work we were not obliged to do.

We were obliged. kagane serves covers with cross-origin-resource-policy: same-origin behind a JavaScript challenge (measured 2026-08-08), so no <img> outside kagane.to can load one under any combination of referrer policy and crossorigin attribute. The first fix for that was a kagane-only proxy applied in the web templates — and it produced issue #47, because the JSON API kept emitting the raw kagane URL and the userscript rendered it into a broken-image glyph. A per-Site exception that only one of two clients knows about is not a fix; it is a bug with a delay on it. Uniformity is the property being bought: every client renders every Cover the same way, and a Site changing its CORP header or its CDN cannot break a client again.

Considered options

Per-Site exceptions, proxying only what must be proxied. Cheapest, and what we had. Rejected: it is what produced #47, and it requires every current and future client to know which Sites are special.

A host allowlist for the outbound fetch, mirroring fetchableSeriesURL. Rejected in favour of destination-class control — see below.

Cover bytes in Postgres bytea, extending the existing covers table. Rejected: covers are immutable blobs served straight to browsers, which is what a filesystem is for. The cost is real and accepted — durability is now two things to back up instead of one, against ADR-0001's grain.

Per-Reader Cover overrides. Rejected, consistent with ADR-0003's rejection of per-Reader title overrides. A Cover is a fact about the Series.

Two deliberate relaxations

Destination control is deny-class, not an allowlist. The outbound fetch requires https, resolves DNS first and refuses loopback, private, link-local and CGNAT addresses, re-checks on every redirect hop, and caps body size and content type. It does not pin a host set, which is what fetchableSeriesURL does for series_url. Cover hosts are CDNs that move: demonicscans.org serves its covers from readermc.org, a host with no visible relationship to the Site. An allowlist would silently stop producing Covers the day a Site switched CDN, and the failure would look like this bug. The resolved-IP check is the load-bearing part; without it, an attacker-controlled page need only publish a DNS name pointing at 127.0.0.1.

The cover route is public, where the kagane proxy was session-gated. An <img> in the userscript panel cannot send a bearer token, and it cannot be given one: the panel's shadow root is mode: "open", so the host page's own JavaScript can read any src we set. A credential in an image URL is a credential handed to a third-party site. The route serves public artwork from public Sites and its path reveals nothing about which Reader holds what. The residual cost is that we can be hot-linked by others.

Consequences

  • The poll's cover prefetch, today guarded by sr.Site != "kagane", applies to every Site in both Libraries. Nothing about Covers is conditioned on kind.
  • Only kagane still needs the CDP browser for its bytes. The other five Sites fetch over plain TLS — including novelfull, whose HTML answers cf-mitigated: challenge while its image paths answer 200 with access-control-allow-origin: * (measured 2026-08-09).
  • Client-side cover scraping is deleted from both userscripts. It could not help: a scraped URL has no render path left, and it is absent exactly when a Series is created — neither comix nor lightnovelworld exposes a cover on a chapter page, which is where a Reader bookmarks mid-read.
  • PUT /bookmarks/{key} still accepts a cover field and ignores it. This extends ADR-0003's "ignored after creation" to "ignored always", and keeps the flat wire contract ADR-0004 requires so installed scripts keep working. The field is therefore permanently inert rather than pending removal, and says so at the decode site.
  • A Cover that fails to load falls back to the placeholder in both clients. The broken-image glyph reported in #47 is not a state we render.
  • The existing kagane covers rows are dropped rather than migrated; that path re-fetches on demand already.
  • Two Sites deserve a note for whoever writes the extractor: asura's .webp cover URL answers Content-Type: image/jpeg, so trust the header; demonic's og:image carries a raw unencoded space and must be percent-encoded before fetching.