One registry entry per Site, and one shared Series-page read for the Poll and the Acquisition #94
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Problem Statement
Adding support for a new Site means finding six unrelated places that compare
the site string, spread across three files in the
latestpackage: the LatestChapter parse switch, the Cover parse switch, the browser-backed Site list, the
fetcher choice, the host pin, and the payload read inside the browser. Nothing
ties them together and nothing fails at compile time, so a Site wired into five
of the six works in development and then misbehaves in production — silently
skipped, or fetched over the wrong route, or parsed against the wrong document.
The same friction shows up a second way. The Poll and the Acquisition both need
the same five steps — gate the address, choose the route, fetch the page, find
the Latest Chapter, find the Cover address — and each has its own copy. The two
copies have already drifted once: cover-byte routing had to be deliberately
pulled out into a shared address-shaped rule to stop them diverging again.
Both problems land on the same person: whoever adds the seventh Site, or changes
how any existing Site is read.
Solution
Give each Site exactly one entry in one registry, and make the Poll and the
Acquisition share the single read they both perform.
A registry entry answers a fixed set of questions about its Site: which hostname
its addresses must carry, how to find the Latest Chapter in a page body, how to
find the Cover address in a page body, and — only for a Site behind a JavaScript
challenge — how to read its payload from a cleared browser tab and how to know
the payload arrived. Everything about how a Site behaves lives in its entry, and
nowhere else.
The Poll and the Acquisition keep their own purposes and their own policies.
Only the read they have in common moves into one module, which returns the two
facts it found and persists nothing.
Recorded in ADR-0009. The domain term for the creation-time read is
Acquisition, now defined in
CONTEXT.md.User Stories
Implementation Decisions
Phase one — the Site registry.
latestpackage, keyed by the stored site string, holding one entry per Site. Six entries: asura, demonic, comix, kagane, novelfull, lightnovelworld.interfacewith a method set. Most Sites differ in one or two answers and three share a single Cover implementation, so a method set would produce near-empty types. A missing answer is a nil value, handled where an unknown Site is already handled.www.or subdomain tolerance — matching the existing pins.Phase two — the shared Series-page read.
Ordering. Phase one lands before phase two, so that the shared read is written against the registry rather than against the switches it replaces.
Testing Decisions
A good test here asserts observable behaviour: what the Poll or the Acquisition
writes to a Series given a page body, which Sites are considered fetchable, and
what happens when a fetch fails. It does not assert the shape of the registry,
the presence of a field, or that a particular function was called.
Out of Scope
latestpackage's parse layer. Noted as a real observation about the test surface, but not required by this work and not attempted here.Further Notes
ADR-0009 records the principle behind the registry's shape and names the failure
mode a future review is likely to hit: an override used by exactly one Site
looks like an inconsistency to collapse, and collapsing it means either widening
the question set for every Site or scattering the shared tab lifecycle. Both
were rejected on evidence.
Three facts constrain the browser side and are easy to lose in a refactor. The
tab is held open across re-reads because a challenge needs several seconds of
live page to solve itself and write clearance into the shared cookie jar —
reading once and closing the tab clears nothing. The browser is on-demand and on
a separate machine, so an unreachable one must degrade exactly as no browser at
all. And the two Sites behind the browser have different payloads: one exposes
its chapter list only through an API that must be called from inside the page,
the other renders its chapters into the served HTML.
The pinning of the three previously unpinned Sites is the only behavioural
change in this work. For one of them the pin is a strict improvement rather than
a restriction: its old domain redirects deep links to the site root, discarding
the path, so a stored address on that domain currently fetches the wrong
document and parses it. After pinning it fails cleanly instead.