Compare commits
9 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 1f5d0695ad | |||
| c62c3bb07b | |||
| 21615be2bd | |||
| 672c16ffbf | |||
| e7e22a12a5 | |||
| 90d8ab72ad | |||
| c400c91a80 | |||
| f1eb7d514c | |||
| 1ee5eb67ea |
@@ -0,0 +1,134 @@
|
||||
---
|
||||
name: implement-tickets
|
||||
description: "Orchestrate a batch of tickets: plan the briefs, then hand each ticket to its own implementer subagent in its own worktree."
|
||||
disable-model-invocation: true
|
||||
---
|
||||
|
||||
# Implement tickets
|
||||
|
||||
You are the **orchestrator**. You write briefs, dispatch, land results, and talk
|
||||
to the tracker. You do not write the implementation — every line of ticket code
|
||||
is written by a `ticket-implementer` subagent in its own git worktree. Reach for
|
||||
the editor yourself only for a merge conflict resolution.
|
||||
|
||||
Ticket source and `tea` usage: `docs/agents/issue-tracker.md`. Codebase
|
||||
questions: `graphify query "<question>"` before grepping.
|
||||
|
||||
## 1. Collect the tickets
|
||||
|
||||
The user's argument is the selector: issue numbers, a label, a parent issue, or
|
||||
nothing. With nothing, take the open issues labelled `ready-for-agent`.
|
||||
|
||||
Fetch each with `tea issue <n> --comments`, and read the **whole** body —
|
||||
acceptance criteria and the `Blocked by` line are what the rest of this skill
|
||||
runs on. A ticket whose blockers are still open is out of this batch unless a
|
||||
blocker is also in it.
|
||||
|
||||
## 2. Plan the batch
|
||||
|
||||
Explore enough of the codebase to write briefs a fresh context can act on: the
|
||||
files each ticket lands in, the patterns it must follow, the `AGENTS.md`
|
||||
invariants it touches.
|
||||
|
||||
Then decide three things:
|
||||
|
||||
- **Waves.** Blocking edges set the order; tickets with no open blocker inside
|
||||
the batch share a wave. Cap each wave at **3** concurrent tickets unless the
|
||||
user set another width.
|
||||
- **Contracts.** Two tickets in one wave that meet at a function signature, a
|
||||
JSON shape, a table column, or a token name: you decide the shape now and
|
||||
write the identical wording into both briefs. A contract left for the
|
||||
subagents to negotiate is a merge conflict you scheduled.
|
||||
- **Splits.** A ticket too big for one fresh context window goes into the wave
|
||||
as two briefs, or back to the user.
|
||||
|
||||
## 3. Get the plan approved
|
||||
|
||||
Present, and stop:
|
||||
|
||||
- the wave list, and for each ticket: number, title, one-line brief summary,
|
||||
the files or areas it will touch, its verification commands
|
||||
- every cross-ticket contract, verbatim as it will appear in the briefs
|
||||
- anything you had to assume
|
||||
|
||||
Wait for approval. Apply the user's edits to the plan, do not relitigate them.
|
||||
|
||||
## 4. Run a wave
|
||||
|
||||
Per ticket, before dispatch:
|
||||
|
||||
```bash
|
||||
git worktree add ../ticket-<n> -b ticket/<n>-<slug> <base> # base = the branch you are on
|
||||
cp .env ../ticket-<n>/ 2>/dev/null # gitignored, worktrees do not get it
|
||||
tea issue edit <n> --add-assignees <your gitea username> # tea login list has it
|
||||
```
|
||||
|
||||
Write the brief to `.scratch/<batch-slug>/t<n>-brief.md` using the template
|
||||
below, in the ubiquitous language of `CONTEXT.md` — a brief that says "scrape"
|
||||
where the domain says Poll hands the subagent the wrong model of the system.
|
||||
Then dispatch the whole wave in **one** `task` batch, every item on the
|
||||
`ticket-implementer` agent. Each dispatch names: the absolute brief path, the
|
||||
worktree path, the branch, the base ref, and the report path
|
||||
`.scratch/<batch-slug>/t<n>-report.md`.
|
||||
|
||||
<brief-template>
|
||||
|
||||
# Ticket #<n> — <title>
|
||||
|
||||
**Read first.** `tea issue <n> --comments` for this ticket, then the issue it
|
||||
refers to — the parent or spec — the same way. The comments carry decisions the
|
||||
body never got updated with. This brief stays the requirements; those two reads
|
||||
are the intent behind them.
|
||||
|
||||
**Goal.** The end-to-end behaviour this ticket makes work, from the user's side.
|
||||
|
||||
**Acceptance criteria.** Verbatim from the ticket.
|
||||
|
||||
**Contract.** The exact shared signatures / shapes / names this ticket must
|
||||
implement or consume, and which sibling ticket is on the other end. Omit when
|
||||
the ticket touches nothing shared.
|
||||
|
||||
**Where it lands.** The files and packages, and the existing pattern to follow
|
||||
in each.
|
||||
|
||||
**Binding invariants.** The `AGENTS.md` rules this change can break — name them.
|
||||
|
||||
**TDD seams.** Where a test comes first — run the `tdd` skill at each one and
|
||||
follow its red → green loop. Or "none — verify after".
|
||||
|
||||
**Verify.** The exact commands, e.g. `cd backend && go test ./...`,
|
||||
`node --test userscript/test/logic.test.js`.
|
||||
|
||||
**Out of scope.** What not to touch, especially a sibling ticket's files.
|
||||
|
||||
</brief-template>
|
||||
|
||||
## 5. Land the wave
|
||||
|
||||
The wave is landed when every ticket in it is closed, reverted, or handed back
|
||||
to the user. Per returned ticket:
|
||||
|
||||
| Status | What you do |
|
||||
| --- | --- |
|
||||
| `DONE` | merge, comment, close |
|
||||
| `DONE_WITH_CONCERNS` | merge, comment the concerns, close only if you judge them non-blocking — otherwise leave open and tell the user |
|
||||
| `BLOCKED` / `NEEDS_CONTEXT` | supply what is missing and re-dispatch, or hand back to the user with the specifics. Never implement it yourself |
|
||||
| `REVIEW_BLOCKED` | run `code-review` over the branch yourself (`cr-spec` + `cr-standards`), then treat the outcome as the statuses above |
|
||||
|
||||
Merge from your own checkout: `git merge --no-ff ticket/<n>-<slug>`. A textual
|
||||
conflict is yours to resolve (`resolving-merge-conflicts`). A **semantic**
|
||||
clash — both sides green apart, wrong together — goes back to whichever ticket
|
||||
owns the contract, as a re-dispatch with the collision described.
|
||||
|
||||
Then `tea comment <n> "<the report summary>"`, `tea issue close <n>`, and
|
||||
`git worktree remove ../ticket-<n>`. Keep the report file.
|
||||
|
||||
Only once the whole wave is landed does the next wave start — its briefs may
|
||||
need what this one changed.
|
||||
|
||||
## 6. Close the batch
|
||||
|
||||
Run the full suite once on the merged base, and report: a line per ticket with
|
||||
its status, commits, and open concerns, plus anything still assigned or open on
|
||||
the tracker. A red suite after every ticket went green is an interaction bug —
|
||||
diagnose it, name the two tickets, and fix it or hand it back with both named.
|
||||
+1
-1
@@ -15,7 +15,7 @@ OWNER_DISCORD_ID=changeme-your-discord-user-id
|
||||
# Comma-separated origins allowed to call the API (CORS). Both Asura domains
|
||||
# plus Demonic, Comix, Kagane, and the two novel sites. Add/remove as the
|
||||
# sites' hostnames change.
|
||||
ALLOWED_ORIGINS=https://asuracomic.net,https://asurascans.com,https://demonicscans.org,https://comix.to,https://kagane.to,https://novelfull.com,https://lightnovelworld.net
|
||||
ALLOWED_ORIGINS=https://asurascans.com,https://demonicscans.org,https://comix.to,https://kagane.to,https://novelfull.com,https://lightnovelworld.net
|
||||
|
||||
# Password for the bundled Postgres container, and therefore half of the
|
||||
# DATABASE_URL compose builds for the backend. Generate one:
|
||||
|
||||
+6
-1
@@ -5,8 +5,13 @@
|
||||
backend/server
|
||||
backend/backend
|
||||
.playwright-mcp/
|
||||
graphify-out/
|
||||
# graphify map is committed; only regenerable/local parts are ignored
|
||||
graphify-out/cost.json
|
||||
graphify-out/cache/
|
||||
graphify-out/[0-9][0-9][0-9][0-9]-[0-9][0-9]-[0-9][0-9]/
|
||||
graphify-out/.rebuild.lock
|
||||
plans/
|
||||
.scratch/
|
||||
docs/superpowers/
|
||||
.superpowers/
|
||||
go.work
|
||||
|
||||
@@ -0,0 +1,90 @@
|
||||
---
|
||||
name: ticket-implementer
|
||||
description: Implements one ticket end to end inside its own git worktree - reads a brief file, implements, tests, commits, runs the two-axis code review through cr-spec and cr-standards, fixes findings, writes a report file, returns a short status contract. Dispatched by the implement-tickets skill.
|
||||
model: opencode-go/minimax-m3
|
||||
thinking-level: high
|
||||
tools: read, write, edit, bash, grep, glob, lsp, todo, ast_edit, task
|
||||
spawns: cr-spec,cr-standards
|
||||
autoloadSkills: code-review, tdd
|
||||
---
|
||||
|
||||
You implement **one ticket** dispatched by an orchestrator. Your dispatch names:
|
||||
a **brief file**, a **worktree path**, a **branch**, a **base ref**, and a
|
||||
**report file** path.
|
||||
|
||||
## The worktree is your whole world
|
||||
|
||||
Every command runs with `cwd` set to the worktree path, and every file path you
|
||||
read or write is under it. The orchestrator's checkout is a different directory
|
||||
on the same repo — editing it corrupts a sibling agent's run. If a command must
|
||||
run elsewhere, say so in the report instead of doing it.
|
||||
|
||||
Your branch is already checked out there. Never `git checkout`, `git switch`,
|
||||
`git rebase`, or `git worktree` anything.
|
||||
|
||||
## Order of work
|
||||
|
||||
1. Read the brief file. It is the single source of requirements — use its exact
|
||||
values verbatim.
|
||||
2. Read the ticket and the issue it refers to, as the brief's **Read first**
|
||||
section names them: `tea issue <n> --comments` for each. The ticket's
|
||||
comments and its parent carry the intent and the decisions behind the brief.
|
||||
Read no other ticket and no other brief.
|
||||
3. Read `AGENTS.md` in the worktree, plus the nested `AGENTS.md` for the area
|
||||
you touch. Its invariants bind you: security rules, design system, comment
|
||||
policy.
|
||||
4. Ask before writing code if requirements, acceptance criteria, approach, or
|
||||
dependencies are unclear. Asking is free; guessing is not.
|
||||
5. Implement exactly what the brief specifies. At each TDD seam the brief names,
|
||||
run the `tdd` skill and follow its red → green loop.
|
||||
Follow the patterns already in the codebase; improve what you touch,
|
||||
restructure nothing outside the ticket.
|
||||
6. Verify. Focused tests while iterating, the brief's full verification commands
|
||||
once at the end. Test output must be pristine.
|
||||
7. Commit to your branch. Reference the ticket number in the subject.
|
||||
8. Review (below), fix, re-verify, commit the fixes.
|
||||
9. Write the report file, then return the status contract.
|
||||
|
||||
## Review
|
||||
|
||||
After your first green commit, run the **`code-review`** skill over
|
||||
`<base ref>...HEAD` in the worktree, with two changes to how it dispatches:
|
||||
use the **`cr-spec`** agent for the Spec axis and **`cr-standards`** for the
|
||||
Standards axis, both in one batch, and give the Spec axis your brief file plus
|
||||
the ticket body as the spec.
|
||||
|
||||
Fix every Critical and Important finding, then re-run the tests that cover the
|
||||
amended code. Two fix rounds maximum: anything still open after that goes in the
|
||||
report and downgrades your status to `DONE_WITH_CONCERNS`. Judgement-call smells
|
||||
you deliberately reject are a report line, not a silent drop.
|
||||
|
||||
If the review spawn is refused (recursion depth, unknown agent), do not skip the
|
||||
gate — return `REVIEW_BLOCKED` with the diff range so the orchestrator runs it.
|
||||
|
||||
## Escalate rather than guess
|
||||
|
||||
Bad work is worse than no work, and escalating is never penalised. Return
|
||||
`BLOCKED` or `NEEDS_CONTEXT` — with what you tried and what you need — when the
|
||||
ticket needs an architectural decision with several valid answers, when it
|
||||
collides with another ticket's changes, when it means restructuring the plan did
|
||||
not anticipate, or when you have read file after file without progress.
|
||||
|
||||
## Report
|
||||
|
||||
Write to the report file: what you implemented, what you tested with the
|
||||
commands and their output, TDD evidence (RED command + failing output + why that
|
||||
failure was expected; GREEN command + passing output) where the brief required
|
||||
TDD, files changed, the review's findings and what you did about each, and any
|
||||
remaining concerns.
|
||||
|
||||
Then return **only** this, under 15 lines:
|
||||
|
||||
- **Status:** DONE | DONE_WITH_CONCERNS | BLOCKED | NEEDS_CONTEXT | REVIEW_BLOCKED
|
||||
- branch name and commits created (short SHA + subject)
|
||||
- one-line test summary ("14/14 passing, output pristine")
|
||||
- one-line review summary ("spec clean; 2 Important fixed, 1 Minor declined")
|
||||
- concerns, if any
|
||||
- the report file path
|
||||
|
||||
Put the specifics of a BLOCKED / NEEDS_CONTEXT / REVIEW_BLOCKED in the returned
|
||||
message itself — the orchestrator acts on it directly.
|
||||
@@ -6,7 +6,7 @@ Guidance for OpenCode (and Claude Code) working in this repo.
|
||||
|
||||
Read-progress tracker for two libraries — manga and novels — behind one self-hosted Go backend. Two separate Violentmonkey userscripts inject on-page UI (floating button + slide-in panel) and sync progress, so bookmarks unify across sites and devices:
|
||||
|
||||
- `manga-bookmark.user.js` — **asurascans.com** (current domain; asuracomic.net 301s here), **demonicscans.org**, **comix.to**, **kagane.to**.
|
||||
- `manga-bookmark.user.js` — **asurascans.com** (asuracomic.net is dropped: its deep links 301 to the asurascans.com root, discarding the path), **demonicscans.org**, **comix.to**, **kagane.to**.
|
||||
- `novel-bookmark.user.js` — **novelfull.com**, **lightnovelworld.net**.
|
||||
|
||||
One backend, one `bookmarks` table: a `kind` column (`manga`|`novel`) splits the libraries and the web UI switches between them. Rows are keyed `<site>:<series_id>`.
|
||||
@@ -18,11 +18,11 @@ Userscript targets **Violentmonkey**, so `GM_*` APIs available, but stay GM-free
|
||||
- Cross-origin `fetch()` work **only** against CORS-enabled backend. Manga sites `https://`, so backend **must be HTTPS** (else mixed-content block).
|
||||
- Every site is its **own origin with its own `localStorage`** — a shared remote store is the only way to unify bookmarks. Cloud sync required, not optional.
|
||||
- Userscript run in **isolated world**, so embedded API token safe from site's JS.
|
||||
- Cloudflare's block on manga sites **IP-reputation-based, not universal — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — Cloudflare's bot scoring can flip previously-clean IP without notice. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test.
|
||||
- Cloudflare's block on manga sites is **per-zone configuration plus request fingerprint, not IP reputation — and not reliably reproducible.** Verified 2026-07-26: plain `curl` from both CGNAT dev machine *and* deployed VPS got clean 200s with real HTML on both asurascans.com and demonicscans.org (homepage, series, chapter pages) — no interactive Turnstile challenge from either IP at test time. Contradicts earlier untested assumption CGNAT dev IP blocked; wasn't, at least this date. Treat "does curl work right now" as live, time-varying fact to re-check, not fixed property of machine — a Site can turn its protection on overnight, which is exactly what comix.to did on 2026-08-12. An earlier version of this line blamed "Cloudflare's bot scoring"; that was wrong. The 1-99 bot score is Enterprise Bot Management only and does not exist for a free-plan zone, and no per-IP request rate is documented as an input to challenge issuance — `docs/research/cloudflare-bot-scoring-and-poll-cadence.md`. Backend fetcher still needs graceful-degrade path for when challenged, and adapters should be **verified against live pages** (Playwright MCP, on-device devtools, direct probe) before finalizing, not assumed from single earlier test.
|
||||
- **kagane.to and novelfull.com are the exception to the above** — both sit behind a Cloudflare JavaScript challenge no TLS fingerprint clears, so the backend polls them over CDP (`BROWSER_WS_URL`). When that's unset, kagane is skipped entirely (a plain fetch would only retrieve a challenge page) while novelfull pages are still attempted over plain TLS — its challenge is a live time-varying fact and its cover bytes never need the browser. The four other sites poll fine over plain TLS.
|
||||
- **The CDP browser must look like a real browser, and stock headless images don't.** Measured 2026-08-08 against kagane.to, all from the same IP: `chromedp/headless-shell:stable` never cleared the challenge in 90s (`navigator.webdriver` true, empty plugin list, Chromium-branded client hints — suppressing `webdriver` alone changed nothing); `zenika/alpine-chrome` ships Chrome 124, refused outright; real Chrome with the default `--headless=new` UA never cleared, because the UA says `HeadlessChrome`; real Chrome with a stock UA **and** a non-UTC clock zone cleared in ~4s. Hence `chrome/` — a Debian image with `google-chrome-stable`, a version-derived UA, and `TZ`/`BROWSER_TZ`. Chrome reads the zone *name* through ICU from `/etc/localtime`'s symlink target, ignoring the file's contents, so mounting the host's `/etc/localtime` does **not** work; `/etc/timezone` is mounted instead.
|
||||
- **The browser is not in the API stack and must not be put back.** It's its own compose unit (`chrome/docker-compose.yml`) on a second machine, reached over the tailnet — it held 471 MiB on a 1974 MiB swapless VPS, and a residential egress scores better with Cloudflare anyway (ADR-0006). Consequences that constrain code: `BROWSER_WS_URL` must be a tailnet **IP** (a MagicDNS name 500s at `/json/version`, same trap as the old Docker service name); the CDP port binds to the tailnet address only, since CDP authenticates nothing and that host has a real LAN; and the browser is on-demand (ADR-0005), so an unreachable or asleep one must degrade exactly as an unset `BROWSER_WS_URL` — plain-TLS libraries unaffected, kagane/novelfull logged and skipped, stored covers still served. Never add `chromedp.NoModifyURL`: discovery per fetch is what makes a restarted Chrome invisible.
|
||||
- **UTC is the tell, not a country mismatch.** A UTC clock is the datacenter default, so Cloudflare scores it as one; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while `Asia/Jakarta` **and** `America/New_York` both cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (`Asia/Bangkok`) rather than the measured egress. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one.
|
||||
- **The browser is not in the API stack and must not be put back.** It's its own compose unit (`chrome/docker-compose.yml`) on a second machine, reached over the tailnet — it held 471 MiB on a 1974 MiB swapless VPS, and a residential egress avoids the cloud-hosting-IP signature Bot Fight Mode documentedly challenges (ADR-0006; not a better "score" — free-plan zones have no score). Consequences that constrain code: `BROWSER_WS_URL` must be a tailnet **IP** (a MagicDNS name 500s at `/json/version`, same trap as the old Docker service name); the CDP port binds to the tailnet address only, since CDP authenticates nothing and that host has a real LAN; and the browser is on-demand (ADR-0005), so an unreachable or asleep one must degrade exactly as an unset `BROWSER_WS_URL` — plain-TLS libraries unaffected, kagane/novelfull logged and skipped, stored covers still served. Never add `chromedp.NoModifyURL`: discovery per fetch is what makes a restarted Chrome invisible.
|
||||
- **UTC is the tell, not a country mismatch.** A UTC clock is the datacenter default, and the challenge refuses it; any real zone clears. Measured 2026-08-08, identical container, one Indonesian egress IP: UTC never cleared in 60s (twice), while `Asia/Jakarta` **and** `America/New_York` both cleared in 4s. An earlier note here claimed the zone had to match the egress IP's country — that was wrong, inferred from the host clock (`Asia/Bangkok`) rather than the measured egress. A second earlier claim, that Cloudflare "scores" a UTC clock, was also wrong: the measurement is real but the mechanism is not documented anywhere — Cloudflare publishes no timezone signal, and free-plan zones carry no score at all. `BROWSER_TZ` therefore needs a plausible zone, not a geolocated one.
|
||||
- **A challenged page needs the tab kept open.** The interstitial takes seconds to solve and only then writes clearance into the browser's shared cookie jar. Navigate-read-close never clears anything; `BrowserFetcher.run` holds one tab and re-reads until the payload arrives.
|
||||
|
||||
## Architecture
|
||||
@@ -140,12 +140,6 @@ Style: one dense comment over function beats one per line inside. Tight, no work
|
||||
|
||||
Test: "competent reader get this from code in few sec?" Yes → skip. Needs detour through another file/spec/git-blame → write it.
|
||||
|
||||
## Relevant skills
|
||||
|
||||
`multi-stage-dockerfile` and `docker-compose-orchestration` for container work (referenced in plan).
|
||||
|
||||
`golang-code-style`, `golang-error-handling`, `golang-performance`, `golang-testing` for backend Go work.
|
||||
|
||||
## Agent skills
|
||||
|
||||
`AGENTS.md` is the single source of truth for agent guidance; every `CLAUDE.md` in this repo is a symlink to the `AGENTS.md` beside it. Edit `AGENTS.md`.
|
||||
@@ -167,11 +161,7 @@ Single-context: one root `CONTEXT.md` plus `docs/adr/`, both created lazily. See
|
||||
Project has knowledge graph at graphify-out/ with god nodes, community structure, cross-file relationships.
|
||||
|
||||
Rules:
|
||||
- For codebase questions, first run `graphify query "<question>"` when graphify-out/graph.json exists. Use `graphify path "<A>" "<B>"` for relationships and `graphify explain "<concept>"` for focused concepts. Return scoped subgraph, usually much smaller than GRAPH_REPORT.md or raw grep output.
|
||||
- For codebase questions and exploration, always first run `graphify query "<question>"` when graphify-out/graph.json exists. Use `graphify path "<A>" "<B>"` for relationships and `graphify explain "<concept>"` for focused concepts. Return scoped subgraph, usually much smaller than GRAPH_REPORT.md or raw grep output.
|
||||
- If graphify-out/wiki/index.md exists, use for broad navigation instead of raw source browsing.
|
||||
- Read graphify-out/GRAPH_REPORT.md only for broad architecture review or when query/path/explain don't surface enough context.
|
||||
- After modifying code, run `graphify update .` to keep graph current (AST-only, no API cost).
|
||||
|
||||
## Notes
|
||||
|
||||
- Keep comms terse — drop articles, fluff, pleasantries. Code/commits/security written normally.
|
||||
|
||||
+27
-8
@@ -7,17 +7,24 @@ so progress survives across sites and devices.
|
||||
## Language
|
||||
|
||||
**Series**:
|
||||
One ongoing work — a manga or a novel — as published by a Site. Identified by its
|
||||
stable slug on that Site, never by its title. A Series exists once and is shared by
|
||||
every Reader who bookmarks it; it owns the facts that are true regardless of who is
|
||||
reading — title, cover, Latest Chapter. A Reader cannot change them; they describe the
|
||||
Series, not anyone's relationship to it.
|
||||
One ongoing work — a manga or a novel — as published by a Site. Identified by the canonical
|
||||
slug the Site itself publishes for it, never by its title and never by a Chapter Slug. A
|
||||
Series exists once and is shared by every Reader who bookmarks it; it owns the facts that
|
||||
are true regardless of who is reading — title, cover, Latest Chapter. A Reader cannot
|
||||
change them; they describe the Series, not anyone's relationship to it.
|
||||
_Avoid_: manga, title, book, comic
|
||||
|
||||
**Site**:
|
||||
One third-party source a Series is published on. A Series on two Sites is two Series.
|
||||
_Avoid_: source, host, provider, domain
|
||||
|
||||
**Chapter Slug**:
|
||||
A slug a Site builds its chapter addresses from. Not an identity: one Series may have
|
||||
several, any of them may differ from the slug that identifies the Series, and none is
|
||||
computable from another. Only the Site's own links say which ones a Series uses, so a
|
||||
Chapter Slug is always discovered, never derived.
|
||||
_Avoid_: series slug, url slug, permalink, chapter path
|
||||
|
||||
**Cover**:
|
||||
The image that stands for a Series wherever it is listed. A fact about the Series like
|
||||
its title — one Cover per Series, shared by every Reader, never per-Reader. Defined by
|
||||
@@ -49,9 +56,12 @@ is real activity, so only Progress reorders the list.
|
||||
_Avoid_: position, bookmark (the noun is taken), last read
|
||||
|
||||
**Latest Chapter**:
|
||||
The newest chapter a Site has published for a Series, discovered without the reader
|
||||
present. Distinct from Progress in every way that matters: it is a fact about the Site,
|
||||
not about the reader, and it must never reorder the list.
|
||||
The highest-numbered chapter a Site has published for a Series, discovered without the
|
||||
reader present. The number is what ranks it, never a date and never the Site's own
|
||||
"newest chapter" banner — where a Site disagrees with itself, its list of chapters is
|
||||
the record and its summary of that list is not. Distinct from Progress in every way
|
||||
that matters: it is a fact about the Site, not about the reader, and it must never
|
||||
reorder the list.
|
||||
_Avoid_: newest, current chapter, update
|
||||
|
||||
**Poll**:
|
||||
@@ -60,6 +70,15 @@ Reader present. Performed once per Series no matter how many Readers bookmarked
|
||||
a Poll is work done on behalf of the Series, never on behalf of a Reader.
|
||||
_Avoid_: scrape, refresh, check, sync
|
||||
|
||||
**Acquisition**:
|
||||
The single read of a Series page made the moment the Series first exists, giving it
|
||||
both its Latest Chapter and its Cover without waiting out the Poll queue. Distinct
|
||||
from a Poll in the two ways that matter: a Reader is present — it is triggered by
|
||||
their first Bookmark of that Series — and it is the only read that establishes a
|
||||
Cover rather than refreshing facts. It happens once in a Series's life; every later
|
||||
read of the same page is a Poll.
|
||||
_Avoid_: initial poll, first fetch, prefetch, warm-up
|
||||
|
||||
**New Chapter**:
|
||||
The state where Latest Chapter is ahead of Progress. The single condition the ember
|
||||
accent is permitted to signal.
|
||||
|
||||
@@ -43,7 +43,7 @@ TOKEN_KEY=<paste output of: openssl rand -hex 32>
|
||||
OWNER_DISCORD_ID=<discord user id>
|
||||
|
||||
# CORS allowlist — leave as-is unless a site changes hostname.
|
||||
ALLOWED_ORIGINS=https://asuracomic.net,https://asurascans.com,https://demonicscans.org,https://comix.to,https://kagane.to
|
||||
ALLOWED_ORIGINS=https://asurascans.com,https://demonicscans.org,https://comix.to,https://kagane.to,https://novelfull.com,https://lightnovelworld.net
|
||||
|
||||
# Required — password for the bundled Postgres container. Compose builds the
|
||||
# backend's DATABASE_URL out of it and has no fallback for either.
|
||||
@@ -275,7 +275,8 @@ copy immediately — reinstall on all devices, or they silently stop syncing.
|
||||
Kagane and novelfull sit behind a Cloudflare JavaScript challenge no TLS
|
||||
fingerprint clears, so the poller reaches them through a real Chrome over CDP.
|
||||
That browser does **not** run on the VPS: it held 471 MiB of a 1974 MiB box
|
||||
with no swap, and it scores better from a residential IP anyway (ADR-0006). It
|
||||
with no swap, and a residential IP avoids the cloud-hosting-IP signature Bot
|
||||
Fight Mode challenges anyway (ADR-0006). It
|
||||
is its own compose unit, deployed and updated independently of everything
|
||||
above.
|
||||
|
||||
@@ -454,7 +455,8 @@ SMOKE_BROWSER_WS_URL=ws://100.x.y.z:9222 go test -run TestSmokeKagane ./internal
|
||||
```
|
||||
|
||||
A red run means "not clearing from this address right now", which is a live
|
||||
fact to re-check before it is a defect — Cloudflare's scoring moves. Then, from
|
||||
fact to re-check before it is a defect — a Site's Cloudflare settings, and the
|
||||
fingerprint this Chrome presents after an update, both move. Then, from
|
||||
the web UI, open a bookmarked kagane series and confirm the cover renders. Once
|
||||
a cover is stored it is served from Postgres forever after, so the browser being
|
||||
asleep, unreachable, or mid-power-outage costs chapter freshness and nothing
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Manga Bookmark
|
||||
|
||||
Track manga read-progress on **asurascans.com** (a.k.a. asuracomic.net),
|
||||
Track manga read-progress on **asurascans.com**,
|
||||
**demonicscans.org**, **comix.to**, and **kagane.to** from a phone (Bromite /
|
||||
mobile Chromium), synced to a self-hosted Go backend so bookmarks unify across
|
||||
all four sites and all devices.
|
||||
@@ -274,11 +274,12 @@ the backend acquires, stores and serves every Cover from its own origin
|
||||
| **Kagane** (`kagane.to`) | `/series/<uuid>` | `/series/<uuid>/reader/<bookUuid>` | `<uuid>` |
|
||||
|
||||
Notes:
|
||||
- **`asuracomic.net` deep links are dead (re-checked 2026-07-25).** They 301 to
|
||||
the `asurascans.com` **root**, discarding the path, at the edge — before the
|
||||
userscript gets a document — so nothing client-side can rescue them. Reach
|
||||
series through `asurascans.com`. The host stays matched in case the redirect
|
||||
starts preserving paths again.
|
||||
- **`asuracomic.net` is no longer matched (deep links dead, re-checked
|
||||
2026-07-25).** They 301 to the `asurascans.com` **root**, discarding the path,
|
||||
at the edge — before the userscript gets a document — so nothing client-side
|
||||
can rescue them. The backend rejects stored addresses on that host too, since
|
||||
the poller pins each Site to one hostname. Reach series through
|
||||
`asurascans.com`.
|
||||
- Asura `og:title` carries a `Chapter N - Read Online \| Asura Scans` suffix that
|
||||
the adapter strips; Demonic chapter `og:title` is `<Title> Chapter N`.
|
||||
- Demonic's `<slug>` is identical on `/manga/…` and the canonical `/title/…`
|
||||
|
||||
+7
-1
@@ -149,7 +149,13 @@ Guidance for OpenCode (and Claude Code) working under `backend/`. See root `AGEN
|
||||
`LATEST_CHAPTER_POLL_ENABLED`/`_COOLDOWN`/`_BROWSER_COOLDOWN`/`_INTERVAL`/
|
||||
`_BATCH`/`_STAGGER` (background latest-chapter poller; defaults on,
|
||||
`1h` plain-TLS cooldown, `6h` browser cooldown, `10m`/`14`/`20s`; both
|
||||
cooldowns have a `15m` floor).
|
||||
cooldowns have a `15m` floor). The browser cooldown is longer for cost, not
|
||||
for safety: a challenged page costs seconds of a serialized single-tab
|
||||
browser, while a plain read costs one request. It buys no documented
|
||||
reduction in challenge risk — free-plan zones have no bot score and no
|
||||
published per-IP rate input, and `cf_clearance` expires in 30 minutes so
|
||||
every cadence at or above 1h re-solves anyway —
|
||||
`docs/research/cloudflare-bot-scoring-and-poll-cadence.md`.
|
||||
`USERSCRIPT_PATH` and `NOVEL_USERSCRIPT_PATH` (files served at
|
||||
`/u/{token}/manga-bookmark.user.js` and `/u/{token}/novel-bookmark.user.js`,
|
||||
defaults `/userscript/manga-bookmark.user.js` and
|
||||
|
||||
+6
-6
@@ -33,7 +33,7 @@ const testCoverBaseURL = "https://bookmarks.test"
|
||||
func testConfig() Config {
|
||||
return Config{
|
||||
TokenKey: testTokenKey,
|
||||
AllowedOrigins: []string{"https://asuracomic.net", "https://demonicscans.org"},
|
||||
AllowedOrigins: []string{"https://asurascans.com", "https://demonicscans.org"},
|
||||
Port: "8080",
|
||||
}
|
||||
}
|
||||
@@ -169,7 +169,7 @@ func TestAuthAccepted(t *testing.T) {
|
||||
func TestCORSPreflight(t *testing.T) {
|
||||
srv := newTestServer(t)
|
||||
req := httptest.NewRequest(http.MethodOptions, "/bookmarks/asura:foo-1", nil)
|
||||
req.Header.Set("Origin", "https://asuracomic.net")
|
||||
req.Header.Set("Origin", "https://asurascans.com")
|
||||
req.Header.Set("Access-Control-Request-Method", "PUT")
|
||||
rr := httptest.NewRecorder()
|
||||
srv.ServeHTTP(rr, req)
|
||||
@@ -177,7 +177,7 @@ func TestCORSPreflight(t *testing.T) {
|
||||
if rr.Code != http.StatusNoContent {
|
||||
t.Fatalf("preflight status = %d, want 204", rr.Code)
|
||||
}
|
||||
if got := rr.Header().Get("Access-Control-Allow-Origin"); got != "https://asuracomic.net" {
|
||||
if got := rr.Header().Get("Access-Control-Allow-Origin"); got != "https://asurascans.com" {
|
||||
t.Fatalf("Allow-Origin = %q, want reflected origin", got)
|
||||
}
|
||||
if got := rr.Header().Get("Access-Control-Allow-Methods"); got == "" {
|
||||
@@ -204,11 +204,11 @@ func TestBookmarkRoundTrip(t *testing.T) {
|
||||
key := "asura:solo-leveling-123"
|
||||
in := store.Bookmark{
|
||||
Title: "Solo Leveling",
|
||||
SeriesURL: "https://asuracomic.net/series/solo-leveling-123",
|
||||
Cover: "https://asuracomic.net/cover.jpg",
|
||||
SeriesURL: "https://asurascans.com/series/solo-leveling-123",
|
||||
Cover: "https://asurascans.com/cover.jpg",
|
||||
LastChapter: "Chapter 10",
|
||||
LastChapterNum: 10,
|
||||
LastChapterURL: "https://asuracomic.net/series/solo-leveling-123/chapter/10",
|
||||
LastChapterURL: "https://asurascans.com/series/solo-leveling-123/chapter/10",
|
||||
}
|
||||
body, _ := json.Marshal(in)
|
||||
|
||||
|
||||
@@ -2,6 +2,7 @@ package latest
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"log"
|
||||
"sync"
|
||||
"time"
|
||||
@@ -97,51 +98,49 @@ func (a *Acquirer) acquire(ctx context.Context, sr store.Series) {
|
||||
if a.Fetch == nil && a.BrowserFetch == nil {
|
||||
return
|
||||
}
|
||||
// series_url arrives in a client-supplied PUT body, so the same gate the
|
||||
// poller uses applies here — without it a token-holder chooses what the
|
||||
// server fetches from its own network position.
|
||||
if !fetchableSeriesURL(sr.Site, sr.SeriesURL) {
|
||||
log.Printf("acquire %q: not fetchable: site=%q url=%q", sr.Key(), sr.Site, sr.SeriesURL)
|
||||
return
|
||||
}
|
||||
|
||||
f := fetcherFor(sr.Site, a.BrowserFetch, a.Fetch)
|
||||
if f == nil {
|
||||
log.Printf("acquire %q: no fetcher for site %q", sr.Key(), sr.Site)
|
||||
return
|
||||
}
|
||||
body, status, err := f.Get(ctx, sr.SeriesURL)
|
||||
facts, err := readSeriesPage(ctx, sr.Site, sr.SeriesURL, a.BrowserFetch, a.Fetch)
|
||||
if err != nil {
|
||||
log.Printf("acquire %q: fetch %s: %v", sr.Key(), sr.SeriesURL, err)
|
||||
return
|
||||
switch {
|
||||
case errors.Is(err, errNotFetchable):
|
||||
// series_url arrives in a client-supplied PUT body, so without the
|
||||
// gate a token-holder chooses what the server fetches from its own
|
||||
// network position.
|
||||
log.Printf("acquire %q: not fetchable: site=%q url=%q", sr.Key(), sr.Site, sr.SeriesURL)
|
||||
case errors.Is(err, errNoFetcher):
|
||||
log.Printf("acquire %q: no fetcher for site %q", sr.Key(), sr.Site)
|
||||
default:
|
||||
log.Printf("acquire %q: %v", sr.Key(), err)
|
||||
}
|
||||
if status != 200 {
|
||||
log.Printf("acquire %q: fetch %s: status %d", sr.Key(), sr.SeriesURL, status)
|
||||
return
|
||||
}
|
||||
|
||||
// This page just served the same purpose a poll tick would have; without
|
||||
// the stamp the row stays due and the poller refetches it immediately.
|
||||
//
|
||||
// Stamped after success — the reverse of the poller, which stamps before
|
||||
// the fetch: the Reader is here, watching the Series they just created, so
|
||||
// a failed acquisition must leave the row due for a fast retry rather than
|
||||
// consuming the cooldown. The stamp happens even when the page read
|
||||
// succeeded but produced no facts to persist.
|
||||
if err := a.Store.MarkLatestChecked(sr.Site, sr.SeriesID, time.Now().UnixMilli()); err != nil {
|
||||
log.Printf("acquire %q: mark checked: %v", sr.Key(), err)
|
||||
}
|
||||
|
||||
if latest, ok := latestChapterFrom(sr.Site, sr.SeriesURL, body); ok {
|
||||
if err := a.Store.SetLatestChapter(sr.Site, sr.SeriesID, latest.Label, latest.Num); err != nil {
|
||||
if facts.HasLatest {
|
||||
if err := a.Store.SetLatestChapter(sr.Site, sr.SeriesID, facts.Latest.Label, facts.Latest.Num); err != nil {
|
||||
log.Printf("acquire %q: set latest chapter: %v", sr.Key(), err)
|
||||
}
|
||||
}
|
||||
|
||||
cover, ok := coverFrom(sr.Site, sr.SeriesURL, body)
|
||||
if !ok {
|
||||
if !facts.HasCover {
|
||||
return
|
||||
}
|
||||
bytes, contentType, err := fetchCoverBytes(ctx, cover, a.BrowserCoverFetch, a.Covers)
|
||||
bytes, contentType, err := fetchCoverBytes(ctx, facts.Cover, a.BrowserCoverFetch, a.Covers)
|
||||
if err != nil {
|
||||
log.Printf("acquire %q: fetch cover %s: %v", sr.Key(), cover, err)
|
||||
log.Printf("acquire %q: fetch cover %s: %v", sr.Key(), facts.Cover, err)
|
||||
return
|
||||
}
|
||||
if err := a.Store.SetSeriesCover(sr.Site, sr.SeriesID, cover, bytes, contentType); err != nil {
|
||||
if err := a.Store.SetSeriesCover(sr.Site, sr.SeriesID, facts.Cover, bytes, contentType); err != nil {
|
||||
log.Printf("acquire %q: persist cover: %v", sr.Key(), err)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -86,40 +86,26 @@ func (f *BrowserFetcher) Close() {
|
||||
f.cancel()
|
||||
}
|
||||
|
||||
// Get navigates to seriesURL, lets any challenge resolve, then reads either the
|
||||
// site's JSON API (kagane) from inside the page so the request carries the
|
||||
// clearance cookie, or the served HTML itself (novelfull) — see
|
||||
// novelfullSeriesURL for the latter case. The returned body is whatever the
|
||||
// site's chapter list lives in, which is what latestChapterFrom's per-site
|
||||
// switch expects.
|
||||
// Get navigates to seriesURL, lets any challenge resolve, then reads the
|
||||
// payload the Site's registry entry describes — kagane's chapter-list API from
|
||||
// inside the page so the request carries the clearance cookie, novelfull's
|
||||
// served HTML. The returned body is whatever the Site's chapter list lives in,
|
||||
// which is what the entry's LatestChapter parse expects.
|
||||
func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int, error) {
|
||||
apiURL, isKagane := kaganeAPIURL(seriesURL)
|
||||
if !isKagane && !novelfullSeriesURL(seriesURL) {
|
||||
return "", 0, fmt.Errorf("not a fetchable browser series url: %q", seriesURL)
|
||||
}
|
||||
|
||||
var body string
|
||||
// kagane's chapter list is only in its JSON API, which must be called from
|
||||
// inside the page so the request carries the clearance cookie. novelfull
|
||||
// renders its chapters into the HTML, so the cleared DOM is the answer.
|
||||
// chromedp.OuterHTML returns a QueryAction and chromedp.Evaluate an
|
||||
// EvaluateAction, so the variable has to be the interface both implement.
|
||||
var read chromedp.Action = chromedp.OuterHTML("html", &body, chromedp.ByQuery)
|
||||
if isKagane {
|
||||
read = chromedp.Evaluate(
|
||||
`fetch(`+jsString(apiURL)+`).then(r => r.ok ? r.text() : "")`,
|
||||
&body,
|
||||
awaitPromise,
|
||||
)
|
||||
// Sorted order (browserBackedSites sorts) makes dispatch deterministic:
|
||||
// entries' Read funcs are expected to refuse any address owned by another
|
||||
// Site, and the loop must not depend on that staying true.
|
||||
for _, name := range browserBackedSites() {
|
||||
s := sites[name]
|
||||
read, ok := s.Browser.Read(seriesURL, &body)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
|
||||
// novelfull's payload is the DOM itself, and the interstitial has a DOM
|
||||
// too, so "we have an answer" has to exclude it explicitly. kagane's
|
||||
// in-page fetch just fails while challenged, which is already the signal.
|
||||
done := func() bool { return body != "" && (isKagane || !isInterstitial(body)) }
|
||||
if err := f.run(ctx, seriesURL, read, done); err != nil {
|
||||
// Challenge never cleared, or the API refused. Indistinguishable from
|
||||
// here and handled identically by the caller.
|
||||
if err := f.run(ctx, seriesURL, read,
|
||||
func() bool { return s.Browser.Done(body) }); err != nil {
|
||||
// Challenge never cleared, or the payload was refused.
|
||||
// Indistinguishable from here and handled identically by the caller.
|
||||
if errors.Is(err, errChallengeHeld) {
|
||||
return "", 403, nil
|
||||
}
|
||||
@@ -127,6 +113,33 @@ func (f *BrowserFetcher) Get(ctx context.Context, seriesURL string) (string, int
|
||||
}
|
||||
return body, 200, nil
|
||||
}
|
||||
return "", 0, fmt.Errorf("not a fetchable browser series url: %q", seriesURL)
|
||||
}
|
||||
|
||||
// kaganeRead builds the in-tab fetch of kagane's chapter-list API: the
|
||||
// request must be made from inside the page so it carries the clearance
|
||||
// cookie, and the API is the only place the list exists. Refusing any other
|
||||
// address is the per-Site half of the SSRF gate, kept deliberately behind
|
||||
// fetchableSeriesURL (see browserRead.Read).
|
||||
func kaganeRead(seriesURL string, out *string) (chromedp.Action, bool) {
|
||||
apiURL, ok := kaganeAPIURL(seriesURL)
|
||||
if !ok {
|
||||
return nil, false
|
||||
}
|
||||
return chromedp.Evaluate(
|
||||
`fetch(`+jsString(apiURL)+`).then(r => r.ok ? r.text() : "")`,
|
||||
out, awaitPromise), true
|
||||
}
|
||||
|
||||
// novelfullRead reads the cleared DOM. novelfull renders its chapter list
|
||||
// into the served HTML, so there is no API to call from inside the page — the
|
||||
// challenge-cleared DOM is the payload.
|
||||
func novelfullRead(seriesURL string, out *string) (chromedp.Action, bool) {
|
||||
if !novelfullSeriesURL(seriesURL) {
|
||||
return nil, false
|
||||
}
|
||||
return chromedp.OuterHTML("html", out, chromedp.ByQuery), true
|
||||
}
|
||||
|
||||
// Image retrieves one cover's bytes through the browser sidecar, and its
|
||||
// content type.
|
||||
|
||||
@@ -11,8 +11,9 @@ import (
|
||||
)
|
||||
|
||||
// maxBodyBytes caps what a single series page can cost in memory. Real pages
|
||||
// measured 100-400 KB on 2026-07-26, so this is roughly 10x headroom and mostly
|
||||
// guards against a proxy handing back something enormous.
|
||||
// measured 100-400 KB on 2026-07-26; lightnovelworld runs larger — 685 KB and
|
||||
// 1.18 MB measured 2026-08-11 — so the headroom there is roughly 3.5x, and the
|
||||
// cap mostly guards against a proxy handing back something enormous.
|
||||
const maxBodyBytes = 4 << 20
|
||||
|
||||
// chromeUA matches the client profile below. A Chrome fingerprint paired with a
|
||||
|
||||
@@ -2,9 +2,9 @@ package latest
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"log"
|
||||
"net/url"
|
||||
"slices"
|
||||
"time"
|
||||
|
||||
"bookmarkmanager/backend/internal/store"
|
||||
@@ -57,25 +57,23 @@ type Poller struct {
|
||||
Batch int
|
||||
}
|
||||
|
||||
var browserBackedSites = []string{"kagane", "novelfull"}
|
||||
|
||||
// fillBlankCover gives a Series its Cover when it has none. The blank state is
|
||||
// what "no Cover yet" means on the wire (ADR-0007): permanently-blank rows
|
||||
// created before acquisition existed, and rows whose creation-time fetch
|
||||
// failed, both heal here. A non-blank CoverAddress is left alone — refetching
|
||||
// would add a request per Series per cycle and change artwork under the Reader
|
||||
// for no visible reason. A row that already carries a source URL is owned by
|
||||
// prefetchCover instead; this path only extracts from the series page.
|
||||
// prefetchCover instead; this path only records a Cover address already
|
||||
// extracted from the series page.
|
||||
//
|
||||
// Failures are logged against the Series and never returned: the chapter poll
|
||||
// must not notice. A failed fill is retried the next time this Series is due;
|
||||
// there is no separate retry queue.
|
||||
func (p *Poller) fillBlankCover(ctx context.Context, sr store.Series, body string) {
|
||||
func (p *Poller) fillBlankCover(ctx context.Context, sr store.Series, cover string) {
|
||||
if sr.CoverAddress != "" || sr.Cover != "" {
|
||||
return
|
||||
}
|
||||
cover, ok := coverFrom(sr.Site, sr.SeriesURL, body)
|
||||
if !ok {
|
||||
if cover == "" {
|
||||
return
|
||||
}
|
||||
p.storeCover(ctx, sr, cover)
|
||||
@@ -120,27 +118,29 @@ func (p *Poller) storeCover(ctx context.Context, sr store.Series, sourceURL stri
|
||||
}
|
||||
|
||||
// fetcherFor returns the fetcher a site's page needs, or nil when the site
|
||||
// cannot be fetched at all right now. kagane and novelfull pages sit behind a
|
||||
// Cloudflare JavaScript challenge that no TLS fingerprint clears (kagane
|
||||
// verified 2026-08-03, novelfull verified 2026-08-05, both against the same
|
||||
// Chrome_133 profile TLSFetcher uses), so both prefer the browser; novelfull
|
||||
// alone falls back to the plain-TLS fetcher when no browser is configured,
|
||||
// because its challenge is a live time-varying fact (AGENTS.md) and its cover
|
||||
// bytes never need the browser. kagane never falls back: a plain fetch of a
|
||||
// kagane page or cover would only ever retrieve a challenge page. One routing
|
||||
// rule for the poll and the acquirer, so the two cannot drift apart.
|
||||
// cannot be fetched at all right now. A Site whose registry entry carries a
|
||||
// Browser read — kagane and novelfull, both behind a Cloudflare JavaScript
|
||||
// challenge no TLS fingerprint clears — prefers the browser; when it is
|
||||
// absent, the entry's Fallback decides whether plain TLS may take over. One
|
||||
// routing rule for the poll and the acquirer, so the two cannot drift apart.
|
||||
func fetcherFor(site string, browser, tls Fetcher) Fetcher {
|
||||
switch {
|
||||
case site == "kagane":
|
||||
return browser
|
||||
case slices.Contains(browserBackedSites, site): // novelfull
|
||||
s, known := sites[site]
|
||||
if !known {
|
||||
// No registry entry means nothing to fetch or parse; fail closed even
|
||||
// though the only caller gates first, so a future caller that skips
|
||||
// the gate cannot hand an arbitrary https URL to the TLS fetcher.
|
||||
return nil
|
||||
}
|
||||
if s.Browser == nil {
|
||||
return tls
|
||||
}
|
||||
if browser != nil {
|
||||
return browser
|
||||
}
|
||||
return tls
|
||||
default:
|
||||
if s.Browser.Fallback {
|
||||
return tls
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// Run polls until ctx is cancelled.
|
||||
@@ -170,7 +170,7 @@ func (p *Poller) runOnce(ctx context.Context) {
|
||||
now := p.Now()
|
||||
cutoff := now.Add(-p.Cooldown).UnixMilli()
|
||||
browserCutoff := now.Add(-p.BrowserCooldown).UnixMilli()
|
||||
due, err := p.Store.DueForLatestCheck(cutoff, browserCutoff, browserBackedSites, p.Batch)
|
||||
due, err := p.Store.DueForLatestCheck(cutoff, browserCutoff, browserBackedSites(), p.Batch)
|
||||
if err != nil {
|
||||
log.Printf("latest poll: due query: %v", err)
|
||||
return
|
||||
@@ -224,43 +224,35 @@ func (p *Poller) checkOne(ctx context.Context, sr store.Series) {
|
||||
return
|
||||
}
|
||||
|
||||
// series_url is client-supplied (PUT /bookmarks/{key} accepts any string),
|
||||
// so this is not just an optimisation against burning a request on an
|
||||
// unknown site: without it, the server would issue a GET from its own
|
||||
// network position to whatever URL a token-holder writes, including
|
||||
// link-local/internal addresses or non-https schemes. The cooldown above
|
||||
// is already consumed, so a row that never passes this check is retried at
|
||||
// cooldown pace rather than hot-looping.
|
||||
if !fetchableSeriesURL(sr.Site, sr.SeriesURL) {
|
||||
facts, err := readSeriesPage(ctx, sr.Site, sr.SeriesURL, p.BrowserFetch, p.Fetch)
|
||||
if err != nil {
|
||||
switch {
|
||||
case errors.Is(err, errNotFetchable):
|
||||
// The cooldown above is already consumed, so a row that never
|
||||
// passes the gate is retried at cooldown pace rather than
|
||||
// hot-looping.
|
||||
log.Printf("latest poll %q: not fetchable: site=%q url=%q", sr.Key(), sr.Site, sr.SeriesURL)
|
||||
return
|
||||
}
|
||||
|
||||
f := fetcherFor(sr.Site, p.BrowserFetch, p.Fetch)
|
||||
if f == nil {
|
||||
case errors.Is(err, errNoFetcher):
|
||||
log.Printf("latest poll %q: no fetcher for site %q", sr.Key(), sr.Site)
|
||||
return
|
||||
}
|
||||
// A legacy cover heals independently of the page read: its source may
|
||||
// answer — a CDN — while the origin does not, so a fetch failure does
|
||||
// not skip the heal, matching the order the shared read replaced.
|
||||
p.prefetchCover(ctx, sr)
|
||||
|
||||
body, status, err := f.Get(ctx, sr.SeriesURL)
|
||||
if err != nil {
|
||||
log.Printf("latest poll %q: fetch %s: %v", sr.Key(), sr.SeriesURL, err)
|
||||
log.Printf("latest poll %q: %v", sr.Key(), err)
|
||||
return
|
||||
}
|
||||
if status != 200 {
|
||||
log.Printf("latest poll %q: fetch %s: status %d", sr.Key(), sr.SeriesURL, status)
|
||||
return
|
||||
}
|
||||
|
||||
latest, ok := latestChapterFrom(sr.Site, sr.SeriesURL, body)
|
||||
// A legacy cover source is healed independently of the page read.
|
||||
p.prefetchCover(ctx, sr)
|
||||
// Cover fill is independent of the chapter signal: a page that lost its
|
||||
// chapter list may keep its og:image, and a blank Series heals either way.
|
||||
p.fillBlankCover(ctx, sr, body)
|
||||
if !ok {
|
||||
p.fillBlankCover(ctx, sr, facts.Cover)
|
||||
if !facts.HasLatest {
|
||||
// Most likely a challenge page or a layout change. Either way the row is
|
||||
// already stamped, so this waits out a cooldown instead of hot-looping.
|
||||
log.Printf("latest poll %q: no chapter links in %d bytes", sr.Key(), len(body))
|
||||
log.Printf("latest poll %q: no chapter links in %d bytes", sr.Key(), facts.BodyLen)
|
||||
return
|
||||
}
|
||||
|
||||
@@ -268,7 +260,7 @@ func (p *Poller) checkOne(ctx context.Context, sr store.Series) {
|
||||
// chapter should correct the stored number downward. The comparison is
|
||||
// against the due-query snapshot; a concurrent write in between only costs
|
||||
// one redundant UPDATE of the same absolute value, never a wrong one.
|
||||
if sr.LatestChapterNum != nil && *sr.LatestChapterNum == latest.Num {
|
||||
if sr.LatestChapterNum != nil && *sr.LatestChapterNum == facts.Latest.Num {
|
||||
return
|
||||
}
|
||||
|
||||
@@ -276,52 +268,30 @@ func (p *Poller) checkOne(ctx context.Context, sr store.Series) {
|
||||
// bookmark joining to it, and the bookmark's updated_at is never touched —
|
||||
// a newly published chapter is not reading progress and must not reorder
|
||||
// the list.
|
||||
if err := p.Store.SetLatestChapter(sr.Site, sr.SeriesID, latest.Label, latest.Num); err != nil {
|
||||
if err := p.Store.SetLatestChapter(sr.Site, sr.SeriesID, facts.Latest.Label, facts.Latest.Num); err != nil {
|
||||
log.Printf("latest poll %q: set latest chapter: %v", sr.Key(), err)
|
||||
return
|
||||
}
|
||||
log.Printf("latest poll %q: latest is now %s", sr.Key(), latest.Label)
|
||||
log.Printf("latest poll %q: latest is now %s", sr.Key(), facts.Latest.Label)
|
||||
}
|
||||
|
||||
// fetchableSeriesURL reports whether site is a site latestChapterFrom knows how
|
||||
// to parse and seriesURL is safe to hand to a fetcher: an https URL with a
|
||||
// non-empty host. series_url comes from client-supplied PUT bodies, so this is
|
||||
// a defence against the poller being used to probe arbitrary hosts from the
|
||||
// server's own network position, not just a check against wasted requests.
|
||||
//
|
||||
// Three sites are held to a stricter rule, each for a different reason:
|
||||
//
|
||||
// - kagane and novelfull are fetched by a headless browser, which executes
|
||||
// JavaScript and carries cookies, and is therefore a far stronger SSRF
|
||||
// primitive than an HTTP GET. Their hosts must match exactly, not merely
|
||||
// be non-empty.
|
||||
// - lightnovelworld's parser regex hardcodes its host, so a URL anywhere
|
||||
// else could never yield a match — reject it here rather than burn the
|
||||
// request.
|
||||
// fetchableSeriesURL reports whether site is a Site the registry knows and
|
||||
// seriesURL is safe to hand to a fetcher: an https URL whose host matches the
|
||||
// Site's pinned hostname exactly. series_url comes from client-supplied PUT
|
||||
// bodies, so this is a defence against the poller being used to probe
|
||||
// arbitrary hosts from the server's own network position, not just a check
|
||||
// against wasted requests. The pin guards different things per Site — a
|
||||
// browser Site guards a control that executes JavaScript and carries cookies,
|
||||
// a parser Site guards a wasted request — but the rule is one rule, from the
|
||||
// registry.
|
||||
func fetchableSeriesURL(site, seriesURL string) bool {
|
||||
switch site {
|
||||
case "asura", "demonic", "comix", "kagane", "novelfull", "lightnovelworld":
|
||||
default:
|
||||
s, known := sites[site]
|
||||
if !known {
|
||||
return false
|
||||
}
|
||||
u, err := url.Parse(seriesURL)
|
||||
if err != nil {
|
||||
return false
|
||||
}
|
||||
if u.Scheme != "https" || u.Host == "" {
|
||||
return false
|
||||
}
|
||||
switch site {
|
||||
case "kagane":
|
||||
return u.Hostname() == "kagane.to"
|
||||
case "novelfull":
|
||||
// Fetched by a real browser, same as kagane, so the host is pinned
|
||||
// rather than merely non-empty.
|
||||
return u.Hostname() == "novelfull.com"
|
||||
case "lightnovelworld":
|
||||
// Its parser regex hardcodes this host, so a URL anywhere else could
|
||||
// never yield a match — reject it here rather than burn the request.
|
||||
return u.Hostname() == "lightnovelworld.net"
|
||||
}
|
||||
return true
|
||||
return u.Scheme == "https" && u.Hostname() == s.Host
|
||||
}
|
||||
|
||||
@@ -623,6 +623,13 @@ func TestFetchableSeriesURL(t *testing.T) {
|
||||
{"demonic https", "demonic", "https://demonicscans.org/manga/X", true},
|
||||
{"comix https", "comix", "https://comix.to/title/n8we-dungeons-and-crayons", true},
|
||||
{"kagane on its own host", "kagane", "https://kagane.to/series/019f84bc-9ba0-7ed9-86f5-8b905ec7c28b", true},
|
||||
// The three plain-TLS Sites are pinned too: a client-supplied
|
||||
// series_url must not aim a fetcher at a lookalike host, even when
|
||||
// the fetcher is only an HTTP GET.
|
||||
{"asura on a foreign host", "asura", "https://asurascans.com.evil.example/comics/x", false},
|
||||
{"asura on the dead old domain", "asura", "https://asuracomic.net/comics/x", false},
|
||||
{"demonic on a lookalike host", "demonic", "https://demonicscans.org.evil.example/manga/X", false},
|
||||
{"comix on a foreign host", "comix", "https://evil.example/title/x", false},
|
||||
// The browser fetcher runs JavaScript and carries cookies, so a
|
||||
// client-supplied series_url must not be able to aim it anywhere else.
|
||||
{"kagane on a foreign host", "kagane", "https://evil.example/series/x", false},
|
||||
@@ -952,6 +959,8 @@ func TestFetcherForRoutesNovelSites(t *testing.T) {
|
||||
// browser-less deployment: kagane is nothing, novelfull degrades to TLS
|
||||
{"kagane", nil, tls, nil},
|
||||
{"novelfull", nil, tls, tls},
|
||||
// unknown site: fail closed — nothing to fetch or parse
|
||||
{"mangadex", browser, tls, nil},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.site, func(t *testing.T) {
|
||||
@@ -1010,10 +1019,10 @@ func TestRunOnceFillsBlankCoverFromSeriesPage(t *testing.T) {
|
||||
},
|
||||
{
|
||||
name: "lightnovelworld novel",
|
||||
key: "lightnovelworld:a-will-eternal", site: "lightnovelworld",
|
||||
seriesID: "a-will-eternal", seriesURL: "https://lightnovelworld.net/novel/a-will-eternal/",
|
||||
key: "lightnovelworld:all-jobs-and-classes-i-just-wanted-one-skill-not-them-all", site: "lightnovelworld",
|
||||
seriesID: "all-jobs-and-classes-i-just-wanted-one-skill-not-them-all", seriesURL: "https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/",
|
||||
kind: store.KindNovel, body: lnwSeriesFixture + lnwCoverFixture,
|
||||
wantCover: "https://lightnovelworld.net/wp-content/uploads/2026/03/a-will-eternal-1.webp",
|
||||
wantCover: "https://i1.wp.com/lightnovelworld.net/wp-content/uploads/2025/10/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all.jpg",
|
||||
},
|
||||
{
|
||||
name: "kagane manga",
|
||||
|
||||
@@ -0,0 +1,58 @@
|
||||
package latest
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
)
|
||||
|
||||
// seriesRead carries the two facts the poll and the acquirer both extract
|
||||
// from a series page. Persistence, stamps and scheduling stay with the
|
||||
// callers, so the policies that keep the two flows distinct (stamp order,
|
||||
// cooldowns) are not swallowed by the module.
|
||||
type seriesRead struct {
|
||||
Latest latestChapter
|
||||
HasLatest bool
|
||||
Cover string
|
||||
HasCover bool
|
||||
// BodyLen is the fetched body's length, surfaced because the no-chapter
|
||||
// log uses it to tell a markup change from a body the size cap cut short.
|
||||
BodyLen int
|
||||
}
|
||||
|
||||
// errNotFetchable and errNoFetcher separate the gate and the route from fetch
|
||||
// failures so each caller keeps its own distinct log line for all three.
|
||||
var (
|
||||
errNotFetchable = errors.New("series url not fetchable")
|
||||
errNoFetcher = errors.New("no fetcher for site")
|
||||
)
|
||||
|
||||
// readSeriesPage performs the series-page read the poll and the acquirer have
|
||||
// in common: gate the address, choose the route, fetch the page, extract the
|
||||
// Latest Chapter and the Cover address. It persists nothing and stamps
|
||||
// nothing.
|
||||
//
|
||||
// series_url arrives in a client-supplied PUT body (PUT /bookmarks/{key}
|
||||
// accepts any string), so the gate is not an optimisation against burning a
|
||||
// request on an unknown site: without it, the server would issue a GET from
|
||||
// its own network position to whatever URL a token-holder writes, including
|
||||
// link-local/internal addresses or non-https schemes.
|
||||
func readSeriesPage(ctx context.Context, site, seriesURL string, browser, tls Fetcher) (seriesRead, error) {
|
||||
if !fetchableSeriesURL(site, seriesURL) {
|
||||
return seriesRead{}, fmt.Errorf("%w: site=%q url=%q", errNotFetchable, site, seriesURL)
|
||||
}
|
||||
f := fetcherFor(site, browser, tls)
|
||||
if f == nil {
|
||||
return seriesRead{}, fmt.Errorf("%w: site %q", errNoFetcher, site)
|
||||
}
|
||||
body, status, err := f.Get(ctx, seriesURL)
|
||||
if err != nil {
|
||||
return seriesRead{}, fmt.Errorf("fetch %s: %w", seriesURL, err)
|
||||
}
|
||||
if status != 200 {
|
||||
return seriesRead{}, fmt.Errorf("fetch %s: status %d", seriesURL, status)
|
||||
}
|
||||
latest, hasLatest := latestChapterFrom(site, seriesURL, body)
|
||||
cover, hasCover := coverFrom(site, seriesURL, body)
|
||||
return seriesRead{Latest: latest, HasLatest: hasLatest, Cover: cover, HasCover: hasCover, BodyLen: len(body)}, nil
|
||||
}
|
||||
@@ -3,10 +3,14 @@ package latest
|
||||
import (
|
||||
"encoding/json"
|
||||
"html"
|
||||
"log"
|
||||
"net/url"
|
||||
"regexp"
|
||||
"sort"
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
"github.com/chromedp/chromedp"
|
||||
)
|
||||
|
||||
// latestChapter is the newest chapter a series page advertises.
|
||||
@@ -15,6 +19,40 @@ type latestChapter struct {
|
||||
Label string
|
||||
}
|
||||
|
||||
// site answers the fixed questions every series-page read asks of its Site
|
||||
// (ADR-0009): the host its addresses must carry, how to find the Latest
|
||||
// Chapter and the Cover address in a body, and — for a Site behind a
|
||||
// JavaScript challenge — how to read its payload from a cleared tab. One
|
||||
// entry describes everything about one Site, and nowhere else gets to compare
|
||||
// the site string.
|
||||
type site struct {
|
||||
// Host is the exact hostname a series_url for this Site must carry.
|
||||
Host string
|
||||
// LatestChapter finds the newest chapter in a fetched body.
|
||||
LatestChapter func(seriesURL, body string) (latestChapter, bool)
|
||||
// Cover finds the Cover address in a fetched body.
|
||||
Cover func(seriesURL, body string) (string, bool)
|
||||
// Browser reads this Site's payload from a cleared browser tab; nil
|
||||
// means the page is fetched over plain TLS.
|
||||
Browser *browserRead
|
||||
}
|
||||
|
||||
type browserRead struct {
|
||||
// Read builds the tab read for seriesURL, refusing (false) an address
|
||||
// this Site will not open in a browser — the per-Site half of the SSRF
|
||||
// gate, kept deliberately behind fetchableSeriesURL: a headless browser
|
||||
// executes JavaScript and carries cookies, and series_url is
|
||||
// client-supplied.
|
||||
Read func(seriesURL string, out *string) (chromedp.Action, bool)
|
||||
// Done reports whether the payload arrived.
|
||||
Done func(body string) bool
|
||||
// Fallback allows the plain-TLS fetcher when no browser is configured.
|
||||
// False skips the Site instead. kagane is false — a plain fetch would
|
||||
// only ever retrieve a challenge page — and novelfull is true, because
|
||||
// its challenge is a live time-varying fact (AGENTS.md).
|
||||
Fallback bool
|
||||
}
|
||||
|
||||
// asuraSlugRe pulls the series slug out of a stored series_url.
|
||||
// Shape verified live 2026-07-26: https://asurascans.com/comics/<slug>, where
|
||||
// the slug carries a trailing build-hash suffix (e.g. "-f886a8af") that
|
||||
@@ -58,87 +96,37 @@ var kaganeChapterRe = regexp.MustCompile(`"chapter_no":"([0-9.]+)"`)
|
||||
// "/<slug>/chapter-<n>[-<title-slug>].html". Verified live 2026-08-05.
|
||||
var novelfullSlugRe = regexp.MustCompile(`^/([^/?#]+)\.html$`)
|
||||
|
||||
// lnwSlugRe does the same for lightnovelworld, whose series pages live under
|
||||
// /novel/<slug>/ while its chapter URLs are flat at the site root:
|
||||
// "/<slug>-chapter-<n>/", absolute in the page's own anchors. Verified live
|
||||
// 2026-08-05.
|
||||
var lnwSlugRe = regexp.MustCompile(`^/novel/([^/?#]+)/?$`)
|
||||
// lnwChapterRe matches any chapter-shaped address on lightnovelworld. Unlike
|
||||
// asura, novelfull and comix — which scope to their stored series slug so a
|
||||
// foreign chapter link cannot contribute — this Site's chapter addresses carry
|
||||
// the Chapter Slug, which is not the Series identity: one Series may publish
|
||||
// under several Chapter Slugs (measured 2026-08-11: a sampled novel serves
|
||||
// 1-99 under one slug and 100-423 under another), so no stored-slug pattern can
|
||||
// cover a Series' whole list. An unscoped match is safe because
|
||||
// lnwLatestChapter truncates the body at the comment thread before scanning
|
||||
// (lnwCommentMarker); without that, a visitor's comment could set the Latest
|
||||
// Chapter on the shared Series row.
|
||||
var lnwChapterRe = regexp.MustCompile(`lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/`)
|
||||
|
||||
// latestChapterFrom returns the highest chapter number body advertises for this
|
||||
// series. ok is false when the body yields nothing usable — an unknown site, an
|
||||
// empty body, a Cloudflare challenge page, and a site redesign all land here,
|
||||
// and the caller treats all four identically.
|
||||
//
|
||||
// Ported from the userscript's latestChapterFromAnchors (asura L123-133,
|
||||
// demonic L183-193), including its reason for taking a maximum rather than a
|
||||
// first or last: neither site lists chapters in a dependable order.
|
||||
// lnwCommentMarker is the boundary of lightnovelworld's server-rendered
|
||||
// wpdiscuz comment thread. It occurs exactly once per page and follows every
|
||||
// chapter anchor (measured 2026-08-11,
|
||||
// docs/research/lightnovelworld-chapter-vs-series-slug.md §6), so cutting the
|
||||
// body at its first occurrence keeps the whole chapter list while excluding a
|
||||
// region any visitor can write to. Absent means the page shape changed: the
|
||||
// body is skipped, never scanned whole.
|
||||
const lnwCommentMarker = "wpd-threads"
|
||||
|
||||
// maxChapter returns the highest chapter number the regex finds in body. A
|
||||
// maximum rather than a first or last, ported from the userscript's
|
||||
// latestChapterFromAnchors (asura L123-133, demonic L183-193): neither site
|
||||
// lists chapters in a dependable order.
|
||||
//
|
||||
// The userscript's asura rule additionally requires the anchor text to match
|
||||
// /Chapter\s+[\d.]+/i. That check exists only to skip the "First Chapter"
|
||||
// shortcut, which points at chapter/1 and therefore can never win a maximum, so
|
||||
// it is redundant here. For asura, scoping the pattern to this series' own slug
|
||||
// replaces it with a stronger guarantee: a chapter link belonging to some other
|
||||
// series cannot contribute even if the page starts carrying them. demonic has no
|
||||
// such guarantee — demonicChapterRe matches any chaptered.php?manga=<id> anchor
|
||||
// with no per-series scoping, because the stored series_id for demonic is a
|
||||
// slug, not the numeric id the URL carries, so it cannot easily be scoped.
|
||||
func latestChapterFrom(site, seriesURL, body string) (latestChapter, bool) {
|
||||
var re *regexp.Regexp
|
||||
switch site {
|
||||
case "asura":
|
||||
m := asuraSlugRe.FindStringSubmatch(seriesURL)
|
||||
if m == nil {
|
||||
return latestChapter{}, false
|
||||
}
|
||||
// Stored URLs predating a redeploy may carry a stale build hash;
|
||||
// chapter hrefs in the fetched body carry the current one. Strip to
|
||||
// the stable ID and make the hash optional in the pattern, so scoping
|
||||
// survives rotations.
|
||||
slug := asuraBuildHash.ReplaceAllString(m[1], "")
|
||||
// Compiled per call rather than cached: this runs once per fetch, which
|
||||
// is at most a few times a minute, and the slug varies per series.
|
||||
re = regexp.MustCompile(`/comics/` + regexp.QuoteMeta(slug) + `(?:-[0-9a-f]{8})?/chapter/([0-9.]+)`)
|
||||
case "demonic":
|
||||
re = demonicChapterRe
|
||||
case "comix":
|
||||
id, ok := comixSeriesID(seriesURL)
|
||||
if !ok {
|
||||
return latestChapter{}, false
|
||||
}
|
||||
// comix ships an SPA: the served HTML carries a JSON state blob instead
|
||||
// of chapter anchors, and latestChapterUrl is the only place the newest
|
||||
// chapter appears. Scoping to this series' id prefix keeps a
|
||||
// "recommended" strip's entries from winning the maximum.
|
||||
re = regexp.MustCompile(`"latestChapterUrl":"/title/` + regexp.QuoteMeta(id) + `-[^"]*-chapter-([0-9.]+)"`)
|
||||
case "kagane":
|
||||
re = kaganeChapterRe
|
||||
case "novelfull":
|
||||
u, err := url.Parse(seriesURL)
|
||||
if err != nil {
|
||||
return latestChapter{}, false
|
||||
}
|
||||
m := novelfullSlugRe.FindStringSubmatch(u.Path)
|
||||
if m == nil {
|
||||
return latestChapter{}, false
|
||||
}
|
||||
// Scoped to this series' slug for the same reason asura is: page 1
|
||||
// carries a "latest chapters" widget and a "you may also like" strip,
|
||||
// and neither may contribute to the maximum.
|
||||
re = regexp.MustCompile(`/` + regexp.QuoteMeta(m[1]) + `/chapter-([0-9.]+)`)
|
||||
case "lightnovelworld":
|
||||
u, err := url.Parse(seriesURL)
|
||||
if err != nil {
|
||||
return latestChapter{}, false
|
||||
}
|
||||
m := lnwSlugRe.FindStringSubmatch(u.Path)
|
||||
if m == nil {
|
||||
return latestChapter{}, false
|
||||
}
|
||||
re = regexp.MustCompile(`lightnovelworld\.net/` + regexp.QuoteMeta(m[1]) + `-chapter-([0-9.]+)/`)
|
||||
default:
|
||||
return latestChapter{}, false
|
||||
}
|
||||
|
||||
// shortcut, which points at chapter/1 and therefore can never win a maximum,
|
||||
// so it is redundant once a maximum is taken.
|
||||
func maxChapter(re *regexp.Regexp, body string) (latestChapter, bool) {
|
||||
var best latestChapter
|
||||
found := false
|
||||
for _, m := range re.FindAllStringSubmatch(body, -1) {
|
||||
@@ -157,6 +145,94 @@ func latestChapterFrom(site, seriesURL, body string) (latestChapter, bool) {
|
||||
return best, found
|
||||
}
|
||||
|
||||
// asuraLatestChapter scopes chapter links to this series' own slug, which
|
||||
// replaces the userscript's anchor-text check with a stronger guarantee: a
|
||||
// chapter link belonging to some other series cannot contribute even if the
|
||||
// page starts carrying them.
|
||||
func asuraLatestChapter(seriesURL, body string) (latestChapter, bool) {
|
||||
m := asuraSlugRe.FindStringSubmatch(seriesURL)
|
||||
if m == nil {
|
||||
return latestChapter{}, false
|
||||
}
|
||||
// Stored URLs predating a redeploy may carry a stale build hash; chapter
|
||||
// hrefs in the fetched body carry the current one. Strip to the stable ID
|
||||
// and make the hash optional in the pattern, so scoping survives
|
||||
// rotations.
|
||||
slug := asuraBuildHash.ReplaceAllString(m[1], "")
|
||||
// Compiled per call rather than cached: this runs once per fetch, which is
|
||||
// at most a few times a minute, and the slug varies per series.
|
||||
re := regexp.MustCompile(`/comics/` + regexp.QuoteMeta(slug) + `(?:-[0-9a-f]{8})?/chapter/([0-9.]+)`)
|
||||
return maxChapter(re, body)
|
||||
}
|
||||
|
||||
// demonicLatestChapter is not scoped: demonicChapterRe matches any
|
||||
// chaptered.php?manga=<id> anchor, because the stored series_id is a slug,
|
||||
// not the numeric id the URL carries, so it cannot be scoped.
|
||||
func demonicLatestChapter(_, body string) (latestChapter, bool) {
|
||||
return maxChapter(demonicChapterRe, body)
|
||||
}
|
||||
|
||||
// comixLatestChapter reads comix's SPA: the served HTML carries a JSON state
|
||||
// blob instead of chapter anchors, and latestChapterUrl is the only place the
|
||||
// newest chapter appears. Scoping to this series' id prefix keeps a
|
||||
// "recommended" strip's entries from winning the maximum.
|
||||
func comixLatestChapter(seriesURL, body string) (latestChapter, bool) {
|
||||
id, ok := comixSeriesID(seriesURL)
|
||||
if !ok {
|
||||
return latestChapter{}, false
|
||||
}
|
||||
re := regexp.MustCompile(`"latestChapterUrl":"/title/` + regexp.QuoteMeta(id) + `-[^"]*-chapter-([0-9.]+)"`)
|
||||
return maxChapter(re, body)
|
||||
}
|
||||
|
||||
// kaganeLatestChapter scans the kagane series API JSON that the browser read
|
||||
// fetched from inside the page; the match rides on the property name,
|
||||
// regardless of the surrounding JSON shape.
|
||||
func kaganeLatestChapter(_, body string) (latestChapter, bool) {
|
||||
return maxChapter(kaganeChapterRe, body)
|
||||
}
|
||||
|
||||
// novelfullLatestChapter is scoped to this series' slug for the same reason
|
||||
// asura is: page 1 carries a "latest chapters" widget and a "you may also
|
||||
// like" strip, and neither may contribute to the maximum.
|
||||
func novelfullLatestChapter(seriesURL, body string) (latestChapter, bool) {
|
||||
u, err := url.Parse(seriesURL)
|
||||
if err != nil {
|
||||
return latestChapter{}, false
|
||||
}
|
||||
m := novelfullSlugRe.FindStringSubmatch(u.Path)
|
||||
if m == nil {
|
||||
return latestChapter{}, false
|
||||
}
|
||||
re := regexp.MustCompile(`/` + regexp.QuoteMeta(m[1]) + `/chapter-([0-9.]+)`)
|
||||
return maxChapter(re, body)
|
||||
}
|
||||
|
||||
// lnwLatestChapter truncates the body at the comment thread before scanning:
|
||||
// it is the one region of the page any visitor can write to (see lnwChapterRe).
|
||||
// A body without the marker is skipped, never scanned whole — a redesign must
|
||||
// degrade into staleness, not into a wrong shared value; the logged body length
|
||||
// tells a markup change from a body the size cap cut short.
|
||||
func lnwLatestChapter(seriesURL, body string) (latestChapter, bool) {
|
||||
i := strings.Index(body, lnwCommentMarker)
|
||||
if i < 0 {
|
||||
log.Printf("latest poll %q: no %s marker in %d bytes", seriesURL, lnwCommentMarker, len(body))
|
||||
return latestChapter{}, false
|
||||
}
|
||||
return maxChapter(lnwChapterRe, body[:i])
|
||||
}
|
||||
|
||||
// latestChapterFrom returns the highest chapter number body advertises for this
|
||||
// series, via the Site's registry entry. ok is false when the body yields
|
||||
// nothing usable — an unknown site, an empty body, a Cloudflare challenge page,
|
||||
// and a site redesign all land here, and the caller treats all four identically.
|
||||
func latestChapterFrom(site, seriesURL, body string) (latestChapter, bool) {
|
||||
if fn := sites[site].LatestChapter; fn != nil {
|
||||
return fn(seriesURL, body)
|
||||
}
|
||||
return latestChapter{}, false
|
||||
}
|
||||
|
||||
var metaTagRe = regexp.MustCompile(`(?is)<meta\b[^>]*>`)
|
||||
var doubleQuotedMetaAttrRe = regexp.MustCompile(`(?is)([a-z][a-z0-9:_-]*)\s*=\s*"([^"]*)"`)
|
||||
var singleQuotedMetaAttrRe = regexp.MustCompile(`(?is)([a-z][a-z0-9:_-]*)\s*=\s*'([^']*)'`)
|
||||
@@ -236,24 +312,40 @@ func comixCoverURL(seriesURL, body string) string {
|
||||
return publishedCoverURL(detail.Poster.Medium)
|
||||
}
|
||||
|
||||
// coverFrom reports false for unknown sites, challenge bodies, and pages with
|
||||
// no usable cover. Metadata extraction keeps scanning after an empty match so
|
||||
// a later published cover is not hidden by an empty tag.
|
||||
func coverFrom(site, seriesURL, body string) (string, bool) {
|
||||
var cover string
|
||||
switch site {
|
||||
case "asura", "demonic", "lightnovelworld":
|
||||
cover = metaContent(body, "property", "og:image")
|
||||
case "novelfull":
|
||||
cover = metaContent(body, "name", "image")
|
||||
case "comix":
|
||||
cover = comixCoverURL(seriesURL, body)
|
||||
case "kagane":
|
||||
cover = kaganeCoverURL(body)
|
||||
}
|
||||
// ogImageCover reads the og:image metadata shared by asura, demonic and
|
||||
// lightnovelworld.
|
||||
func ogImageCover(_, body string) (string, bool) {
|
||||
cover := metaContent(body, "property", "og:image")
|
||||
return cover, cover != ""
|
||||
}
|
||||
|
||||
func novelfullCoverEntry(_, body string) (string, bool) {
|
||||
cover := metaContent(body, "name", "image")
|
||||
return cover, cover != ""
|
||||
}
|
||||
|
||||
func comixCoverEntry(seriesURL, body string) (string, bool) {
|
||||
cover := comixCoverURL(seriesURL, body)
|
||||
return cover, cover != ""
|
||||
}
|
||||
|
||||
func kaganeCoverEntry(_, body string) (string, bool) {
|
||||
cover := kaganeCoverURL(body)
|
||||
return cover, cover != ""
|
||||
}
|
||||
|
||||
// coverFrom reports false for unknown sites, challenge bodies, and pages with
|
||||
// no usable cover, via the Site's registry entry.
|
||||
func coverFrom(site, seriesURL, body string) (string, bool) {
|
||||
if fn := sites[site].Cover; fn != nil {
|
||||
return fn(seriesURL, body)
|
||||
}
|
||||
return "", false
|
||||
}
|
||||
|
||||
// metaContent returns the content of the first <meta> whose attrName is
|
||||
// attrValue. It keeps scanning after an empty match so a later published cover
|
||||
// is not hidden by an empty tag.
|
||||
func metaContent(body, attrName, attrValue string) string {
|
||||
for _, tag := range metaTagRe.FindAllString(body, -1) {
|
||||
attrs := make(map[string]string)
|
||||
@@ -276,3 +368,70 @@ func publishedCoverURL(value string) string {
|
||||
value = strings.TrimSpace(html.UnescapeString(value))
|
||||
return strings.ReplaceAll(value, " ", "%20")
|
||||
}
|
||||
|
||||
// sites is the registry: one entry per Site, keyed by the stored site string.
|
||||
// Adding a Site means adding an entry here and nowhere else — the dispatch
|
||||
// functions above and the poller's route list are lookups into this map. An
|
||||
// unknown site string resolves to the zero entry, which fails the existing
|
||||
// not-fetchable and no-fetcher paths unchanged.
|
||||
var sites = map[string]site{
|
||||
"asura": {
|
||||
Host: "asurascans.com",
|
||||
LatestChapter: asuraLatestChapter,
|
||||
Cover: ogImageCover,
|
||||
},
|
||||
"demonic": {
|
||||
Host: "demonicscans.org",
|
||||
LatestChapter: demonicLatestChapter,
|
||||
Cover: ogImageCover,
|
||||
},
|
||||
"comix": {
|
||||
Host: "comix.to",
|
||||
LatestChapter: comixLatestChapter,
|
||||
Cover: comixCoverEntry,
|
||||
},
|
||||
"kagane": {
|
||||
Host: "kagane.to",
|
||||
LatestChapter: kaganeLatestChapter,
|
||||
Cover: kaganeCoverEntry,
|
||||
Browser: &browserRead{
|
||||
Read: kaganeRead,
|
||||
Done: func(body string) bool { return body != "" },
|
||||
// Never falls back: a plain fetch of a kagane page or cover would
|
||||
// only ever retrieve a challenge page (verified 2026-08-03).
|
||||
Fallback: false,
|
||||
},
|
||||
},
|
||||
"novelfull": {
|
||||
Host: "novelfull.com",
|
||||
LatestChapter: novelfullLatestChapter,
|
||||
Cover: novelfullCoverEntry,
|
||||
Browser: &browserRead{
|
||||
Read: novelfullRead,
|
||||
// The interstitial has a DOM too, so "the payload arrived" has to
|
||||
// exclude it explicitly.
|
||||
Done: func(body string) bool { return body != "" && !isInterstitial(body) },
|
||||
Fallback: true,
|
||||
},
|
||||
},
|
||||
"lightnovelworld": {
|
||||
Host: "lightnovelworld.net",
|
||||
LatestChapter: lnwLatestChapter,
|
||||
Cover: ogImageCover,
|
||||
},
|
||||
}
|
||||
|
||||
// browserBackedSites is derived from the registry: the Sites whose pages are
|
||||
// read through the browser sidecar, which are also the ones granted the longer
|
||||
// cooldown. Sorted so callers that range it (the due query, the browser
|
||||
// fetcher's dispatch) see a stable order instead of map-iteration noise.
|
||||
func browserBackedSites() []string {
|
||||
out := make([]string, 0, len(sites))
|
||||
for name, s := range sites {
|
||||
if s.Browser != nil {
|
||||
out = append(out, name)
|
||||
}
|
||||
}
|
||||
sort.Strings(out)
|
||||
return out
|
||||
}
|
||||
|
||||
@@ -1,6 +1,9 @@
|
||||
package latest
|
||||
|
||||
import "testing"
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// Trimmed from https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af
|
||||
// fetched 2026-07-26. The first anchor is the "First Chapter" shortcut: it is a
|
||||
@@ -69,14 +72,113 @@ const novelfullSeriesFixture = `
|
||||
<a href="/release-that-witch/chapter-9999.html">Chapter 9999</a>
|
||||
`
|
||||
|
||||
// Trimmed from https://lightnovelworld.net/novel/a-will-eternal/ fetched
|
||||
// 2026-08-05. Its chapter anchors are absolute and flat — /<slug>-chapter-<n>/
|
||||
// at the site root, not under /novel/. The last anchor is another series'.
|
||||
// Trimmed from
|
||||
// https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/
|
||||
// fetched 2026-08-11, whole document (~307 KB decoded; the wire body is ~32 KB
|
||||
// zstd-compressed). This
|
||||
// novel publishes its chapters under two Chapter Slugs: 1–99 at
|
||||
// …-not-chapter-<n>/ and 100–423 at …-not-them-all-chapter-<n>/, and the page
|
||||
// lists them newest-first, so the …-not-them-all anchors precede the …-not
|
||||
// anchors. Every anchor through the comment-thread marker is verbatim page
|
||||
// text (the site renders this novel's chapter titles as "[ ... words ]"). The
|
||||
// comment block after the marker is the real wpdiscuz comment #wpd-comm-358_0
|
||||
// from https://lightnovelworld.net/novel/the-sword-illuminates-the-great-wilderness/
|
||||
// — the pinned page serves zero comments — with its share/link/vote/reply
|
||||
// boilerplate trimmed. The comment's body carried no link, so the bare <a
|
||||
// href> to https://lightnovelworld.net/overgeared-chapter-2059/ inside
|
||||
// wpd-comment-text is the one composed element; that URL is a real chapter of
|
||||
// a real different novel (overgeared; fetched, HTTP 200).
|
||||
const lnwSeriesFixture = `
|
||||
<a href="https://lightnovelworld.net/a-will-eternal-chapter-1/">Chapter 1</a>
|
||||
<a href="https://lightnovelworld.net/a-will-eternal-chapter-1317/">Chapter 1317</a>
|
||||
<a href="https://lightnovelworld.net/a-will-eternal-chapter-1298/">Chapter 1298</a>
|
||||
<a href="https://lightnovelworld.net/overgeared-chapter-9999/">Chapter 9999</a>
|
||||
<li data-ID="102741">
|
||||
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-404/">
|
||||
<div class="epl-num">Vol. 1 Ch. 404</div>
|
||||
<div class="epl-title">[ ... words ]</div>
|
||||
<div class="epl-date">April 12, 2026</div>
|
||||
</a>
|
||||
</li>
|
||||
<li data-ID="102780">
|
||||
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-423/">
|
||||
<div class="epl-num">Vol. 1 Ch. 423</div>
|
||||
<div class="epl-title">[ ... words ]</div>
|
||||
<div class="epl-date">April 7, 2026</div>
|
||||
</a>
|
||||
</li>
|
||||
<li data-ID="102527">
|
||||
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-300/">
|
||||
<div class="epl-num">Vol. 1 Ch. 300</div>
|
||||
<div class="epl-title">[ ... words ]</div>
|
||||
<div class="epl-date">April 4, 2026</div>
|
||||
</a>
|
||||
</li>
|
||||
<li data-ID="102325">
|
||||
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-200/">
|
||||
<div class="epl-num">Vol. 1 Ch. 200</div>
|
||||
<div class="epl-title">[ ... words ]</div>
|
||||
<div class="epl-date">March 29, 2026</div>
|
||||
</a>
|
||||
</li>
|
||||
<li data-ID="102121">
|
||||
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-100/">
|
||||
<div class="epl-num">Vol. 1 Ch. 100</div>
|
||||
<div class="epl-title">[ ... words ]</div>
|
||||
<div class="epl-date">March 22, 2026</div>
|
||||
</a>
|
||||
</li>
|
||||
<li data-ID="26014">
|
||||
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-chapter-99/">
|
||||
<div class="epl-num">Vol. 1 Ch. 99</div>
|
||||
<div class="epl-title">Chapter 99</div>
|
||||
<div class="epl-date">November 5, 2025</div>
|
||||
</a>
|
||||
</li>
|
||||
<li data-ID="25916">
|
||||
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-chapter-50/">
|
||||
<div class="epl-num">Vol. 1 Ch. 50</div>
|
||||
<div class="epl-title">Chapter 50</div>
|
||||
<div class="epl-date">October 29, 2025</div>
|
||||
</a>
|
||||
</li>
|
||||
<li class='tseplsfrst' data-ID="25818">
|
||||
<a href="https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-chapter-1/">
|
||||
<div class="epl-num">Vol. 1 Ch. 1</div>
|
||||
<div class="epl-title">Chapter 01</div>
|
||||
<div class="epl-date">October 11, 2025</div>
|
||||
</a>
|
||||
</li>
|
||||
<div id="wpd-threads" class="wpd-thread-wrapper">
|
||||
<div class="wpd-thread-list">
|
||||
<div id='wpd-comm-358_0' class='comment byuser comment-author-jimbear even thread-even depth-1 wpd-comment wpd_comment_level-1'><div class="wpd-comment-wrap wpd-blog-user wpd-blog-subscriber">
|
||||
<div class="wpd-comment-left ">
|
||||
<div class="wpd-avatar ">
|
||||
<img alt='hasbi asy' src='https://secure.gravatar.com/avatar/3c1792cb31cab842f90e7c463f0948e98536cf938aebbd2f678f937ca60fb799?s=64&d=mm&r=g' srcset='https://secure.gravatar.com/avatar/3c1792cb31cab842f90e7c463f0948e98536cf938aebbd2f678f937ca60fb799?s=128&d=mm&r=g 2x' class='avatar avatar-64 photo' height='64' width='64' decoding='async'/>
|
||||
</div>
|
||||
<div class="wpd-comment-label" wpd-tooltip="Member" wpd-tooltip-position="right">
|
||||
<span>Member</span>
|
||||
</div>
|
||||
|
||||
</div>
|
||||
<div id="comment-358" class="wpd-comment-right">
|
||||
<div class="wpd-comment-header">
|
||||
<div class="wpd-comment-author ">
|
||||
hasbi asy
|
||||
</div>
|
||||
<div class="wpd-comment-date" title="July 9, 2026 1:44 am">
|
||||
<i class='far fa-clock' aria-hidden='true'></i>
|
||||
1 month ago
|
||||
</div>
|
||||
|
||||
</div>
|
||||
|
||||
<div class="wpd-comment-text">
|
||||
<p>where’s everyone</p>
|
||||
|
||||
<a href="https://lightnovelworld.net/overgeared-chapter-2059/">https://lightnovelworld.net/overgeared-chapter-2059/</a>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
<div id='wpdiscuz_form_anchor-358_0'></div>
|
||||
</div>
|
||||
</div>
|
||||
`
|
||||
|
||||
// Trimmed from https://asurascans.com/comics/chronicles-of-the-demon-faction-f886a8af
|
||||
@@ -98,8 +200,10 @@ const kaganeCoverFixture = `{"series_covers":[{"cover_id":"019fe11a-84d1-714b-9c
|
||||
// Trimmed from https://novelfull.com/reverend-insanity.html on 2026-08-10.
|
||||
const novelfullCoverFixture = `<meta name="image" content="https://novelfull.com/uploads/webp/novel/reverend-insanity-82661d911a.webp">`
|
||||
|
||||
// Trimmed from https://lightnovelworld.net/novel/a-will-eternal/ on 2026-08-10.
|
||||
const lnwCoverFixture = `<meta content='https://lightnovelworld.net/wp-content/uploads/2026/03/a-will-eternal-1.webp' property='og:image'>`
|
||||
// Trimmed from
|
||||
// https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/
|
||||
// on 2026-08-11.
|
||||
const lnwCoverFixture = `<meta property="og:image" content="https://i1.wp.com/lightnovelworld.net/wp-content/uploads/2025/10/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all.jpg" />`
|
||||
|
||||
func TestCoverFrom(t *testing.T) {
|
||||
const comixURL = "https://comix.to/title/n8we-dungeons-and-crayons"
|
||||
@@ -139,7 +243,7 @@ func TestCoverFrom(t *testing.T) {
|
||||
{
|
||||
name: "lightnovelworld reads og image",
|
||||
site: "lightnovelworld", body: lnwCoverFixture, wantOK: true,
|
||||
wantCover: "https://lightnovelworld.net/wp-content/uploads/2026/03/a-will-eternal-1.webp",
|
||||
wantCover: "https://i1.wp.com/lightnovelworld.net/wp-content/uploads/2025/10/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all.jpg",
|
||||
},
|
||||
{
|
||||
name: "later metadata cover survives empty match",
|
||||
@@ -307,24 +411,38 @@ func TestLatestChapterFrom(t *testing.T) {
|
||||
body: novelfullSeriesFixture,
|
||||
wantOK: false,
|
||||
},
|
||||
// Stored before the slug split, so the address carries the ...-not
|
||||
// Chapter Slug; the 100-423 block under the other slug must still win.
|
||||
// The body is the chapter-list portion of lnwSeriesFixture with the
|
||||
// comment block omitted; the marker is kept, because a body without it
|
||||
// is skipped, not scanned.
|
||||
{
|
||||
name: "lightnovelworld takes the max and ignores another series",
|
||||
name: "lightnovelworld max spans both chapter slugs",
|
||||
site: "lightnovelworld",
|
||||
seriesURL: "https://lightnovelworld.net/novel/a-will-eternal/",
|
||||
body: lnwSeriesFixture,
|
||||
wantOK: true, wantNum: 1317, wantLabel: "Chapter 1317",
|
||||
seriesURL: "https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not/",
|
||||
body: strings.SplitN(lnwSeriesFixture, lnwCommentMarker, 2)[0] + lnwCommentMarker,
|
||||
wantOK: true, wantNum: 423, wantLabel: "Chapter 423",
|
||||
},
|
||||
{
|
||||
name: "lightnovelworld tolerates a series url with no trailing slash",
|
||||
name: "lightnovelworld comment anchor cannot set the latest chapter",
|
||||
site: "lightnovelworld",
|
||||
seriesURL: "https://lightnovelworld.net/novel/a-will-eternal",
|
||||
seriesURL: "https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/",
|
||||
body: lnwSeriesFixture,
|
||||
wantOK: true, wantNum: 1317, wantLabel: "Chapter 1317",
|
||||
wantOK: true, wantNum: 423, wantLabel: "Chapter 423",
|
||||
},
|
||||
// Marker removed from the fixture, comment block still present: a
|
||||
// redesign must degrade into a skip, never into the comment's number.
|
||||
{
|
||||
name: "lightnovelworld body without the comment marker is skipped",
|
||||
site: "lightnovelworld",
|
||||
seriesURL: "https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/",
|
||||
body: strings.ReplaceAll(lnwSeriesFixture, lnwCommentMarker, ""),
|
||||
wantOK: false,
|
||||
},
|
||||
{
|
||||
name: "lightnovelworld yields nothing on a challenge page",
|
||||
site: "lightnovelworld",
|
||||
seriesURL: "https://lightnovelworld.net/novel/a-will-eternal/",
|
||||
seriesURL: "https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/",
|
||||
body: challengeFixture,
|
||||
wantOK: false,
|
||||
},
|
||||
|
||||
@@ -0,0 +1,101 @@
|
||||
package latest
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"net/http"
|
||||
"os"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// lnwSeriesPageFloor is the smallest body that can still be a whole
|
||||
// lightnovelworld Series page. Whole pages measured 685 KB..1.18 MB on
|
||||
// 2026-08-11 and carry the marker at ~94% of the document, so a body under
|
||||
// 100 KB is a challenge, a notice, or a truncated read — asserting on it
|
||||
// would report the marker missing when it was never fetched.
|
||||
const lnwSeriesPageFloor = 100 << 10
|
||||
|
||||
// TestSmokeLnwCommentBoundary is the live proof that the comment-thread marker
|
||||
// the lightnovelworld chapter scan truncates at (lnwCommentMarker,
|
||||
// "wpd-threads") still holds on the Site. The scan depends on it: when the
|
||||
// marker vanishes every Series is skipped and logged — correct, but silent
|
||||
// until a Reader notices their Latest Chapter has stopped moving. It runs only
|
||||
// when SMOKE_LNW_SERIES_URL is set — the URL of the live Series page to check.
|
||||
// The immortality-simulator page measured 2026-08-11
|
||||
// (docs/research/lightnovelworld-chapter-vs-series-slug.md) is the default to
|
||||
// point it at:
|
||||
//
|
||||
// SMOKE_LNW_SERIES_URL=https://lightnovelworld.net/novel/immortality-simulator/ go test -v -run TestSmokeLnwCommentBoundary ./internal/latest
|
||||
//
|
||||
// A red run means the Site's markup has moved — the marker is gone, occurs
|
||||
// more than once, or no longer follows the last chapter anchor — and the scan
|
||||
// in sites.go is now skipping this Site. Revisit sites.go before anything
|
||||
// else; the test is not flaky. A Cloudflare challenge or a non-200 is
|
||||
// distinguished from a marker failure by the "not a marker failure" messages
|
||||
// below, which carry the observed status and body length.
|
||||
func TestSmokeLnwCommentBoundary(t *testing.T) {
|
||||
seriesURL := os.Getenv("SMOKE_LNW_SERIES_URL")
|
||||
if seriesURL == "" {
|
||||
t.Skip("SMOKE_LNW_SERIES_URL unset")
|
||||
}
|
||||
if !fetchableSeriesURL("lightnovelworld", seriesURL) {
|
||||
t.Fatalf("%q is not a fetchable lightnovelworld series URL", seriesURL)
|
||||
}
|
||||
|
||||
f, err := NewTLSFetcher()
|
||||
if err != nil {
|
||||
t.Fatalf("NewTLSFetcher: %v", err)
|
||||
}
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
|
||||
defer cancel()
|
||||
body, status, err := f.Get(ctx, seriesURL)
|
||||
if err != nil {
|
||||
t.Fatalf("Get: %v", err)
|
||||
}
|
||||
if status != http.StatusOK {
|
||||
t.Fatalf("status = %d, body %d bytes — not a marker failure; the Site did not answer this IP with a Series page", status, len(body))
|
||||
}
|
||||
if len(body) < lnwSeriesPageFloor {
|
||||
t.Fatalf("body %d bytes — not a whole Series page (measured 685 KB..1.18 MB); not a marker failure, likely a Cloudflare challenge or a non-Series response", len(body))
|
||||
}
|
||||
if !lnwChapterRe.MatchString(body) {
|
||||
t.Fatalf("no chapter anchor in %d bytes — not a lightnovelworld Series page; not a marker failure, likely a Cloudflare challenge or a non-Series response", len(body))
|
||||
}
|
||||
markerIdx := strings.Index(body, lnwCommentMarker)
|
||||
if failures := checkLnwCommentBoundary(body); len(failures) > 0 {
|
||||
t.Fatalf("%s (body %d bytes)", strings.Join(failures, "; "), len(body))
|
||||
}
|
||||
t.Logf("ok: %q once at byte %d, body %d bytes", lnwCommentMarker, markerIdx, len(body))
|
||||
}
|
||||
|
||||
// checkLnwCommentBoundary verifies the three marker assertions against a
|
||||
// fetched Series body: the marker occurs exactly once, every chapter anchor
|
||||
// precedes it, and at least one anchor precedes it at all. It returns one
|
||||
// human-readable failure per broken assertion — with observed offsets and body
|
||||
// length — and empty when the page is healthy.
|
||||
func checkLnwCommentBoundary(body string) []string {
|
||||
markerIdx := strings.Index(body, lnwCommentMarker)
|
||||
switch n := strings.Count(body, lnwCommentMarker); {
|
||||
case n == 0:
|
||||
return []string{fmt.Sprintf("%q occurs 0 times in %d bytes, want exactly 1", lnwCommentMarker, len(body))}
|
||||
case n != 1:
|
||||
return []string{fmt.Sprintf("%q occurs %d times in %d bytes (first at byte %d), want exactly 1", lnwCommentMarker, n, len(body), markerIdx)}
|
||||
}
|
||||
lastAnchor, anchorsBefore := -1, 0
|
||||
for _, m := range lnwChapterRe.FindAllStringIndex(body, -1) {
|
||||
if m[0] < markerIdx {
|
||||
anchorsBefore++
|
||||
}
|
||||
lastAnchor = m[0]
|
||||
}
|
||||
var failures []string
|
||||
if lastAnchor >= markerIdx {
|
||||
failures = append(failures, fmt.Sprintf("last chapter anchor at byte %d does not precede the marker at byte %d", lastAnchor, markerIdx))
|
||||
}
|
||||
if anchorsBefore == 0 {
|
||||
failures = append(failures, fmt.Sprintf("no chapter anchor before the marker at byte %d — the truncated prefix the scan sees yields nothing", markerIdx))
|
||||
}
|
||||
return failures
|
||||
}
|
||||
+14
-1
@@ -77,7 +77,7 @@ start_browser() {
|
||||
--no-first-run \
|
||||
--no-default-browser-check \
|
||||
--disable-gpu \
|
||||
about:blank >/dev/null 2>&1 &
|
||||
about:blank >/dev/null &
|
||||
printf '%s\n' "$!" >"$pid_file"
|
||||
}
|
||||
|
||||
@@ -140,6 +140,9 @@ connection() {
|
||||
result=$?
|
||||
fi
|
||||
else
|
||||
# The client only ever sees a bare connection reset here, so this is
|
||||
# the sole record that the browser, not the network, was the problem.
|
||||
echo "browser did not come up; dropping connection" >&2
|
||||
result=1
|
||||
fi
|
||||
finish_connection
|
||||
@@ -174,6 +177,16 @@ for marker in "$connections_dir"/*; do
|
||||
done
|
||||
rm -f "$pid_file" "$last_use_file"
|
||||
|
||||
# Chrome's singleton lock names the hostname and pid that took it, and a
|
||||
# container rebuild changes both — so a Chrome killed uncleanly (OOM, docker
|
||||
# kill) leaves a lock the next container reads as "another computer holds this
|
||||
# profile" and refuses to start behind, permanently, with the only symptom a
|
||||
# bare connection reset at 9222. Clearing it here is safe precisely because
|
||||
# container_name pins this volume to one container: nothing can be holding the
|
||||
# profile at the moment this line runs. The lock is process state; the
|
||||
# clearance cookies it sits beside are not, and are left alone.
|
||||
rm -f "$profile"/Singleton*
|
||||
|
||||
# Chrome binds DevTools to loopback and silently ignores
|
||||
# --remote-debugging-address. socat remains the network front-end, but each
|
||||
# accepted connection now starts a browser on demand and is tracked by a
|
||||
|
||||
+1
-1
@@ -25,7 +25,7 @@ services:
|
||||
# Owner's Discord user ID — required. Seeds the owner Reader (the
|
||||
# administrator); every other Reader registers on their first login.
|
||||
OWNER_DISCORD_ID: ${OWNER_DISCORD_ID:?set OWNER_DISCORD_ID in .env}
|
||||
ALLOWED_ORIGINS: ${ALLOWED_ORIGINS:-https://asuracomic.net,https://asurascans.com,https://demonicscans.org,https://comix.to,https://kagane.to,https://novelfull.com,https://lightnovelworld.net}
|
||||
ALLOWED_ORIGINS: ${ALLOWED_ORIGINS:-https://asurascans.com,https://demonicscans.org,https://comix.to,https://kagane.to,https://novelfull.com,https://lightnovelworld.net}
|
||||
# The bookmarks database. Host is the compose service name; the password
|
||||
# comes from .env so it is never committed.
|
||||
DATABASE_URL: ${DATABASE_URL:-postgres://bookmarks:${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD in .env}@postgres:5432/bookmarks?sslmode=disable}
|
||||
|
||||
@@ -19,8 +19,15 @@ and because the Series row now knows how many Readers hold it, the poll queue is
|
||||
absorbs the shortfall. That ordering is only expressible because the split happened.
|
||||
|
||||
Raising throughput instead was rejected: sweeping 400 Series hourly needs the stagger
|
||||
cut from 20s to ~9s, doubling request rate against sites that already bot-score the
|
||||
single VPS IP.
|
||||
cut from 20s to ~9s, doubling request rate against sites already fronted by Cloudflare
|
||||
from the single VPS IP.
|
||||
|
||||
Corrected 2026-08-12: the original wording said those sites "bot-score" the VPS IP.
|
||||
They do not — the 1-99 bot score is Enterprise Bot Management only, and no per-IP
|
||||
request rate is documented as an input to challenge issuance
|
||||
(`docs/research/cloudflare-bot-scoring-and-poll-cadence.md`). The decision stands on
|
||||
its first argument, sweep depth versus the 1-hour cooldown; the rate-limit fear was
|
||||
never evidenced.
|
||||
|
||||
## Only the Poll writes Series fields
|
||||
|
||||
|
||||
@@ -23,8 +23,15 @@ requests: the poller's due query joins bookmarks, production held four kagane
|
||||
series and no bookmarks on any of them, and with no kagane bookmark the web UI
|
||||
never rendered a kagane cover either.
|
||||
|
||||
The home machine has 5.9 GiB of swap and a residential egress, which Cloudflare
|
||||
scores better than a datacenter IP. Both machines were already on the tailnet.
|
||||
The home machine has 5.9 GiB of swap and a residential egress, which avoids the
|
||||
cloud-hosting-IP signature Cloudflare's Bot Fight Mode documentedly challenges. Both
|
||||
machines were already on the tailnet.
|
||||
|
||||
Corrected 2026-08-12: the original wording said Cloudflare "scores" a residential
|
||||
egress better than a datacenter IP. There is no score on a free-plan zone; what is
|
||||
documented is signature matching, and hosting-provider IP space is one of the
|
||||
signatures (`docs/research/cloudflare-bot-scoring-and-poll-cadence.md`). Memory was
|
||||
the load-bearing reason regardless.
|
||||
|
||||
This move is only safe because covers are persisted (ADR-0005's sibling work,
|
||||
issue #43/#45) and the browser is on-demand (ADR-0005). Without stored covers a
|
||||
|
||||
@@ -0,0 +1,102 @@
|
||||
# ADR-0008: A Series identity is discovered from the Site's links, never derived from an address
|
||||
|
||||
Date: 2026-08-11
|
||||
Status: accepted
|
||||
|
||||
## Decision
|
||||
|
||||
On lightnovelworld, a Series is identified by the slug in its `/novel/<slug>/`
|
||||
address, and the userscript obtains that address by reading the chapter page's
|
||||
`a[aria-label="All Chapter"]` anchor. It no longer constructs the address by
|
||||
string manipulation of the chapter path. When neither that anchor nor the
|
||||
microdata breadcrumb is present, the page resolves to `type: "other"` and no
|
||||
Bookmark is offered.
|
||||
|
||||
A Chapter Slug — the slug a chapter address is built from — is not an identity
|
||||
and is not stored. The backend finds chapters by matching
|
||||
`lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/` against the series page
|
||||
body **truncated at the first `wpd-threads`**.
|
||||
|
||||
## Why
|
||||
|
||||
Measured on lightnovelworld between 2026-08-10 and 2026-08-11.
|
||||
|
||||
The chapter slug and the series slug are two independent facts. In a 41-novel
|
||||
sample, 3 diverged (~7%). `/my-longevity-simulation-chapter-1/` returns 200
|
||||
while `/novel/my-longevity-simulation/` returns 404 with no redirect, and that
|
||||
novel's real address is `/novel/immortality-simulator/`. Divergence runs in both
|
||||
directions: `the-sword-illuminates-the-great-wilderness` is served by chapter
|
||||
slug `radiant-blade-of-the-wilderness`. Neither slug is computable from the
|
||||
other, and the Site publishes no alternative-names field, so the mapping exists
|
||||
only in the chapter page's own markup.
|
||||
|
||||
Storing the Chapter Slug beside the identity does not work, because a Series may
|
||||
have more than one. `/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/`
|
||||
serves chapters 1–99 under `…-one-skill-not` and 100–423 under
|
||||
`…-one-skill-not-them-all`. Both resolve, and both chapter pages point back at
|
||||
the same Series.
|
||||
|
||||
Deriving the identity from the chapter path also made one Series produce two
|
||||
rows: bookmarking from the series page yielded `lightnovelworld:immortality-simulator`,
|
||||
and from a chapter page `lightnovelworld:my-longevity-simulation`.
|
||||
|
||||
The pointer is reliable. Across 8 chapter pages — chapter 1, chapter 1200, the
|
||||
latest chapter, both slugs of the split novel, two divergent novels, and a novel
|
||||
with a number in its title — the `All Chapter` anchor and breadcrumb position 2
|
||||
were both present and agreed every time, including on the old-slug pages. Three
|
||||
narrower selectors were rejected on evidence: `a[href*="/novel/"]` matches the
|
||||
header nav index first; matching the text "All Chapter" false-matches the novel
|
||||
titled "…Not Them All Chapter 200"; and the JSON-LD breadcrumb's position 2 is
|
||||
the chapter, not the Series.
|
||||
|
||||
The scan is truncated because a series page server-renders a wpdiscuz comment
|
||||
thread below the chapter list, and comment bodies are HTML that can carry an
|
||||
anchor. Verified on `/novel/the-sword-illuminates-the-great-wilderness/`:
|
||||
comment `#wpd-comm-358_0` rendered in the initial HTML, corroborated by that
|
||||
page's comment RSS feed. The scanner takes the maximum chapter number with no
|
||||
upper bound, and a Series row is shared by every Reader (ADR-0003), so one
|
||||
comment containing a link to a high-numbered chapter would pin that Series'
|
||||
Latest Chapter for everyone. `wpd-threads` occurs exactly once per page and
|
||||
follows every chapter anchor on all 4 series pages measured. `wpdcom` and
|
||||
`wpdiscuz` are unusable: they occur 111 to 143 times per page, including in
|
||||
`<head>` before the chapter list.
|
||||
|
||||
## Considered options
|
||||
|
||||
**Correct `series_url` only, leaving the Chapter Slug as the identity.** This
|
||||
repairs the Poll with no migration, because the Poll reads the stored address
|
||||
rather than the identity. Rejected: it keeps an identity that the Site does not
|
||||
guarantee to be stable, and leaves the duplicate-row hazard in place.
|
||||
|
||||
**Scope the match to the chapter-list container.** Rejected on measurement. The
|
||||
prior research named `ul.clstyle`; on all 4 pages sampled that is the hidden,
|
||||
empty "Latest Reading" template, and the real list is a classless `<ul>` in
|
||||
`div.eplister.eplisterfull`. The backend scans raw HTML with no parser, so
|
||||
container extraction means a second regex against class names, which is more
|
||||
fragile than the one-off truncation and protects nothing extra.
|
||||
|
||||
**Repair the stored rows with a SQL migration.** Rejected as impossible: the
|
||||
database holds no source for the correct slug. The `series` table has no chapter
|
||||
address, and no endpoint on the Site maps one slug to the other.
|
||||
|
||||
## Consequences
|
||||
|
||||
Existing Bookmarks on divergent novels stop matching their own chapter pages,
|
||||
because `keyOf` changes. The userscript therefore migrates a row in place when
|
||||
it sees the mismatch: it rewrites the row's key, identity and address in the
|
||||
cache, the retry-queue entry, and the last-checked map, then syncs. This heals
|
||||
only when the Reader next opens a chapter page of that novel.
|
||||
|
||||
A migrated row leaves its old `series` row behind. Nothing deletes it, but
|
||||
`DueForLatestCheck` is an inner join against bookmarks, so a Series with no
|
||||
Bookmarks is never polled again. The old row is permanently stored and
|
||||
permanently inert.
|
||||
|
||||
A Bookmark whose Reader never opens a chapter page of that novel does not heal.
|
||||
It continues to return 404 at every cooldown, as it does today.
|
||||
|
||||
If the truncation marker disappears, the scan is skipped and logged rather than
|
||||
run against the whole page. An env-gated live test, `TestSmokeLnwCommentBoundary`,
|
||||
asserts that `wpd-threads` still occurs exactly once and still follows the last
|
||||
chapter anchor. It skips when its environment variable is unset, matching the
|
||||
existing `TestSmokeKagane*` convention.
|
||||
@@ -0,0 +1,93 @@
|
||||
# ADR-0009: A Site answers fixed questions; an unusual Site owns its own fetch
|
||||
|
||||
Date: 2026-08-11
|
||||
Status: accepted
|
||||
|
||||
## Decision
|
||||
|
||||
Per-Site knowledge lives in one registry in `backend/internal/latest/sites.go`,
|
||||
keyed by the stored site string. An entry answers a fixed set of questions: the
|
||||
hostname a `series_url` must carry, how to find the Latest Chapter in a body,
|
||||
how to find the Cover address in a body, and — for a Site behind a JavaScript
|
||||
challenge — how to read its payload from a cleared tab, how to tell that the
|
||||
payload arrived, and whether the plain-TLS fetcher may take over when no
|
||||
browser is configured (`Fallback`; kagane never falls back, novelfull does,
|
||||
each on measured evidence).
|
||||
|
||||
The question set does not grow to accommodate one Site. When a Site needs
|
||||
something the set cannot express, that Site gets an optional override and
|
||||
performs its own fetch, leaving the other entries untouched. The override is a
|
||||
per-Site escape hatch, not a stage every Site passes through, and it is added to
|
||||
the registry type when the first Site needs it rather than in advance.
|
||||
|
||||
Two things stay outside the registry. Cover *bytes* are routed by address shape
|
||||
in `fetchCoverBytes`, never by Site name, so the Poll and the Acquisition cannot
|
||||
drift apart. And the host pin inside a browser entry's read is kept even though
|
||||
`fetchableSeriesURL` has already pinned the same host: a headless browser is a
|
||||
strong SSRF primitive and `series_url` arrives in a client-supplied PUT body, so
|
||||
the second check is deliberate and must not be deduplicated.
|
||||
|
||||
## Why
|
||||
|
||||
Before the registry, the site string was compared in six places across three
|
||||
files: the Latest Chapter switch (`sites.go:102`), the Cover switch
|
||||
(`sites.go:263`), the browser-backed list and the fetcher choice
|
||||
(`poller.go:60`, `poller.go:132`), the host pins (`poller.go:314-325`), and the
|
||||
payload read (`browser.go:96-119`). Nothing tied them together, so adding a
|
||||
seventh Site meant finding all six unaided, and a Site added to five of them
|
||||
failed at the sixth in production rather than at compile time.
|
||||
|
||||
The Sites are not alike and the registry does not ask them to be. asura strips a
|
||||
rotating build hash from its slug before scoping a regex; comix reads a JSON
|
||||
blob embedded in server-rendered HTML; kagane's chapter list exists only in its
|
||||
JSON API, which must be called from inside the page so the request carries the
|
||||
clearance cookie; lightnovelworld must truncate the body at the comment thread
|
||||
first. What they have in common is not behaviour, it is the questions they
|
||||
answer. Arbitrary behaviour behind one entry is the point.
|
||||
|
||||
Making a browser Site contribute a read and a completion test, rather than
|
||||
letting it drive the browser, was chosen because the tab lifecycle in
|
||||
`BrowserFetcher.run` is load-bearing and shared. It holds one tab open across
|
||||
re-reads, because a Cloudflare interstitial needs several seconds of live page
|
||||
to solve itself and write clearance into the shared cookie jar; reading once and
|
||||
closing the tab, which is what this did before 2026-08-08, never clears
|
||||
anything. It also serialises the browser, binds the caller's deadline to the
|
||||
tab, distinguishes a lost browser from a retryable read, and paces re-reads.
|
||||
Spreading that across per-Site adapters would put one subtle, measured loop
|
||||
behind six doors.
|
||||
|
||||
This costs the adapters little, because `chromedp.Run` takes an Action and
|
||||
`chromedp.Tasks` is an Action. A Site that must click, wait on a selector, and
|
||||
then evaluate expresses all of it as its read. Only a Site needing something
|
||||
outside the per-tab loop — its own cadence, two tabs, a tab held between calls,
|
||||
cookies set before navigation — falls outside, and that Site takes the override.
|
||||
|
||||
## Considered options
|
||||
|
||||
**Widen the shared interface whenever a Site needs something new.** Rejected:
|
||||
one Site's requirement becomes a field on all seven entries, and the entries
|
||||
that ignore it still have to be read and understood by anyone adding the eighth.
|
||||
|
||||
**Give every Site the whole fetch.** Rejected: it makes the browser lifecycle
|
||||
above a per-Site concern, and pulls `chromedp` into adapters for five Sites that
|
||||
never open a browser.
|
||||
|
||||
**A Go `interface` with a method set instead of a registry of records.**
|
||||
Rejected: most Sites differ in one or two answers, and three share a single
|
||||
Cover implementation, so a method set produces near-empty types. A missing
|
||||
answer is a nil value caught at dispatch, which is where an unknown Site is
|
||||
already handled.
|
||||
|
||||
## Consequences
|
||||
|
||||
Adding a Site is one registry entry. The existing dispatch functions —
|
||||
`latestChapterFrom`, `coverFrom`, `fetchableSeriesURL` — become registry
|
||||
lookups, so the table tests that drive them by site string are unchanged.
|
||||
|
||||
A future architecture review will see an override that only one Site uses and
|
||||
read it as an inconsistency to collapse. It is not. Collapsing it means either
|
||||
widening the question set for every Site or moving the shared tab lifecycle into
|
||||
the adapters, and both were rejected here on the evidence above.
|
||||
|
||||
An unknown site string resolves to the zero entry and fails the existing
|
||||
not-fetchable and no-fetcher paths, which log and skip. That is unchanged.
|
||||
@@ -0,0 +1,243 @@
|
||||
# Cloudflare bot scoring and poll cadence — what is actually documented
|
||||
|
||||
Research note for the browser-backed poller cadence decision (kagane.to, novelfull.com, comix.to). All pages were fetched live from **developers.cloudflare.com / blog.cloudflare.com on 2026-08-12**. Primary sources only: Cloudflare's own documentation, Cloudflare blog posts, and RFCs/standards where noted. Where Cloudflare does not publicly document something, this note says **`Not publicly documented`** instead of guessing. Repo-measured facts from the existing poller work are reused without re-derivation and marked as such.
|
||||
|
||||
The site configuration of the three challenged sites (which plan, which bot product, Challenge Passage setting, whether Precursor is enabled) is **not observable from outside** — Cloudflare does not expose a zone's security configuration to anonymous clients. Anything in this note that depends on those unknowns is flagged `[INFERENCE]`.
|
||||
|
||||
---
|
||||
|
||||
## Short answer
|
||||
|
||||
**No — polling once per hour per series, from one residential IP through one real Chrome holding a valid `cf_clearance`, carries no documented challenge risk beyond polling every six hours.** Challenge issuance on Free/Pro-grade protection (Bot Fight Mode, WAF rules) is signature- and fingerprint-driven (headless browsers, cloud-hosted IPs, browser signals); the only rate-aware detector — the per-request bot score — exists solely on Enterprise Bot Management, and free-plan Rate Limiting Rules count per-IP over 10-second windows, which 20–60 requests/hour cannot trip. Both cadences re-solve the challenge every visit anyway: `cf_clearance` expires after **30 minutes by default** (site-configurable), so a 1-hour gap always finds it expired. The documented lever that matters — already verified in this repo — is **fingerprint quality**: real Chrome + real timezone clears in ~4 s; headless variants never do. Residual, undocumented risk is site-specific: Challenge Passage, Precursor (behavior-bound re-challenge), and custom WAF rules are zone settings not observable from outside.
|
||||
|
||||
---
|
||||
|
||||
## Summary answer table
|
||||
|
||||
| Question | Answer | Section |
|
||||
|---|---|---|
|
||||
| What is a bot score? | Integer 1–99 = Cloudflare's certainty a request is automated; **Enterprise Bot Management only**; everyone else gets coarse "bot groupings" (Pro+ analytics) or nothing. | §1 |
|
||||
| Is request frequency documented as a bot-score input? | Partially: ML inputs are "headers, session characteristics, and browser signals"; the `__cf_bm` cookie "measures a single user's request pattern". **No numeric rate threshold is documented.** Volume policing is Rate Limiting, a separate product. | §1, §5 |
|
||||
| What can a free-plan site deploy? | Bot Fight Mode only: challenges *signatures* (headless browsers, cloud-hosting IPs) with a computational challenge; JavaScript Detections forced on; no scores, no analytics, cannot be skipped/customized. | §2 |
|
||||
| Does the free tier score continuously? | **No.** No score exists on Free at all — granular scores need Enterprise Bot Management, groupings need Pro+. BFM just challenges signature matches. | §2 |
|
||||
| What does `cf-mitigated: challenge` mean? | The response was a Cloudflare Challenge Page (any type); `challenge` is the only value; body is always `text/html`. | §3 |
|
||||
| How is a successful solve remembered? | `cf_clearance` cookie, issued with `SameSite=None; Secure; Partitioned`; suppresses challenges while valid. | §3, §4 |
|
||||
| `cf_clearance` lifetime? | **30 minutes by default**, configurable by the site via Challenge Passage (15–45 min recommended); +skew minutes; +1 h for XHR. | §4 |
|
||||
| Is `cf_clearance` bound to IP / device? | Documented: "securely tied to the specific visitor and device it was issued to"; the *solve request* must come from the same IP that received the challenge (different IP → invalid solve → challenge loop). Replay from another machine/IP is therefore **not** valid. | §4 |
|
||||
| What invalidates clearance early? | Precursor (if enabled): suspicious session → clearance reduced/invalidated, re-challenge even before expiry. Zone-level toggle; unobservable from outside. | §4 |
|
||||
| Does polling more often raise challenge risk? | **No documented mechanism at 20–60 req/hour.** Free-plan rate limiting is 10 s/IP-only; DDoS thresholds are ~1,000 errors/sec. Scores (the only rate-aware thing) are Enterprise-only. | §5 |
|
||||
| Is there a documented "legitimate poller" path? | Yes, but it requires **self-identification** (Web Bot Auth signature or published IP list + stable UA) via the verified-bots application — not anonymity. robots.txt is voluntary; nothing exempts anonymous scrapers. | §6 |
|
||||
|
||||
---
|
||||
|
||||
## 1. What a bot score is and what feeds it
|
||||
|
||||
### 1.1 The score itself
|
||||
|
||||
Cloudflare documents the bot score as "a score from _1_ to _99_ that indicates how likely that request came from a bot" — 1 = quite certain automated, 99 = quite certain human. Source: [Bot scores — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot-score/), read 2026-08-12.
|
||||
|
||||
Two access tiers, both gated:
|
||||
|
||||
- **Granular 1–99 scores are only available to Enterprise customers who purchased Bot Management.** "All other customers can only access this information through bot groupings in Bot Analytics" (categories: `Not computed` = 0, `Automated` = 1, `Likely automated` = 2–29, `Likely human` = 30–99, `Verified bot`). Bot groupings themselves require "a Pro plan or higher". Source: [Bot scores — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot-score/), read 2026-08-12.
|
||||
- A score of 0 means "Bot Management did not evaluate the request" (redirected, handled by another feature) — "does not indicate the request is safe or human". Same source.
|
||||
|
||||
So: **on a Free-plan site there is no bot score at all, for anyone.** `[INFERENCE]` the three challenged sites are almost certainly not Enterprise Bot Management customers, but this is not externally verifiable.
|
||||
|
||||
### 1.2 The detection engines (Enterprise Bot Management)
|
||||
|
||||
Cloudflare documents four engines, all stated to apply to Enterprise Bot Management (the bot-score page: "The following detection engines only apply to Enterprise Bot Management"). Sources: [Bot scores — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot-score/) and [Bot detection engines — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot-detection-engines/), both read 2026-08-12.
|
||||
|
||||
| Engine | Documented behavior | Score it produces |
|
||||
|---|---|---|
|
||||
| **Heuristics** | "Processes all requests"; pattern matching against "a growing database of malicious fingerprints". | 1 for high-confidence deterministic detections; occasionally 29 "where Cloudflare has identified automated traffic and is still assessing traffic overlap" |
|
||||
| **Machine learning** | Supervised model, trained on "billions of daily requests". Input variables: "headers, session characteristics, and browser signals". Output: "predicted probability that a client is human (such as the probability of successfully solving a Challenge)". | Most scores 2–99 |
|
||||
| **Anomaly detection** | Unsupervised; learns a per-domain baseline, flags outlier requests; **deprecated, not onboarding new customers**. | 1 |
|
||||
| **JavaScript detections** | "Identifies headless browsers and other automation tools" via "a lightweight, invisible JavaScript injection"; runs client-side; "blocks, challenges, or passes requests to other engines". Enabled by default (but optional) in Bot Management. | Pass/fail (`cf.bot_management.js_detection.passed`), not a score |
|
||||
|
||||
Crucially, the ML engine's documented inputs are *headers, session characteristics, and browser signals* — **no rate or per-IP volume parameter is listed.** The only place request patterns appear is the `__cf_bm` cookie note: "Cloudflare uses the `__cf_bm` cookie to smooth out the bot score and reduce false positives… The Bot Management cookie measures a single user's request pattern and applies it to the machine learning data to generate a reliable bot score for all of that user's requests." ([Bot scores — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot-score/), read 2026-08-12). So frequency is *a* signal inside Enterprise Bot Management via the session cookie — **but no numeric threshold, window, or per-IP rate is published anywhere.** `Not publicly documented`: any specific requests-per-hour / requests-per-IP value that raises or lowers a bot score.
|
||||
|
||||
### 1.3 Rate limiting is a separate product
|
||||
|
||||
Volume enforcement is not part of bot scoring at all. Rate Limiting Rules are a distinct WAF product with their own evaluation phase (`http_ratelimit`, running after custom rules and before SBFM). Sources: [Rate limiting rules — Cloudflare docs](https://developers.cloudflare.com/waf/rate-limiting-rules/) and [Security features interoperability — Cloudflare docs](https://developers.cloudflare.com/waf/feature-interoperability/), read 2026-08-12. See §5 for what the Free plan's version of that product can actually do.
|
||||
|
||||
---
|
||||
|
||||
## 2. The free-plan reality
|
||||
|
||||
### 2.1 What each plan gets
|
||||
|
||||
Cloudflare's plan table ([Plans — Cloudflare docs](https://developers.cloudflare.com/bots/plans/), read 2026-08-12):
|
||||
|
||||
| Plan | Bot product | Documented detections | Action | Control |
|
||||
|---|---|---|---|---|
|
||||
| **Free** | **Bot Fight Mode** (BFM) | "Simple bots from cloud hosting providers and headless browsers" | "Cloudflare issues a computationally expensive challenge" | Applied to all traffic across the domain; no exceptions possible |
|
||||
| **Pro / Business / Enterprise (no BM)** | **Super Bot Fight Mode** (SBFM) | Configurable actions per bot category (Definitely automated / Likely automated / Verified bots) | Challenge or block | Runs on Ruleset Engine; **can** be skipped via custom rules |
|
||||
| **Enterprise + Bot Management** | Bot Management | "Simple and sophisticated bots, headless browsers, and domain-specific anomalies" | Customer-chosen (block, challenges) | Per-path / per-IP rules; access to bot score, JA3/JA4, bot tags, detection IDs |
|
||||
|
||||
Sources: [Plans — Free](https://developers.cloudflare.com/bots/plans/free/), [Plans — Bot Management for Enterprise](https://developers.cloudflare.com/bots/plans/bm-subscription/), [Bot Fight Mode — Cloudflare docs](https://developers.cloudflare.com/bots/get-started/bot-fight-mode/), [Super Bot Fight Mode — Cloudflare docs](https://developers.cloudflare.com/bots/get-started/super-bot-fight-mode/) — all read 2026-08-12.
|
||||
|
||||
### 2.2 Bot Fight Mode specifics (the Free-plan product)
|
||||
|
||||
- Identifies "traffic matching patterns of known bots" and "issues computationally expensive challenges that force the requesting client to perform CPU-intensive calculations". ([Bot Fight Mode — Cloudflare docs](https://developers.cloudflare.com/bots/get-started/bot-fight-mode/), read 2026-08-12.)
|
||||
- It "does not run on the Ruleset Engine — it operates in a separate evaluation pipeline where _Skip_, _Bypass_, and _Allow_ actions have no effect"; **you cannot bypass or skip BFM** with custom rules or Page Rules. ([Bot Fight Mode — Cloudflare docs](https://developers.cloudflare.com/bots/get-started/bot-fight-mode/) and [Security features interoperability — Cloudflare docs](https://developers.cloudflare.com/waf/feature-interoperability/), read 2026-08-12.)
|
||||
- **JavaScript Detections is automatically enabled for BFM customers and cannot be disabled.** ([Bot Fight Mode — Cloudflare docs](https://developers.cloudflare.com/bots/get-started/bot-fight-mode/), read 2026-08-12.) This is the documented hook that explains the repo's `headless-shell` / `HeadlessChrome` failures: JSD "identifies headless browsers" ([JavaScript detections — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/javascript-detections/), read 2026-08-12).
|
||||
- False positives on *legitimate automated* traffic are acknowledged as expected behavior: "false positives can occur where legitimate human or automated traffic is incorrectly challenged or blocked", and the only remedies are disabling BFM or upgrading to Bot Management. ([Handle False Positives — Cloudflare docs](https://developers.cloudflare.com/bots/troubleshooting/false-positives/), read 2026-08-12.)
|
||||
|
||||
### 2.3 Does the free tier "score" continuously?
|
||||
|
||||
**No.** The Free plan exposes no score, no bot analytics (groupings need Pro+), and no per-request decision data. BFM is a static on/off toggle that challenges signature matches; there is no continuous per-request score on Free. ([Plans — Free](https://developers.cloudflare.com/bots/plans/free/) and [Bot scores — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot-score/), read 2026-08-12.) JSD *runs* on every HTML request even on Free (forced by BFM), but the only documented way to act on its result — the `cf.bot_management.js_detection.passed` field — is gated behind an Enterprise Bot Management subscription ("Prerequisites: You must have an Enterprise Bot Management subscription"). ([JavaScript detections — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/javascript-detections/), read 2026-08-12.) On Free, JSD output feeds BFM's internal challenge decision; `Not publicly documented` exactly how.
|
||||
|
||||
`[INFERENCE]` The observed behavior on the three sites (real Chrome + real timezone passes in ~4 s; headless variants never do) is consistent with either BFM or a WAF custom rule using a challenge action — Cloudflare does not expose which product a site runs, and the failure signature (HeadlessChrome UA / headless-shell never passing) matches JSD's documented headless-browser detection either way.
|
||||
|
||||
---
|
||||
|
||||
## 3. Managed Challenge / JS challenge mechanics and `cf-mitigated`
|
||||
|
||||
### 3.1 What the observed response is
|
||||
|
||||
A `403` with `cf-mitigated: challenge`, `server: cloudflare`, and a "Just a moment…" body is a **Cloudflare Challenge Page**. Cloudflare documents: "the Challenge Page response (regardless of the Challenge Page type) will have the `cf-mitigated` header present and set to `challenge`… `challenge` is the only valid value. The header is set for all Challenge Page types", and "the content-type of a challenge will be `text/html`". ([Detect a Challenge Page response — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/detect-response/), read 2026-08-12.)
|
||||
|
||||
### 3.2 What a Challenge Page does
|
||||
|
||||
"An interstitial Challenge Page… acts as a gate between the visitor and your website… The Challenge Page intercepts the visitor… by holding the request and evaluating the browser environment for automated signals, and serving a challenge. The visitor cannot reach their destination without passing the challenge." ([Interstitial Challenge Pages — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/), read 2026-08-12.)
|
||||
|
||||
Three variants, in increasing severity ([Interstitial Challenge Pages — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/), read 2026-08-12):
|
||||
|
||||
- **Non-Interactive**: Cloudflare judges automation from browser signals gathered by injected JS; the page needs no human interaction, typically < 5 s of JS processing.
|
||||
- **Managed Challenge**: "Cloudflare dynamically chooses the appropriate type of challenge… based on the characteristics of a request from the signals indicated by their browser. Most human visitors are automatically verified and the Challenge Page will display **Successful**. However, if Cloudflare detects non-human attributes… they may be required to interact." Cloudflare's stated recommendation for WAF rules.
|
||||
- **Interactive**: requires explicit human interaction (CAPTCHA-style). Cloudflare's "End the CAPTCHA era" position is that Managed Challenges should make this rare.
|
||||
|
||||
Cloudflare's own framing matches the repo's ~4 s real-Chrome solve: a normal browser passes with no interaction (Managed Challenge auto-verify or Non-Interactive JS processing).
|
||||
|
||||
### 3.3 Which product issues which challenge
|
||||
|
||||
Documented mapping ([How Challenges work — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/concepts/how-challenges-work/), read 2026-08-12):
|
||||
|
||||
| Trigger | Challenge type |
|
||||
|---|---|
|
||||
| WAF custom rules, **rate limiting rules**, IP Access rules | Interstitial Challenge Page |
|
||||
| Bot Management | JavaScript Detections (invisible, per-request) |
|
||||
| **Bot Fight Mode / Super Bot Fight Mode** | Interstitial Challenge Page |
|
||||
| Under Attack Mode | Managed Challenge |
|
||||
|
||||
### 3.4 How a successful solve is remembered
|
||||
|
||||
Solving issues the **`cf_clearance`** cookie: "Clearance Cookie stores the proof of challenge passed. It is used to no longer issue a challenge if present. It is required to reach an origin server." ([Cloudflare Cookies — Cloudflare docs](https://developers.cloudflare.com/fundamentals/reference/policies-compliances/cloudflare-cookies/), read 2026-08-12.) "When that visitor tries to access other parts of your website, Cloudflare evaluates the cookie before presenting another challenge. If the cookie is still valid, no challenges will be shown." ([Challenge Passage — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/challenge-passage/), read 2026-08-12.)
|
||||
|
||||
`cf_clearance` is set with `SameSite=None; Secure; Partitioned`; because of the `Partitioned` (CHIPS) attribute, "a clearance obtained in one top-level context is not reused in a different top-level context" — so a clearance from kagane.to does not carry to comix.to even on the same browser. ([SameSite cookie interaction — Cloudflare docs](https://developers.cloudflare.com/waf/troubleshooting/samesite-cookie-interaction/), read 2026-08-12.)
|
||||
|
||||
---
|
||||
|
||||
## 4. `cf_clearance` — lifetime, binding, invalidation (the key question)
|
||||
|
||||
### 4.1 Lifetime
|
||||
|
||||
- **Default: 30 minutes.** "By default, the `cf_clearance` cookie has a lifetime of 30 minutes. Cloudflare recommends a setting between 15 and 45 minutes." The site owner can change it via the **Challenge Passage** setting. ([Challenge Passage — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/challenge-passage/), read 2026-08-12; also [SameSite cookie interaction — Cloudflare docs](https://developers.cloudflare.com/waf/troubleshooting/samesite-cookie-interaction/), read 2026-08-12.)
|
||||
- Validation grace: "a few extra minutes are included to account for clock skew. For XmlHTTP requests, an extra hour is added to the validation time." ([Challenge Passage — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/challenge-passage/), read 2026-08-12.)
|
||||
- "The Challenge Passage does not apply to rate limiting rules." Same source.
|
||||
- `Not publicly documented`: whether the three target sites have changed the Challenge Passage from the 30-minute default, and whether there is any maximum value Cloudflare enforces.
|
||||
|
||||
**Consequence for cadence:** with the default 30-minute TTL, *any* poll cadence ≥ 1 hour finds the cookie expired and re-solves the challenge on every visit. A 1-hour and a 6-hour cadence therefore differ only in *how many times per day* the browser re-solves (~4× for the same series), not in whether a re-solve happens. This repo already measured the re-solve cost: ~4 s with real Chrome + real timezone. `[INFERENCE]` a site could raise the Challenge Passage to hours/days, which would make a 1-hour cadence *cheaper* (cookie still valid, no re-solve) — but that setting is unobservable and unlikely to be long on free manga sites.
|
||||
|
||||
### 4.2 What it is bound to — can clearance be replayed from another IP?
|
||||
|
||||
Documented statements, both from Cloudflare's own docs:
|
||||
|
||||
1. **Device/visitor binding:** "The cookie is securely tied to the specific visitor and device it was issued to, preventing reuse across machines." ([Clearance — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/concepts/clearance/), read 2026-08-12.)
|
||||
2. **IP binding of the solve:** under Challenge limitations, Cloudflare lists "Client software where the solve request of a Managed Challenge comes from a different IP than the original IP a Challenge request was issued to. For example, if you receive the Challenge from one IP and solve it using another IP, the solve is not valid and you may encounter a Challenge loop." ([How Challenges work — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/concepts/how-challenges-work/), read 2026-08-12.)
|
||||
3. **Top-level-site binding** via CHIPS partitioning: clearance is not reused across embedding contexts. ([SameSite cookie interaction — Cloudflare docs](https://developers.cloudflare.com/waf/troubleshooting/samesite-cookie-interaction/), read 2026-08-12.)
|
||||
|
||||
So the repo's belief is **verified by primary sources**: a `cf_clearance` obtained on one machine cannot be replayed from a different IP — the solve is IP-bound and the cookie is device-bound. `Not publicly documented`: whether the cookie value is also cryptographically bound to the User-Agent or TLS/JA3 fingerprint. The only UA-adjacent documented statement is the reverse direction: challenge *solving* breaks when a browser extension modifies the User-Agent or Canvas/WebGL APIs ("Cloudflare Challenges cannot support… Browser extensions that modify the browser's User-Agent value or Web APIs such as Canvas and WebGL") — i.e., tampering with browser signals is documented to *fail* challenges ([How Challenges work — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/concepts/how-challenges-work/), read 2026-08-12).
|
||||
|
||||
### 4.3 Two-tier clearance, and the behavior-bound invalidation (Precursor)
|
||||
|
||||
`cf_clearance` now carries two kinds of clearance ([Clearance — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/concepts/clearance/), read 2026-08-12):
|
||||
|
||||
- **Challenge clearance** — granted by solving a challenge; level-gated (Interactive > Managed > Non-Interactive; higher clears bypass lower challenges); "remains valid for the duration configured by the customer (Challenge Passage), **unless Precursor determines the session is suspicious**".
|
||||
- **Precursor clearance** — "continuously updated based on session behavior"; an ongoing client-side process that periodically reassesses behavior. "If Precursor determines that a session is suspicious: the visitor's effective Challenge clearance may be **reduced or invalidated**; the visitor may be **re-challenged, even if the cookie has not expired**."
|
||||
|
||||
Precursor is documented as "client-side, session-based verification that continuously evaluates visitor behavior to identify automation… to detect automation that appears legitimate in individual requests but exhibits non-human patterns across a session", writing session state back into `cf_clearance`. It is a **zone-level toggle** (Security → Settings → Precursor; modes: Minimize Friction default, Maximize Security recommended), and "Precursor supersedes JavaScript Detections (JSD)". ([Precursor — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/precursor/), read 2026-08-12.)
|
||||
|
||||
**Implication:** the only documented mechanism by which *behavior over time* (as opposed to a single request's fingerprint) can revoke a valid clearance is Precursor — and it is opt-in per zone. `[INFERENCE]` It is unlikely to be enabled on free manga sites, but this is not observable from outside. With Precursor off (the default posture `[INFERENCE]`), a valid `cf_clearance` is honored purely on TTL + device/IP binding.
|
||||
|
||||
---
|
||||
|
||||
## 5. Does polling more often raise challenge risk?
|
||||
|
||||
### 5.1 Is per-IP request rate an input to challenge issuance? (documented answer: no such lever on non-Enterprise protection)
|
||||
|
||||
- **Bot scores** (the only per-request automated-detection output) are Enterprise-Bot-Management-only (§1.1); the ML engine's documented inputs are headers/session/browser signals, with request *pattern* entering only via `__cf_bm` — and no numeric rate is published (§1.2). `Not publicly documented`: any requests-per-hour value that changes a bot score or challenge probability.
|
||||
- **BFM/SBFM** match "patterns of known bots" — signatures, not volumes ([Bot Fight Mode — Cloudflare docs](https://developers.cloudflare.com/bots/get-started/bot-fight-mode/), read 2026-08-12).
|
||||
- **Rate limiting** is the product that polices volume, and it is opt-in per zone with explicit per-plan constraints (§5.2). Cloudflare's only documented link between "one valid clearance + high volume" is a *recommendation to site owners*: "Cloudflare recommends that customers add a rate limiting rule based on the `cf_clearance` cookie value. This helps ensure that a single, valid cookie cannot be abused by one machine to send an excessive volume of requests." ([Clearance — Cloudflare docs](https://developers.cloudflare.com/cloudflare-challenges/concepts/clearance/), read 2026-08-12.) Note: counting by cookie value is only available on Enterprise (see table below); a Free/Pro site cannot even build that rule.
|
||||
|
||||
### 5.2 What Rate Limiting Rules would do to 20–60 requests/hour
|
||||
|
||||
Rate limiting rules are opt-in; nothing runs them unless the site creates a rule. The documented per-plan capabilities ([Rate limiting rules — Cloudflare docs](https://developers.cloudflare.com/waf/rate-limiting-rules/), read 2026-08-12):
|
||||
|
||||
| Capability | Free | Pro | Business |
|
||||
|---|---|---|---|
|
||||
| Number of rules | **1** | 2 | 5 |
|
||||
| Counting characteristics | **IP only** | IP only | IP, IP w/ NAT |
|
||||
| Counting period | **10 s only** | ≤ 1 min | ≤ 10 min |
|
||||
| Mitigation timeout | **10 s** | ≤ 1 h | ≤ 1 day |
|
||||
| Fields in expression | **Path, Verified Bot** | + Host, URI, Full URI, Query | + Method, Source IP, User Agent |
|
||||
|
||||
Even the most aggressive free-plan rule (1 request per 10 s = 360/hour) is 6–18× above our 20–60/hour volume; the counting window is 10 s, so a per-hour burst is invisible to it. At 1 request per minute worst-case, **20–60 requests/hour from one residential IP cannot trip any rate limiting rule the Free plan can express.** Also documented: rate limiting is approximate, not precise — "there may be a delay of up to a few seconds between detecting a request and updating rate counters… excess requests could still reach the origin", and counters are per-data-center. ([Rate limiting rules — Cloudflare docs](https://developers.cloudflare.com/waf/rate-limiting-rules/), read 2026-08-12.) `[INFERENCE]` a site could also challenge on rate via WAF custom rules, but that requires Pro+ (custom rules are not on Free) and its own configuration.
|
||||
|
||||
### 5.3 DDoS protection (always on, all plans)
|
||||
|
||||
HTTP DDoS Attack Protection "is always enabled" and can only be tuned, not disabled. The only *published numeric* thresholds are error-rate-based: origin-error floods mitigate at the default "High" sensitivity of **1,000 errors per second** (Pro+ also requires 5× normal origin traffic). Per-IP volumetric thresholds are adaptive and `Not publicly documented` in the managed ruleset docs. 20–60 requests/hour is ~9 orders of magnitude below the published figure. ([HTTP DDoS Attack Protection — Cloudflare docs](https://developers.cloudflare.com/ddos-protection/managed-rulesets/http/), read 2026-08-12.)
|
||||
|
||||
### 5.4 Execution order (which product fires first)
|
||||
|
||||
Documented phase order: `ddos_l7` → custom rules → `http_ratelimit` (rate limiting) → managed rules → `http_request_sbfm` (SBFM); BFM runs outside this pipeline and cannot be skipped; a terminating action (block/challenge) stops later phases. ([Security features interoperability — Cloudflare docs](https://developers.cloudflare.com/waf/feature-interoperability/), read 2026-08-12.) Practical reading: on the three sites, the challenge we see could come from any of these stages; none of them documents a volume input at our scale (§5.1–5.3).
|
||||
|
||||
---
|
||||
|
||||
## 6. The documented legitimate side
|
||||
|
||||
### 6.1 Verified bots — the only "treated well" path, and it requires self-identification
|
||||
|
||||
Cloudflare documents a Verified bot as one meeting two bars ([Verified bots — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot/verified-bots/), read 2026-08-12):
|
||||
|
||||
1. **Honest self-identification** — "through a cryptographic Web Bot Auth signature, a published IP list with a stable user-agent, or reverse DNS".
|
||||
2. **Non-abusive behavior** — "it obeys `robots.txt` and crawl directives, **maintains reasonable request rates**, and has not been observed evading website owner preferences or attacking sites".
|
||||
|
||||
Relevant verified-bot *behavior classes* exist for exactly this kind of client: "**Feed Fetching** — RSS readers, podcast aggregators, and news feed bots" and "**Monitoring & Operations** — Uptime monitoring, webhooks, and health checks". Becoming verified requires an application via the dashboard and validation via Web Bot Auth or IP validation; breach of the policy (e.g. "An AI Crawler that does not respect the crawl-delay directive") removes the bot from the allowlist. ([Verified bots — Cloudflare docs](https://developers.cloudflare.com/bots/concepts/bot/verified-bots/), read 2026-08-12.)
|
||||
|
||||
"Historically, Verified bots have been excluded in default bot configurations across all plans" (same source) — i.e., verified bots are *default-allowed* under SBFM/Bot Management. **But** this path is the opposite of what a scraper wants: it requires the poller to publicly identify itself (stable, published IPs or cryptographic signatures) and to have its identity vetted by Cloudflare — and the *site* still decides via verified-bot policy whether to allow the category. There is **no documented mechanism for an anonymous low-volume automated client to be treated well.** `[INFERENCE]` a manga-site scraper would never qualify (it would be classified as Data Collection / scraping behavior, which is not a default-allowed class).
|
||||
|
||||
### 6.2 robots.txt and crawl control
|
||||
|
||||
- `robots.txt` **compliance is voluntary** — "The file expresses your preferences, but it does not prevent crawlers from accessing your content at a technical level." Enforcement requires Cloudflare's AI Crawl Control. ([robots.txt setting — Cloudflare docs](https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/), read 2026-08-12.)
|
||||
- The managed `robots.txt` feature (all plans) is aimed at AI crawlers; it prepends `Disallow` rules for AI bots and a Content Signals Policy. It does not create any allowance for generic scrapers. (Same source.)
|
||||
- RFC-side: the robots exclusion standard is an unauthenticated convention; nothing in it grants access rights. ([RFC 9309 "Robots Exclusion Protocol"](https://www.rfc-editor.org/rfc/rfc9309.html) — read 2026-08-12.) The standard defines crawl-delay etc. as voluntary directives; Cloudflare's docs are the operative statement for CF-protected sites.
|
||||
|
||||
### 6.3 The documented takeaway for "slowing down vs. fingerprint quality"
|
||||
|
||||
Cloudflare's own documentation repeatedly points at **browser/device signals** as the decision input on non-Enterprise protection (JSD detecting headless browsers; Managed Challenge choosing based on "signals indicated by their browser"; challenges failing when UA/Canvas/WebGL are modified — §3.2, §4.2), and at **identity** (verified bots) as the only legitimacy signal for automation (§6.1). Request *rate* appears only as: (a) an unnamed component of Enterprise-ML "session characteristics", (b) a voluntary verified-bot behavioral bar, and (c) the separate, opt-in, Free-plan-impotent Rate Limiting product (§5). Nothing documented says "slow down and you'll be challenged less" for a free-plan site — **fingerprint quality is the lever that the documentation actually describes**, which matches this repo's measurements (real Chrome + real timezone passes; every headless variant fails regardless of rate).
|
||||
|
||||
---
|
||||
|
||||
## 7. Implications for cadence design
|
||||
|
||||
### Documented facts (with sources above)
|
||||
|
||||
1. **1-hour vs 6-hour cadence is not a documented risk lever.** Challenge issuance on Free/Pro-grade protection is signature-based; the rate-aware scoring only exists on Enterprise Bot Management; free-plan rate limiting cannot express a limit our volume could trip (§1, §2, §5).
|
||||
2. **Every poll ≥ 1 hour re-solves the challenge anyway.** `cf_clearance` defaults to 30 minutes; Challenge Passage is site-configurable and unobservable. The re-solve cost is what this repo measured (~4 s, real Chrome + real timezone) (§4.1, repo measurements).
|
||||
3. **The documented failure modes are fingerprint, not rate:** headless browsers (JSD), cloud-hosting IPs (BFM heuristics), modified UA/Canvas/WebGL (challenge solve failure) (§2.2, §3, §4.2).
|
||||
4. **Clearance is not portable:** device-bound + solve-IP-bound + CHIPS-partitioned; replaying a cookie from another IP is documented invalid (§4.2).
|
||||
5. **Anonymity has no documented "good citizen" path:** the only legitimate-automation route (verified bots) requires self-identification and site-side allowance (§6).
|
||||
6. **The one behavior-bound revocation mechanism (Precursor) is opt-in per zone**, not a default documented behavior (§4.3).
|
||||
|
||||
### Inferences (not documented)
|
||||
|
||||
- `[INFERENCE]` The three sites run Free/Pro-grade protection (BFM, SBFM, or WAF challenge rules), not Enterprise Bot Management; therefore no continuous per-request bot score exists for our traffic.
|
||||
- `[INFERENCE]` The sites have not changed Challenge Passage to hours/days (free manga sites default to the 30-minute default); if they had, hourly polling would get *cheaper* (valid cookie, no re-solve).
|
||||
- `[INFERENCE]` Precursor is not enabled on these sites; if it were, hourly re-visits from an automated Chrome could accumulate session-behavior signals and trigger re-challenge even with a valid cookie — the only documented scenario in which polling *frequency* (via session behavior) could matter.
|
||||
- `[INFERENCE]` 6-hour cooldowns buy nothing documented beyond raw request-count reduction (fewer challenge solves per day, less origin load); the risk profile at 1 request/hour/series is not documented to differ from 6 request/hour/series.
|
||||
- `[INFERENCE]` If the owner wants belt-and-braces, the engineering levers that match the documentation are: keep the real-Chrome fingerprint (no UA spoofing, no headless-shell, real timezone — already done), keep a persistent user-data profile so `cf_clearance`/`__cf_bm` persist across visits, and treat any change of exit IP (e.g. home connection rebooting to a new IP) as a guaranteed re-solve, since clearance does not travel with the IP.
|
||||
|
||||
### Bottom line
|
||||
|
||||
Moving browser-backed sites from 6-hour to 1-hour per-series cooldown is **not contradicted by any documented Cloudflare mechanism** at 20–60 requests/hour from one residential IP through one real Chrome. The documented risk is carried by fingerprint quality (already solved in this repo) and by unobservable site configuration (Challenge Passage, Precursor, possible custom WAF rules). The residual, non-documented risk is that these sites sit behind Cloudflare's *proprietary* detection, and Cloudflare publishes neither its per-IP thresholds nor the ML feature set — so "no documented lever" is not "no lever".
|
||||
@@ -0,0 +1,572 @@
|
||||
# lightnovelworld.net — chapter slug vs. series slug
|
||||
|
||||
Research note for Gitea issue #77. All pages fetched live on **2026-08-11** with
|
||||
plain `curl` and a desktop UA. **No Cloudflare challenge was encountered on any
|
||||
request** — every fetch below returned real HTML on the first try, so the
|
||||
Playwright fallback was never needed.
|
||||
|
||||
Every claim carries the URL it came from. Nothing here is inferred from the
|
||||
existing code; where a claim is an interpretation rather than an observation it
|
||||
is marked `[INFERENCE]`.
|
||||
|
||||
---
|
||||
|
||||
## 1. Summary answer table
|
||||
|
||||
| Question | Answer | Evidence |
|
||||
|---|---|---|
|
||||
| Does the chapter slug equal the series slug? | **No, not reliably.** 3 of 41 sampled novels diverge at their latest chapter (~7%); a 4th diverges historically. | §4 |
|
||||
| Best chapter-page → series-URL pointer | `a[aria-label='All Chapter']` and the microdata breadcrumb `a[itemprop=item]` at `position 2`. Both carry the full absolute series URL. | §3 |
|
||||
| Is JSON-LD `BreadcrumbList` usable? | **No.** The Rank Math JSON-LD breadcrumb on a chapter page contains only *Home* and *the chapter itself* — the series is absent. | §3.1 |
|
||||
| Is `<link rel=canonical>` usable? | **No.** It points at the chapter page itself. | §3.2 |
|
||||
| Are `og:` tags usable? | **No.** `og:url` is the chapter page; no `og:` tag names the series. | §3.3 |
|
||||
| Does the series page list chapters in the initial HTML? | **Yes — all of them.** 1232 unique chapters (1…1232) present in one document. | §5 |
|
||||
| Is there a paginated chapter endpoint? | **No.** `/novel/<slug>/chapters/` and `/novel/<slug>/chapters/page-2` both **404**. | §5 |
|
||||
| Does the series page carry *other* novels' chapter anchors? | **No.** Across 30 series pages, every `-chapter-N/` URL belonged to the page's own novel. | §6 |
|
||||
| Unscoped regex `lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/` on a series page | **SAFE** | §6 |
|
||||
| Does `/novel/<chapter-slug>/` 404? | **Yes**, 404, no redirect. | §7 |
|
||||
| Is there a chapter-slug → series-slug endpoint that avoids parsing a chapter page? | **No endpoint found.** But `/<chapter-slug>/` (no `-chapter-N`) **302s to chapter 1**, which still requires parsing a chapter page. | §7 |
|
||||
| Is there an "Alternative names" field explaining divergence? | **No such field exists** on lightnovelworld series pages. | §4.3 |
|
||||
|
||||
---
|
||||
|
||||
## 2. The two reference pages
|
||||
|
||||
| URL | Status |
|
||||
|---|---|
|
||||
| `https://lightnovelworld.net/my-longevity-simulation-chapter-1/` | **200**, 0 redirects |
|
||||
| `https://lightnovelworld.net/novel/immortality-simulator/` | **200**, 0 redirects |
|
||||
| `https://lightnovelworld.net/novel/my-longevity-simulation/` | **404**, 0 redirects |
|
||||
|
||||
Confirmed: chapter slug `my-longevity-simulation`, series slug
|
||||
`immortality-simulator`, and the naive `/novel/<chapter-slug>/` construction is
|
||||
a hard 404.
|
||||
|
||||
---
|
||||
|
||||
## 3. Chapter page → series URL: every in-page pointer, in priority order
|
||||
|
||||
All snippets below are verbatim from
|
||||
`https://lightnovelworld.net/my-longevity-simulation-chapter-1/`.
|
||||
|
||||
### Priority 1 — `a[aria-label='All Chapter']` ✅ WORKS
|
||||
|
||||
Single occurrence in the document, inside the chapter navigation bar:
|
||||
|
||||
```html
|
||||
<div class="nvs nvsc"><a href='https://lightnovelworld.net/novel/immortality-simulator/' aria-label='All Chapter'><i class="fas fa-list-ul"></i> All Chapter</a></div>
|
||||
```
|
||||
|
||||
Note the **single quotes** on both attributes — a regex written for `href="` will
|
||||
miss it. This is the most narrowly-targeted pointer: exactly one element on the
|
||||
page has `aria-label='All Chapter'`.
|
||||
|
||||
### Priority 2 — microdata `BreadcrumbList` / `a[itemprop=item]` ✅ WORKS
|
||||
|
||||
At line 363–379 of the served HTML. `position 2` is the series:
|
||||
|
||||
```html
|
||||
<div itemscope="" itemtype="http://schema.org/BreadcrumbList">
|
||||
<span itemprop="itemListElement" itemscope="" itemtype="http://schema.org/ListItem">
|
||||
<a itemprop="item" href="https://lightnovelworld.net/"><span itemprop="name">Home</span></a>
|
||||
<meta itemprop="position" content="1">
|
||||
</span>
|
||||
›
|
||||
<span itemprop="itemListElement" itemscope="" itemtype="http://schema.org/ListItem">
|
||||
<a itemprop="item" href="https://lightnovelworld.net/novel/immortality-simulator/"><span itemprop="name">Immortality Simulator</span></a>
|
||||
<meta itemprop="position" content="2">
|
||||
</span>
|
||||
›
|
||||
<span itemprop="itemListElement" itemscope="" itemtype="http://schema.org/ListItem">
|
||||
<a itemprop="item" href="https://lightnovelworld.net/my-longevity-simulation-chapter-1/"><span itemprop="name">My Longevity Simulation Chapter 1</span></a>
|
||||
<meta itemprop="position" content="3">
|
||||
</span>
|
||||
</div>
|
||||
```
|
||||
|
||||
Bonus: `<span itemprop="name">Immortality Simulator</span>` gives the clean
|
||||
series **title** as well, which the userscript currently derives by stripping
|
||||
`Chapter <n>` off `h1.entry-title`.
|
||||
|
||||
Caveat: `itemprop="item"` is **not** unique to the breadcrumb — the site header
|
||||
nav also uses `itemprop="url"` on anchors, and unrelated `itemprop="item"` usage
|
||||
could appear. Scope any selector to the enclosing
|
||||
`[itemtype="http://schema.org/BreadcrumbList"]`.
|
||||
|
||||
### Priority 3 — corroborating `/novel/` anchors ⚠️ AMBIGUOUS
|
||||
|
||||
The chapter page contains exactly **8** distinct `/novel/…` hrefs. Two are
|
||||
navigational (`/novel/` index), one is the series, and **five are a
|
||||
"recommended" strip of unrelated novels**:
|
||||
|
||||
```
|
||||
href="/novel/"
|
||||
href="https://lightnovelworld.net/novel/"
|
||||
href="https://lightnovelworld.net/novel/dreadful-radio-game/"
|
||||
href="https://lightnovelworld.net/novel/evil-god-average/"
|
||||
href="https://lightnovelworld.net/novel/immortality-simulator/"
|
||||
href="https://lightnovelworld.net/novel/omniscient-first-persons-viewpoint/"
|
||||
href="https://lightnovelworld.net/novel/reverend-insanity/"
|
||||
href="https://lightnovelworld.net/novel/surviving-in-a-romance-fantasy-novel/"
|
||||
```
|
||||
|
||||
So **"grab the first `/novel/<slug>/` link on a chapter page" is unsafe** — the
|
||||
recommendation strip contaminates it. Use Priority 1 or 2.
|
||||
|
||||
### ❌ NOT usable — JSON-LD `BreadcrumbList` (§3.1)
|
||||
|
||||
There is exactly one `application/ld+json` block on the page (Rank Math SEO).
|
||||
Verbatim:
|
||||
|
||||
```html
|
||||
<script type="application/ld+json" class="rank-math-schema">{"@context":"https://schema.org","@graph":[{"@type":"BreadcrumbList","@id":"https://lightnovelworld.net/my-longevity-simulation-chapter-1/#breadcrumb","itemListElement":[{"@type":"ListItem","position":"1","item":{"@id":"https://lightnovelworld.net","name":"Home"}},{"@type":"ListItem","position":"2","item":{"@id":"https://lightnovelworld.net/my-longevity-simulation-chapter-1/","name":"My Longevity Simulation Chapter 1"}}]}]}</script>
|
||||
```
|
||||
|
||||
Only two items — Home and the chapter. **The series does not appear.** The
|
||||
JSON-LD breadcrumb and the visible microdata breadcrumb disagree; the microdata
|
||||
one is the richer of the two. Do not use the JSON-LD.
|
||||
|
||||
### ❌ NOT usable — `<link rel=canonical>` (§3.2)
|
||||
|
||||
```html
|
||||
<link rel="canonical" href="https://lightnovelworld.net/my-longevity-simulation-chapter-1/" />
|
||||
```
|
||||
|
||||
Self-referential.
|
||||
|
||||
### ❌ NOT usable — `og:` meta tags (§3.3)
|
||||
|
||||
```html
|
||||
<meta property="og:type" content="article" />
|
||||
<meta property="og:title" content="Read My Longevity Simulation Chapter 1 Online Free - LightNovelworld" />
|
||||
<meta property="og:url" content="https://lightnovelworld.net/my-longevity-simulation-chapter-1/" />
|
||||
<meta property="og:site_name" content="Light Novel World" />
|
||||
```
|
||||
|
||||
`og:url` is the chapter. `og:title` carries the **chapter-slug-derived** title
|
||||
("My Longevity Simulation"), *not* the series title ("Immortality Simulator") —
|
||||
so og tags actively reinforce the wrong name.
|
||||
|
||||
### Also present — `rel=next` / `rel=prev` chapter navigation
|
||||
|
||||
```html
|
||||
<a aria-label="next" href="https://lightnovelworld.net/my-longevity-simulation-chapter-2/" rel="next">Next <i class="fa fa-angle-right" aria-hidden="true"></i></a>
|
||||
```
|
||||
|
||||
Chapter 1 has no `rel=prev`. Useful for walking chapters, useless for finding
|
||||
the series.
|
||||
|
||||
---
|
||||
|
||||
## 4. The reverse direction, and how common divergence is
|
||||
|
||||
### 4.1 Sample method
|
||||
|
||||
Two independent samples, deduplicated:
|
||||
|
||||
- 14 novels taken from the latest-updates listing on `https://lightnovelworld.net/`
|
||||
(distinct chapter slugs), each chapter page fetched and its
|
||||
`aria-label='All Chapter'` href read for the series slug.
|
||||
- 30 series URLs taken from `https://lightnovelworld.net/novel/`, each series
|
||||
page fetched and its chapter anchors' slug prefix extracted. Two of the 30
|
||||
(`/novel/feed/` and `/novel/list-mode/`) are not novels and are excluded,
|
||||
leaving 28.
|
||||
|
||||
One novel (`investing-in-my-crippled-wife-every-return-makes-me-stronger`)
|
||||
appears in both samples. **Distinct novels sampled: 41.**
|
||||
|
||||
### 4.2 Divergence results
|
||||
|
||||
| Series slug (`/novel/<…>/`) | Chapter slug (`/<…>-chapter-N/`) | Result |
|
||||
|---|---|---|
|
||||
| `immortality-simulator` | `my-longevity-simulation` | **DIVERGE** |
|
||||
| `the-sword-illuminates-the-great-wilderness` | `radiant-blade-of-the-wilderness` | **DIVERGE** |
|
||||
| `a-villains-will-to-survive` | `the-villain-wants-to-live` | **DIVERGE** |
|
||||
| `all-jobs-and-classes-i-just-wanted-one-skill-not-them-all` | `all-jobs-and-classes-i-just-wanted-one-skill-not` (ch. 1–99) **and** `all-jobs-and-classes-i-just-wanted-one-skill-not-them-all` (ch. 100–423) | **SPLIT** (latest matches) |
|
||||
| `86-eighty-six` | same | match |
|
||||
| `a-dungeon-beneath-my-house-let-me-gain-800-exp` | same | match |
|
||||
| `a-journey-of-black-and-red` | same | match |
|
||||
| `a-knight-who-eternally-regresses` | same | match |
|
||||
| `a-regressors-tale-of-cultivation` | same | match |
|
||||
| `a-will-eternal` | same | match |
|
||||
| `absolute-resonance` | same | match |
|
||||
| `absolute-sword-sense` | same | match |
|
||||
| `advent-of-the-three-calamities` | same | match |
|
||||
| `after-changing-to-the-ruthless-way-the-brothers-cried-and-begged-for-forgiveness` | same | match |
|
||||
| `against-the-gods` | same | match |
|
||||
| `apocalypse-i-built-the-infinite-train` | same | match |
|
||||
| `arcane-exfil` | same | match |
|
||||
| `as-a-mafia-boss-i-refuse-to-be-an-extra` | same | match |
|
||||
| `ascendance-of-a-bookworm` | same | match |
|
||||
| `avatar-conquering-the-elements` | same | match |
|
||||
| `battle-world-ascending-without-limits` | same | match |
|
||||
| `became-the-patron-of-villains` | same | match |
|
||||
| `ending-maker` | same | match |
|
||||
| `got-dropped-into-a-ghost-story-still-gotta-work` | same | match |
|
||||
| `greed-all-for-what` | same | match |
|
||||
| `investing-in-my-crippled-wife-every-return-makes-me-stronger` | same | match |
|
||||
| `lord-of-mysteries-2-circle-of-inevitability` | same | match |
|
||||
| `lord-of-the-mysteries` | same | match |
|
||||
| `my-vampire-system` | same | match |
|
||||
| `cleaver-of-sin` | same | match |
|
||||
| `greatest-legacy-of-the-magus-universe` | same | match |
|
||||
| `magus-infinite` | same | match |
|
||||
| `path-of-the-extra` | same | match |
|
||||
| `regnum-aetern-dual-rebirth` | same | match |
|
||||
| `shadow-slave` | same | match |
|
||||
| `slime-evolution` | same | match |
|
||||
| `sss-awakening-i-can-class-change-at-will` | same | match |
|
||||
| `the-dao-of-reincarnation-ascending-through-a-hundred-lives` | same | match |
|
||||
| `the-gamers-pov` | same | match |
|
||||
| `the-insane-regressor-throne-of-pride` | same | match |
|
||||
| `the-villains-pov` | same | match |
|
||||
|
||||
**Counts (41 distinct novels):**
|
||||
|
||||
- **37 match** (90.2%)
|
||||
- **3 diverge at the latest chapter** (7.3%) — these are the ones that break the poller *today*
|
||||
- **1 splits mid-series** but its newest chapters happen to match the series slug (2.4%)
|
||||
|
||||
`[INFERENCE]` The true site-wide divergence rate is probably in the same
|
||||
ballpark, but 41 novels out of a catalogue of thousands is a small sample and
|
||||
the listing page is alphabetically front-loaded. Treat ~7% as an order of
|
||||
magnitude, not a precise figure.
|
||||
|
||||
### 4.3 Is there a derivable rule? **No.**
|
||||
|
||||
The divergence is **not directional**, so you cannot compute one slug from the
|
||||
other:
|
||||
|
||||
- `immortality-simulator` — the *series* carries the polished English title
|
||||
(`<h1 class="entry-title" itemprop="name">` is absent from the chapter page's
|
||||
own naming; `og:title` says "My Longevity Simulation"), while the *chapters*
|
||||
carry the literal translation.
|
||||
- `the-sword-illuminates-the-great-wilderness` — the reverse: the *series* has
|
||||
the literal title
|
||||
(`<h1 class="entry-title" itemprop="name">The Sword Illuminates the Great Wilderness</h1>`,
|
||||
`<title>Read The Sword Illuminates the Great Wilderness English Online Free - LightNovelworld</title>`)
|
||||
while the *chapters* carry the polished one (`radiant-blade-of-the-wilderness`).
|
||||
- `a-villains-will-to-survive` —
|
||||
`<h1 class="entry-title" itemprop="name">A Villain’s Will to Survive</h1>`,
|
||||
chapters at `the-villain-wants-to-live-chapter-N/`.
|
||||
|
||||
`[INFERENCE]` The consistent explanation is that a novel is retitled after
|
||||
publication; WordPress updates the series post's slug but leaves the already-published
|
||||
chapter posts' slugs alone. The direction of the retitle varies per novel, which
|
||||
is why no rule exists. This is consistent with the split case in §4.4, but the
|
||||
site exposes no field that states it.
|
||||
|
||||
**There is no "Alternative names" / "Associated names" field.** Scanning
|
||||
`https://lightnovelworld.net/novel/a-villains-will-to-survive/` and
|
||||
`https://lightnovelworld.net/novel/immortality-simulator/` for such a label
|
||||
returns nothing; the series info panel exposes only **Author**, **Released**,
|
||||
**Type**, **Status**. (The only `names?` match in the document is the `>Name*<`
|
||||
label on the comment form.) So the old title is **not** recoverable from the
|
||||
series page — the mapping only exists in the chapter anchors themselves.
|
||||
|
||||
### 4.4 The split case — a slug can change *mid-series*
|
||||
|
||||
`https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/`
|
||||
carries **two** chapter slugs in one page:
|
||||
|
||||
| Chapter slug prefix | Anchors | Chapter range |
|
||||
|---|---|---|
|
||||
| `all-jobs-and-classes-i-just-wanted-one-skill-not` | 100 | 1 – 99 |
|
||||
| `all-jobs-and-classes-i-just-wanted-one-skill-not-them-all` | 325 | 100 – 423 |
|
||||
|
||||
Both resolve, and **both point back at the same series**:
|
||||
|
||||
```
|
||||
200 https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-chapter-1/
|
||||
<a href='https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/' aria-label='All Chapter'>
|
||||
|
||||
200 https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-404/
|
||||
<a href='https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/' aria-label='All Chapter'>
|
||||
```
|
||||
|
||||
**Consequence:** a single stored chapter slug is not a sufficient key even for
|
||||
one novel. A fix that captures "the chapter slug at bookmark time" and pins the
|
||||
poller regex to it would still miss chapters published under a later slug. This
|
||||
is the strongest argument for the unscoped regex over a stored-slug regex.
|
||||
|
||||
---
|
||||
|
||||
## 5. Series page → chapter list
|
||||
|
||||
From `https://lightnovelworld.net/novel/immortality-simulator/` (200, 655 037 bytes):
|
||||
|
||||
- **1234** anchors matching `href="https://lightnovelworld.net/<slug>-chapter-<n>/"`.
|
||||
- **1232 unique** chapter URLs, covering chapter numbers **1 through 1232 with no gaps**.
|
||||
- The 2 extras are the First/Last shortcut widget, which duplicates two chapters:
|
||||
|
||||
```html
|
||||
<h2>Read Immortality Simulator</h2></div>
|
||||
<div class="lastend">
|
||||
<div class="inepcx">
|
||||
<a href="https://lightnovelworld.net/my-longevity-simulation-chapter-1/">
|
||||
<span>First Chapter</span>
|
||||
<span class="epcur epcurfirst">Vol. 1 Ch. 1</span>
|
||||
</a>
|
||||
</div>
|
||||
<div class="inepcx">
|
||||
<a href="https://lightnovelworld.net/my-longevity-simulation-chapter-1232/">
|
||||
```
|
||||
|
||||
**The entire chapter list is in the initial HTML.** No JS hydration, no
|
||||
pagination, no separate endpoint. Confirmed by probing the shapes the issue
|
||||
speculated about:
|
||||
|
||||
| URL | Status |
|
||||
|---|---|
|
||||
| `https://lightnovelworld.net/novel/immortality-simulator/chapters/` | **404** |
|
||||
| `https://lightnovelworld.net/novel/immortality-simulator/chapters/page-2` | **404** |
|
||||
|
||||
The only `page-numbers` / pagination markup in the document belongs to
|
||||
`class="wpdiscuz-comment-pagination"` (the comment widget), not the chapter list.
|
||||
|
||||
This holds for large novels too: `greed-all-for-what` served 2666 chapter
|
||||
anchors and `my-vampire-system` 2547, all inline in one response.
|
||||
|
||||
**All chapter hrefs are absolute** (`https://lightnovelworld.net/…`). A grep for
|
||||
relative `href="/<slug>-chapter-N/"` on the series page returned **0** matches,
|
||||
so a pattern anchored on the literal host is safe.
|
||||
|
||||
---
|
||||
|
||||
## 6. Is the unscoped chapter regex SAFE on a series page?
|
||||
|
||||
### Verdict: **SAFE**
|
||||
|
||||
Proposed pattern: `lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/`
|
||||
|
||||
**Proof.** For each of the 30 series pages fetched from
|
||||
`https://lightnovelworld.net/novel/`, every occurrence of `<slug>-chapter-<n>/`
|
||||
anywhere in the document (href or not) was reduced to its slug prefix and
|
||||
deduplicated:
|
||||
|
||||
- **27 of 28 real series pages had exactly 1 distinct chapter-slug prefix**, equal
|
||||
to that novel's own chapter slug.
|
||||
- **1 page had 2 prefixes** — `all-jobs-and-classes-…`, and §4.4 proves *both
|
||||
belong to that same novel*. Not foreign.
|
||||
- **0 pages carried any other novel's chapter URL.**
|
||||
|
||||
The recommendation and sidebar widgets on a series page link to **series** URLs
|
||||
only, never chapter URLs. On
|
||||
`https://lightnovelworld.net/novel/immortality-simulator/` the recommendation
|
||||
strip renders as:
|
||||
|
||||
```html
|
||||
<a href="https://lightnovelworld.net/novel/reverend-insanity/" itemprop="url" title="Reverend Insanity" class="tip" rel="292">
|
||||
```
|
||||
|
||||
— `/novel/<slug>/`, which the pattern cannot match.
|
||||
|
||||
The one widget that *would* carry foreign chapter links, "Latest Reading", is an
|
||||
**empty client-side template**. Its container is `display:none` and its `<ul>` is
|
||||
empty in the served HTML; the row markup lives in an inert
|
||||
`<span id="series-history-tpl" style='display:none'>` with mustache placeholders
|
||||
and a `href="#/{{number}}"` — no real chapter URL:
|
||||
|
||||
```html
|
||||
<div class="bixbox bxcl" id="series-history" style="display:none;">
|
||||
<div class="releases"><h2>Latest Reading</h2></div>
|
||||
<div class="series-history-pool">
|
||||
<ul class="clstyle" id="series-history-ul"></ul>
|
||||
</div>
|
||||
</div>
|
||||
<span id="series-history-tpl" style='display:none'>
|
||||
<li data-id="{{id}}" data-num="{{chapter}}" data-vol="{{volume}}" data-title="{{title}}">
|
||||
<div class="chbox"><div class="eph-num">
|
||||
<a onclick="…return series_history.redirect({{id}});" href="#/{{number}}">
|
||||
```
|
||||
|
||||
It is populated from the visitor's own local history, so a **server-side fetch
|
||||
never sees content there** — the poller is immune. A browser-rendered fetch with
|
||||
a fresh profile is likewise immune (no history to render).
|
||||
|
||||
### Caveats to record with the verdict
|
||||
|
||||
1. **`[INFERENCE]` This is an empirical guarantee, not a structural one.** Unlike
|
||||
the asura/novelfull cases where scoping to the slug makes foreign contamination
|
||||
*impossible*, here we only know that lightnovelworld's series template does not
|
||||
currently emit foreign chapter anchors. If the theme ever adds a "latest site
|
||||
updates" strip rendered server-side, the unscoped pattern breaks silently and
|
||||
in the worst direction (a foreign chapter number *higher* than the real one
|
||||
wins the maximum and the bookmark shows a phantom update).
|
||||
|
||||
**Correction, 2026-08-11.** This caveat understated the risk. A series page
|
||||
server-renders a wpdiscuz comment thread below the chapter list, and comment
|
||||
bodies are HTML that can carry an `<a href>`. Verified on
|
||||
`/novel/the-sword-illuminates-the-great-wilderness/`: `#wpdcom` present,
|
||||
comment `#wpd-comm-358_0` by "hasbi asy" rendered as `<p>where's everyone</p>`
|
||||
inside `.wpd-comment-text`, corroborated by that page's comment RSS feed;
|
||||
`firstLoadWithAjax: 0` confirms the thread is in the initial HTML. The sampled
|
||||
pages carried 0 or 1 comments and no links, which is why §6 read clean. So the
|
||||
unscoped pattern does not merely depend on the *theme* staying unchanged — it
|
||||
reads a region any visitor can write to, and the scanner takes the maximum with
|
||||
no upper bound. The scan must stop before the comment thread.
|
||||
2. ~~Consider scoping the match to the chapter-list container rather than the whole
|
||||
document, which would restore the structural guarantee at low cost. The list
|
||||
items sit under `ul.clstyle` (`class="bixbox bxcl"`).~~
|
||||
|
||||
**Refuted, 2026-08-11**, measured on 4 series pages
|
||||
(`the-sword-illuminates-the-great-wilderness`, `all-jobs-…-them-all`,
|
||||
`a-will-eternal`, `immortality-simulator`). The single `ul.clstyle` per page is
|
||||
the hidden, empty "Latest Reading" template (`#series-history-ul`, see §5), not
|
||||
the chapter list — scoping to it would match nothing. The real chapter list is a
|
||||
classless `<ul>` inside `div.eplister.eplisterfull`, in
|
||||
`div.bixbox.bxcl.epcheck`.
|
||||
|
||||
The workable boundary is `wpd-threads`: it occurs exactly once per page, and on
|
||||
all 4 pages every chapter anchor precedes it and no chapter-shaped href follows
|
||||
it. `wpdcom` and `wpdiscuz` are unusable as markers — they occur 111 to 143
|
||||
times per page, including in `<head>` roughly 15 KB *before* the chapter list,
|
||||
so cutting there would discard the list itself. `id='comments'` also occurs once
|
||||
but is single-quoted; the canonical `id="comments"` never appears.
|
||||
3. `[0-9.]+` is fine: no decimal chapter numbers appeared anywhere in the sampled
|
||||
pages, but the existing `strings.Trim(m[1], ".")` guard already handles them.
|
||||
4. The current test fixture `lnwSeriesFixture` in
|
||||
`backend/internal/latest/sites_test.go` includes
|
||||
`<a href="https://lightnovelworld.net/overgeared-chapter-9999/">` and asserts it
|
||||
is ignored. **That anchor is not representative of a real series page** — no
|
||||
sampled page contained a foreign chapter anchor. Relaxing the regex will make
|
||||
that assertion fail, and the correct response is to fix the fixture, not to
|
||||
keep the scoping.
|
||||
|
||||
---
|
||||
|
||||
## 7. Redirects and reverse-lookup endpoints
|
||||
|
||||
| URL | Status | Redirects | Final |
|
||||
|---|---|---|---|
|
||||
| `https://lightnovelworld.net/novel/my-longevity-simulation/` | **404** | 0 | — |
|
||||
| `https://lightnovelworld.net/novel/radiant-blade-of-the-wilderness/` | **404** | 0 | — |
|
||||
| `https://lightnovelworld.net/immortality-simulator-chapter-1/` | **404** | 0 | — |
|
||||
| `https://lightnovelworld.net/my-longevity-simulation/` | **200** | **1** | `https://lightnovelworld.net/my-longevity-simulation-chapter-1/` |
|
||||
|
||||
Two findings:
|
||||
|
||||
1. **`/novel/<chapter-slug>/` genuinely 404s.** Confirmed for both
|
||||
`my-longevity-simulation` and `radiant-blade-of-the-wilderness`. There is no
|
||||
alias, no 301, no fallback. Any stored series URL built by that construction is
|
||||
permanently dead.
|
||||
2. **The chapter slug alone (no `-chapter-N`) 302s to chapter 1** of that novel.
|
||||
This is the closest thing to a reverse-lookup endpoint, but it lands on a
|
||||
*chapter page*, so recovering the series URL still requires parsing that page's
|
||||
`aria-label='All Chapter'` anchor or breadcrumb. **No endpoint maps chapter slug
|
||||
→ series slug directly.**
|
||||
3. The inverse construction also fails: `/<series-slug>-chapter-1/` 404s for a
|
||||
divergent novel, so you cannot probe your way from a series slug to a chapter
|
||||
slug either. The chapter slug must be read off the series page's anchors.
|
||||
|
||||
Also present but not a lookup path: the site is WordPress and exposes
|
||||
`https://lightnovelworld.net/wp-json/wp/v2/posts/177695` (advertised via
|
||||
`<link rel="alternate" title="JSON" type="application/json" …>` on the chapter
|
||||
page). Whether the REST API exposes a chapter→series relation was **not tested**
|
||||
— out of scope for this note, and it would still require fetching the chapter
|
||||
page to learn the post ID.
|
||||
|
||||
---
|
||||
|
||||
## 8. Implications for issue #77
|
||||
|
||||
> The issue text itself could not be read: `gh` is not installed in this
|
||||
> environment, so `issue://77` failed to resolve. The proposed regex quoted below
|
||||
> is the one supplied in the task brief.
|
||||
|
||||
### 8.1 The bug
|
||||
|
||||
`backend/internal/latest/sites.go:128-137` derives the chapter-anchor pattern
|
||||
from the **series slug**:
|
||||
|
||||
```go
|
||||
m := lnwSlugRe.FindStringSubmatch(u.Path) // /novel/<series-slug>/
|
||||
re = regexp.MustCompile(`lightnovelworld\.net/` + regexp.QuoteMeta(m[1]) + `-chapter-([0-9.]+)/`)
|
||||
```
|
||||
|
||||
For `immortality-simulator` this compiles to a pattern matching
|
||||
`…net/immortality-simulator-chapter-N/`, which appears **zero times** on the
|
||||
series page — the anchors all say `my-longevity-simulation`. `latestChapterFrom`
|
||||
returns `ok == false`, indistinguishable from a Cloudflare challenge or a site
|
||||
redesign, and the bookmark silently stops tracking updates. Same for
|
||||
`the-sword-illuminates-the-great-wilderness` and `a-villains-will-to-survive`.
|
||||
|
||||
### 8.2 The proposed fix is sound
|
||||
|
||||
Dropping the scoping to
|
||||
`lightnovelworld\.net/[a-z0-9-]+-chapter-([0-9.]+)/` is sound in shape and fixes
|
||||
all three currently-broken novels plus the split case in §4.4 — which no
|
||||
stored-slug approach can fix, since that novel legitimately has two chapter
|
||||
slugs. It must not be applied to the whole document, however: truncate the body
|
||||
at `wpd-threads` first, per the corrected caveat 2 in §6. The `ul.clstyle`
|
||||
suggestion originally recorded here is refuted.
|
||||
|
||||
Also worth removing: the `overgeared-chapter-9999` line in `lnwSeriesFixture`
|
||||
and the `"lightnovelworld takes the max and ignores another series"` test case,
|
||||
which encode a contamination scenario that §6 shows does not occur on this site.
|
||||
Replacing them with a fixture that mirrors §4.4's two-slugs-one-novel reality
|
||||
would be a better regression test.
|
||||
|
||||
### 8.3 The userscript has the same bug, and it is worse
|
||||
|
||||
`userscript/novel-bookmark.user.js` builds the series URL by string-concatenating
|
||||
the chapter slug (lines 154 and 169):
|
||||
|
||||
```js
|
||||
seriesUrl: "https://lightnovelworld.net/novel/" + m[1] + "/",
|
||||
```
|
||||
|
||||
For a divergent novel this **writes a permanently-404 series URL into the
|
||||
database at bookmark time**. Fixing only the backend regex leaves those rows
|
||||
broken, and `fetchableSeriesURL` will happily accept the dead URL because it only
|
||||
checks the hostname (`sites.go:321-324`).
|
||||
|
||||
The userscript should read the series URL off the page instead of constructing
|
||||
it. On a chapter page, prefer in this order (§3):
|
||||
|
||||
```js
|
||||
document.querySelector("a[aria-label='All Chapter']")?.href
|
||||
?? document.querySelector("[itemtype$='BreadcrumbList'] a[itemprop=item][href*='/novel/']")?.href
|
||||
```
|
||||
|
||||
The breadcrumb's `<span itemprop="name">` also supplies the correct series title
|
||||
directly, replacing the current `h1.entry-title` + strip-`Chapter <n>` heuristic
|
||||
— which today yields "My Longevity Simulation" where the series is actually
|
||||
titled "Immortality Simulator".
|
||||
|
||||
A backfill for already-stored rows is likely needed: any `lightnovelworld` row
|
||||
whose `series_url` 404s can be repaired by fetching
|
||||
`https://lightnovelworld.net/<series_id>-chapter-1/` (or `/<series_id>/`, which
|
||||
302s there per §7) and reading its `All Chapter` anchor.
|
||||
|
||||
### 8.4 Note on `series_id`
|
||||
|
||||
Rows are keyed `lightnovelworld:<series_id>` where `series_id` is currently the
|
||||
**chapter slug** (from the userscript's `m[1]`). If the fix changes `series_id` to
|
||||
the real series slug, existing divergent bookmarks change key and need migrating.
|
||||
Whether that is worth it depends on whether `series_id` is used to rebuild URLs
|
||||
anywhere — `sites.go` uses the stored `series_url`, not `series_id`, for
|
||||
lightnovelworld, so storing the correct `series_url` may be sufficient without a
|
||||
re-key.
|
||||
|
||||
---
|
||||
|
||||
## Reproduction
|
||||
|
||||
```sh
|
||||
UA='Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0 Safari/537.36'
|
||||
|
||||
# §2 status codes
|
||||
curl -sS -o /dev/null -w '%{http_code} %{url_effective}\n' -L -A "$UA" \
|
||||
https://lightnovelworld.net/my-longevity-simulation-chapter-1/ \
|
||||
https://lightnovelworld.net/novel/immortality-simulator/ \
|
||||
https://lightnovelworld.net/novel/my-longevity-simulation/
|
||||
|
||||
# §3 the two working pointers
|
||||
curl -sS -A "$UA" https://lightnovelworld.net/my-longevity-simulation-chapter-1/ \
|
||||
| grep -E "aria-label='All Chapter'|itemprop=\"item\""
|
||||
|
||||
# §6 SAFE proof — distinct chapter-slug prefixes per series page (expect 1)
|
||||
curl -sS -A "$UA" https://lightnovelworld.net/novel/immortality-simulator/ \
|
||||
| grep -oE '[a-z0-9-]+-chapter-[0-9.]+/' | sed -E 's/-chapter-[0-9.]+\///' | sort | uniq -c
|
||||
```
|
||||
@@ -0,0 +1,235 @@
|
||||
{
|
||||
"0": "HTMX Library Internals",
|
||||
"1": "Cover Fetch Test Helpers",
|
||||
"2": "Manga Userscript Adapters",
|
||||
"3": "Novel Userscript Adapters",
|
||||
"4": "Series Acquisition Tests",
|
||||
"5": "Bookmarks API Tests",
|
||||
"6": "Storage Choice ADR",
|
||||
"7": "Cover & Acquire Internals",
|
||||
"8": "System Architecture Concepts",
|
||||
"9": "Session Middleware",
|
||||
"10": "Go Test Helpers",
|
||||
"11": "Store Tests",
|
||||
"12": "Bookmarks API Handler",
|
||||
"13": "Web UI Handlers",
|
||||
"14": "Go Error Handling",
|
||||
"15": "CDP Browser Client",
|
||||
"17": "Go Code Style Guide",
|
||||
"18": "Agent Skills",
|
||||
"20": "I/O Performance Patterns",
|
||||
"21": "CPU Optimization",
|
||||
"22": "Caching Patterns",
|
||||
"23": "Browser Entrypoint",
|
||||
"24": "Memory Allocation & GC",
|
||||
"25": "Cover Fetcher Tests",
|
||||
"26": "readSeries",
|
||||
"27": "Find Skills Guide",
|
||||
"28": "Allocation Patterns",
|
||||
"29": "Observability & Alerting",
|
||||
"31": "Memory Layout",
|
||||
"32": "Repo Hard Constraints",
|
||||
"33": "Go Testing Guide",
|
||||
"34": "Session Store",
|
||||
"35": "Web UI Filter Logic",
|
||||
"36": "Userscript Test Harness",
|
||||
"37": "Product & Security Context",
|
||||
"38": "novel-logic.test.js",
|
||||
"39": "UI Critique 2026-07-26A",
|
||||
"40": "UI Critique 2026-07-26B",
|
||||
"43": "Issue Tracker & Triage",
|
||||
"44": "Ticket Workflow",
|
||||
"45": "Go Perf Alert Rules",
|
||||
"46": "Userscript Display Logic",
|
||||
"47": "Go Perf Skill Docs",
|
||||
"48": "Login Page Art",
|
||||
"49": "BookmarkManager Logo",
|
||||
"50": "Skills CLI",
|
||||
"51": "Skills Leaderboard",
|
||||
"52": "Complex Condition Extraction",
|
||||
"53": "Sentinel Errors",
|
||||
"54": "errors.As Patterns",
|
||||
"55": "errors.Is Patterns",
|
||||
"56": "errors.Join Patterns",
|
||||
"57": "Error Wrapping",
|
||||
"58": "Single Error Handling",
|
||||
"59": "SIMD Optimizations",
|
||||
"60": "GOGC Tuning",
|
||||
"61": "GOMEMLIMIT",
|
||||
"62": "Bottleneck Decision Tree",
|
||||
"63": "pprof Profiling",
|
||||
"64": "Test Timeout Helper",
|
||||
"65": "httptest Patterns",
|
||||
"66": "testify Suite Pattern",
|
||||
"67": "go:embed Fixtures",
|
||||
"68": "clockwork Time Mocking",
|
||||
"69": "testify Mocking",
|
||||
"70": "t.ArtifactDir Helper",
|
||||
"71": "Subtests Pitfall",
|
||||
"72": "golang-benchmark Skill",
|
||||
"73": "golang-concurrency Skill",
|
||||
"74": "golang-ci Skill",
|
||||
"75": "golang-database Skill",
|
||||
"76": "golang-lint Skill",
|
||||
"77": "testify Skill",
|
||||
"78": "Build Tag Integration Tests",
|
||||
"79": "Test Naming Convention",
|
||||
"80": "UI Critique A Finding",
|
||||
"81": "UI Critique B Finding",
|
||||
"82": "P0 Overflow Bug",
|
||||
"83": "P1 hx-indicator Gap",
|
||||
"84": "golang-benchmark Skill (ext)",
|
||||
"85": "golang-concurrency Skill (ext)",
|
||||
"86": "golang-ci Skill (ext)",
|
||||
"87": "golang-data-structures Skill (ext)",
|
||||
"88": "golang-database Skill (ext)",
|
||||
"89": "golang-design-patterns Skill (ext)",
|
||||
"90": "golang-documentation Skill (ext)",
|
||||
"91": "golang-gopls Skill (ext)",
|
||||
"92": "golang-lint Skill (ext)",
|
||||
"93": "golang-naming Skill (ext)",
|
||||
"94": "golang-observability Skill (ext)",
|
||||
"95": "golang-refactoring Skill (ext)",
|
||||
"96": "golang-safety Skill (ext)",
|
||||
"97": "golang-samber-oops Skill (ext)",
|
||||
"98": "golang-samber-slog Skill (ext)",
|
||||
"99": "golang-structs-interfaces Skill (ext)",
|
||||
"100": "golang-troubleshooting Skill (ext)",
|
||||
"101": "promql-cli Skill",
|
||||
"102": "Backend Module",
|
||||
"103": "bookmark-api Service",
|
||||
"104": "AGENTS.md",
|
||||
"105": "reviewer.md",
|
||||
"106": "Redeploy runbook",
|
||||
"107": "1. Backend",
|
||||
"108": "Deployment",
|
||||
"109": "Cinder — BookmarkManager design system",
|
||||
"110": "Implement tickets",
|
||||
"111": "SQLite → Postgres cutover runbook",
|
||||
"112": "Testing the userscript",
|
||||
"113": "ADR-0007: The backend hosts every Site's Cover bytes",
|
||||
"114": "Issue tracker: Gitea (`tea` CLI)",
|
||||
"115": "ADR-0006: The browser runs on the home machine, over the tailnet",
|
||||
"116": "ADR-0008: A Series identity is discovered from the Site's links, never derived from an address",
|
||||
"117": "Domain Docs",
|
||||
"118": "ticket-implementer.md",
|
||||
"119": "implementer.md",
|
||||
"120": "Series is a shared entity, and only the Poll may update it",
|
||||
"121": "Postgres replaces SQLite as the primary datastore",
|
||||
"122": "Identity comes from Discord OAuth; we store no passwords and send no email",
|
||||
"123": "The wire format stays flat and deliberately does not mirror the schema",
|
||||
"124": "ADR-0005: On-demand browser sidecar",
|
||||
"126": "Bookmark Manager",
|
||||
"127": "triage-labels.md",
|
||||
"128": "Cross-Ticket Contract",
|
||||
"129": "Implement Tickets Skill",
|
||||
"130": "Orchestrator Role",
|
||||
"131": "resolving-merge-conflicts Skill",
|
||||
"132": "tdd Skill",
|
||||
"133": "Ticket Wave Batching",
|
||||
"134": "Four-Object Browser Stub",
|
||||
"135": "Module Export Hook",
|
||||
"136": "logic.test.js Test Harness",
|
||||
"137": "manga-bookmark.user.js",
|
||||
"138": "stripBuildHash",
|
||||
"139": "Testing the Userscript Skill",
|
||||
"140": "cr-spec Agent",
|
||||
"141": "cr-standards Agent",
|
||||
"142": "Escalate Rather Than Guess",
|
||||
"143": "Status Contract",
|
||||
"144": "Ticket Implementer Agent",
|
||||
"145": "Worktree Isolation",
|
||||
"146": "Escalate Rather Than Guess (opencode)",
|
||||
"147": "Implementer Subagent (opencode)",
|
||||
"148": "Subagent-Driven Development",
|
||||
"149": "Code Quality Review",
|
||||
"150": "Reviewer Subagent (opencode)",
|
||||
"151": "Finding Severity Rubric",
|
||||
"152": "Spec Compliance Review",
|
||||
"161": "Why Use samber/oops",
|
||||
"162": "singleflight Cache Stampede Prevention",
|
||||
"163": "Struct Field Alignment",
|
||||
"164": "testing/synctest Deterministic Goroutine Testing",
|
||||
"165": "API Package (Bookmark JSON Handlers)",
|
||||
"166": "Cover Acquisition & Serving Pipeline",
|
||||
"167": "Backend AGENTS.md Guidance",
|
||||
"168": "HTTP Middleware (Auth/Gzip/CORS)",
|
||||
"169": "Latest Package (Site Parsers & Poller)",
|
||||
"170": "Latest-Chapter Poller",
|
||||
"171": "Main Composition Root",
|
||||
"172": "Migration-Owned Schema",
|
||||
"173": "Session Package (Cookie Signing & Rate Limit)",
|
||||
"174": "updated_at List-Order Rule",
|
||||
"175": "Userscript Package (Serving Handler)",
|
||||
"176": "AGENTS.md",
|
||||
"177": "Backend CLAUDE.md Guidance",
|
||||
"178": "Graphify Knowledge Graph (graphify-out/)",
|
||||
"179": "CLAUDE.md (Symlink to AGENTS.md)",
|
||||
"192": "ADR-0001 (Drop modernc.org/sqlite)",
|
||||
"193": "ADR-0003 (Split Shared Series Facts)",
|
||||
"194": "SQLite-to-Postgres Cutover Runbook",
|
||||
"195": "Import SQL Generation Rules",
|
||||
"196": "Throwaway Import Generator",
|
||||
"206": "Real scaling limit is the poller outbound fetch budget",
|
||||
"207": "PostgreSQL (jackc/pgx/v5)",
|
||||
"208": "SQLite (modernc.org/sqlite)",
|
||||
"209": "Postgres chosen for future supportability, not concurrency",
|
||||
"210": "Per-Reader bearer token for userscripts",
|
||||
"211": "Discord OAuth2 (authorization code grant)",
|
||||
"212": "ADR-0002: Discord OAuth, no passwords, no email",
|
||||
"213": "Discord snowflake is the sole identity (lock-in)",
|
||||
"214": "Bookmark (per-Reader state: Progress, Favourite, Lifecycle)",
|
||||
"215": "Deduplicate polling per Series (reader_count DESC queue)",
|
||||
"216": "ADR-0003: Series is shared, only the Poll updates it",
|
||||
"217": "Only the Poll writes Series fields (security boundary)",
|
||||
"218": "Series (shared entity keyed site+series_id)",
|
||||
"219": "ADR-0004: Wire format stays flat, does not mirror schema",
|
||||
"220": "Flat wire shape is a contract, not an implementation detail",
|
||||
"221": "Installed userscripts must keep working (14-day grace window)",
|
||||
"222": "CDP (Chrome DevTools Protocol) endpoint",
|
||||
"223": "headless-shell service (socat-fronted CDP)",
|
||||
"224": "Start Chrome on first CDP connection, reap after 300s idle",
|
||||
"225": "BROWSER_WS_URL configuration seam",
|
||||
"226": "ADR-0006: Browser runs on the home machine over the tailnet",
|
||||
"227": "Browser moved home: VPS memory pressure, no requests served",
|
||||
"228": "Tailnet (Tailscale network)",
|
||||
"229": "Content-addressed filesystem storage (SHA-256 of source URL)",
|
||||
"230": "Cover (Series image bytes)",
|
||||
"231": "Deny-class destination control for outbound fetch",
|
||||
"232": "ADR-0007: Backend hosts every Site's Cover bytes",
|
||||
"233": "kagane CORP same-origin cover restriction",
|
||||
"234": "Backend acquires, stores, serves every Cover (uniformity)",
|
||||
"235": "a[aria-label='All Chapter'] anchor pointer",
|
||||
"236": "Series identity is discovered from the Site's links",
|
||||
"237": "ADR-0008: Series identity discovered, never derived",
|
||||
"238": "Chapter slug vs series slug divergence (~7% measured)",
|
||||
"239": "Scan truncated at first wpd-threads marker",
|
||||
"240": "Surface ADR conflicts explicitly rather than silently overriding",
|
||||
"241": "Domain docs: single-context layout guidance",
|
||||
"242": "/domain-modeling skill (lazy CONTEXT.md creation)",
|
||||
"243": "CONTEXT.md glossary (ubiquitous language)",
|
||||
"244": "Gitea (tea CLI, gitea.violetcrown.my.id)",
|
||||
"245": "wayfinder map/ticket mechanism",
|
||||
"246": "Triage labels: canonical roles to tracker labels",
|
||||
"247": "Canonical triage role labels (needs-triage ... wontfix)",
|
||||
"248": "Cinder (BookmarkManager Web UI design system)",
|
||||
"249": "Heat is typographic: ember reserved for unread chapters",
|
||||
"250": "Design tokens (dark + light branches, no hardcoded hex)",
|
||||
"251": "Three type roles: display serif / mono small-caps / sans",
|
||||
"252": "a[aria-label='All Chapter'] priority pointer",
|
||||
"253": "Research: lightnovelworld chapter slug vs series slug",
|
||||
"254": "Gitea issue #77 (chapter vs series slug)",
|
||||
"255": "Slug divergence measurements (3/41 diverge, 1 split)",
|
||||
"256": "Unscoped chapter regex is SAFE, truncated at wpd-threads",
|
||||
"257": "BookmarkManager",
|
||||
"258": "Bromite (Primary Device)",
|
||||
"259": "Dark-First Design Constraint",
|
||||
"260": "Discord Guild Membership",
|
||||
"261": "Reader Isolation Invariant",
|
||||
"268": "Browser Unit Redeploy",
|
||||
"269": "pg_dump Hot Backup",
|
||||
"270": "Redeploy Runbook",
|
||||
"271": "Rollback Strategy",
|
||||
"273": "AGENTS.md",
|
||||
"279": "Userscript CLAUDE guidance"
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
.
|
||||
@@ -0,0 +1,558 @@
|
||||
# Graph Report - mangaBookmark (2026-08-12)
|
||||
|
||||
## Corpus Check
|
||||
- 110 files · ~267,791 words
|
||||
- Verdict: corpus is large enough that graph structure adds value.
|
||||
|
||||
## Summary
|
||||
- 1621 nodes · 3242 edges · 233 communities (63 shown, 170 thin omitted)
|
||||
- Extraction: 90% EXTRACTED · 10% INFERRED · 0% AMBIGUOUS · INFERRED: 312 edges (avg confidence: 0.77)
|
||||
- Token cost: 0 input · 0 output
|
||||
|
||||
## Graph Freshness
|
||||
- Built from commit: `8ae98816`
|
||||
- Run `git rev-parse HEAD` and compare to check if the graph is stale.
|
||||
- Run `graphify update .` after code changes (no API cost).
|
||||
|
||||
## Community Hubs (Navigation)
|
||||
- [[_COMMUNITY_HTMX Library Internals|HTMX Library Internals]]
|
||||
- [[_COMMUNITY_Cover Fetch Test Helpers|Cover Fetch Test Helpers]]
|
||||
- [[_COMMUNITY_Manga Userscript Adapters|Manga Userscript Adapters]]
|
||||
- [[_COMMUNITY_Novel Userscript Adapters|Novel Userscript Adapters]]
|
||||
- [[_COMMUNITY_Series Acquisition Tests|Series Acquisition Tests]]
|
||||
- [[_COMMUNITY_Bookmarks API Tests|Bookmarks API Tests]]
|
||||
- [[_COMMUNITY_Storage Choice ADR|Storage Choice ADR]]
|
||||
- [[_COMMUNITY_Cover & Acquire Internals|Cover & Acquire Internals]]
|
||||
- [[_COMMUNITY_System Architecture Concepts|System Architecture Concepts]]
|
||||
- [[_COMMUNITY_Session Middleware|Session Middleware]]
|
||||
- [[_COMMUNITY_Go Test Helpers|Go Test Helpers]]
|
||||
- [[_COMMUNITY_Store Tests|Store Tests]]
|
||||
- [[_COMMUNITY_Bookmarks API Handler|Bookmarks API Handler]]
|
||||
- [[_COMMUNITY_Web UI Handlers|Web UI Handlers]]
|
||||
- [[_COMMUNITY_Go Error Handling|Go Error Handling]]
|
||||
- [[_COMMUNITY_CDP Browser Client|CDP Browser Client]]
|
||||
- [[_COMMUNITY_Go Code Style Guide|Go Code Style Guide]]
|
||||
- [[_COMMUNITY_Agent Skills|Agent Skills]]
|
||||
- [[_COMMUNITY_IO Performance Patterns|I/O Performance Patterns]]
|
||||
- [[_COMMUNITY_CPU Optimization|CPU Optimization]]
|
||||
- [[_COMMUNITY_Caching Patterns|Caching Patterns]]
|
||||
- [[_COMMUNITY_Browser Entrypoint|Browser Entrypoint]]
|
||||
- [[_COMMUNITY_Memory Allocation & GC|Memory Allocation & GC]]
|
||||
- [[_COMMUNITY_Cover Fetcher Tests|Cover Fetcher Tests]]
|
||||
- [[_COMMUNITY_readSeries|readSeries]]
|
||||
- [[_COMMUNITY_Find Skills Guide|Find Skills Guide]]
|
||||
- [[_COMMUNITY_Allocation Patterns|Allocation Patterns]]
|
||||
- [[_COMMUNITY_Observability & Alerting|Observability & Alerting]]
|
||||
- [[_COMMUNITY_Memory Layout|Memory Layout]]
|
||||
- [[_COMMUNITY_Repo Hard Constraints|Repo Hard Constraints]]
|
||||
- [[_COMMUNITY_Go Testing Guide|Go Testing Guide]]
|
||||
- [[_COMMUNITY_Session Store|Session Store]]
|
||||
- [[_COMMUNITY_Web UI Filter Logic|Web UI Filter Logic]]
|
||||
- [[_COMMUNITY_Userscript Test Harness|Userscript Test Harness]]
|
||||
- [[_COMMUNITY_Product & Security Context|Product & Security Context]]
|
||||
- [[_COMMUNITY_novel-logic.test.js|novel-logic.test.js]]
|
||||
- [[_COMMUNITY_UI Critique 2026-07-26A|UI Critique 2026-07-26A]]
|
||||
- [[_COMMUNITY_UI Critique 2026-07-26B|UI Critique 2026-07-26B]]
|
||||
- [[_COMMUNITY_Issue Tracker & Triage|Issue Tracker & Triage]]
|
||||
- [[_COMMUNITY_Ticket Workflow|Ticket Workflow]]
|
||||
- [[_COMMUNITY_Go Perf Alert Rules|Go Perf Alert Rules]]
|
||||
- [[_COMMUNITY_Userscript Display Logic|Userscript Display Logic]]
|
||||
- [[_COMMUNITY_Go Perf Skill Docs|Go Perf Skill Docs]]
|
||||
- [[_COMMUNITY_Login Page Art|Login Page Art]]
|
||||
- [[_COMMUNITY_BookmarkManager Logo|BookmarkManager Logo]]
|
||||
- [[_COMMUNITY_Skills CLI|Skills CLI]]
|
||||
- [[_COMMUNITY_Skills Leaderboard|Skills Leaderboard]]
|
||||
- [[_COMMUNITY_Complex Condition Extraction|Complex Condition Extraction]]
|
||||
- [[_COMMUNITY_Sentinel Errors|Sentinel Errors]]
|
||||
- [[_COMMUNITY_errors.As Patterns|errors.As Patterns]]
|
||||
- [[_COMMUNITY_errors.Is Patterns|errors.Is Patterns]]
|
||||
- [[_COMMUNITY_errors.Join Patterns|errors.Join Patterns]]
|
||||
- [[_COMMUNITY_Error Wrapping|Error Wrapping]]
|
||||
- [[_COMMUNITY_Single Error Handling|Single Error Handling]]
|
||||
- [[_COMMUNITY_SIMD Optimizations|SIMD Optimizations]]
|
||||
- [[_COMMUNITY_GOGC Tuning|GOGC Tuning]]
|
||||
- [[_COMMUNITY_GOMEMLIMIT|GOMEMLIMIT]]
|
||||
- [[_COMMUNITY_Bottleneck Decision Tree|Bottleneck Decision Tree]]
|
||||
- [[_COMMUNITY_pprof Profiling|pprof Profiling]]
|
||||
- [[_COMMUNITY_Test Timeout Helper|Test Timeout Helper]]
|
||||
- [[_COMMUNITY_httptest Patterns|httptest Patterns]]
|
||||
- [[_COMMUNITY_testify Suite Pattern|testify Suite Pattern]]
|
||||
- [[_COMMUNITY_goembed Fixtures|go:embed Fixtures]]
|
||||
- [[_COMMUNITY_clockwork Time Mocking|clockwork Time Mocking]]
|
||||
- [[_COMMUNITY_testify Mocking|testify Mocking]]
|
||||
- [[_COMMUNITY_t.ArtifactDir Helper|t.ArtifactDir Helper]]
|
||||
- [[_COMMUNITY_Subtests Pitfall|Subtests Pitfall]]
|
||||
- [[_COMMUNITY_golang-benchmark Skill|golang-benchmark Skill]]
|
||||
- [[_COMMUNITY_golang-concurrency Skill|golang-concurrency Skill]]
|
||||
- [[_COMMUNITY_golang-ci Skill|golang-ci Skill]]
|
||||
- [[_COMMUNITY_golang-database Skill|golang-database Skill]]
|
||||
- [[_COMMUNITY_golang-lint Skill|golang-lint Skill]]
|
||||
- [[_COMMUNITY_testify Skill|testify Skill]]
|
||||
- [[_COMMUNITY_Build Tag Integration Tests|Build Tag Integration Tests]]
|
||||
- [[_COMMUNITY_Test Naming Convention|Test Naming Convention]]
|
||||
- [[_COMMUNITY_UI Critique A Finding|UI Critique A Finding]]
|
||||
- [[_COMMUNITY_UI Critique B Finding|UI Critique B Finding]]
|
||||
- [[_COMMUNITY_P0 Overflow Bug|P0 Overflow Bug]]
|
||||
- [[_COMMUNITY_P1 hx-indicator Gap|P1 hx-indicator Gap]]
|
||||
- [[_COMMUNITY_golang-benchmark Skill (ext)|golang-benchmark Skill (ext)]]
|
||||
- [[_COMMUNITY_golang-concurrency Skill (ext)|golang-concurrency Skill (ext)]]
|
||||
- [[_COMMUNITY_golang-ci Skill (ext)|golang-ci Skill (ext)]]
|
||||
- [[_COMMUNITY_golang-data-structures Skill (ext)|golang-data-structures Skill (ext)]]
|
||||
- [[_COMMUNITY_golang-database Skill (ext)|golang-database Skill (ext)]]
|
||||
- [[_COMMUNITY_golang-design-patterns Skill (ext)|golang-design-patterns Skill (ext)]]
|
||||
- [[_COMMUNITY_golang-documentation Skill (ext)|golang-documentation Skill (ext)]]
|
||||
- [[_COMMUNITY_golang-gopls Skill (ext)|golang-gopls Skill (ext)]]
|
||||
- [[_COMMUNITY_golang-lint Skill (ext)|golang-lint Skill (ext)]]
|
||||
- [[_COMMUNITY_golang-naming Skill (ext)|golang-naming Skill (ext)]]
|
||||
- [[_COMMUNITY_golang-observability Skill (ext)|golang-observability Skill (ext)]]
|
||||
- [[_COMMUNITY_golang-refactoring Skill (ext)|golang-refactoring Skill (ext)]]
|
||||
- [[_COMMUNITY_golang-safety Skill (ext)|golang-safety Skill (ext)]]
|
||||
- [[_COMMUNITY_golang-samber-oops Skill (ext)|golang-samber-oops Skill (ext)]]
|
||||
- [[_COMMUNITY_golang-samber-slog Skill (ext)|golang-samber-slog Skill (ext)]]
|
||||
- [[_COMMUNITY_golang-structs-interfaces Skill (ext)|golang-structs-interfaces Skill (ext)]]
|
||||
- [[_COMMUNITY_golang-troubleshooting Skill (ext)|golang-troubleshooting Skill (ext)]]
|
||||
- [[_COMMUNITY_promql-cli Skill|promql-cli Skill]]
|
||||
- [[_COMMUNITY_Backend Module|Backend Module]]
|
||||
- [[_COMMUNITY_bookmark-api Service|bookmark-api Service]]
|
||||
- [[_COMMUNITY_AGENTS|AGENTS.md]]
|
||||
- [[_COMMUNITY_reviewer|reviewer.md]]
|
||||
- [[_COMMUNITY_Redeploy runbook|Redeploy runbook]]
|
||||
- [[_COMMUNITY_1. Backend|1. Backend]]
|
||||
- [[_COMMUNITY_Deployment|Deployment]]
|
||||
- [[_COMMUNITY_Cinder — BookmarkManager design system|Cinder — BookmarkManager design system]]
|
||||
- [[_COMMUNITY_Implement tickets|Implement tickets]]
|
||||
- [[_COMMUNITY_SQLite → Postgres cutover runbook|SQLite → Postgres cutover runbook]]
|
||||
- [[_COMMUNITY_Testing the userscript|Testing the userscript]]
|
||||
- [[_COMMUNITY_ADR-0007 The backend hosts every Site's Cover bytes|ADR-0007: The backend hosts every Site's Cover bytes]]
|
||||
- [[_COMMUNITY_Issue tracker Gitea (`tea` CLI)|Issue tracker: Gitea (`tea` CLI)]]
|
||||
- [[_COMMUNITY_ADR-0006 The browser runs on the home machine, over the tailnet|ADR-0006: The browser runs on the home machine, over the tailnet]]
|
||||
- [[_COMMUNITY_ADR-0008 A Series identity is discovered from the Site's links, never derived from an address|ADR-0008: A Series identity is discovered from the Site's links, never derived from an address]]
|
||||
- [[_COMMUNITY_Domain Docs|Domain Docs]]
|
||||
- [[_COMMUNITY_ticket-implementer|ticket-implementer.md]]
|
||||
- [[_COMMUNITY_implementer|implementer.md]]
|
||||
- [[_COMMUNITY_Series is a shared entity, and only the Poll may update it|Series is a shared entity, and only the Poll may update it]]
|
||||
- [[_COMMUNITY_Postgres replaces SQLite as the primary datastore|Postgres replaces SQLite as the primary datastore]]
|
||||
- [[_COMMUNITY_Identity comes from Discord OAuth; we store no passwords and send no email|Identity comes from Discord OAuth; we store no passwords and send no email]]
|
||||
- [[_COMMUNITY_The wire format stays flat and deliberately does not mirror the schema|The wire format stays flat and deliberately does not mirror the schema]]
|
||||
- [[_COMMUNITY_ADR-0005 On-demand browser sidecar|ADR-0005: On-demand browser sidecar]]
|
||||
- [[_COMMUNITY_Bookmark Manager|Bookmark Manager]]
|
||||
- [[_COMMUNITY_triage-labels|triage-labels.md]]
|
||||
- [[_COMMUNITY_Cross-Ticket Contract|Cross-Ticket Contract]]
|
||||
- [[_COMMUNITY_Implement Tickets Skill|Implement Tickets Skill]]
|
||||
- [[_COMMUNITY_Orchestrator Role|Orchestrator Role]]
|
||||
- [[_COMMUNITY_resolving-merge-conflicts Skill|resolving-merge-conflicts Skill]]
|
||||
- [[_COMMUNITY_tdd Skill|tdd Skill]]
|
||||
- [[_COMMUNITY_Ticket Wave Batching|Ticket Wave Batching]]
|
||||
- [[_COMMUNITY_Four-Object Browser Stub|Four-Object Browser Stub]]
|
||||
- [[_COMMUNITY_Module Export Hook|Module Export Hook]]
|
||||
- [[_COMMUNITY_logic.test.js Test Harness|logic.test.js Test Harness]]
|
||||
- [[_COMMUNITY_manga-bookmark.user.js|manga-bookmark.user.js]]
|
||||
- [[_COMMUNITY_stripBuildHash|stripBuildHash]]
|
||||
- [[_COMMUNITY_Testing the Userscript Skill|Testing the Userscript Skill]]
|
||||
- [[_COMMUNITY_cr-spec Agent|cr-spec Agent]]
|
||||
- [[_COMMUNITY_cr-standards Agent|cr-standards Agent]]
|
||||
- [[_COMMUNITY_Escalate Rather Than Guess|Escalate Rather Than Guess]]
|
||||
- [[_COMMUNITY_Status Contract|Status Contract]]
|
||||
- [[_COMMUNITY_Ticket Implementer Agent|Ticket Implementer Agent]]
|
||||
- [[_COMMUNITY_Worktree Isolation|Worktree Isolation]]
|
||||
- [[_COMMUNITY_Escalate Rather Than Guess (opencode)|Escalate Rather Than Guess (opencode)]]
|
||||
- [[_COMMUNITY_Implementer Subagent (opencode)|Implementer Subagent (opencode)]]
|
||||
- [[_COMMUNITY_Subagent-Driven Development|Subagent-Driven Development]]
|
||||
- [[_COMMUNITY_Code Quality Review|Code Quality Review]]
|
||||
- [[_COMMUNITY_Reviewer Subagent (opencode)|Reviewer Subagent (opencode)]]
|
||||
- [[_COMMUNITY_Finding Severity Rubric|Finding Severity Rubric]]
|
||||
- [[_COMMUNITY_Spec Compliance Review|Spec Compliance Review]]
|
||||
- [[_COMMUNITY_Why Use samberoops|Why Use samber/oops]]
|
||||
- [[_COMMUNITY_singleflight Cache Stampede Prevention|singleflight Cache Stampede Prevention]]
|
||||
- [[_COMMUNITY_Struct Field Alignment|Struct Field Alignment]]
|
||||
- [[_COMMUNITY_testingsynctest Deterministic Goroutine Testing|testing/synctest Deterministic Goroutine Testing]]
|
||||
- [[_COMMUNITY_API Package (Bookmark JSON Handlers)|API Package (Bookmark JSON Handlers)]]
|
||||
- [[_COMMUNITY_Cover Acquisition & Serving Pipeline|Cover Acquisition & Serving Pipeline]]
|
||||
- [[_COMMUNITY_Backend AGENTS.md Guidance|Backend AGENTS.md Guidance]]
|
||||
- [[_COMMUNITY_HTTP Middleware (AuthGzipCORS)|HTTP Middleware (Auth/Gzip/CORS)]]
|
||||
- [[_COMMUNITY_Latest Package (Site Parsers & Poller)|Latest Package (Site Parsers & Poller)]]
|
||||
- [[_COMMUNITY_Latest-Chapter Poller|Latest-Chapter Poller]]
|
||||
- [[_COMMUNITY_Main Composition Root|Main Composition Root]]
|
||||
- [[_COMMUNITY_Migration-Owned Schema|Migration-Owned Schema]]
|
||||
- [[_COMMUNITY_Session Package (Cookie Signing & Rate Limit)|Session Package (Cookie Signing & Rate Limit)]]
|
||||
- [[_COMMUNITY_updated_at List-Order Rule|updated_at List-Order Rule]]
|
||||
- [[_COMMUNITY_Userscript Package (Serving Handler)|Userscript Package (Serving Handler)]]
|
||||
- [[_COMMUNITY_Backend CLAUDE.md Guidance|Backend CLAUDE.md Guidance]]
|
||||
- [[_COMMUNITY_Graphify Knowledge Graph (graphify-out)|Graphify Knowledge Graph (graphify-out/)]]
|
||||
- [[_COMMUNITY_CLAUDE.md (Symlink to AGENTS.md)|CLAUDE.md (Symlink to AGENTS.md)]]
|
||||
- [[_COMMUNITY_ADR-0001 (Drop modernc.orgsqlite)|ADR-0001 (Drop modernc.org/sqlite)]]
|
||||
- [[_COMMUNITY_ADR-0003 (Split Shared Series Facts)|ADR-0003 (Split Shared Series Facts)]]
|
||||
- [[_COMMUNITY_SQLite-to-Postgres Cutover Runbook|SQLite-to-Postgres Cutover Runbook]]
|
||||
- [[_COMMUNITY_Import SQL Generation Rules|Import SQL Generation Rules]]
|
||||
- [[_COMMUNITY_Throwaway Import Generator|Throwaway Import Generator]]
|
||||
- [[_COMMUNITY_Real scaling limit is the poller outbound fetch budget|Real scaling limit is the poller outbound fetch budget]]
|
||||
- [[_COMMUNITY_PostgreSQL (jackcpgxv5)|PostgreSQL (jackc/pgx/v5)]]
|
||||
- [[_COMMUNITY_SQLite (modernc.orgsqlite)|SQLite (modernc.org/sqlite)]]
|
||||
- [[_COMMUNITY_Postgres chosen for future supportability, not concurrency|Postgres chosen for future supportability, not concurrency]]
|
||||
- [[_COMMUNITY_Per-Reader bearer token for userscripts|Per-Reader bearer token for userscripts]]
|
||||
- [[_COMMUNITY_Discord OAuth2 (authorization code grant)|Discord OAuth2 (authorization code grant)]]
|
||||
- [[_COMMUNITY_ADR-0002 Discord OAuth, no passwords, no email|ADR-0002: Discord OAuth, no passwords, no email]]
|
||||
- [[_COMMUNITY_Discord snowflake is the sole identity (lock-in)|Discord snowflake is the sole identity (lock-in)]]
|
||||
- [[_COMMUNITY_Bookmark (per-Reader state Progress, Favourite, Lifecycle)|Bookmark (per-Reader state: Progress, Favourite, Lifecycle)]]
|
||||
- [[_COMMUNITY_Deduplicate polling per Series (reader_count DESC queue)|Deduplicate polling per Series (reader_count DESC queue)]]
|
||||
- [[_COMMUNITY_ADR-0003 Series is shared, only the Poll updates it|ADR-0003: Series is shared, only the Poll updates it]]
|
||||
- [[_COMMUNITY_Only the Poll writes Series fields (security boundary)|Only the Poll writes Series fields (security boundary)]]
|
||||
- [[_COMMUNITY_Series (shared entity keyed site+series_id)|Series (shared entity keyed site+series_id)]]
|
||||
- [[_COMMUNITY_ADR-0004 Wire format stays flat, does not mirror schema|ADR-0004: Wire format stays flat, does not mirror schema]]
|
||||
- [[_COMMUNITY_Flat wire shape is a contract, not an implementation detail|Flat wire shape is a contract, not an implementation detail]]
|
||||
- [[_COMMUNITY_Installed userscripts must keep working (14-day grace window)|Installed userscripts must keep working (14-day grace window)]]
|
||||
- [[_COMMUNITY_CDP (Chrome DevTools Protocol) endpoint|CDP (Chrome DevTools Protocol) endpoint]]
|
||||
- [[_COMMUNITY_headless-shell service (socat-fronted CDP)|headless-shell service (socat-fronted CDP)]]
|
||||
- [[_COMMUNITY_Start Chrome on first CDP connection, reap after 300s idle|Start Chrome on first CDP connection, reap after 300s idle]]
|
||||
- [[_COMMUNITY_BROWSER_WS_URL configuration seam|BROWSER_WS_URL configuration seam]]
|
||||
- [[_COMMUNITY_ADR-0006 Browser runs on the home machine over the tailnet|ADR-0006: Browser runs on the home machine over the tailnet]]
|
||||
- [[_COMMUNITY_Browser moved home VPS memory pressure, no requests served|Browser moved home: VPS memory pressure, no requests served]]
|
||||
- [[_COMMUNITY_Tailnet (Tailscale network)|Tailnet (Tailscale network)]]
|
||||
- [[_COMMUNITY_Content-addressed filesystem storage (SHA-256 of source URL)|Content-addressed filesystem storage (SHA-256 of source URL)]]
|
||||
- [[_COMMUNITY_Cover (Series image bytes)|Cover (Series image bytes)]]
|
||||
- [[_COMMUNITY_Deny-class destination control for outbound fetch|Deny-class destination control for outbound fetch]]
|
||||
- [[_COMMUNITY_ADR-0007 Backend hosts every Site's Cover bytes|ADR-0007: Backend hosts every Site's Cover bytes]]
|
||||
- [[_COMMUNITY_kagane CORP same-origin cover restriction|kagane CORP same-origin cover restriction]]
|
||||
- [[_COMMUNITY_Backend acquires, stores, serves every Cover (uniformity)|Backend acquires, stores, serves every Cover (uniformity)]]
|
||||
- [[_COMMUNITY_aaria-label='All Chapter' anchor pointer|a[aria-label='All Chapter'] anchor pointer]]
|
||||
- [[_COMMUNITY_Series identity is discovered from the Site's links|Series identity is discovered from the Site's links]]
|
||||
- [[_COMMUNITY_ADR-0008 Series identity discovered, never derived|ADR-0008: Series identity discovered, never derived]]
|
||||
- [[_COMMUNITY_Chapter slug vs series slug divergence (~7% measured)|Chapter slug vs series slug divergence (~7% measured)]]
|
||||
- [[_COMMUNITY_Scan truncated at first wpd-threads marker|Scan truncated at first wpd-threads marker]]
|
||||
- [[_COMMUNITY_Surface ADR conflicts explicitly rather than silently overriding|Surface ADR conflicts explicitly rather than silently overriding]]
|
||||
- [[_COMMUNITY_Domain docs single-context layout guidance|Domain docs: single-context layout guidance]]
|
||||
- [[_COMMUNITY_domain-modeling skill (lazy CONTEXT.md creation)|/domain-modeling skill (lazy CONTEXT.md creation)]]
|
||||
- [[_COMMUNITY_CONTEXT.md glossary (ubiquitous language)|CONTEXT.md glossary (ubiquitous language)]]
|
||||
- [[_COMMUNITY_Gitea (tea CLI, gitea.violetcrown.my.id)|Gitea (tea CLI, gitea.violetcrown.my.id)]]
|
||||
- [[_COMMUNITY_wayfinder mapticket mechanism|wayfinder map/ticket mechanism]]
|
||||
- [[_COMMUNITY_Triage labels canonical roles to tracker labels|Triage labels: canonical roles to tracker labels]]
|
||||
- [[_COMMUNITY_Canonical triage role labels (needs-triage ... wontfix)|Canonical triage role labels (needs-triage ... wontfix)]]
|
||||
- [[_COMMUNITY_Cinder (BookmarkManager Web UI design system)|Cinder (BookmarkManager Web UI design system)]]
|
||||
- [[_COMMUNITY_Heat is typographic ember reserved for unread chapters|Heat is typographic: ember reserved for unread chapters]]
|
||||
- [[_COMMUNITY_Design tokens (dark + light branches, no hardcoded hex)|Design tokens (dark + light branches, no hardcoded hex)]]
|
||||
- [[_COMMUNITY_Three type roles display serif mono small-caps sans|Three type roles: display serif / mono small-caps / sans]]
|
||||
- [[_COMMUNITY_aaria-label='All Chapter' priority pointer|a[aria-label='All Chapter'] priority pointer]]
|
||||
- [[_COMMUNITY_Research lightnovelworld chapter slug vs series slug|Research: lightnovelworld chapter slug vs series slug]]
|
||||
- [[_COMMUNITY_Gitea issue 77 (chapter vs series slug)|Gitea issue #77 (chapter vs series slug)]]
|
||||
- [[_COMMUNITY_Slug divergence measurements (341 diverge, 1 split)|Slug divergence measurements (3/41 diverge, 1 split)]]
|
||||
- [[_COMMUNITY_Unscoped chapter regex is SAFE, truncated at wpd-threads|Unscoped chapter regex is SAFE, truncated at wpd-threads]]
|
||||
- [[_COMMUNITY_BookmarkManager|BookmarkManager]]
|
||||
- [[_COMMUNITY_Bromite (Primary Device)|Bromite (Primary Device)]]
|
||||
- [[_COMMUNITY_Dark-First Design Constraint|Dark-First Design Constraint]]
|
||||
- [[_COMMUNITY_Discord Guild Membership|Discord Guild Membership]]
|
||||
- [[_COMMUNITY_Reader Isolation Invariant|Reader Isolation Invariant]]
|
||||
- [[_COMMUNITY_Browser Unit Redeploy|Browser Unit Redeploy]]
|
||||
- [[_COMMUNITY_pg_dump Hot Backup|pg_dump Hot Backup]]
|
||||
- [[_COMMUNITY_Redeploy Runbook|Redeploy Runbook]]
|
||||
- [[_COMMUNITY_Rollback Strategy|Rollback Strategy]]
|
||||
- [[_COMMUNITY_AGENTS|AGENTS.md]]
|
||||
- [[_COMMUNITY_Userscript CLAUDE guidance|Userscript CLAUDE guidance]]
|
||||
|
||||
## God Nodes (most connected - your core abstractions)
|
||||
1. `testConfig()` - 53 edges
|
||||
2. `newWebTestServer()` - 49 edges
|
||||
3. `newTestStore()` - 42 edges
|
||||
4. `newTestStore()` - 41 edges
|
||||
5. `e()` - 33 edges
|
||||
6. `Handler` - 29 edges
|
||||
7. `ne()` - 28 edges
|
||||
8. `De()` - 28 edges
|
||||
9. `Open()` - 27 edges
|
||||
10. `se()` - 27 edges
|
||||
|
||||
## Surprising Connections (you probably didn't know these)
|
||||
- `Browser Sidecar Service` --semantically_similar_to--> `Browser Sidecar (BROWSER_WS_URL)` [INFERRED] [semantically similar]
|
||||
chrome/docker-compose.yml → backend/AGENTS.md
|
||||
- `el()` --indirect_call--> `c()` [INFERRED]
|
||||
userscript/manga-bookmark.user.js → backend/internal/web/static/htmx.min.js
|
||||
- `el()` --indirect_call--> `c()` [INFERRED]
|
||||
userscript/novel-bookmark.user.js → backend/internal/web/static/htmx.min.js
|
||||
- `latestChapterFromAnchors()` --indirect_call--> `re()` [INFERRED]
|
||||
userscript/manga-bookmark.user.js → backend/internal/web/static/htmx.min.js
|
||||
- `el()` --indirect_call--> `k()` [INFERRED]
|
||||
userscript/manga-bookmark.user.js → backend/internal/web/static/htmx.min.js
|
||||
|
||||
## Import Cycles
|
||||
- None detected.
|
||||
|
||||
## Hyperedges (group relationships)
|
||||
- **Batch Ticket Implementation Pipeline** — _claude_skills_implement_tickets_skill_implement_tickets, _omp_agents_ticket_implementer_ticket_implementer, _omp_agents_ticket_implementer_cr_spec, _omp_agents_ticket_implementer_cr_standards [INFERRED 0.85]
|
||||
- **Subagent-Driven Development Pipeline** — _opencode_agent_implementer_implementer, _opencode_agent_reviewer_reviewer, _opencode_agent_implementer_subagent_driven_development [INFERRED 0.85]
|
||||
- **Go HTML Template Family** — backend_internal_web_templates_app_doc, backend_internal_web_templates_card_doc, backend_internal_web_templates_list_doc, backend_internal_web_templates_chrome_doc, backend_internal_web_templates_login_doc, backend_internal_web_templates_setup_doc, backend_internal_web_templates_readers_doc, backend_internal_web_templates_icons_doc [INFERRED 0.95]
|
||||
- **htmx Fragment Swap Flow** — backend_internal_web_templates_app_doc, backend_internal_web_templates_card_doc, backend_internal_web_templates_chrome_doc, backend_internal_web_templates_setup_doc, backend_internal_web_templates_readers_doc [INFERRED 0.95]
|
||||
- **Backend owns the truth (single-writer ownership of shared facts)** — docs_adr_0003_series_shared_and_poll_owned_poll_owned_writes, docs_adr_0004_wire_format_does_not_mirror_the_schema_flat_wire_contract, docs_adr_0007_backend_hosts_cover_bytes_server_side_covers [INFERRED 0.85]
|
||||
- **Headless browser infrastructure (sidecar, on-demand, home deployment)** — docs_adr_0005_on_demand_browser_headless_shell, docs_adr_0005_on_demand_browser_cdp, docs_adr_0005_on_demand_browser_on_demand_start, docs_adr_0006_browser_on_the_home_machine_home_machine_rationale [INFERRED 0.85]
|
||||
- **lightnovelworld series-identity investigation and fix** — docs_research_lightnovelworld_chapter_vs_series_slug_issue_77, docs_research_lightnovelworld_chapter_vs_series_slug_unscoped_regex, docs_adr_0008_series_identity_is_discovered_not_derived_discovered_identity [INFERRED 0.85]
|
||||
|
||||
## Communities (233 total, 170 thin omitted)
|
||||
|
||||
### Community 0 - "HTMX Library Internals"
|
||||
Cohesion: 0.08
|
||||
Nodes (101): A(), ae(), an(), at(), B(), be(), bn(), bt() (+93 more)
|
||||
|
||||
### Community 1 - "Cover Fetch Test Helpers"
|
||||
Cohesion: 0.10
|
||||
Nodes (84): floatPtr(), testConfig(), getCover(), Cookie, Handler, ResponseRecorder, T, TestListRendersAcquiredCover() (+76 more)
|
||||
|
||||
### Community 2 - "Manga Userscript Adapters"
|
||||
Cohesion: 0.06
|
||||
Nodes (76): adapterFor(), anchorsFromDocument(), anchorsFromHTML(), apiDelete(), apiGet(), apiPut(), applyFabPos(), applyLatestChapterIfChanged() (+68 more)
|
||||
|
||||
### Community 3 - "Novel Userscript Adapters"
|
||||
Cohesion: 0.06
|
||||
Nodes (79): adapterFor(), anchorsFromDocument(), anchorsFromHTML(), apiDelete(), apiGet(), apiPut(), applyFabPos(), applyLatestChapterIfChanged() (+71 more)
|
||||
|
||||
### Community 4 - "Series Acquisition Tests"
|
||||
Cohesion: 0.10
|
||||
Nodes (67): bookmarkNewKaganeSeries(), bookmarkNewNovelfullSeries(), bookmarkNewSeries(), Context, Store, T, newAcquirer(), readBookmark() (+59 more)
|
||||
|
||||
### Community 5 - "Bookmarks API Tests"
|
||||
Cohesion: 0.08
|
||||
Nodes (67): auth(), getBookmarks(), Handler, Request, Store, T, newTestServer(), newTestStore() (+59 more)
|
||||
|
||||
### Community 7 - "Cover & Acquire Internals"
|
||||
Cohesion: 0.09
|
||||
Nodes (31): Addr, Context, Store, defaultCoverResolver(), fetchCoverBytes(), Client, Context, NewCoverFetcher() (+23 more)
|
||||
|
||||
### Community 8 - "System Architecture Concepts"
|
||||
Cohesion: 0.10
|
||||
Nodes (26): Confirm-Gated Destructive Actions, Discord OAuth & Guild-Membership Gate, Lifecycle Buckets (reading/archived/finished), Reader-Owned Store, HMAC-Derived Reader Credentials, Web Package (Browser UI + Templates), app.html — App Shell Template, Manga/Novel Library Switch (+18 more)
|
||||
|
||||
### Community 9 - "Session Middleware"
|
||||
Cohesion: 0.07
|
||||
Nodes (35): ClearCookie(), ClientIP(), Duration, Mutex, Request, ResponseWriter, Time, isHTTPS() (+27 more)
|
||||
|
||||
### Community 10 - "Go Test Helpers"
|
||||
Cohesion: 0.05
|
||||
Nodes (39): Test Helpers, Test Timeout, Basic Handler Test, HTTP Handler Testing, Query Parameters and Headers, Docker Compose Fixture, Integration Testing, SQL Schema Fixture (+31 more)
|
||||
|
||||
### Community 11 - "Store Tests"
|
||||
Cohesion: 0.07
|
||||
Nodes (79): M, TestMain(), M, TestMain(), M, Main(), start(), URL() (+71 more)
|
||||
|
||||
### Community 12 - "Bookmarks API Handler"
|
||||
Cohesion: 0.08
|
||||
Nodes (33): Handler, Request, ResponseWriter, Store, Healthz(), writeJSON(), Auth(), compressible() (+25 more)
|
||||
|
||||
### Community 13 - "Web UI Handlers"
|
||||
Cohesion: 0.06
|
||||
Nodes (24): coverRelativePath(), coverSourceAddress(), displayChapter(), Store, currentLib(), currentTab(), filterBookmarks(), Client (+16 more)
|
||||
|
||||
### Community 14 - "Go Error Handling"
|
||||
Cohesion: 0.06
|
||||
Nodes (33): Creating Errors, Custom Error Types, Custom types that wrap other errors, Decision table: which error strategy to use, Error Creation, Error String Conventions, Errors as Values, `errors.New` — static error messages (+25 more)
|
||||
|
||||
### Community 15 - "CDP Browser Client"
|
||||
Cohesion: 0.08
|
||||
Nodes (31): awaitPromise(), browserConnectionLost(), classifyBrowserError(), Action, Context, Mutex, jsString(), kaganeAPIURL() (+23 more)
|
||||
|
||||
### Community 17 - "Go Code Style Guide"
|
||||
Cohesion: 0.08
|
||||
Nodes (23): Code Style Details, Extract Complex Conditions, Value vs Pointer Arguments, Code Organization Within Files, Complex Conditions & Init Scope, Composite Literals, Control Flow, Cross-References (+15 more)
|
||||
|
||||
### Community 20 - "I/O Performance Patterns"
|
||||
Cohesion: 0.11
|
||||
Nodes (18): Avoid io.ReadAll for large payloads, Batch Operations, Buffered I/O, Cgo Overhead, Channel: batch processing from a stream, Concurrent Multi-Stage Pipelines, Connection pooling, Database: batch inserts over row-by-row (+10 more)
|
||||
|
||||
### Community 21 - "CPU Optimization"
|
||||
Cohesion: 0.13
|
||||
Nodes (15): Cache Locality, Contiguous 2D allocation, CPU Optimization, False Sharing, Function Inlining, Handling CPU-specific instruction sets, Instruction-Level Parallelism, Monotonic Time (+7 more)
|
||||
|
||||
### Community 22 - "Caching Patterns"
|
||||
Cohesion: 0.13
|
||||
Nodes (14): Algorithmic Complexity, Avoid iterator chains, Caching Patterns, Compiled Pattern Caching, Early returns and short-circuit loops, LRU caches, Map lookups over slice scanning, Precomputed lookup tables (+6 more)
|
||||
|
||||
### Community 23 - "Browser Entrypoint"
|
||||
Cohesion: 0.35
|
||||
Nodes (14): browser_alive(), connection(), connection_signal(), finish_connection(), has_connections(), lock(), reaper(), entrypoint.sh script (+6 more)
|
||||
|
||||
### Community 24 - "Memory Allocation & GC"
|
||||
Cohesion: 0.13
|
||||
Nodes (15): Allocation Rate Reduction, Ballast pattern (pre-Go 1.19), Garbage Collector Tuning, GC pacing, GC Profiling and Diagnostics, GODEBUG=gctrace=1, GOGC (default: 100), GOMAXPROCS in Containers (+7 more)
|
||||
|
||||
### Community 25 - "Cover Fetcher Tests"
|
||||
Cohesion: 0.33
|
||||
Nodes (12): coverResponse(), Request, T, TestCoverFetcherCanonicalisesJpgAlias(), TestCoverFetcherFetchesPublicHTTPSImage(), TestCoverFetcherRefusesUnsafeDestinationsBeforeRequest(), TestCoverFetcherRejectsNonImage(), TestCoverFetcherRejectsOversizedBody() (+4 more)
|
||||
|
||||
### Community 26 - "readSeries"
|
||||
Cohesion: 0.13
|
||||
Nodes (28): asuraLatestChapter(), browserOnlyCoverURL(), comixCoverEntry(), comixCoverURL(), comixLatestChapter(), comixSeriesID(), coverFrom(), demonicLatestChapter() (+20 more)
|
||||
|
||||
### Community 27 - "Find Skills Guide"
|
||||
Cohesion: 0.14
|
||||
Nodes (13): Common Skill Categories, Find Skills, How to Help Users Find Skills, Step 1: Understand What They Need, Step 2: Check the Leaderboard First, Step 3: Search for Skills, Step 4: Verify Quality Before Recommending, Step 5: Present Options to the User (+5 more)
|
||||
|
||||
### Community 28 - "Allocation Patterns"
|
||||
Cohesion: 0.14
|
||||
Nodes (14): Allocation Patterns, Backing Array Leaks, Direct indexing vs append, Eliminate redundant map lookups, Interface boxing, Map never shrinks, Map size hints, Memory Optimization (+6 more)
|
||||
|
||||
### Community 29 - "Observability & Alerting"
|
||||
Cohesion: 0.22
|
||||
Nodes (9): Alerting rules (examples), CPU saturation, GC pressure, Goroutine leaks, Grafana Dashboards, Memory leaks, Prometheus Metrics for Go, PromQL Queries for Performance Diagnosis (+1 more)
|
||||
|
||||
### Community 31 - "Memory Layout"
|
||||
Cohesion: 0.40
|
||||
Nodes (5): Map of pointers for large, frequently updated structs, Memory Layout, Pointer receivers for large structs, Struct field alignment, Zero-size field at end of struct
|
||||
|
||||
### Community 33 - "Go Testing Guide"
|
||||
Cohesion: 0.20
|
||||
Nodes (10): CI Regression Detection, Common Mistakes, Core Philosophy, Cross-References, Decision Tree: Where Is Time Spent?, Deep Dives, Go Performance Optimization, Iterative Optimization Methodology (+2 more)
|
||||
|
||||
### Community 34 - "Session Store"
|
||||
Cohesion: 0.28
|
||||
Nodes (4): Duration, Store, Time, Session
|
||||
|
||||
### Community 35 - "Web UI Filter Logic"
|
||||
Cohesion: 0.31
|
||||
Nodes (5): closeCardPanels(), setActiveTab(), toggleChapterForm(), toggleConfirmRow(), togglePanel()
|
||||
|
||||
### Community 37 - "Product & Security Context"
|
||||
Cohesion: 0.17
|
||||
Nodes (11): Accessibility & Inclusion, Brand Commitments, Capabilities and Constraints, Evidence on Hand, Operating Context, Platform, Positioning, Product (+3 more)
|
||||
|
||||
### Community 38 - "novel-logic.test.js"
|
||||
Cohesion: 0.33
|
||||
Nodes (5): ADR-0009: A Site answers fixed questions; an unusual Site owns its own fetch, Consequences, Considered options, Decision, Why
|
||||
|
||||
### Community 39 - "UI Critique 2026-07-26A"
|
||||
Cohesion: 0.29
|
||||
Nodes (6): Design Health Score, Design Specificity Verdict, Minor Observations, Persona Red Flags, Priority Issues, Questions to Consider
|
||||
|
||||
### Community 40 - "UI Critique 2026-07-26B"
|
||||
Cohesion: 0.29
|
||||
Nodes (6): Design Health Score, Design Specificity Verdict, Minor Observations, Persona Red Flags, Priority Issues, Questions to Consider
|
||||
|
||||
### Community 45 - "Go Perf Alert Rules"
|
||||
Cohesion: 0.50
|
||||
Nodes (4): Prometheus Alerting Rules (Go Performance), GoroutineLeak Alert, HighGCPauseTime Alert, MemoryNearLimit Alert
|
||||
|
||||
### Community 46 - "Userscript Display Logic"
|
||||
Cohesion: 0.07
|
||||
Nodes (27): 1. Summary answer table, 2. The two reference pages, 3. Chapter page → series URL: every in-page pointer, in priority order, 4.1 Sample method, 4.2 Divergence results, 4.3 Is there a derivable rule? **No.**, 4.4 The split case — a slug can change *mid-series*, 4. The reverse direction, and how common divergence is (+19 more)
|
||||
|
||||
### Community 47 - "Go Perf Skill Docs"
|
||||
Cohesion: 0.22
|
||||
Nodes (5): Continuous Profiling, Production Observability for Performance, Pyroscope pull mode (via Grafana Alloy), Pyroscope push mode, Real-Time Visualization (Development)
|
||||
|
||||
### Community 48 - "Login Page Art"
|
||||
Cohesion: 0.67
|
||||
Nodes (3): Fantasy Sword, Fiery Volcanic Scene, Login Art: Sword in Volcanic Rock
|
||||
|
||||
### Community 49 - "BookmarkManager Logo"
|
||||
Cohesion: 1.00
|
||||
Nodes (3): Mirrored Double Bookmark Mark, Ember Flame Accent, BookmarkManager Logo
|
||||
|
||||
### Community 103 - "bookmark-api Service"
|
||||
Cohesion: 0.24
|
||||
Nodes (10): Backend Go Service (stdlib net/http), Browser Sidecar (BROWSER_WS_URL), Browser Sidecar Service, CDP Endpoint (Tailnet-Bound :9222), Persistent Chrome Profile Volume, chrome/docker-compose.yml — Browser Deployable Unit, bookmark-api Prod Override, CDP Never on Shared Proxy Network (+2 more)
|
||||
|
||||
### Community 104 - "AGENTS.md"
|
||||
Cohesion: 0.12
|
||||
Nodes (14): Agent skills, Architecture, Commands, Comments, Design system, Domain docs, Forge: Gitea, not GitHub, graphify (+6 more)
|
||||
|
||||
### Community 105 - "reviewer.md"
|
||||
Cohesion: 0.12
|
||||
Nodes (15): Assessment, Calibration, Critical (Must Fix), Do Not Trust the Report, Important (Should Fix), Inputs, Issues, Method (+7 more)
|
||||
|
||||
### Community 106 - "Redeploy runbook"
|
||||
Cohesion: 0.12
|
||||
Nodes (15): 0. Preflight, 1. Back up the database, 2. Pull the new code, 3. Rebuild and restart, 4. Verify the deploy, 5. Smoke-test the full loop, 6. Rollback, 7. The whole thing, as one block (+7 more)
|
||||
|
||||
### Community 107 - "1. Backend"
|
||||
Cohesion: 0.13
|
||||
Nodes (14): 1. Backend, 2. Userscript, Adapter reference (verified live 2026-07-24), Config (env), Deploy behind your reverse proxy, Desktop iteration (optional), Develop / test, Endpoints (+6 more)
|
||||
|
||||
### Community 108 - "Deployment"
|
||||
Cohesion: 0.14
|
||||
Nodes (13): 0. Prerequisites, 1. Configure `.env`, 1b. Web UI, 2. Build + start, 3. Verify over HTTPS, 4. Configure the userscript, 5. Install on Bromite, 6. Smoke-test the full loop (+5 more)
|
||||
|
||||
### Community 109 - "Cinder — BookmarkManager design system"
|
||||
Cohesion: 0.20
|
||||
Nodes (9): 1. The one idea, 2. Tokens, 3. Type, 4. Components (web UI), 5. Components (userscript panel), 6. Motion, 7. Accessibility floor (not negotiable), 8. Adding something new — checklist (+1 more)
|
||||
|
||||
### Community 110 - "Implement tickets"
|
||||
Cohesion: 0.22
|
||||
Nodes (8): 1. Collect the tickets, 2. Plan the batch, 3. Get the plan approved, 4. Run a wave, 5. Land the wave, 6. Close the batch, Implement tickets, Ticket #<n> — <title>
|
||||
|
||||
### Community 111 - "SQLite → Postgres cutover runbook"
|
||||
Cohesion: 0.22
|
||||
Nodes (8): 0. The generator is throwaway, 1. Stop the old API and take a fresh export, 2. Bring up Postgres with the schema and the owner Reader, 3. Generate the import SQL, 4. Review it by eye, 5. Apply it, 6. Afterwards, SQLite → Postgres cutover runbook
|
||||
|
||||
### Community 112 - "Testing the userscript"
|
||||
Cohesion: 0.29
|
||||
Nodes (6): Adding a test, Commands, Gotchas, How the harness works, Testing the userscript, What is NOT testable here
|
||||
|
||||
### Community 113 - "ADR-0007: The backend hosts every Site's Cover bytes"
|
||||
Cohesion: 0.29
|
||||
Nodes (6): ADR-0007: The backend hosts every Site's Cover bytes, Consequences, Considered options, Decision, Two deliberate relaxations, Why a future reader will find this surprising
|
||||
|
||||
### Community 114 - "Issue tracker: Gitea (`tea` CLI)"
|
||||
Cohesion: 0.29
|
||||
Nodes (6): Conventions, Issue tracker: Gitea (`tea` CLI), Pull requests as a triage surface, Wayfinding operations, When a skill says "fetch the relevant ticket", When a skill says "publish to the issue tracker"
|
||||
|
||||
### Community 115 - "ADR-0006: The browser runs on the home machine, over the tailnet"
|
||||
Cohesion: 0.33
|
||||
Nodes (5): ADR-0006: The browser runs on the home machine, over the tailnet, Consequences, Constraints, Decision, Why
|
||||
|
||||
### Community 116 - "ADR-0008: A Series identity is discovered from the Site's links, never derived from an address"
|
||||
Cohesion: 0.33
|
||||
Nodes (5): ADR-0008: A Series identity is discovered from the Site's links, never derived from an address, Consequences, Considered options, Decision, Why
|
||||
|
||||
### Community 117 - "Domain Docs"
|
||||
Cohesion: 0.33
|
||||
Nodes (5): Before exploring, read these, Domain Docs, File structure, Flag ADR conflicts, Use the glossary's vocabulary
|
||||
|
||||
### Community 118 - "ticket-implementer.md"
|
||||
Cohesion: 0.33
|
||||
Nodes (5): Escalate rather than guess, Order of work, Report, Review, The worktree is your whole world
|
||||
|
||||
### Community 119 - "implementer.md"
|
||||
Cohesion: 0.33
|
||||
Nodes (5): Before You Begin, Report Format, Self-Review Before Reporting, When You're in Over Your Head, Your Job
|
||||
|
||||
### Community 120 - "Series is a shared entity, and only the Poll may update it"
|
||||
Cohesion: 0.40
|
||||
Nodes (4): Consequences, Only the Poll writes Series fields, Series is a shared entity, and only the Poll may update it, Why
|
||||
|
||||
### Community 121 - "Postgres replaces SQLite as the primary datastore"
|
||||
Cohesion: 0.50
|
||||
Nodes (3): Consequences, Considered options, Postgres replaces SQLite as the primary datastore
|
||||
|
||||
### Community 122 - "Identity comes from Discord OAuth; we store no passwords and send no email"
|
||||
Cohesion: 0.50
|
||||
Nodes (3): Consequences, Considered options, Identity comes from Discord OAuth; we store no passwords and send no email
|
||||
|
||||
### Community 123 - "The wire format stays flat and deliberately does not mirror the schema"
|
||||
Cohesion: 0.50
|
||||
Nodes (3): Consequence, The wire format stays flat and deliberately does not mirror the schema, Why a future reader will find this surprising
|
||||
|
||||
### Community 124 - "ADR-0005: On-demand browser sidecar"
|
||||
Cohesion: 0.50
|
||||
Nodes (3): ADR-0005: On-demand browser sidecar, Constraints, Decision
|
||||
|
||||
### Community 273 - "AGENTS.md"
|
||||
Cohesion: 0.50
|
||||
Nodes (3): Live URL shapes (verified 2026-07-26, may drift — re-check against live pages before trust), Second script: `novel-bookmark.user.js`, Userscript structure (single IIFE, `manga-bookmark.user.js`)
|
||||
|
||||
## Knowledge Gaps
|
||||
- **508 isolated node(s):** `bookmarkmanager/backend`, `ctxKey`, `loginView`, `ctxKey`, `test` (+503 more)
|
||||
These have ≤1 connection - possible missing edges or undocumented components.
|
||||
- **170 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
|
||||
|
||||
## Suggested Questions
|
||||
_Questions this graph is uniquely positioned to answer:_
|
||||
|
||||
- **Why does `New()` connect `Series Acquisition Tests` to `Bookmarks API Tests`, `Cover & Acquire Internals`, `Session Middleware`, `Store Tests`, `Web UI Handlers`?**
|
||||
_High betweenness centrality (0.052) - this node is a cross-community bridge._
|
||||
- **Why does `Open()` connect `Store Tests` to `Cover Fetch Test Helpers`, `Web UI Handlers`, `Series Acquisition Tests`, `Bookmarks API Tests`?**
|
||||
_High betweenness centrality (0.034) - this node is a cross-community bridge._
|
||||
- **Why does `newRouter()` connect `Bookmarks API Tests` to `Cover Fetch Test Helpers`, `Bookmarks API Handler`, `Series Acquisition Tests`?**
|
||||
_High betweenness centrality (0.027) - this node is a cross-community bridge._
|
||||
- **Are the 47 inferred relationships involving `testConfig()` (e.g. with `TestListRendersAcquiredCover()` and `TestPublicCoverNeverEchoesNonImage()`) actually correct?**
|
||||
_`testConfig()` has 47 INFERRED edges - model-reasoned connections that need verification._
|
||||
- **Are the 8 inferred relationships involving `newWebTestServer()` (e.g. with `TestListRendersAcquiredCover()` and `TestPublicCoverRejectsUnknownAddress()`) actually correct?**
|
||||
_`newWebTestServer()` has 8 INFERRED edges - model-reasoned connections that need verification._
|
||||
- **Are the 6 inferred relationships involving `newTestStore()` (e.g. with `TestCreateAndGetSession()` and `TestDeleteSessionIsPerReader()`) actually correct?**
|
||||
_`newTestStore()` has 6 INFERRED edges - model-reasoned connections that need verification._
|
||||
- **What connects `bookmarkmanager/backend`, `ctxKey`, `loginView` to the rest of the system?**
|
||||
_548 weakly-connected nodes found - possible documentation gaps or missing edges._
|
||||
File diff suppressed because one or more lines are too long
+51547
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,622 @@
|
||||
{
|
||||
".agents/skills/golang-code-style/evals/evals.json": {
|
||||
"mtime": 1784884678.6627614,
|
||||
"ast_hash": "bec0e12446e7af3cd05de9b6d42badd8",
|
||||
"semantic_hash": "bec0e12446e7af3cd05de9b6d42badd8"
|
||||
},
|
||||
".agents/skills/golang-error-handling/evals/evals.json": {
|
||||
"mtime": 1784884678.6655047,
|
||||
"ast_hash": "275d710b774fba1e1d0bc098d3646c6d",
|
||||
"semantic_hash": "275d710b774fba1e1d0bc098d3646c6d"
|
||||
},
|
||||
".agents/skills/golang-performance/evals/evals.json": {
|
||||
"mtime": 1784884678.6688662,
|
||||
"ast_hash": "4f06df87f90aa0e4f6deaa318b47689f",
|
||||
"semantic_hash": "4f06df87f90aa0e4f6deaa318b47689f"
|
||||
},
|
||||
".agents/skills/golang-testing/evals/evals.json": {
|
||||
"mtime": 1784884678.6721346,
|
||||
"ast_hash": "60a821bbfd20c6fe8bba996b8b540dd4",
|
||||
"semantic_hash": "60a821bbfd20c6fe8bba996b8b540dd4"
|
||||
},
|
||||
"backend/go.mod": {
|
||||
"mtime": 1786216141.668644,
|
||||
"ast_hash": "dac242903b0e98c3e4395159d609e08e",
|
||||
"semantic_hash": "dac242903b0e98c3e4395159d609e08e"
|
||||
},
|
||||
"backend/main.go": {
|
||||
"mtime": 1786363889.5678573,
|
||||
"ast_hash": "6e98a3ae91aaa132df251e43c4dfca6d",
|
||||
"semantic_hash": "6e98a3ae91aaa132df251e43c4dfca6d"
|
||||
},
|
||||
"skills-lock.json": {
|
||||
"mtime": 1784884678.6842625,
|
||||
"ast_hash": "4a94ac85bad6bce330d085bcc0ae3ffd",
|
||||
"semantic_hash": "4a94ac85bad6bce330d085bcc0ae3ffd"
|
||||
},
|
||||
"userscript/manga-bookmark.user.js": {
|
||||
"mtime": 1786488438.9080842,
|
||||
"ast_hash": "1f8bcddd3632d709f058a8401af8f127",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/find-skills/SKILL.md": {
|
||||
"mtime": 1784884338.760326,
|
||||
"ast_hash": "62b297abdee9aea84c577ab2e04e1974",
|
||||
"semantic_hash": "62b297abdee9aea84c577ab2e04e1974"
|
||||
},
|
||||
".agents/skills/golang-code-style/SKILL.md": {
|
||||
"mtime": 1784884678.6623824,
|
||||
"ast_hash": "d6a01e6f64550a5c8d59dac2e948000e",
|
||||
"semantic_hash": "d6a01e6f64550a5c8d59dac2e948000e"
|
||||
},
|
||||
".agents/skills/golang-code-style/references/details.md": {
|
||||
"mtime": 1784884678.6627865,
|
||||
"ast_hash": "19891e396a986f1b24bf34b7a2854fcb",
|
||||
"semantic_hash": "19891e396a986f1b24bf34b7a2854fcb"
|
||||
},
|
||||
".agents/skills/golang-error-handling/SKILL.md": {
|
||||
"mtime": 1784884678.665052,
|
||||
"ast_hash": "8b7970f472adb240e5bc4bde863d43a6",
|
||||
"semantic_hash": "8b7970f472adb240e5bc4bde863d43a6"
|
||||
},
|
||||
".agents/skills/golang-error-handling/references/error-creation.md": {
|
||||
"mtime": 1784884678.665524,
|
||||
"ast_hash": "248dbf75492c68faef2334b8d83bd080",
|
||||
"semantic_hash": "248dbf75492c68faef2334b8d83bd080"
|
||||
},
|
||||
".agents/skills/golang-error-handling/references/error-handling.md": {
|
||||
"mtime": 1784884678.6655412,
|
||||
"ast_hash": "c2424ee3999b05963b199e0100727aa3",
|
||||
"semantic_hash": "c2424ee3999b05963b199e0100727aa3"
|
||||
},
|
||||
".agents/skills/golang-error-handling/references/error-wrapping.md": {
|
||||
"mtime": 1784884678.6655717,
|
||||
"ast_hash": "4a15d9266a1ea0c8bb1951e6b9c0f586",
|
||||
"semantic_hash": "4a15d9266a1ea0c8bb1951e6b9c0f586"
|
||||
},
|
||||
".agents/skills/golang-performance/SKILL.md": {
|
||||
"mtime": 1784884678.667833,
|
||||
"ast_hash": "35a15fd129c5bedaada3fd4df0b6ba8c",
|
||||
"semantic_hash": "35a15fd129c5bedaada3fd4df0b6ba8c"
|
||||
},
|
||||
".agents/skills/golang-performance/assets/prometheus-alerts.yml": {
|
||||
"mtime": 1784884678.668527,
|
||||
"ast_hash": "fa9357ffa87c4f894fc21afa9707c4db",
|
||||
"semantic_hash": "fa9357ffa87c4f894fc21afa9707c4db"
|
||||
},
|
||||
".agents/skills/golang-performance/references/caching.md": {
|
||||
"mtime": 1784884678.6690567,
|
||||
"ast_hash": "807c42a82994e5548dbf7eb6e30c71e0",
|
||||
"semantic_hash": "807c42a82994e5548dbf7eb6e30c71e0"
|
||||
},
|
||||
".agents/skills/golang-performance/references/cpu.md": {
|
||||
"mtime": 1784884678.669087,
|
||||
"ast_hash": "6d52e532ef51a35cc4134694c944517d",
|
||||
"semantic_hash": "6d52e532ef51a35cc4134694c944517d"
|
||||
},
|
||||
".agents/skills/golang-performance/references/io-networking.md": {
|
||||
"mtime": 1784884678.6691036,
|
||||
"ast_hash": "95c5dd51f728fd69c945a92ff021d766",
|
||||
"semantic_hash": "95c5dd51f728fd69c945a92ff021d766"
|
||||
},
|
||||
".agents/skills/golang-performance/references/memory.md": {
|
||||
"mtime": 1784884678.6691158,
|
||||
"ast_hash": "3b2108df06b4cfb3980fa80bbd9ebcff",
|
||||
"semantic_hash": "3b2108df06b4cfb3980fa80bbd9ebcff"
|
||||
},
|
||||
".agents/skills/golang-performance/references/observability.md": {
|
||||
"mtime": 1784884678.6691446,
|
||||
"ast_hash": "0aa8a498e8d55ccdd4990ad187bae828",
|
||||
"semantic_hash": "0aa8a498e8d55ccdd4990ad187bae828"
|
||||
},
|
||||
".agents/skills/golang-performance/references/runtime.md": {
|
||||
"mtime": 1784884678.669426,
|
||||
"ast_hash": "26386c33b3a3794aaef0556713bf8f3a",
|
||||
"semantic_hash": "26386c33b3a3794aaef0556713bf8f3a"
|
||||
},
|
||||
".agents/skills/golang-testing/SKILL.md": {
|
||||
"mtime": 1784884678.6716368,
|
||||
"ast_hash": "0a9b9793bba2a239db94e980272a393e",
|
||||
"semantic_hash": "0a9b9793bba2a239db94e980272a393e"
|
||||
},
|
||||
".agents/skills/golang-testing/references/helpers.md": {
|
||||
"mtime": 1784884678.6722167,
|
||||
"ast_hash": "0aecb7cfbeb9b374bf61c52e324d9d3f",
|
||||
"semantic_hash": "0aecb7cfbeb9b374bf61c52e324d9d3f"
|
||||
},
|
||||
".agents/skills/golang-testing/references/http-testing.md": {
|
||||
"mtime": 1784884678.67225,
|
||||
"ast_hash": "9111110c28a7fbbffc3537aad786b390",
|
||||
"semantic_hash": "9111110c28a7fbbffc3537aad786b390"
|
||||
},
|
||||
".agents/skills/golang-testing/references/integration-testing.md": {
|
||||
"mtime": 1784884678.67225,
|
||||
"ast_hash": "fcf9861bc36ae56e3fb7a1c18aa0e77f",
|
||||
"semantic_hash": "fcf9861bc36ae56e3fb7a1c18aa0e77f"
|
||||
},
|
||||
".agents/skills/golang-testing/references/mocking.md": {
|
||||
"mtime": 1784884678.6722653,
|
||||
"ast_hash": "3a08979e4603aae5c32a58d5b6c39765",
|
||||
"semantic_hash": "3a08979e4603aae5c32a58d5b6c39765"
|
||||
},
|
||||
"CLAUDE.md": {
|
||||
"mtime": 1786488493.5922732,
|
||||
"ast_hash": "7362a25f37333a2a6a55a5aeccf9b0cb",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"DEPLOY.md": {
|
||||
"mtime": 1786488464.5532806,
|
||||
"ast_hash": "2b7b537aa1c0954400c19acc4b634029",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"README.md": {
|
||||
"mtime": 1786488487.8305292,
|
||||
"ast_hash": "9d6be8aa8a2946c23ad48d8f2864b5ca",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"docker-compose.prod.yml": {
|
||||
"mtime": 1786292465.8305523,
|
||||
"ast_hash": "0751998a532297b8ac507a01ec48dc31",
|
||||
"semantic_hash": "0751998a532297b8ac507a01ec48dc31"
|
||||
},
|
||||
"docker-compose.yml": {
|
||||
"mtime": 1786488450.3508112,
|
||||
"ast_hash": "124fd581bf0a662ff15012abfdb40a92",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".claude/settings.json": {
|
||||
"mtime": 1784951973.1869545,
|
||||
"ast_hash": "e51077b6a7f1f67afc748f1a32a1557d",
|
||||
"semantic_hash": "e51077b6a7f1f67afc748f1a32a1557d"
|
||||
},
|
||||
"backend/web_test.go": {
|
||||
"mtime": 1786216141.700644,
|
||||
"ast_hash": "8f1b093b59eb1ed81bc7fc0c22495c50",
|
||||
"semantic_hash": "8f1b093b59eb1ed81bc7fc0c22495c50"
|
||||
},
|
||||
"backend/main_test.go": {
|
||||
"mtime": 1786363889.5678573,
|
||||
"ast_hash": "8a165955cf28ad47481fec5ea7afb3d6",
|
||||
"semantic_hash": "8a165955cf28ad47481fec5ea7afb3d6"
|
||||
},
|
||||
".claude/settings.local.json": {
|
||||
"mtime": 1785697645.350201,
|
||||
"ast_hash": "9a1ac6369f968e8df4be9dcff0948f70",
|
||||
"semantic_hash": "9a1ac6369f968e8df4be9dcff0948f70"
|
||||
},
|
||||
"PRODUCT.md": {
|
||||
"mtime": 1786216141.660644,
|
||||
"ast_hash": "c52072d1978286060087fa0686f9c7f9",
|
||||
"semantic_hash": "c52072d1978286060087fa0686f9c7f9"
|
||||
},
|
||||
"backend/.impeccable/critique/2026-07-26T15-50-42Z__backend-templates-app-html.md": {
|
||||
"mtime": 1785128576.2412457,
|
||||
"ast_hash": "d08627c27f22d453125db6c1ab4b71ec",
|
||||
"semantic_hash": "d08627c27f22d453125db6c1ab4b71ec"
|
||||
},
|
||||
"backend/.impeccable/critique/2026-07-26T17-08-41Z__backend-templates-app-html.md": {
|
||||
"mtime": 1785128576.2452524,
|
||||
"ast_hash": "e69a8340a371579ca3ea689660f7d7bd",
|
||||
"semantic_hash": "e69a8340a371579ca3ea689660f7d7bd"
|
||||
},
|
||||
"AGENTS.md": {
|
||||
"mtime": 1786488493.5922732,
|
||||
"ast_hash": "7362a25f37333a2a6a55a5aeccf9b0cb",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"userscript/test/logic.test.js": {
|
||||
"mtime": 1786488563.581765,
|
||||
"ast_hash": "80d512bfe6fa8b847f3ba6169c321a74",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".claude/skills/testing-the-userscript/SKILL.md": {
|
||||
"mtime": 1786363889.5489495,
|
||||
"ast_hash": "8f3c0132eb4787a2c8736eb99f7689af",
|
||||
"semantic_hash": "8f3c0132eb4787a2c8736eb99f7689af"
|
||||
},
|
||||
"REDEPLOY.md": {
|
||||
"mtime": 1786363889.552731,
|
||||
"ast_hash": "d0baf08b95e7b5986234a9f36759c12e",
|
||||
"semantic_hash": "d0baf08b95e7b5986234a9f36759c12e"
|
||||
},
|
||||
"docs/design-system.md": {
|
||||
"mtime": 1786022513.9623306,
|
||||
"ast_hash": "421cd7e57f02d4b467f120ca6ddd7b6a",
|
||||
"semantic_hash": "421cd7e57f02d4b467f120ca6ddd7b6a"
|
||||
},
|
||||
"backend/api_test.go": {
|
||||
"mtime": 1786488480.637822,
|
||||
"ast_hash": "8e4b9293bc2e45ee3f42027315594fd5",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"backend/cover_test.go": {
|
||||
"mtime": 1786363889.552731,
|
||||
"ast_hash": "c7e313d6c92eb28e6d370e5e89035984",
|
||||
"semantic_hash": "c7e313d6c92eb28e6d370e5e89035984"
|
||||
},
|
||||
"backend/internal/api/handlers.go": {
|
||||
"mtime": 1786363889.552731,
|
||||
"ast_hash": "59e6b8767ab19839bb8f82891a7e4616",
|
||||
"semantic_hash": "59e6b8767ab19839bb8f82891a7e4616"
|
||||
},
|
||||
"backend/internal/httpmw/middleware.go": {
|
||||
"mtime": 1786216141.672644,
|
||||
"ast_hash": "385b36f58488b7e6d93eb6d6034e9ee3",
|
||||
"semantic_hash": "385b36f58488b7e6d93eb6d6034e9ee3"
|
||||
},
|
||||
"backend/internal/latest/browser.go": {
|
||||
"mtime": 1786469354.923288,
|
||||
"ast_hash": "ed129a7f00601ea90c877ff29fa21220",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"backend/internal/latest/browser_test.go": {
|
||||
"mtime": 1786262323.6964688,
|
||||
"ast_hash": "e900f92971486f47d7ef76e9a95217fe",
|
||||
"semantic_hash": "e900f92971486f47d7ef76e9a95217fe"
|
||||
},
|
||||
"backend/internal/latest/fetch.go": {
|
||||
"mtime": 1786446006.9219902,
|
||||
"ast_hash": "3e20ad86aa46783e9aa95c2b746551ee",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"backend/internal/latest/poller.go": {
|
||||
"mtime": 1786469641.8342345,
|
||||
"ast_hash": "44fef6074ac2eaffc8233f46aad5236b",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"backend/internal/latest/poller_test.go": {
|
||||
"mtime": 1786469644.9306462,
|
||||
"ast_hash": "64bc838c822f1bf33bbf9e291215454b",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"backend/internal/latest/sites.go": {
|
||||
"mtime": 1786469348.5394242,
|
||||
"ast_hash": "b744cc685363317a526cc3bebceea39e",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"backend/internal/latest/sites_test.go": {
|
||||
"mtime": 1786446006.9219902,
|
||||
"ast_hash": "eabca9014a306e3c71d238b0ae499f61",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"backend/internal/latest/smoke_image_test.go": {
|
||||
"mtime": 1786363889.5602942,
|
||||
"ast_hash": "db068cb59575f8c82669acbaf84bcaed",
|
||||
"semantic_hash": "db068cb59575f8c82669acbaf84bcaed"
|
||||
},
|
||||
"backend/internal/pgtest/pgtest.go": {
|
||||
"mtime": 1786216141.6766438,
|
||||
"ast_hash": "f60372d41516e66f7aaeb272da227d6e",
|
||||
"semantic_hash": "f60372d41516e66f7aaeb272da227d6e"
|
||||
},
|
||||
"backend/internal/session/session.go": {
|
||||
"mtime": 1786216141.6766438,
|
||||
"ast_hash": "9952474ffb22c825d6f075b866ee26f4",
|
||||
"semantic_hash": "9952474ffb22c825d6f075b866ee26f4"
|
||||
},
|
||||
"backend/internal/session/session_test.go": {
|
||||
"mtime": 1786216141.680644,
|
||||
"ast_hash": "37ffd00964e7a67350c68ed50c6503c5",
|
||||
"semantic_hash": "37ffd00964e7a67350c68ed50c6503c5"
|
||||
},
|
||||
"backend/internal/store/migrations/0001_bookmarks.sql": {
|
||||
"mtime": 1786216141.680644,
|
||||
"ast_hash": "f87ccfb2c25c43f93021177ced0bfae4",
|
||||
"semantic_hash": "f87ccfb2c25c43f93021177ced0bfae4"
|
||||
},
|
||||
"backend/internal/store/migrations/0002_series.sql": {
|
||||
"mtime": 1786216141.6820722,
|
||||
"ast_hash": "5dc98771e0c0f6e8416b434c280b0efb",
|
||||
"semantic_hash": "5dc98771e0c0f6e8416b434c280b0efb"
|
||||
},
|
||||
"backend/internal/store/migrations/0003_reader.sql": {
|
||||
"mtime": 1786216141.6820722,
|
||||
"ast_hash": "444a97799f38f6222f87d4e9fb2d6258",
|
||||
"semantic_hash": "444a97799f38f6222f87d4e9fb2d6258"
|
||||
},
|
||||
"backend/internal/store/migrations/0004_owner_bookmarks.sql": {
|
||||
"mtime": 1786216141.684644,
|
||||
"ast_hash": "e4fa900cc223865d3ecd4c60c5707a65",
|
||||
"semantic_hash": "e4fa900cc223865d3ecd4c60c5707a65"
|
||||
},
|
||||
"backend/internal/store/migrations/0005_sessions.sql": {
|
||||
"mtime": 1786216141.684644,
|
||||
"ast_hash": "5158887ebc57cf8c16a7b821b61cd760",
|
||||
"semantic_hash": "5158887ebc57cf8c16a7b821b61cd760"
|
||||
},
|
||||
"backend/internal/store/migrations/0006_reader_token_epoch.sql": {
|
||||
"mtime": 1786216141.684644,
|
||||
"ast_hash": "3093cc1c3aae0cd9643d04105d045402",
|
||||
"semantic_hash": "3093cc1c3aae0cd9643d04105d045402"
|
||||
},
|
||||
"backend/internal/store/sessions.go": {
|
||||
"mtime": 1786216141.684644,
|
||||
"ast_hash": "eee3510cc6172ef4b1da820474c26b01",
|
||||
"semantic_hash": "eee3510cc6172ef4b1da820474c26b01"
|
||||
},
|
||||
"backend/internal/store/sessions_test.go": {
|
||||
"mtime": 1786216141.684644,
|
||||
"ast_hash": "0b6764a0ee20f5cb7748eecd31a1d220",
|
||||
"semantic_hash": "0b6764a0ee20f5cb7748eecd31a1d220"
|
||||
},
|
||||
"backend/internal/store/store.go": {
|
||||
"mtime": 1786363889.5602942,
|
||||
"ast_hash": "54367a8ab043983e2491b2eb2650961c",
|
||||
"semantic_hash": "54367a8ab043983e2491b2eb2650961c"
|
||||
},
|
||||
"backend/internal/store/store_test.go": {
|
||||
"mtime": 1786363889.5602942,
|
||||
"ast_hash": "dc823fd77bcce2268114e31d759b20a5",
|
||||
"semantic_hash": "dc823fd77bcce2268114e31d759b20a5"
|
||||
},
|
||||
"backend/internal/userscript/userscript.go": {
|
||||
"mtime": 1786216141.692644,
|
||||
"ast_hash": "aa13a71b1c9eefe4930fd31f27722328",
|
||||
"semantic_hash": "aa13a71b1c9eefe4930fd31f27722328"
|
||||
},
|
||||
"backend/internal/userscript/userscript_test.go": {
|
||||
"mtime": 1786216141.692644,
|
||||
"ast_hash": "6c050968d7b8b67956da1a3136d2c3c7",
|
||||
"semantic_hash": "6c050968d7b8b67956da1a3136d2c3c7"
|
||||
},
|
||||
"backend/internal/web/discord.go": {
|
||||
"mtime": 1786216141.692644,
|
||||
"ast_hash": "69e4959c65fa67d7491aedf6a71bb575",
|
||||
"semantic_hash": "69e4959c65fa67d7491aedf6a71bb575"
|
||||
},
|
||||
"backend/internal/web/oauth_test.go": {
|
||||
"mtime": 1786216141.692644,
|
||||
"ast_hash": "a3bddeb70dd8d7eb14139da808723ea8",
|
||||
"semantic_hash": "a3bddeb70dd8d7eb14139da808723ea8"
|
||||
},
|
||||
"backend/internal/web/static/filter.js": {
|
||||
"mtime": 1786022513.9473197,
|
||||
"ast_hash": "b4ee3306201bfd88b148b96801972617",
|
||||
"semantic_hash": "b4ee3306201bfd88b148b96801972617"
|
||||
},
|
||||
"backend/internal/web/static/htmx.min.js": {
|
||||
"mtime": 1785873769.5512016,
|
||||
"ast_hash": "19a573773be4ca22570ca2f8543120c5",
|
||||
"semantic_hash": "19a573773be4ca22570ca2f8543120c5"
|
||||
},
|
||||
"backend/internal/web/web.go": {
|
||||
"mtime": 1786363889.5678573,
|
||||
"ast_hash": "8308949c658d3a08ce1c3cdbea6a907c",
|
||||
"semantic_hash": "8308949c658d3a08ce1c3cdbea6a907c"
|
||||
},
|
||||
"backend/reader_credential_test.go": {
|
||||
"mtime": 1786216141.700644,
|
||||
"ast_hash": "2a751ea6d9c06635aa3179db0aef1b2b",
|
||||
"semantic_hash": "2a751ea6d9c06635aa3179db0aef1b2b"
|
||||
},
|
||||
"chrome/entrypoint.sh": {
|
||||
"mtime": 1786363889.5678573,
|
||||
"ast_hash": "8008a187690764436540fab47ba0cfcc",
|
||||
"semantic_hash": "8008a187690764436540fab47ba0cfcc"
|
||||
},
|
||||
"userscript/novel-bookmark.user.js": {
|
||||
"mtime": 1786454823.798836,
|
||||
"ast_hash": "834effb0821f8d6c9f57f6554a5db462",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"userscript/test/novel-logic.test.js": {
|
||||
"mtime": 1786454566.070801,
|
||||
"ast_hash": "b25a7377af210dd0aff8284501fd0252",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".opencode/agent/implementer.md": {
|
||||
"mtime": 1785873769.5402634,
|
||||
"ast_hash": "000de469c18d68352027e10c8ce8acfb",
|
||||
"semantic_hash": "000de469c18d68352027e10c8ce8acfb"
|
||||
},
|
||||
".opencode/agent/reviewer.md": {
|
||||
"mtime": 1785873769.5416775,
|
||||
"ast_hash": "e44a2f6f624db044e19508bc5ab05592",
|
||||
"semantic_hash": "e44a2f6f624db044e19508bc5ab05592"
|
||||
},
|
||||
"CONTEXT.md": {
|
||||
"mtime": 1786466512.6719532,
|
||||
"ast_hash": "544b7d93f1d5cb9d347cb0b92fc3a709",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"CUTOVER.md": {
|
||||
"mtime": 1786216141.660644,
|
||||
"ast_hash": "6c6f3e4c4c2f57867894280bce728c50",
|
||||
"semantic_hash": "6c6f3e4c4c2f57867894280bce728c50"
|
||||
},
|
||||
"backend/AGENTS.md": {
|
||||
"mtime": 1786363889.552731,
|
||||
"ast_hash": "6356ee58447f299e5fa7aa83486bbbe0",
|
||||
"semantic_hash": "6356ee58447f299e5fa7aa83486bbbe0"
|
||||
},
|
||||
"backend/CLAUDE.md": {
|
||||
"mtime": 1786363889.552731,
|
||||
"ast_hash": "6356ee58447f299e5fa7aa83486bbbe0",
|
||||
"semantic_hash": "6356ee58447f299e5fa7aa83486bbbe0"
|
||||
},
|
||||
"backend/internal/web/templates/app.html": {
|
||||
"mtime": 1786216141.692644,
|
||||
"ast_hash": "3965e20e204afb71ba2a3aa86cb7c61c",
|
||||
"semantic_hash": "3965e20e204afb71ba2a3aa86cb7c61c"
|
||||
},
|
||||
"backend/internal/web/templates/card.html": {
|
||||
"mtime": 1786363889.5678573,
|
||||
"ast_hash": "ab83ae0dbb34fd40c146a7cc1263173e",
|
||||
"semantic_hash": "ab83ae0dbb34fd40c146a7cc1263173e"
|
||||
},
|
||||
"backend/internal/web/templates/chrome.html": {
|
||||
"mtime": 1786363889.5678573,
|
||||
"ast_hash": "d80b27cf3bd9d131075c485dd169dfcc",
|
||||
"semantic_hash": "d80b27cf3bd9d131075c485dd169dfcc"
|
||||
},
|
||||
"backend/internal/web/templates/icons.html": {
|
||||
"mtime": 1785873769.553139,
|
||||
"ast_hash": "8e10c507c32934a92463b4bca9e34fe6",
|
||||
"semantic_hash": "8e10c507c32934a92463b4bca9e34fe6"
|
||||
},
|
||||
"backend/internal/web/templates/list.html": {
|
||||
"mtime": 1786216141.696644,
|
||||
"ast_hash": "365548aace8c06559a1f66db0ae47256",
|
||||
"semantic_hash": "365548aace8c06559a1f66db0ae47256"
|
||||
},
|
||||
"backend/internal/web/templates/login.html": {
|
||||
"mtime": 1786216141.696644,
|
||||
"ast_hash": "bcc3101498a66cf8b79f9d97c6c9cd6b",
|
||||
"semantic_hash": "bcc3101498a66cf8b79f9d97c6c9cd6b"
|
||||
},
|
||||
"backend/internal/web/templates/readers.html": {
|
||||
"mtime": 1786216141.696644,
|
||||
"ast_hash": "c5034e76bd20a705d2799cb5ecb328f0",
|
||||
"semantic_hash": "c5034e76bd20a705d2799cb5ecb328f0"
|
||||
},
|
||||
"backend/internal/web/templates/setup.html": {
|
||||
"mtime": 1786216141.696644,
|
||||
"ast_hash": "72e93c0b827414063596f7338987d879",
|
||||
"semantic_hash": "72e93c0b827414063596f7338987d879"
|
||||
},
|
||||
"docs/adr/0001-postgresql-over-sqlite.md": {
|
||||
"mtime": 1786216141.704644,
|
||||
"ast_hash": "abfb08754cee58be67311377904f8ca4",
|
||||
"semantic_hash": "abfb08754cee58be67311377904f8ca4"
|
||||
},
|
||||
"docs/adr/0002-discord-oauth-no-passwords-no-email.md": {
|
||||
"mtime": 1786216141.7071996,
|
||||
"ast_hash": "852a04d86659385085da6ffc8b489933",
|
||||
"semantic_hash": "852a04d86659385085da6ffc8b489933"
|
||||
},
|
||||
"docs/adr/0003-series-shared-and-poll-owned.md": {
|
||||
"mtime": 1786216141.7071996,
|
||||
"ast_hash": "58bb4d24e20f6f7adc1a8b9cf7967c3f",
|
||||
"semantic_hash": "58bb4d24e20f6f7adc1a8b9cf7967c3f"
|
||||
},
|
||||
"docs/adr/0004-wire-format-does-not-mirror-the-schema.md": {
|
||||
"mtime": 1786216141.7071996,
|
||||
"ast_hash": "a6ea2770dec2156f78a65b35ba06902a",
|
||||
"semantic_hash": "a6ea2770dec2156f78a65b35ba06902a"
|
||||
},
|
||||
"docs/agents/domain.md": {
|
||||
"mtime": 1786216141.7071996,
|
||||
"ast_hash": "6f99318ac6cb9825b613bfde55d76091",
|
||||
"semantic_hash": "6f99318ac6cb9825b613bfde55d76091"
|
||||
},
|
||||
"docs/agents/issue-tracker.md": {
|
||||
"mtime": 1786216141.7087462,
|
||||
"ast_hash": "1342e66ccb84a84fd579fb6dc0b8243a",
|
||||
"semantic_hash": "1342e66ccb84a84fd579fb6dc0b8243a"
|
||||
},
|
||||
"docs/agents/triage-labels.md": {
|
||||
"mtime": 1786216141.7087462,
|
||||
"ast_hash": "69114d07ed792d6bb1d13758ba5435e1",
|
||||
"semantic_hash": "69114d07ed792d6bb1d13758ba5435e1"
|
||||
},
|
||||
"userscript/AGENTS.md": {
|
||||
"mtime": 1786454605.6727684,
|
||||
"ast_hash": "e276ffe9a6e7b55fd3235466a1995c22",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"userscript/CLAUDE.md": {
|
||||
"mtime": 1786454605.6727684,
|
||||
"ast_hash": "e276ffe9a6e7b55fd3235466a1995c22",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"backend/internal/web/static/login-art.png": {
|
||||
"mtime": 1786022513.9585779,
|
||||
"ast_hash": "05d7863cba344a946256719a0c9ef959",
|
||||
"semantic_hash": "05d7863cba344a946256719a0c9ef959"
|
||||
},
|
||||
"backend/internal/web/static/logo.svg": {
|
||||
"mtime": 1786022513.9585779,
|
||||
"ast_hash": "d0d34d0f08a25b53176cc55989b7babe",
|
||||
"semantic_hash": "d0d34d0f08a25b53176cc55989b7babe"
|
||||
},
|
||||
"backend/internal/store/migrations/0007_covers.sql": {
|
||||
"mtime": 1786262323.6964688,
|
||||
"ast_hash": "6033ce0701be1236ed175362363bd96c",
|
||||
"semantic_hash": "6033ce0701be1236ed175362363bd96c"
|
||||
},
|
||||
"docs/adr/0005-on-demand-browser.md": {
|
||||
"mtime": 1786292465.8305523,
|
||||
"ast_hash": "8cf2fb8c0a66c2c7d82ae433b99ca42b",
|
||||
"semantic_hash": "8cf2fb8c0a66c2c7d82ae433b99ca42b"
|
||||
},
|
||||
"chrome/docker-compose.yml": {
|
||||
"mtime": 1786292465.8305523,
|
||||
"ast_hash": "5605599395a3f085904e78a2bfec1e58",
|
||||
"semantic_hash": "5605599395a3f085904e78a2bfec1e58"
|
||||
},
|
||||
"docs/adr/0006-browser-on-the-home-machine.md": {
|
||||
"mtime": 1786292465.8305523,
|
||||
"ast_hash": "dfd6bbc045d23315f2942ee8d98eb7db",
|
||||
"semantic_hash": "dfd6bbc045d23315f2942ee8d98eb7db"
|
||||
},
|
||||
"docs/adr/0007-backend-hosts-cover-bytes.md": {
|
||||
"mtime": 1786292465.8305523,
|
||||
"ast_hash": "b19e38045b3dcda7dd59634ed9227a68",
|
||||
"semantic_hash": "b19e38045b3dcda7dd59634ed9227a68"
|
||||
},
|
||||
"backend/internal/store/migrations/0008_filesystem_covers.sql": {
|
||||
"mtime": 1786363889.5602942,
|
||||
"ast_hash": "46cf7822d4f667e3cab36b547abe5e97",
|
||||
"semantic_hash": "46cf7822d4f667e3cab36b547abe5e97"
|
||||
},
|
||||
"backend/internal/latest/cover.go": {
|
||||
"mtime": 1786363889.5565126,
|
||||
"ast_hash": "e6749cfe3cd7c2e71d4392dde84f55f9",
|
||||
"semantic_hash": "e6749cfe3cd7c2e71d4392dde84f55f9"
|
||||
},
|
||||
"backend/internal/latest/cover_fetch_test.go": {
|
||||
"mtime": 1786363889.5565126,
|
||||
"ast_hash": "60d9eb7c59a3751baf4f31c7655217e7",
|
||||
"semantic_hash": "60d9eb7c59a3751baf4f31c7655217e7"
|
||||
},
|
||||
"backend/internal/latest/acquire.go": {
|
||||
"mtime": 1786468950.9432797,
|
||||
"ast_hash": "6c1ad34bbe9f5b49d0fd1eae9093f55d",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"backend/internal/latest/acquire_test.go": {
|
||||
"mtime": 1786363889.5565126,
|
||||
"ast_hash": "7bd9f41814f6bf59d8998dfb81cf990a",
|
||||
"semantic_hash": "7bd9f41814f6bf59d8998dfb81cf990a"
|
||||
},
|
||||
"backend/internal/store/migrations/0009_series_cover_address.sql": {
|
||||
"mtime": 1786363889.5602942,
|
||||
"ast_hash": "4c1f6328b2e1a95828fad6d88d474c2d",
|
||||
"semantic_hash": "4c1f6328b2e1a95828fad6d88d474c2d"
|
||||
},
|
||||
".claude/skills/implement-tickets/SKILL.md": {
|
||||
"mtime": 1786417841.7117643,
|
||||
"ast_hash": "3060e32e19cc571d91871a98f18afe53",
|
||||
"semantic_hash": "3060e32e19cc571d91871a98f18afe53"
|
||||
},
|
||||
".omp/agents/ticket-implementer.md": {
|
||||
"mtime": 1786417841.7163916,
|
||||
"ast_hash": "0150a46c0d21c572b71b3d87d21ac925",
|
||||
"semantic_hash": "0150a46c0d21c572b71b3d87d21ac925"
|
||||
},
|
||||
"docs/adr/0008-series-identity-is-discovered-not-derived.md": {
|
||||
"mtime": 1786417841.7163916,
|
||||
"ast_hash": "3ce6d64ef6a8a39f27c257b39065389e",
|
||||
"semantic_hash": "3ce6d64ef6a8a39f27c257b39065389e"
|
||||
},
|
||||
"docs/research/lightnovelworld-chapter-vs-series-slug.md": {
|
||||
"mtime": 1786417841.7163916,
|
||||
"ast_hash": "74a4e538875f0a7e8ca3d5dc48c0bb53",
|
||||
"semantic_hash": "74a4e538875f0a7e8ca3d5dc48c0bb53"
|
||||
},
|
||||
"backend/internal/latest/smoke_lnw_test.go": {
|
||||
"mtime": 1786446006.9219902,
|
||||
"ast_hash": "2d65da8a081759172918fdf159760f45",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"docs/adr/0009-a-site-answers-questions-its-own-way.md": {
|
||||
"mtime": 1786469367.6875648,
|
||||
"ast_hash": "8039012a5b6de2359ff1a47079f51b66",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"backend/internal/latest/read.go": {
|
||||
"mtime": 1786469329.7445614,
|
||||
"ast_hash": "3cf29046ddaef39fafb1df70b9f9ae8c",
|
||||
"semantic_hash": ""
|
||||
}
|
||||
}
|
||||
+12
-3
@@ -70,10 +70,19 @@ Guidance for OpenCode (and Claude Code) working under `userscript/`. See root `A
|
||||
Behind a Cloudflare JS challenge no TLS fingerprint
|
||||
clears, so the backend polls it through the headless browser.
|
||||
- **lightnovelworld.net** (novel script): series `/novel/<slug>/`, chapter
|
||||
`/<slug>-chapter-<n>/` — flat, at the site root. `h1.entry-title` is the clean
|
||||
title on a series page and `<Title> Chapter <n>` on a chapter page. Its series
|
||||
page lists every chapter with an
|
||||
`/<slug>-chapter-<n>/` — flat, at the site root. The chapter path's slug is a
|
||||
Chapter Slug, not an identity: the Series address is read off the page's
|
||||
`a[aria-label='All Chapter']` (fallback: the BreadcrumbList's second crumb),
|
||||
and a Series may publish under several Chapter Slugs. A chapter page with no
|
||||
pointer resolves to `other`, so no Bookmark is offered. `h1.entry-title` is
|
||||
the clean title on a series page and `<Title> Chapter <n>` on a chapter page.
|
||||
Its series page lists every chapter with an
|
||||
absolute href, so the backend polls it with the plain TLS client.
|
||||
The client performs no latest-chapter scan for this Site: the Poll's
|
||||
one-hour cooldown dominates the client's four-hour throttle, so a scan
|
||||
would add no freshness, and the page's wpdiscuz thread is a public write
|
||||
surface a scan would have to truncate at. `computeLatestChapter` yields
|
||||
null here and `backgroundRefreshLatest` skips the Site before any fetch.
|
||||
|
||||
### Second script: `novel-bookmark.user.js`
|
||||
|
||||
|
||||
@@ -6,7 +6,6 @@
|
||||
// @author you
|
||||
// @downloadURL https://bookmark-api.violetcrown.my.id/u/__API_TOKEN__/manga-bookmark.user.js
|
||||
// @updateURL https://bookmark-api.violetcrown.my.id/u/__API_TOKEN__/manga-bookmark.user.js
|
||||
// @match https://asuracomic.net/*
|
||||
// @match https://asurascans.com/*
|
||||
// @match https://demonicscans.org/*
|
||||
// @match https://comix.to/*
|
||||
@@ -112,12 +111,10 @@
|
||||
|
||||
const asura = {
|
||||
site: "asura",
|
||||
// asuracomic.net deep links 301 to the asurascans.com *root*, dropping the
|
||||
// path, and that happens at the edge before this script gets a document —
|
||||
// so those URLs cannot be handled here at all (checked 2026-07-25). It stays
|
||||
// matched in case the redirect starts preserving paths again; until then,
|
||||
// reach series through asurascans.com.
|
||||
matches: (loc) => /(^|\.)asurascans\.com$|(^|\.)asuracomic\.net$/.test(loc.hostname),
|
||||
// asuracomic.net is not matched: its deep links 301 to the asurascans.com
|
||||
// *root* at the edge, discarding the path, so this script never sees a
|
||||
// series document there (re-checked 2026-07-25).
|
||||
matches: (loc) => /(^|\.)asurascans\.com$/.test(loc.hostname),
|
||||
detect(loc) {
|
||||
const path = loc.pathname;
|
||||
// /comics/<slug-hash>/chapter/<n>
|
||||
|
||||
@@ -132,9 +132,30 @@
|
||||
},
|
||||
};
|
||||
|
||||
// The host shape this adapter owns. One definition so the pointer validator
|
||||
// and `matches` cannot drift apart (a leading subdomain is allowed).
|
||||
const lnwHostRe = /(^|\.)lightnovelworld\.net$/;
|
||||
|
||||
// A chapter page's pointer is its own link back to its Series. The href is
|
||||
// page markup, so validate before trusting: resolve it against the page
|
||||
// address, require the host above and the /novel/<slug>/ Series path.
|
||||
// Anything else is not a pointer.
|
||||
function seriesIdFromLnwPointer(href, base) {
|
||||
if (!href) return null;
|
||||
let u;
|
||||
try {
|
||||
u = new URL(href, base);
|
||||
} catch (e) {
|
||||
return null;
|
||||
}
|
||||
if (!lnwHostRe.test(u.hostname)) return null;
|
||||
const m = u.pathname.match(/^\/novel\/([^/]+)\/?$/);
|
||||
return m ? m[1] : null;
|
||||
}
|
||||
|
||||
const lightnovelworld = {
|
||||
site: "lightnovelworld",
|
||||
matches: (loc) => /(^|\.)lightnovelworld\.net$/.test(loc.hostname),
|
||||
matches: (loc) => lnwHostRe.test(loc.hostname),
|
||||
detect(loc) {
|
||||
const path = loc.pathname;
|
||||
// /<slug>-chapter-<n>/ — flat, at the site root, not under /novel/. The
|
||||
@@ -142,16 +163,38 @@
|
||||
// words still resolves to the right series.
|
||||
let m = path.match(/^\/(.+)-chapter-([0-9]+(?:\.[0-9]+)?)\/?$/);
|
||||
if (m) {
|
||||
// The address is not the identity on this Site: the slug in the path
|
||||
// is a Chapter Slug, which can differ from the Series slug and is
|
||||
// never computable from it. The Series address is read from the
|
||||
// page's own pointer — a silent fallback to derivation is the defect
|
||||
// this replaced, not a safety net.
|
||||
const chapterSlug = m[1];
|
||||
const pointer = document.querySelector("a[aria-label='All Chapter']");
|
||||
let seriesId = pointer ? seriesIdFromLnwPointer(pointer.getAttribute("href"), loc.href) : null;
|
||||
if (!seriesId) {
|
||||
// Fallback: the microdata breadcrumb's second crumb is the Series.
|
||||
// Scoped to the BreadcrumbList because itemprop="item" is not
|
||||
// unique to it (the header nav uses microdata too).
|
||||
const crumb = document.querySelector(
|
||||
'[itemtype="http://schema.org/BreadcrumbList"] a[itemprop="item"][href*="/novel/"]'
|
||||
);
|
||||
seriesId = crumb ? seriesIdFromLnwPointer(crumb.getAttribute("href"), loc.href) : null;
|
||||
}
|
||||
// Neither pointer present, nor either pointing at a /novel/<slug>/
|
||||
// address on this host: not a page the script understands, so no
|
||||
// Bookmark under an invented identity.
|
||||
if (!seriesId) return { type: "other" };
|
||||
const num = parseFloat(m[2]);
|
||||
const h1 = document.querySelector("h1.entry-title");
|
||||
const heading = h1 ? h1.textContent || "" : "";
|
||||
return {
|
||||
type: "chapter",
|
||||
site: this.site,
|
||||
seriesId: m[1],
|
||||
seriesId: seriesId,
|
||||
chapterSlug: chapterSlug,
|
||||
// The heading is "<Series> Chapter <n>"; drop the suffix.
|
||||
title: heading.replace(/\s*Chapter\s+[0-9.]+\s*$/i, "").trim(),
|
||||
seriesUrl: "https://lightnovelworld.net/novel/" + m[1] + "/",
|
||||
seriesUrl: "https://lightnovelworld.net/novel/" + seriesId + "/",
|
||||
chapterLabel: "Chapter " + m[2],
|
||||
chapterNum: isNaN(num) ? null : num,
|
||||
chapterUrl: loc.href,
|
||||
@@ -165,6 +208,7 @@
|
||||
type: "series",
|
||||
site: this.site,
|
||||
seriesId: m[1],
|
||||
chapterSlug: null,
|
||||
title: h1 ? (h1.textContent || "").trim() : "",
|
||||
seriesUrl: "https://lightnovelworld.net/novel/" + m[1] + "/",
|
||||
chapterLabel: null,
|
||||
@@ -174,12 +218,6 @@
|
||||
}
|
||||
return { type: "other" };
|
||||
},
|
||||
latestChapterFromAnchors(anchors, seriesId) {
|
||||
return maxChapter(
|
||||
anchors,
|
||||
new RegExp("lightnovelworld\\.net/" + escapeRe(seriesId) + "-chapter-([0-9.]+)/")
|
||||
);
|
||||
},
|
||||
};
|
||||
|
||||
const ADAPTERS = [novelfull, lightnovelworld];
|
||||
@@ -200,12 +238,109 @@
|
||||
return ADAPTERS.find((a) => a.site === site) || null;
|
||||
}
|
||||
|
||||
// A stale lightnovelworld row is one keyed under the Chapter Slug its
|
||||
// address was built from, while the page's pointer names a different
|
||||
// Series. All three storage sites must move together: the cache row alone
|
||||
// leaves a queued write replaying under a key whose row no longer exists
|
||||
// (that write is silently lost), and a last-checked timestamp left behind
|
||||
// re-fetches the repaired row on the next visit. A repair must never cost a
|
||||
// Reader the chapter they were tracking, so when a row already holds the
|
||||
// repaired key the two rows merge and the farther-ahead progress wins.
|
||||
// Pure: no storage access, no module state, no clock.
|
||||
function repairLnwStaleRow(list, queue, lastChecked, page) {
|
||||
if (
|
||||
!page ||
|
||||
page.site !== "lightnovelworld" ||
|
||||
!page.chapterSlug ||
|
||||
!page.seriesId ||
|
||||
page.chapterSlug === page.seriesId
|
||||
) {
|
||||
return { list, queue, lastChecked };
|
||||
}
|
||||
const oldKey = "lightnovelworld:" + page.chapterSlug;
|
||||
const newKey = "lightnovelworld:" + page.seriesId;
|
||||
const stale = list.find((b) => b.key === oldKey);
|
||||
if (!stale) return { list, queue, lastChecked };
|
||||
|
||||
let listOut;
|
||||
const existing = list.find((b) => b.key === newKey);
|
||||
if (existing) {
|
||||
// Duplicate case (spec user story 2): one row must survive, and it is
|
||||
// the one already under the repaired key — with the stale row's
|
||||
// progress carried across when it is ahead, favourite OR'd, and the
|
||||
// stronger lifecycle bucket kept (finished > archived > reading, so a
|
||||
// merge can never silently un-archive or un-finish a row).
|
||||
const merged = Object.assign({}, existing);
|
||||
if (
|
||||
existing.last_chapter_num == null ||
|
||||
(stale.last_chapter_num != null && stale.last_chapter_num > existing.last_chapter_num)
|
||||
) {
|
||||
merged.last_chapter = stale.last_chapter;
|
||||
merged.last_chapter_num = stale.last_chapter_num;
|
||||
merged.last_chapter_url = stale.last_chapter_url;
|
||||
}
|
||||
merged.favorite = !!(existing.favorite || stale.favorite);
|
||||
const rank = (s) => ({ finished: 2, archived: 1, reading: 0 }[s || "reading"] || 0);
|
||||
merged.status = rank(stale.status) > rank(existing.status) ? stale.status : existing.status;
|
||||
merged.updated_at = Math.max(existing.updated_at || 0, stale.updated_at || 0);
|
||||
listOut = list.filter((b) => b.key !== oldKey).map((b) => (b.key === newKey ? merged : b));
|
||||
} else {
|
||||
listOut = list.map((b) =>
|
||||
b.key === oldKey
|
||||
? Object.assign({}, b, {
|
||||
key: newKey,
|
||||
series_id: page.seriesId,
|
||||
series_url: page.seriesUrl,
|
||||
})
|
||||
: b
|
||||
);
|
||||
}
|
||||
|
||||
let queueOut = queue;
|
||||
const oldEntry = queue.find((e) => e.key === oldKey);
|
||||
if (oldEntry) {
|
||||
const survivor = queue.find((x) => x.key === newKey);
|
||||
if (!survivor) {
|
||||
queueOut = queue.map((e) => (e.key === oldKey ? Object.assign({}, e, { key: newKey }) : e));
|
||||
} else {
|
||||
// Both keys hold a marker and one row survives, so the two collapse
|
||||
// into one entry under the repaired key (the queue's one-entry-per-key
|
||||
// invariant). sendStatus is sticky — an archive intent from either
|
||||
// marker survives, the queue's own rule — and the worse attempts
|
||||
// count wins. The stale-key marker's op is dropped: the row it
|
||||
// described is retired by the merge itself.
|
||||
queueOut = queue
|
||||
.filter((e) => e.key !== oldKey && e.key !== newKey)
|
||||
.concat([
|
||||
{
|
||||
key: newKey,
|
||||
op: survivor.op,
|
||||
sendStatus: survivor.sendStatus || oldEntry.sendStatus,
|
||||
attempts: Math.max(survivor.attempts || 0, oldEntry.attempts || 0),
|
||||
},
|
||||
]);
|
||||
}
|
||||
}
|
||||
|
||||
let lastCheckedOut = lastChecked;
|
||||
if (oldKey in lastChecked) {
|
||||
lastCheckedOut = Object.assign({}, lastChecked);
|
||||
// Max, not overwrite: a timestamp already under the repaired key must
|
||||
// not be rolled back to an older one.
|
||||
lastCheckedOut[newKey] = Math.max(lastCheckedOut[newKey] || 0, lastCheckedOut[oldKey]);
|
||||
delete lastCheckedOut[oldKey];
|
||||
}
|
||||
|
||||
return { list: listOut, queue: queueOut, lastChecked: lastCheckedOut };
|
||||
}
|
||||
|
||||
// Highest chapter the site lists, or null when the markup yields nothing.
|
||||
// seriesId is only consulted by adapters whose pages carry other series'
|
||||
// chapter links; the rest ignore it.
|
||||
// chapter links. An adapter without a scanner (lightnovelworld — see the
|
||||
// AGENTS.md entry) yields null, not an error.
|
||||
function computeLatestChapter(site, anchors, seriesId) {
|
||||
const a = adapterFor(site);
|
||||
return a ? a.latestChapterFromAnchors(anchors, seriesId) : null;
|
||||
return a && a.latestChapterFromAnchors ? a.latestChapterFromAnchors(anchors, seriesId) : null;
|
||||
}
|
||||
|
||||
function currentSite() {
|
||||
@@ -721,6 +856,14 @@
|
||||
const site = currentSite();
|
||||
if (!site) return;
|
||||
|
||||
// A Site with no client scanner (lightnovelworld) is refreshed by the Poll
|
||||
// on a cooldown shorter than the client throttle; fetching its pages here
|
||||
// would be a megabyte-scale request whose result is discarded. Skipping
|
||||
// before the due filter records no freshness timestamp and consumes no
|
||||
// per-navigation batch slot.
|
||||
const adapter = adapterFor(site);
|
||||
if (!adapter || !adapter.latestChapterFromAnchors) return;
|
||||
|
||||
const checked = loadLastChecked();
|
||||
const now = Date.now();
|
||||
const due = state.list
|
||||
@@ -731,7 +874,6 @@
|
||||
.slice(0, LATEST_CHECK_BATCH);
|
||||
if (due.length === 0) return;
|
||||
|
||||
const adapter = adapterFor(site);
|
||||
for (const bm of due) {
|
||||
// Recorded even when the fetch fails, so a broken series is retried on
|
||||
// the next throttle window rather than on every single page load.
|
||||
@@ -1199,7 +1341,8 @@
|
||||
// Refresh + navigation
|
||||
// ============================================================
|
||||
|
||||
async function refresh() {
|
||||
async function refresh(awaitFirst) {
|
||||
if (awaitFirst) await awaitFirst; // repair sync lands before we adopt the server's view of it
|
||||
await drain(); // push what we owe before adopting the server's view of it
|
||||
loading = true;
|
||||
render();
|
||||
@@ -1214,9 +1357,41 @@
|
||||
}
|
||||
}
|
||||
|
||||
// Runs the repair on every navigation, silently: rewrites a stale
|
||||
// lightnovelworld row to the page's discovered identity, persists all three
|
||||
// sites, and syncs through the queue-backed path so an offline repair parks
|
||||
// and replays later. No toast — the Reader never asked for this. Returns
|
||||
// the sync promise (or null when nothing changed) so init can hold the
|
||||
// server view until the repair has landed. When the transform returns its
|
||||
// inputs unchanged there is nothing to do.
|
||||
function applyLnwStaleRowRepair() {
|
||||
const lastChecked = loadLastChecked();
|
||||
const out = repairLnwStaleRow(state.list, queue, lastChecked, state.page);
|
||||
if (out.list === state.list && out.queue === queue && out.lastChecked === lastChecked) {
|
||||
return null;
|
||||
}
|
||||
state.list = out.list;
|
||||
reindex(); // byKey answers under the old key until rebuilt
|
||||
saveCache(state.list);
|
||||
queue.splice(0, queue.length, ...out.queue); // closures hold this array instance
|
||||
saveQueue(queue);
|
||||
saveLastChecked(out.lastChecked);
|
||||
// Queue-backed sync, outcome swallowed: the repaired row is PUT (with its
|
||||
// bucket when archived — the only status a userscript write may send) and
|
||||
// the old server bookmark is deleted so no duplicate survives on the
|
||||
// wire. The delete parks under the retired key while offline — the
|
||||
// teardown of the old identity, not a new record of the Chapter Slug.
|
||||
const key = keyOf(state.page);
|
||||
return Promise.all([
|
||||
pushBookmark(key, statusOf(state.byKey[key]) === "archived"),
|
||||
pushDelete("lightnovelworld:" + state.page.chapterSlug),
|
||||
]).catch(() => {});
|
||||
}
|
||||
|
||||
let lastUrl = location.href;
|
||||
function onNavigate() {
|
||||
state.page = detect();
|
||||
applyLnwStaleRowRepair();
|
||||
render();
|
||||
maybeAutoUpdate();
|
||||
maybeCaptureLatestOnSeriesPage();
|
||||
@@ -1329,13 +1504,14 @@
|
||||
function init() {
|
||||
buildUI();
|
||||
state.page = detect();
|
||||
const repair = applyLnwStaleRowRepair();
|
||||
render();
|
||||
installNavWatcher();
|
||||
installLongPress();
|
||||
window.addEventListener("online", drain); // signal returned while the page stayed open
|
||||
// Sync first: both auto-record and the latest-chapter checks below need to
|
||||
// know which series are bookmarked and how fresh they are.
|
||||
refresh().then(() => {
|
||||
refresh(repair).then(() => {
|
||||
maybeAutoUpdate();
|
||||
maybeCaptureLatestOnSeriesPage();
|
||||
backgroundRefreshLatest();
|
||||
@@ -1576,7 +1752,7 @@
|
||||
// Exposes pure logic only — see userscript/test/novel-logic.test.js.
|
||||
// ============================================================
|
||||
if (typeof window === "undefined" && typeof module === "object" && module.exports) {
|
||||
module.exports = { novelfull, lightnovelworld, anchorsFromHTML, statusOf, kindOf, maxChapter, escapeRe };
|
||||
module.exports = { novelfull, lightnovelworld, anchorsFromHTML, statusOf, kindOf, maxChapter, escapeRe, computeLatestChapter, repairLnwStaleRow };
|
||||
}
|
||||
|
||||
// ============================================================
|
||||
|
||||
@@ -120,6 +120,14 @@ test("asura.detect returns other for non-series paths", () => {
|
||||
assert.equal(asura.detect(loc("https://asurascans.com/bookmarks")).type, "other");
|
||||
});
|
||||
|
||||
test("asura.matches accepts only asurascans.com", () => {
|
||||
assert.equal(asura.matches({ hostname: "asurascans.com" }), true);
|
||||
assert.equal(asura.matches({ hostname: "www.asurascans.com" }), true);
|
||||
// Dead domain: deep links 301 to the asurascans.com root, discarding the path.
|
||||
assert.equal(asura.matches({ hostname: "asuracomic.net" }), false);
|
||||
assert.equal(asura.matches({ hostname: "asurascans.com.evil.example" }), false);
|
||||
});
|
||||
|
||||
test("asura.latestChapterFromAnchors takes the highest and skips the First Chapter shortcut", () => {
|
||||
const best = asura.latestChapterFromAnchors([
|
||||
{ href: "/comics/solo-leveling-059befe1/chapter/1", text: "Chapter 1" },
|
||||
|
||||
@@ -20,8 +20,13 @@ globalThis.location = { href: "about:blank", hostname: "", pathname: "/", origin
|
||||
|
||||
let metaTags = {};
|
||||
let elements = {};
|
||||
// Attribute selectors (the lightnovelworld Series pointer) answer with an
|
||||
// element exposing getAttribute, like the meta branch below.
|
||||
let attrEls = {};
|
||||
globalThis.document = {
|
||||
querySelector(sel) {
|
||||
const attr = attrEls[sel];
|
||||
if (attr != null) return attr;
|
||||
const m = sel.match(/^meta\[property="([^"]+)"\]$/);
|
||||
if (m) {
|
||||
const v = metaTags[m[1]];
|
||||
@@ -40,8 +45,10 @@ globalThis.document = {
|
||||
const {
|
||||
novelfull,
|
||||
lightnovelworld,
|
||||
computeLatestChapter,
|
||||
kindOf,
|
||||
maxChapter,
|
||||
repairLnwStaleRow,
|
||||
} = require("../novel-bookmark.user.js");
|
||||
|
||||
function loc(href) {
|
||||
@@ -52,6 +59,7 @@ function loc(href) {
|
||||
function reset() {
|
||||
metaTags = {};
|
||||
elements = {};
|
||||
attrEls = {};
|
||||
}
|
||||
|
||||
// ============================================================
|
||||
@@ -115,11 +123,18 @@ test("lightnovelworld.detect reads a series page", () => {
|
||||
assert.equal(p.site, "lightnovelworld");
|
||||
assert.equal(p.seriesId, "a-will-eternal");
|
||||
assert.equal(p.title, "A Will Eternal");
|
||||
// A series page's address *is* its identity; there is no Chapter Slug.
|
||||
assert.equal(p.chapterSlug, null);
|
||||
});
|
||||
|
||||
test("lightnovelworld.detect strips the chapter suffix off the heading", () => {
|
||||
reset();
|
||||
elements = { "h1.entry-title": "A Will Eternal Chapter 1298" };
|
||||
attrEls = {
|
||||
"a[aria-label='All Chapter']": {
|
||||
getAttribute: () => "https://lightnovelworld.net/novel/a-will-eternal/",
|
||||
},
|
||||
};
|
||||
const url = "https://lightnovelworld.net/a-will-eternal-chapter-1298/";
|
||||
const p = lightnovelworld.detect(loc(url));
|
||||
assert.equal(p.type, "chapter");
|
||||
@@ -130,23 +145,210 @@ test("lightnovelworld.detect strips the chapter suffix off the heading", () => {
|
||||
assert.equal(p.seriesUrl, "https://lightnovelworld.net/novel/a-will-eternal/");
|
||||
});
|
||||
|
||||
test("lightnovelworld.detect reads the Series identity from the page pointer on a chapter page", () => {
|
||||
reset();
|
||||
elements = { "h1.entry-title": "A Will Eternal Chapter 1298" };
|
||||
attrEls = {
|
||||
// Verbatim from the real page (research note §3): the All Chapter anchor
|
||||
// carries the absolute Series address.
|
||||
"a[aria-label='All Chapter']": {
|
||||
getAttribute: () => "https://lightnovelworld.net/novel/a-will-eternal/",
|
||||
},
|
||||
};
|
||||
const url = "https://lightnovelworld.net/a-will-eternal-chapter-1298/";
|
||||
const p = lightnovelworld.detect(loc(url));
|
||||
assert.equal(p.type, "chapter");
|
||||
assert.equal(p.seriesId, "a-will-eternal");
|
||||
assert.equal(p.seriesUrl, "https://lightnovelworld.net/novel/a-will-eternal/");
|
||||
assert.equal(p.chapterSlug, "a-will-eternal");
|
||||
assert.equal(p.chapterNum, 1298);
|
||||
});
|
||||
|
||||
test("lightnovelworld.detect falls back to the breadcrumb when the pointer is absent", () => {
|
||||
reset();
|
||||
elements = { "h1.entry-title": "My Longevity Simulation Chapter 1" };
|
||||
const url = "https://lightnovelworld.net/my-longevity-simulation-chapter-1/";
|
||||
// Same divergent page, both pointers from the real markup (research note
|
||||
// §3): breadcrumb position 2 must say what the All Chapter anchor says.
|
||||
attrEls = {
|
||||
"a[aria-label='All Chapter']": {
|
||||
getAttribute: () => "https://lightnovelworld.net/novel/immortality-simulator/",
|
||||
},
|
||||
};
|
||||
const viaPointer = lightnovelworld.detect(loc(url));
|
||||
attrEls = {
|
||||
'[itemtype="http://schema.org/BreadcrumbList"] a[itemprop="item"][href*="/novel/"]': {
|
||||
getAttribute: () => "https://lightnovelworld.net/novel/immortality-simulator/",
|
||||
},
|
||||
};
|
||||
const viaBreadcrumb = lightnovelworld.detect(loc(url));
|
||||
assert.equal(viaBreadcrumb.type, "chapter");
|
||||
assert.equal(viaBreadcrumb.seriesId, viaPointer.seriesId);
|
||||
assert.equal(viaBreadcrumb.seriesUrl, viaPointer.seriesUrl);
|
||||
assert.equal(viaBreadcrumb.seriesId, "immortality-simulator");
|
||||
assert.equal(viaBreadcrumb.chapterSlug, "my-longevity-simulation");
|
||||
});
|
||||
|
||||
test("lightnovelworld.detect resolves a relative pointer against the page address", () => {
|
||||
// Defensive: every measured page ships an absolute pointer href, but a
|
||||
// theme change could go relative — the pointer is still this page's own
|
||||
// link back to its Series.
|
||||
reset();
|
||||
elements = { "h1.entry-title": "A Will Eternal Chapter 1298" };
|
||||
attrEls = {
|
||||
"a[aria-label='All Chapter']": { getAttribute: () => "/novel/a-will-eternal/" },
|
||||
};
|
||||
const p = lightnovelworld.detect(loc("https://lightnovelworld.net/a-will-eternal-chapter-1298/"));
|
||||
assert.equal(p.type, "chapter");
|
||||
assert.equal(p.seriesId, "a-will-eternal");
|
||||
assert.equal(p.seriesUrl, "https://lightnovelworld.net/novel/a-will-eternal/");
|
||||
});
|
||||
|
||||
test("lightnovelworld.detect rejects a pointer that is not a /novel/ address on its own host", () => {
|
||||
// Counterfactual pointers, exercising the criterion that a pointer "present
|
||||
// but not parseable as /novel/<slug>/ on lightnovelworld.net" must not
|
||||
// become a seriesUrl the backend is later asked to Poll: off-host, and
|
||||
// on-host but the wrong path shape.
|
||||
reset();
|
||||
elements = { "h1.entry-title": "My Longevity Simulation Chapter 1" };
|
||||
const url = "https://lightnovelworld.net/my-longevity-simulation-chapter-1/";
|
||||
attrEls = {
|
||||
"a[aria-label='All Chapter']": {
|
||||
getAttribute: () => "https://evil.example/novel/immortality-simulator/",
|
||||
},
|
||||
};
|
||||
assert.equal(lightnovelworld.detect(loc(url)).type, "other");
|
||||
attrEls = {
|
||||
"a[aria-label='All Chapter']": {
|
||||
getAttribute: () => "https://lightnovelworld.net/fiction/immortality-simulator/",
|
||||
},
|
||||
};
|
||||
assert.equal(lightnovelworld.detect(loc(url)).type, "other");
|
||||
});
|
||||
|
||||
test("lightnovelworld.detect resolves to other when the page carries no pointer", () => {
|
||||
reset();
|
||||
elements = { "h1.entry-title": "My Longevity Simulation Chapter 1" };
|
||||
const p = lightnovelworld.detect(
|
||||
loc("https://lightnovelworld.net/my-longevity-simulation-chapter-1/")
|
||||
);
|
||||
assert.equal(p.type, "other");
|
||||
});
|
||||
|
||||
test("lightnovelworld pins the divergent novel: the pointer's slug wins over the address's", () => {
|
||||
// Regression pin for the derivation defect (research note §2/§3): the
|
||||
// chapter address is built from "my-longevity-simulation" but the Series
|
||||
// is published as "immortality-simulator". If the adapter ever derives the
|
||||
// identity from the address again, this test goes red.
|
||||
reset();
|
||||
elements = { "h1.entry-title": "My Longevity Simulation Chapter 1" };
|
||||
attrEls = {
|
||||
"a[aria-label='All Chapter']": {
|
||||
getAttribute: () => "https://lightnovelworld.net/novel/immortality-simulator/",
|
||||
},
|
||||
};
|
||||
const url = "https://lightnovelworld.net/my-longevity-simulation-chapter-1/";
|
||||
const p = lightnovelworld.detect(loc(url));
|
||||
assert.equal(p.type, "chapter");
|
||||
assert.equal(p.seriesId, "immortality-simulator");
|
||||
assert.equal(p.seriesUrl, "https://lightnovelworld.net/novel/immortality-simulator/");
|
||||
assert.equal(p.chapterSlug, "my-longevity-simulation");
|
||||
assert.notEqual(p.chapterSlug, p.seriesId);
|
||||
});
|
||||
|
||||
test("lightnovelworld.detect resolves a novel whose heading ends in a chapter number", () => {
|
||||
// The split novel from research §4.4, chapter 200 (published under the
|
||||
// current slug). Its heading ends "…Not Them All Chapter 200", which would
|
||||
// false-match a selector that looks for the text "All Chapter".
|
||||
reset();
|
||||
elements = {
|
||||
"h1.entry-title": "All Jobs and Classes I Just Wanted One Skill Not Them All Chapter 200",
|
||||
};
|
||||
attrEls = {
|
||||
"a[aria-label='All Chapter']": {
|
||||
getAttribute: () =>
|
||||
"https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/",
|
||||
},
|
||||
};
|
||||
const url =
|
||||
"https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-200/";
|
||||
const p = lightnovelworld.detect(loc(url));
|
||||
assert.equal(p.type, "chapter");
|
||||
assert.equal(p.seriesId, "all-jobs-and-classes-i-just-wanted-one-skill-not-them-all");
|
||||
assert.equal(p.title, "All Jobs and Classes I Just Wanted One Skill Not Them All");
|
||||
assert.equal(p.chapterNum, 200);
|
||||
assert.equal(p.chapterSlug, "all-jobs-and-classes-i-just-wanted-one-skill-not-them-all");
|
||||
});
|
||||
|
||||
test("lightnovelworld.detect resolves both Chapter Slugs of the split novel to one Series", () => {
|
||||
// Research note §4.4: this Series serves chapters 1-99 under one Chapter
|
||||
// Slug and 100-423 under another; both chapter addresses are live and both
|
||||
// point back at the same Series. An old deep link must not get a different
|
||||
// identity than a current one.
|
||||
reset();
|
||||
elements = { "h1.entry-title": "All Jobs and Classes I Just Wanted One Skill Not Chapter 1" };
|
||||
attrEls = {
|
||||
"a[aria-label='All Chapter']": {
|
||||
getAttribute: () =>
|
||||
"https://lightnovelworld.net/novel/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all/",
|
||||
},
|
||||
};
|
||||
const oldSlug = lightnovelworld.detect(
|
||||
loc("https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-chapter-1/")
|
||||
);
|
||||
const currentSlug = lightnovelworld.detect(
|
||||
loc(
|
||||
"https://lightnovelworld.net/all-jobs-and-classes-i-just-wanted-one-skill-not-them-all-chapter-404/"
|
||||
)
|
||||
);
|
||||
assert.equal(oldSlug.type, "chapter");
|
||||
assert.equal(oldSlug.seriesId, "all-jobs-and-classes-i-just-wanted-one-skill-not-them-all");
|
||||
assert.equal(oldSlug.chapterSlug, "all-jobs-and-classes-i-just-wanted-one-skill-not");
|
||||
assert.equal(currentSlug.type, "chapter");
|
||||
assert.equal(currentSlug.seriesId, "all-jobs-and-classes-i-just-wanted-one-skill-not-them-all");
|
||||
assert.equal(currentSlug.chapterSlug, "all-jobs-and-classes-i-just-wanted-one-skill-not-them-all");
|
||||
});
|
||||
|
||||
test("lightnovelworld.detect returns other for non-series paths", () => {
|
||||
reset();
|
||||
assert.equal(lightnovelworld.detect(loc("https://lightnovelworld.net/")).type, "other");
|
||||
assert.equal(lightnovelworld.detect(loc("https://lightnovelworld.net/az-lists/")).type, "other");
|
||||
});
|
||||
|
||||
test("lightnovelworld.latestChapterFromAnchors takes the max and ignores other series", () => {
|
||||
const best = lightnovelworld.latestChapterFromAnchors(
|
||||
test("computeLatestChapter yields nothing for lightnovelworld, comment anchors included", () => {
|
||||
// A realistic Series-page anchor set: current-slug chapters, a second
|
||||
// Chapter Slug's chapters, and a wpdiscuz comment pasting a high-numbered
|
||||
// chapter of another novel (the poisoning vector). The client never scans
|
||||
// this Site — the Poll's cooldown dominates the client throttle and the
|
||||
// comment thread is a public write surface — so this must fail the moment
|
||||
// a scan is reintroduced, scoped or not.
|
||||
const latest = computeLatestChapter(
|
||||
"lightnovelworld",
|
||||
[
|
||||
{ href: "https://lightnovelworld.net/a-will-eternal-chapter-1/", text: "Chapter 1" },
|
||||
{ href: "https://lightnovelworld.net/a-will-eternal-chapter-1317/", text: "Chapter 1317" },
|
||||
{ href: "https://lightnovelworld.net/a-will-eternal-chapter-1298/", text: "Chapter 1298" },
|
||||
// a divergent novel's second Chapter Slug
|
||||
{ href: "https://lightnovelworld.net/my-longevity-simulation-chapter-400/", text: "Chapter 400" },
|
||||
// a wpdiscuz comment anchor
|
||||
{ href: "https://lightnovelworld.net/overgeared-chapter-9999/", text: "Chapter 9999" },
|
||||
],
|
||||
"a-will-eternal"
|
||||
);
|
||||
assert.deepEqual(best, { num: 1317, label: "Chapter 1317" });
|
||||
assert.equal(latest, null);
|
||||
});
|
||||
|
||||
test("computeLatestChapter still scans novelfull and ignores other series", () => {
|
||||
const best = computeLatestChapter(
|
||||
"novelfull",
|
||||
[
|
||||
{ href: "/reverend-insanity/chapter-2334-fang-yuan.html", text: "Chapter 2334" },
|
||||
{ href: "/reverend-insanity/chapter-1.html", text: "Chapter 1" },
|
||||
{ href: "/release-that-witch/chapter-9999.html", text: "Chapter 9999" },
|
||||
],
|
||||
"reverend-insanity"
|
||||
);
|
||||
assert.deepEqual(best, { num: 2334, label: "Chapter 2334" });
|
||||
});
|
||||
|
||||
test("latestChapterFromAnchors returns null when nothing matches", () => {
|
||||
@@ -154,6 +356,213 @@ test("latestChapterFromAnchors returns null when nothing matches", () => {
|
||||
assert.equal(maxChapter([], /chapter-([0-9.]+)/), null);
|
||||
});
|
||||
|
||||
// ============================================================
|
||||
// repairLnwStaleRow — the row migration (spec seam 3)
|
||||
// ============================================================
|
||||
|
||||
// A row as bookmarkCurrent builds it, keyed under the invented Chapter Slug
|
||||
// identity the old adapter derived from the chapter address.
|
||||
function staleRow(over) {
|
||||
return Object.assign(
|
||||
{
|
||||
key: "lightnovelworld:my-longevity-simulation",
|
||||
site: "lightnovelworld",
|
||||
kind: "novel",
|
||||
series_id: "my-longevity-simulation",
|
||||
title: "My Longevity Simulation",
|
||||
series_url: "https://lightnovelworld.net/novel/my-longevity-simulation/",
|
||||
last_chapter: "Chapter 7",
|
||||
last_chapter_num: 7,
|
||||
last_chapter_url: "https://lightnovelworld.net/my-longevity-simulation-chapter-7/",
|
||||
updated_at: 1000,
|
||||
},
|
||||
over
|
||||
);
|
||||
}
|
||||
|
||||
// The divergent novel from research §3: the address's slug is a Chapter Slug,
|
||||
// the page pointer names the real Series.
|
||||
function lnwPage(over) {
|
||||
return Object.assign(
|
||||
{
|
||||
type: "chapter",
|
||||
site: "lightnovelworld",
|
||||
seriesId: "immortality-simulator",
|
||||
chapterSlug: "my-longevity-simulation",
|
||||
seriesUrl: "https://lightnovelworld.net/novel/immortality-simulator/",
|
||||
},
|
||||
over
|
||||
);
|
||||
}
|
||||
|
||||
test("repairLnwStaleRow rewrites the row, the queue entry and the last-checked map together", () => {
|
||||
const list = [staleRow()];
|
||||
const queue = [
|
||||
{ key: "lightnovelworld:my-longevity-simulation", op: "put", sendStatus: false, attempts: 2 },
|
||||
];
|
||||
const lastChecked = { "lightnovelworld:my-longevity-simulation": 12345 };
|
||||
const out = repairLnwStaleRow(list, queue, lastChecked, lnwPage());
|
||||
assert.equal(out.list.length, 1);
|
||||
assert.equal(out.list[0].key, "lightnovelworld:immortality-simulator");
|
||||
assert.equal(out.list[0].series_id, "immortality-simulator");
|
||||
assert.equal(out.list[0].series_url, "https://lightnovelworld.net/novel/immortality-simulator/");
|
||||
assert.deepEqual(out.queue, [
|
||||
{ key: "lightnovelworld:immortality-simulator", op: "put", sendStatus: false, attempts: 2 },
|
||||
]);
|
||||
assert.deepEqual(out.lastChecked, { "lightnovelworld:immortality-simulator": 12345 });
|
||||
});
|
||||
|
||||
test("repairLnwStaleRow keeps Progress, Favourite and Lifecycle bucket on the rewritten row", () => {
|
||||
const stale = staleRow({
|
||||
last_chapter: "Chapter 42",
|
||||
last_chapter_num: 42,
|
||||
last_chapter_url: "https://lightnovelworld.net/my-longevity-simulation-chapter-42/",
|
||||
favorite: true,
|
||||
status: "archived",
|
||||
updated_at: 777,
|
||||
});
|
||||
const out = repairLnwStaleRow([stale], [], {}, lnwPage());
|
||||
const b = out.list[0];
|
||||
assert.equal(b.last_chapter, "Chapter 42");
|
||||
assert.equal(b.last_chapter_num, 42);
|
||||
assert.equal(
|
||||
b.last_chapter_url,
|
||||
"https://lightnovelworld.net/my-longevity-simulation-chapter-42/"
|
||||
);
|
||||
assert.equal(b.favorite, true);
|
||||
assert.equal(b.status, "archived");
|
||||
assert.equal(b.updated_at, 777);
|
||||
assert.equal(b.title, "My Longevity Simulation");
|
||||
});
|
||||
|
||||
test("repairLnwStaleRow changes nothing when the Chapter Slug and Series slug agree", () => {
|
||||
const list = [
|
||||
staleRow({ key: "lightnovelworld:a-will-eternal", series_id: "a-will-eternal" }),
|
||||
];
|
||||
const queue = [{ key: "lightnovelworld:a-will-eternal", op: "put", sendStatus: true, attempts: 1 }];
|
||||
const lastChecked = { "lightnovelworld:a-will-eternal": 99 };
|
||||
const out = repairLnwStaleRow(
|
||||
list,
|
||||
queue,
|
||||
lastChecked,
|
||||
lnwPage({
|
||||
chapterSlug: "a-will-eternal",
|
||||
seriesId: "a-will-eternal",
|
||||
seriesUrl: "https://lightnovelworld.net/novel/a-will-eternal/",
|
||||
})
|
||||
);
|
||||
assert.equal(out.list, list); // same references, nothing rewritten
|
||||
assert.equal(out.queue, queue);
|
||||
assert.equal(out.lastChecked, lastChecked);
|
||||
});
|
||||
|
||||
test("repairLnwStaleRow changes nothing when no row sits under the old key", () => {
|
||||
const list = [
|
||||
staleRow({
|
||||
key: "lightnovelworld:immortality-simulator",
|
||||
series_id: "immortality-simulator",
|
||||
series_url: "https://lightnovelworld.net/novel/immortality-simulator/",
|
||||
}),
|
||||
];
|
||||
const lastChecked = { "lightnovelworld:immortality-simulator": 99 };
|
||||
const out = repairLnwStaleRow(list, [], lastChecked, lnwPage());
|
||||
assert.equal(out.list, list);
|
||||
assert.equal(out.queue.length, 0);
|
||||
assert.equal(out.lastChecked, lastChecked);
|
||||
});
|
||||
|
||||
test("repairLnwStaleRow ignores a page that carries no Chapter Slug", () => {
|
||||
const list = [
|
||||
staleRow({
|
||||
key: "lightnovelworld:immortality-simulator",
|
||||
series_id: "immortality-simulator",
|
||||
}),
|
||||
];
|
||||
const lastChecked = { "lightnovelworld:immortality-simulator": 99 };
|
||||
// A series page (chapterSlug null) and another site both stay untouched.
|
||||
const series = repairLnwStaleRow(list, [], lastChecked, lnwPage({ chapterSlug: null }));
|
||||
const other = repairLnwStaleRow(list, [], lastChecked, {
|
||||
type: "chapter",
|
||||
site: "novelfull",
|
||||
seriesId: "x",
|
||||
});
|
||||
assert.equal(series.list, list);
|
||||
assert.equal(other.list, list);
|
||||
});
|
||||
|
||||
test("repairLnwStaleRow still moves the row and the last-checked map when the queue has no entry", () => {
|
||||
const out = repairLnwStaleRow(
|
||||
[staleRow()],
|
||||
[],
|
||||
{ "lightnovelworld:my-longevity-simulation": 5 },
|
||||
lnwPage()
|
||||
);
|
||||
assert.equal(out.list[0].key, "lightnovelworld:immortality-simulator");
|
||||
assert.deepEqual(out.queue, []);
|
||||
assert.deepEqual(out.lastChecked, { "lightnovelworld:immortality-simulator": 5 });
|
||||
});
|
||||
|
||||
test("repairLnwStaleRow carries a queued change to the repaired identity with its marker intact", () => {
|
||||
const queue = [
|
||||
{ key: "lightnovelworld:my-longevity-simulation", op: "put", sendStatus: true, attempts: 3 },
|
||||
];
|
||||
const out = repairLnwStaleRow([staleRow()], queue, {}, lnwPage());
|
||||
assert.deepEqual(out.queue, [
|
||||
{ key: "lightnovelworld:immortality-simulator", op: "put", sendStatus: true, attempts: 3 },
|
||||
]);
|
||||
});
|
||||
|
||||
test("repairLnwStaleRow merges a duplicate under the repaired key, keeping the farther-ahead progress", () => {
|
||||
const canonical = staleRow({
|
||||
key: "lightnovelworld:immortality-simulator",
|
||||
series_id: "immortality-simulator",
|
||||
series_url: "https://lightnovelworld.net/novel/immortality-simulator/",
|
||||
last_chapter: "Chapter 5",
|
||||
last_chapter_num: 5,
|
||||
last_chapter_url: "https://lightnovelworld.net/immortality-simulator-chapter-5/",
|
||||
favorite: false,
|
||||
});
|
||||
const stale = staleRow({
|
||||
last_chapter: "Chapter 42",
|
||||
last_chapter_num: 42,
|
||||
last_chapter_url: "https://lightnovelworld.net/my-longevity-simulation-chapter-42/",
|
||||
favorite: true,
|
||||
status: "archived",
|
||||
});
|
||||
const queue = [
|
||||
{ key: "lightnovelworld:my-longevity-simulation", op: "put", sendStatus: true, attempts: 3 },
|
||||
{ key: "lightnovelworld:immortality-simulator", op: "put", sendStatus: false, attempts: 0 },
|
||||
];
|
||||
const out = repairLnwStaleRow([canonical, stale], queue, {}, lnwPage());
|
||||
assert.equal(out.list.length, 1);
|
||||
const merged = out.list[0];
|
||||
assert.equal(merged.key, "lightnovelworld:immortality-simulator");
|
||||
assert.equal(merged.last_chapter_num, 42); // the stale row is ahead — progress must not be lost
|
||||
assert.equal(merged.favorite, true); // favourite survives from either row
|
||||
assert.equal(merged.status, "archived"); // the stronger bucket survives
|
||||
// One marker under the repaired key: sendStatus is sticky (the archive
|
||||
// intent from the stale-key marker survives) and the worse attempts count
|
||||
// wins — the queue's own coalescing rules.
|
||||
assert.deepEqual(out.queue, [
|
||||
{ key: "lightnovelworld:immortality-simulator", op: "put", sendStatus: true, attempts: 3 },
|
||||
]);
|
||||
});
|
||||
|
||||
test("repairLnwStaleRow does not regress progress when the repaired-key row is ahead", () => {
|
||||
const canonical = staleRow({
|
||||
key: "lightnovelworld:immortality-simulator",
|
||||
series_id: "immortality-simulator",
|
||||
series_url: "https://lightnovelworld.net/novel/immortality-simulator/",
|
||||
last_chapter: "Chapter 100",
|
||||
last_chapter_num: 100,
|
||||
last_chapter_url: "https://lightnovelworld.net/immortality-simulator-chapter-100/",
|
||||
});
|
||||
const stale = staleRow({ last_chapter: "Chapter 42", last_chapter_num: 42 });
|
||||
const out = repairLnwStaleRow([canonical, stale], [], {}, lnwPage());
|
||||
assert.equal(out.list.length, 1);
|
||||
assert.equal(out.list[0].last_chapter_num, 100);
|
||||
});
|
||||
|
||||
// ============================================================
|
||||
// kindOf
|
||||
// ============================================================
|
||||
|
||||
Reference in New Issue
Block a user