Run on Postgres with today's schema #28

Merged
sulthan merged 3 commits from feat/db-change into main 2026-08-08 06:52:21 +07:00
5 changed files with 229 additions and 0 deletions
Showing only changes of commit 7a0c190ebe - Show all commits
+69
View File
@@ -0,0 +1,69 @@
# Bookmark Manager
Read-progress tracker for serialised fiction. A reader browses third-party manga and
novel sites; userscripts capture where they got to and sync it to a self-hosted backend,
so progress survives across sites and devices.
## Language
**Series**:
One ongoing work — a manga or a novel — as published by a Site. Identified by its
stable slug on that Site, never by its title. A Series exists once and is shared by
every Reader who bookmarks it; it owns the facts that are true regardless of who is
reading — title, cover, Latest Chapter. A Reader cannot change them; they describe the
Series, not anyone's relationship to it.
_Avoid_: manga, title, book, comic
**Site**:
One third-party source a Series is published on. A Series on two Sites is two Series.
_Avoid_: source, host, provider, domain
**Reader**:
A person with their own Progress. Exactly one per set of credentials, so there is no
separate "account" concept to model — the credential belongs to the Reader.
_Avoid_: user, account, member, subscriber
**Bookmark**:
One Reader's tracked relationship with one Series, holding only what differs between
Readers: Progress, Favourite, Lifecycle bucket. Facts about the Series itself belong
to the Series, not here.
_Avoid_: entry, item, record, subscription
**Library**:
One of the two halves of the collection — manga or novel — selected by a Bookmark's
`kind`. The web UI and the userscripts each address exactly one Library at a time.
Not a per-person concept: "everything one person has bookmarked" is a different idea
and must not be called a Library.
_Avoid_: section, tab, category
**Progress**:
The furthest chapter a reader has actually read in a Series. Only a change in Progress
is real activity, so only Progress reorders the list.
_Avoid_: position, bookmark (the noun is taken), last read
**Latest Chapter**:
The newest chapter a Site has published for a Series, discovered without the reader
present. Distinct from Progress in every way that matters: it is a fact about the Site,
not about the reader, and it must never reorder the list.
_Avoid_: newest, current chapter, update
**Poll**:
The backend's own check of a Site for a Series's Latest Chapter, made without the
Reader present. Performed once per Series no matter how many Readers bookmarked it —
a Poll is work done on behalf of the Series, never on behalf of a Reader.
_Avoid_: scrape, refresh, check, sync
**New Chapter**:
The state where Latest Chapter is ahead of Progress. The single condition the ember
accent is permitted to signal.
_Avoid_: unread, update available
**Lifecycle bucket**:
Which of three mutually exclusive states a Bookmark sits in — reading, archived, or
finished. A Bookmark is in exactly one. Orthogonal to being a favourite.
_Avoid_: state, status (as a domain word), list
**Favourite**:
A reader's manual pin on a Bookmark. Orthogonal to the Lifecycle bucket, and never a
reason to reorder the list.
_Avoid_: starred, pinned, priority
+40
View File
@@ -0,0 +1,40 @@
# Postgres replaces SQLite as the primary datastore
Status: accepted
The project is moving from a single-reader tracker to a service published to a community,
so we replaced `modernc.org/sqlite` with Postgres (`jackc/pgx/v5`, still pure Go, so
`CGO_ENABLED=0` and the distroless image are unaffected). The deciding reason is future
supportability — managed hosting, a datastore that survives the app outgrowing one box —
**not** concurrency, which was measured and found to be a non-issue.
## Considered options
**Stay on SQLite.** Benchmarked against the real store at 1,500 rows (≈50 readers × 30
series): ~9,700 upserts/sec single-writer, plateauing at ~780/sec under 8–50 concurrent
writers, with `List()` holding 90–103 calls/sec under continuous write load. Projected
real load at 50 readers is ~0.02 writes/sec — roughly four and a half orders of magnitude
of headroom. `SetMaxOpenConns(1)` serialises writes but was shown not to starve reads;
an apparent read collapse traced to row count and per-row scanning, not lock contention.
SQLite would have worked. It was rejected for where the project is going, not for what
it does today.
**Postgres.** Chosen. Migrating is cheapest now — 29 rows in one table — and gets
materially harder once there are live readers and a multi-tenant schema.
## Consequences
- Every statement in `internal/store` is rewritten: `?` → `$N`, `IS NOT` →
`IS DISTINCT FROM` (this one is load-bearing; it implements the `updated_at`
ordering rule), `pragma_table_info` → `information_schema.columns`,
`INTEGER`/`REAL` → `bigint`/`double precision`, `favorite` int-as-bool → `boolean`.
- ~75 tests currently get a free isolated database from `t.TempDir()`. They now need a
live server, which makes Docker a hard prerequisite for `go test ./...`. This is the
permanent cost of the decision and the main reason it was close.
- Backups get worse, not better: `VACUUM INTO` produced one self-contained file;
restoring now means `pg_dump`/`pg_restore`, a role, and a password.
- A second stateful container joins the VPS alongside the existing headless-shell.
- **Postgres does not address the real scaling limit.** At batch 14 per 10-minute tick
the poller checks at most 84 series/hour; 50 readers × 30 series is 1,500 bookmarks,
an 18-hour sweep against a configured 1-hour cooldown. That ceiling is an outbound
fetch budget and is fixed by deduplicating polls per Series, not by the datastore.
@@ -0,0 +1,44 @@
# Identity comes from Discord OAuth; we store no passwords and send no email
Status: accepted
The service is being published to a community that already lives on Discord, and we have
no transactional email infrastructure. Rather than build email verification and password
reset to get accounts, Readers sign in with Discord OAuth2 (authorization code grant,
`identify` + `guilds.members.read`), and guild membership replaces both the invite gate
and the email-verification step. No password is ever stored and no mail is ever sent.
## Considered options
**Email + password with invite codes, no verification.** Viable and dependency-free:
an invite code proves community membership, which is what email verification was
standing in for anyway. Rejected because it still requires password hashing, a manual
admin-driven reset path, and a credential store — all of which Discord removes.
**Email + password with a transactional provider** (Resend, Brevo). Rejected as
premature: it builds verification and self-serve reset before anyone has asked for them,
and adds deliverability as an operational concern.
**Discord OAuth.** Chosen. It is less code than either alternative — no hashing, no
reset flow, no invite table — and the authorization question ("is this person in my
community?") is answered by the same call that answers the authentication question.
## Consequences
- **Availability is now coupled to Discord.** If Discord's OAuth endpoint is down,
nobody can start a new session. Existing sessions are unaffected, which bounds the
blast radius.
- **Identity is a Discord snowflake.** Migrating off Discord later means re-identifying
every Reader, because we hold no other credential for them. This is the lock-in the
decision buys, and it is the reason this ADR exists.
- **`guilds.members.read` is checked at login, not continuously.** Someone who leaves
the guild keeps their session until it expires. Acceptable; revocation is a session
delete, not an architectural change.
- **The userscripts cannot use OAuth.** They run in an isolated world on third-party
pages with no redirect surface, so they keep a bearer token — now issued per Reader by
the backend rather than a single shared `API_TOKEN` literal. OAuth gates the web UI;
the web UI is where a Reader obtains their personal userscript.
- `WEB_PASSWORD` disappears, and with it the session HMAC key derivation
(`sha256(API_TOKEN | WEB_PASSWORD | …)`), which needs a replacement secret.
- Seeding the first Reader during migration requires knowing the owner's Discord user
ID up front — a stable snowflake, copied from the Discord client.
@@ -0,0 +1,44 @@
# Series is a shared entity, and only the Poll may update it
Status: accepted
Facts about a Series that are true regardless of who is reading — title, cover, canonical
URL, Latest Chapter — moved off the Bookmark onto a shared `series` row keyed
`(site, series_id)`. A Bookmark now holds only what differs between Readers: Progress,
Favourite, Lifecycle bucket. Fifty Readers tracking one Series produce fifty Bookmarks
and one Series, so the Series is polled once rather than fifty times.
## Why
The poller checks at most 84 series/hour (batch 14 per 10-minute tick). With ~50 Readers
holding ~30 Series each, polling per Bookmark means a 1,500-item sweep — roughly 18 hours
against a configured 1-hour cooldown, quietly breaking the New Chapter signal that is the
product's reason to exist. Deduplicating to distinct Series cuts the sweep several-fold,
and because the Series row now knows how many Readers hold it, the poll queue is ordered
`reader_count DESC, latest_checked_at ASC` — popular Series stay fresh and the long tail
absorbs the shortfall. That ordering is only expressible because the split happened.
Raising throughput instead was rejected: sweeping 400 Series hourly needs the stagger
cut from 20s to ~9s, doubling request rate against sites that already bot-score the
single VPS IP.
## Only the Poll writes Series fields
A client may supply `title`, `cover` and `series_url` only when creating a Series nobody
has bookmarked yet. After that, client-supplied values are ignored; only the backend's
own fetch updates them.
This is a security boundary, not tidiness. Those values are scraped from third-party
pages, which `AGENTS.md` requires be treated as attacker-controlled. Before the split, a
hostile or compromised site could corrupt exactly one Reader's row. After it, the same
write lands on a row every Reader sees — one Reader's browser becomes a write path into
everyone else's UI, and a cover URL can point anywhere. The backend's own fetch is the
higher-trust source: its network, its parser, no third-party JavaScript in the path.
## Consequences
- `Store.Upsert` decomposes one incoming flat body across two tables and enforces the
ownership rule at that seam.
- The `updated_at` ordering rule stays on the Bookmark, where Progress lives. Unchanged.
- Per-Reader title overrides are deliberately not supported; they would reintroduce the
duplication this removes.
@@ -0,0 +1,32 @@
# The wire format stays flat and deliberately does not mirror the schema
Status: accepted
Storage splits a tracked series across two tables (ADR-0003), but `GET /bookmarks` and
`PUT /bookmarks/{key}` keep emitting and accepting one **flat** JSON object with `title`,
`cover`, `last_chapter` and `latest_chapter` as siblings — exactly the shape they had
when there was one table. The server joins on the way out and decomposes on the way in.
## Why a future reader will find this surprising
The obvious move after splitting a table is to nest the JSON to match. Don't "fix" this.
**A nested payload would have broken every installed userscript instantly.** Scripts read
`b.title` directly; moving it to `b.series.title` yields `undefined` — no error, just
blank rows and a New Chapter signal that silently reports nothing forever. Because
Violentmonkey updates roughly once a day per device, the migration relies on old scripts
continuing to work during a 14-day grace window. A nested format and that grace window
are mutually exclusive.
**It is also the better contract independently of compatibility.** A client rendering one
row needs the title and the reading position together; nesting exports the re-stitching
to every browser to mirror a decision about disk layout it should not know about. Keeping
them separate lets storage change again later without a client release — which is the
whole reason this ADR is worth the paragraph.
## Consequence
The flat shape is a contract, not an implementation detail. Changing the storage schema
must not change it. It follows the rule already in force for `updated_at`: the server
owns the truth and returns the row **as stored**, and clients adopt the response rather
than their own payload.