feat(backend)!: run on Postgres with a migration-owned schema (#28)

Swap modernc.org/sqlite for jackc/pgx/v5 with no observable change:
same endpoints, same wire format, same updated_at ordering rule.

The schema now comes from numbered SQL embedded in the binary and
applied on startup, one transaction each, recorded in
schema_migrations. That replaces two pieces of SQLite-era machinery,
both deleted rather than ported: the column probing (Postgres has ADD
COLUMN IF NOT EXISTS, and there is no legacy database left to probe)
and the Asura key rewrite, which has run clean on every start for
months now that the userscripts strip build hashes before writing. Its
regexp survives as latest.asuraBuildHash, where the poller still needs
it to scope chapter links to a series whose slug carries a rotating
hash.

Types get real: favorite is a boolean, chapter numbers double
precision, timestamps stay unix-ms bigint. SQLite's null-safe IS NOT
becomes IS DISTINCT FROM, which is what implements the rule that only
reading progress reorders a list. Inside COALESCE/NULLIF the status
and kind parameters need an explicit ::text -- there is no target
column to infer from and Postgres refuses to guess.

Tests lose their free t.TempDir() database, so Docker is now a hard
prerequisite for `go test ./...`: internal/pgtest starts one
postgres:17-alpine per test binary and hands each test a database of
its own.

Also lands CONTEXT.md and the four ADRs written while scoping #18.

BREAKING CHANGE: DB_PATH is retired for DATABASE_URL, which is
required and has no default. Compose gains a postgres service on an
internal network with its own volume; POSTGRES_PASSWORD joins .env.
The old bookmarks-data volume is deliberately left undeclared so
`docker compose down -v` cannot take the pre-migration database with
it. main is not deployable until #25 and #26 land.

Closes #20

Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
This commit was merged in pull request #28.
This commit is contained in:
2026-08-08 06:52:20 +07:00
committed by sulthan
parent b9f9aea82c
commit 08749df050
31 changed files with 961 additions and 966 deletions
+40
View File
@@ -0,0 +1,40 @@
# Postgres replaces SQLite as the primary datastore
Status: accepted
The project is moving from a single-reader tracker to a service published to a community,
so we replaced `modernc.org/sqlite` with Postgres (`jackc/pgx/v5`, still pure Go, so
`CGO_ENABLED=0` and the distroless image are unaffected). The deciding reason is future
supportability — managed hosting, a datastore that survives the app outgrowing one box —
**not** concurrency, which was measured and found to be a non-issue.
## Considered options
**Stay on SQLite.** Benchmarked against the real store at 1,500 rows (≈50 readers × 30
series): ~9,700 upserts/sec single-writer, plateauing at ~780/sec under 8–50 concurrent
writers, with `List()` holding 90–103 calls/sec under continuous write load. Projected
real load at 50 readers is ~0.02 writes/sec — roughly four and a half orders of magnitude
of headroom. `SetMaxOpenConns(1)` serialises writes but was shown not to starve reads;
an apparent read collapse traced to row count and per-row scanning, not lock contention.
SQLite would have worked. It was rejected for where the project is going, not for what
it does today.
**Postgres.** Chosen. Migrating is cheapest now — 29 rows in one table — and gets
materially harder once there are live readers and a multi-tenant schema.
## Consequences
- Every statement in `internal/store` is rewritten: `?` → `$N`, `IS NOT` →
`IS DISTINCT FROM` (this one is load-bearing; it implements the `updated_at`
ordering rule), `pragma_table_info` → `information_schema.columns`,
`INTEGER`/`REAL` → `bigint`/`double precision`, `favorite` int-as-bool → `boolean`.
- ~75 tests currently get a free isolated database from `t.TempDir()`. They now need a
live server, which makes Docker a hard prerequisite for `go test ./...`. This is the
permanent cost of the decision and the main reason it was close.
- Backups get worse, not better: `VACUUM INTO` produced one self-contained file;
restoring now means `pg_dump`/`pg_restore`, a role, and a password.
- A second stateful container joins the VPS alongside the existing headless-shell.
- **Postgres does not address the real scaling limit.** At batch 14 per 10-minute tick
the poller checks at most 84 series/hour; 50 readers × 30 series is 1,500 bookmarks,
an 18-hour sweep against a configured 1-hour cooldown. That ceiling is an outbound
fetch budget and is fixed by deduplicating polls per Series, not by the datastore.
@@ -0,0 +1,44 @@
# Identity comes from Discord OAuth; we store no passwords and send no email
Status: accepted
The service is being published to a community that already lives on Discord, and we have
no transactional email infrastructure. Rather than build email verification and password
reset to get accounts, Readers sign in with Discord OAuth2 (authorization code grant,
`identify` + `guilds.members.read`), and guild membership replaces both the invite gate
and the email-verification step. No password is ever stored and no mail is ever sent.
## Considered options
**Email + password with invite codes, no verification.** Viable and dependency-free:
an invite code proves community membership, which is what email verification was
standing in for anyway. Rejected because it still requires password hashing, a manual
admin-driven reset path, and a credential store — all of which Discord removes.
**Email + password with a transactional provider** (Resend, Brevo). Rejected as
premature: it builds verification and self-serve reset before anyone has asked for them,
and adds deliverability as an operational concern.
**Discord OAuth.** Chosen. It is less code than either alternative — no hashing, no
reset flow, no invite table — and the authorization question ("is this person in my
community?") is answered by the same call that answers the authentication question.
## Consequences
- **Availability is now coupled to Discord.** If Discord's OAuth endpoint is down,
nobody can start a new session. Existing sessions are unaffected, which bounds the
blast radius.
- **Identity is a Discord snowflake.** Migrating off Discord later means re-identifying
every Reader, because we hold no other credential for them. This is the lock-in the
decision buys, and it is the reason this ADR exists.
- **`guilds.members.read` is checked at login, not continuously.** Someone who leaves
the guild keeps their session until it expires. Acceptable; revocation is a session
delete, not an architectural change.
- **The userscripts cannot use OAuth.** They run in an isolated world on third-party
pages with no redirect surface, so they keep a bearer token — now issued per Reader by
the backend rather than a single shared `API_TOKEN` literal. OAuth gates the web UI;
the web UI is where a Reader obtains their personal userscript.
- `WEB_PASSWORD` disappears, and with it the session HMAC key derivation
(`sha256(API_TOKEN | WEB_PASSWORD | …)`), which needs a replacement secret.
- Seeding the first Reader during migration requires knowing the owner's Discord user
ID up front — a stable snowflake, copied from the Discord client.
@@ -0,0 +1,44 @@
# Series is a shared entity, and only the Poll may update it
Status: accepted
Facts about a Series that are true regardless of who is reading — title, cover, canonical
URL, Latest Chapter — moved off the Bookmark onto a shared `series` row keyed
`(site, series_id)`. A Bookmark now holds only what differs between Readers: Progress,
Favourite, Lifecycle bucket. Fifty Readers tracking one Series produce fifty Bookmarks
and one Series, so the Series is polled once rather than fifty times.
## Why
The poller checks at most 84 series/hour (batch 14 per 10-minute tick). With ~50 Readers
holding ~30 Series each, polling per Bookmark means a 1,500-item sweep — roughly 18 hours
against a configured 1-hour cooldown, quietly breaking the New Chapter signal that is the
product's reason to exist. Deduplicating to distinct Series cuts the sweep several-fold,
and because the Series row now knows how many Readers hold it, the poll queue is ordered
`reader_count DESC, latest_checked_at ASC` — popular Series stay fresh and the long tail
absorbs the shortfall. That ordering is only expressible because the split happened.
Raising throughput instead was rejected: sweeping 400 Series hourly needs the stagger
cut from 20s to ~9s, doubling request rate against sites that already bot-score the
single VPS IP.
## Only the Poll writes Series fields
A client may supply `title`, `cover` and `series_url` only when creating a Series nobody
has bookmarked yet. After that, client-supplied values are ignored; only the backend's
own fetch updates them.
This is a security boundary, not tidiness. Those values are scraped from third-party
pages, which `AGENTS.md` requires be treated as attacker-controlled. Before the split, a
hostile or compromised site could corrupt exactly one Reader's row. After it, the same
write lands on a row every Reader sees — one Reader's browser becomes a write path into
everyone else's UI, and a cover URL can point anywhere. The backend's own fetch is the
higher-trust source: its network, its parser, no third-party JavaScript in the path.
## Consequences
- `Store.Upsert` decomposes one incoming flat body across two tables and enforces the
ownership rule at that seam.
- The `updated_at` ordering rule stays on the Bookmark, where Progress lives. Unchanged.
- Per-Reader title overrides are deliberately not supported; they would reintroduce the
duplication this removes.
@@ -0,0 +1,32 @@
# The wire format stays flat and deliberately does not mirror the schema
Status: accepted
Storage splits a tracked series across two tables (ADR-0003), but `GET /bookmarks` and
`PUT /bookmarks/{key}` keep emitting and accepting one **flat** JSON object with `title`,
`cover`, `last_chapter` and `latest_chapter` as siblings — exactly the shape they had
when there was one table. The server joins on the way out and decomposes on the way in.
## Why a future reader will find this surprising
The obvious move after splitting a table is to nest the JSON to match. Don't "fix" this.
**A nested payload would have broken every installed userscript instantly.** Scripts read
`b.title` directly; moving it to `b.series.title` yields `undefined` — no error, just
blank rows and a New Chapter signal that silently reports nothing forever. Because
Violentmonkey updates roughly once a day per device, the migration relies on old scripts
continuing to work during a 14-day grace window. A nested format and that grace window
are mutually exclusive.
**It is also the better contract independently of compatibility.** A client rendering one
row needs the title and the reading position together; nesting exports the re-stitching
to every browser to mirror a decision about disk layout it should not know about. Keeping
them separate lets storage change again later without a client release — which is the
whole reason this ADR is worth the paragraph.
## Consequence
The flat shape is a contract, not an implementation detail. Changing the storage schema
must not change it. It follows the rule already in force for `updated_at`: the server
owns the truth and returns the row **as stored**, and clients adopt the response rather
than their own payload.
+47
View File
@@ -0,0 +1,47 @@
# Domain Docs
How the engineering skills should consume this repo's domain documentation when exploring the
codebase. Layout: **single-context** — one `CONTEXT.md` plus `docs/adr/` at the repo root.
## Before exploring, read these
- **`CONTEXT.md`** at the repo root — the glossary / ubiquitous language.
- **`docs/adr/`** — read ADRs that touch the area you're about to work in.
If any of these files don't exist, **proceed silently**. Don't flag their absence; don't suggest
creating them upfront. The `/domain-modeling` skill (reached via `/grill-with-docs` and
`/improve-codebase-architecture`) creates them lazily when terms or decisions actually get resolved.
Neither exists yet in this repo. The existing `AGENTS.md` / `CLAUDE.md` and `docs/design-system.md`
carry the current architecture and design law — read those regardless.
## File structure
```
/
├── CONTEXT.md
├── docs/adr/
│ ├── 0001-....md
│ └── 0002-....md
├── backend/
└── userscript/
```
If this repo ever splits into genuinely separate contexts, add a root `CONTEXT-MAP.md` pointing at
one `CONTEXT.md` per context and update this file.
## Use the glossary's vocabulary
When your output names a domain concept (in an issue title, a refactor proposal, a hypothesis, a
test name), use the term as defined in `CONTEXT.md`. Don't drift to synonyms the glossary
explicitly avoids.
If the concept you need isn't in the glossary yet, that's a signal — either you're inventing
language the project doesn't use (reconsider) or there's a real gap (note it for
`/domain-modeling`).
## Flag ADR conflicts
If your output contradicts an existing ADR, surface it explicitly rather than silently overriding:
> _Contradicts ADR-0002 (…) — but worth reopening because…_
+60
View File
@@ -0,0 +1,60 @@
# Issue tracker: Gitea (`tea` CLI)
Issues and specs for this repo live as issues on the self-hosted Gitea instance
`gitea.violetcrown.my.id` (repo `sulthan/mangaBookmark`). **`gh` does not work here** — use
[`tea`](https://gitea.com/gitea/tea) for everything past plain git. Auth lives in `tea login`,
not a `GH_TOKEN` env var. `tea` infers the repo from the local clone's `origin`.
`tea` prints rendered boxes rather than plain text; pass `--output json` (or `-o json`) when a
skill needs to parse the result.
## Conventions
- **Create an issue**: `tea issue create --title "..." --description "..."` (`--labels`,
`--assignees` optional). Multi-line bodies: pass the body through a shell variable or heredoc.
- **Read an issue**: `tea issue <number> --comments` (add `-o json` for machine-readable output).
- **List issues**: `tea issue list --state open -o json --fields index,title,body,labels,state,author`;
filter with `--labels "..."`, `--state open|closed|all`, `--assignee`, `--keyword`.
- **Comment**: `tea comment <number> "..."` (alias of `tea comments add`).
- **Apply / remove labels**: `tea issue edit <number> --add-labels "..."` / `--remove-labels "..."`.
Labels must exist first — see `tea labels list` / `tea labels create --name "..." --color "#rrggbb"`.
- **Close**: `tea issue close <number>` (comment separately with `tea comment`; `close` takes no
`--comment` flag).
## Pull requests as a triage surface
**PRs as a request surface: no.** _(Set to `yes` if this repo treats external PRs as feature
requests; `/triage` reads this flag.)_
When set to `yes`, PRs run through the same labels and states as issues, using the `tea pr`
equivalents: `tea pr <number> --comments`, `tea pr list --state open -o json`,
`tea pr create --head <branch> --base main --title "..." --description "..."`, `tea comment`,
`tea pr close`. Gitea shares one index space across issues and PRs, so a bare `#42` may be either
— resolve with `tea pr 42` and fall back to `tea issue 42`.
## When a skill says "publish to the issue tracker"
Create a Gitea issue with `tea issue create`.
## When a skill says "fetch the relevant ticket"
Run `tea issue <number> --comments`.
## Wayfinding operations
Used by `/wayfinder`. The **map** is a single issue; **tickets** are child issues.
- **Map**: one issue labelled `wayfinder:map` holding the Notes / Decisions-so-far / Fog body.
`tea issue create --labels wayfinder:map --title "..." --description "..."`.
- **Child ticket**: an issue labelled `wayfinder:<type>` (`research`/`prototype`/`grilling`/`task`)
with `Part of #<map>` as the first body line, and a task-list entry in the map body. `tea` has no
sub-issue command, so the task list plus the `Part of` line is the canonical link.
- **Blocking**: a `Blocked by: #<n>, #<n>` line at the top of the child body. Gitea's native issue
dependencies exist in the API but `tea` does not expose them; the body line is the source of
truth. A ticket is unblocked when every listed blocker is closed.
- **Frontier query**: `tea issue list --state open -o json` scoped to the map's task list; drop any
ticket with an open blocker or an assignee; first in map order wins.
- **Claim**: `tea issue edit <n> --add-assignees <your-username>` — the session's first write.
(`tea` has no `@me` shorthand; use the Gitea username from `tea login list`.)
- **Resolve**: `tea comment <n> "<answer>"`, then `tea issue close <n>`, then append a context
pointer to the map's Decisions-so-far via `tea issue edit <map> --description "..."`.
+20
View File
@@ -0,0 +1,20 @@
# Triage Labels
The skills speak in terms of five canonical triage roles. This file maps those roles to the actual
label strings used in this repo's issue tracker (Gitea — see `docs/agents/issue-tracker.md`).
| Label in mattpocock/skills | Label in our tracker | Meaning |
| -------------------------- | -------------------- | ---------------------------------------- |
| `needs-triage` | `needs-triage` | Maintainer needs to evaluate this issue |
| `needs-info` | `needs-info` | Waiting on reporter for more information |
| `ready-for-agent` | `ready-for-agent` | Fully specified, ready for an AFK agent |
| `ready-for-human` | `ready-for-human` | Requires human implementation |
| `wontfix` | `wontfix` | Will not be actioned |
When a skill mentions a role (e.g. "apply the AFK-ready triage label"), use the corresponding label
string from this table.
Gitea will not auto-create labels on `tea issue edit --add-labels`; create a missing one first with
`tea labels create --name "<label>" --color "#rrggbb"`.
Edit the right-hand column to match whatever vocabulary you actually use.