feat(backend)!: run on Postgres with a migration-owned schema (#28)
Swap modernc.org/sqlite for jackc/pgx/v5 with no observable change: same endpoints, same wire format, same updated_at ordering rule. The schema now comes from numbered SQL embedded in the binary and applied on startup, one transaction each, recorded in schema_migrations. That replaces two pieces of SQLite-era machinery, both deleted rather than ported: the column probing (Postgres has ADD COLUMN IF NOT EXISTS, and there is no legacy database left to probe) and the Asura key rewrite, which has run clean on every start for months now that the userscripts strip build hashes before writing. Its regexp survives as latest.asuraBuildHash, where the poller still needs it to scope chapter links to a series whose slug carries a rotating hash. Types get real: favorite is a boolean, chapter numbers double precision, timestamps stay unix-ms bigint. SQLite's null-safe IS NOT becomes IS DISTINCT FROM, which is what implements the rule that only reading progress reorders a list. Inside COALESCE/NULLIF the status and kind parameters need an explicit ::text -- there is no target column to infer from and Postgres refuses to guess. Tests lose their free t.TempDir() database, so Docker is now a hard prerequisite for `go test ./...`: internal/pgtest starts one postgres:17-alpine per test binary and hands each test a database of its own. Also lands CONTEXT.md and the four ADRs written while scoping #18. BREAKING CHANGE: DB_PATH is retired for DATABASE_URL, which is required and has no default. Compose gains a postgres service on an internal network with its own volume; POSTGRES_PASSWORD joins .env. The old bookmarks-data volume is deliberately left undeclared so `docker compose down -v` cannot take the pre-migration database with it. main is not deployable until #25 and #26 land. Closes #20 Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com> Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
This commit was merged in pull request #28.
This commit is contained in:
@@ -0,0 +1,40 @@
|
||||
# Postgres replaces SQLite as the primary datastore
|
||||
|
||||
Status: accepted
|
||||
|
||||
The project is moving from a single-reader tracker to a service published to a community,
|
||||
so we replaced `modernc.org/sqlite` with Postgres (`jackc/pgx/v5`, still pure Go, so
|
||||
`CGO_ENABLED=0` and the distroless image are unaffected). The deciding reason is future
|
||||
supportability — managed hosting, a datastore that survives the app outgrowing one box —
|
||||
**not** concurrency, which was measured and found to be a non-issue.
|
||||
|
||||
## Considered options
|
||||
|
||||
**Stay on SQLite.** Benchmarked against the real store at 1,500 rows (≈50 readers × 30
|
||||
series): ~9,700 upserts/sec single-writer, plateauing at ~780/sec under 8–50 concurrent
|
||||
writers, with `List()` holding 90–103 calls/sec under continuous write load. Projected
|
||||
real load at 50 readers is ~0.02 writes/sec — roughly four and a half orders of magnitude
|
||||
of headroom. `SetMaxOpenConns(1)` serialises writes but was shown not to starve reads;
|
||||
an apparent read collapse traced to row count and per-row scanning, not lock contention.
|
||||
SQLite would have worked. It was rejected for where the project is going, not for what
|
||||
it does today.
|
||||
|
||||
**Postgres.** Chosen. Migrating is cheapest now — 29 rows in one table — and gets
|
||||
materially harder once there are live readers and a multi-tenant schema.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Every statement in `internal/store` is rewritten: `?` → `$N`, `IS NOT` →
|
||||
`IS DISTINCT FROM` (this one is load-bearing; it implements the `updated_at`
|
||||
ordering rule), `pragma_table_info` → `information_schema.columns`,
|
||||
`INTEGER`/`REAL` → `bigint`/`double precision`, `favorite` int-as-bool → `boolean`.
|
||||
- ~75 tests currently get a free isolated database from `t.TempDir()`. They now need a
|
||||
live server, which makes Docker a hard prerequisite for `go test ./...`. This is the
|
||||
permanent cost of the decision and the main reason it was close.
|
||||
- Backups get worse, not better: `VACUUM INTO` produced one self-contained file;
|
||||
restoring now means `pg_dump`/`pg_restore`, a role, and a password.
|
||||
- A second stateful container joins the VPS alongside the existing headless-shell.
|
||||
- **Postgres does not address the real scaling limit.** At batch 14 per 10-minute tick
|
||||
the poller checks at most 84 series/hour; 50 readers × 30 series is 1,500 bookmarks,
|
||||
an 18-hour sweep against a configured 1-hour cooldown. That ceiling is an outbound
|
||||
fetch budget and is fixed by deduplicating polls per Series, not by the datastore.
|
||||
@@ -0,0 +1,44 @@
|
||||
# Identity comes from Discord OAuth; we store no passwords and send no email
|
||||
|
||||
Status: accepted
|
||||
|
||||
The service is being published to a community that already lives on Discord, and we have
|
||||
no transactional email infrastructure. Rather than build email verification and password
|
||||
reset to get accounts, Readers sign in with Discord OAuth2 (authorization code grant,
|
||||
`identify` + `guilds.members.read`), and guild membership replaces both the invite gate
|
||||
and the email-verification step. No password is ever stored and no mail is ever sent.
|
||||
|
||||
## Considered options
|
||||
|
||||
**Email + password with invite codes, no verification.** Viable and dependency-free:
|
||||
an invite code proves community membership, which is what email verification was
|
||||
standing in for anyway. Rejected because it still requires password hashing, a manual
|
||||
admin-driven reset path, and a credential store — all of which Discord removes.
|
||||
|
||||
**Email + password with a transactional provider** (Resend, Brevo). Rejected as
|
||||
premature: it builds verification and self-serve reset before anyone has asked for them,
|
||||
and adds deliverability as an operational concern.
|
||||
|
||||
**Discord OAuth.** Chosen. It is less code than either alternative — no hashing, no
|
||||
reset flow, no invite table — and the authorization question ("is this person in my
|
||||
community?") is answered by the same call that answers the authentication question.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Availability is now coupled to Discord.** If Discord's OAuth endpoint is down,
|
||||
nobody can start a new session. Existing sessions are unaffected, which bounds the
|
||||
blast radius.
|
||||
- **Identity is a Discord snowflake.** Migrating off Discord later means re-identifying
|
||||
every Reader, because we hold no other credential for them. This is the lock-in the
|
||||
decision buys, and it is the reason this ADR exists.
|
||||
- **`guilds.members.read` is checked at login, not continuously.** Someone who leaves
|
||||
the guild keeps their session until it expires. Acceptable; revocation is a session
|
||||
delete, not an architectural change.
|
||||
- **The userscripts cannot use OAuth.** They run in an isolated world on third-party
|
||||
pages with no redirect surface, so they keep a bearer token — now issued per Reader by
|
||||
the backend rather than a single shared `API_TOKEN` literal. OAuth gates the web UI;
|
||||
the web UI is where a Reader obtains their personal userscript.
|
||||
- `WEB_PASSWORD` disappears, and with it the session HMAC key derivation
|
||||
(`sha256(API_TOKEN | WEB_PASSWORD | …)`), which needs a replacement secret.
|
||||
- Seeding the first Reader during migration requires knowing the owner's Discord user
|
||||
ID up front — a stable snowflake, copied from the Discord client.
|
||||
@@ -0,0 +1,44 @@
|
||||
# Series is a shared entity, and only the Poll may update it
|
||||
|
||||
Status: accepted
|
||||
|
||||
Facts about a Series that are true regardless of who is reading — title, cover, canonical
|
||||
URL, Latest Chapter — moved off the Bookmark onto a shared `series` row keyed
|
||||
`(site, series_id)`. A Bookmark now holds only what differs between Readers: Progress,
|
||||
Favourite, Lifecycle bucket. Fifty Readers tracking one Series produce fifty Bookmarks
|
||||
and one Series, so the Series is polled once rather than fifty times.
|
||||
|
||||
## Why
|
||||
|
||||
The poller checks at most 84 series/hour (batch 14 per 10-minute tick). With ~50 Readers
|
||||
holding ~30 Series each, polling per Bookmark means a 1,500-item sweep — roughly 18 hours
|
||||
against a configured 1-hour cooldown, quietly breaking the New Chapter signal that is the
|
||||
product's reason to exist. Deduplicating to distinct Series cuts the sweep several-fold,
|
||||
and because the Series row now knows how many Readers hold it, the poll queue is ordered
|
||||
`reader_count DESC, latest_checked_at ASC` — popular Series stay fresh and the long tail
|
||||
absorbs the shortfall. That ordering is only expressible because the split happened.
|
||||
|
||||
Raising throughput instead was rejected: sweeping 400 Series hourly needs the stagger
|
||||
cut from 20s to ~9s, doubling request rate against sites that already bot-score the
|
||||
single VPS IP.
|
||||
|
||||
## Only the Poll writes Series fields
|
||||
|
||||
A client may supply `title`, `cover` and `series_url` only when creating a Series nobody
|
||||
has bookmarked yet. After that, client-supplied values are ignored; only the backend's
|
||||
own fetch updates them.
|
||||
|
||||
This is a security boundary, not tidiness. Those values are scraped from third-party
|
||||
pages, which `AGENTS.md` requires be treated as attacker-controlled. Before the split, a
|
||||
hostile or compromised site could corrupt exactly one Reader's row. After it, the same
|
||||
write lands on a row every Reader sees — one Reader's browser becomes a write path into
|
||||
everyone else's UI, and a cover URL can point anywhere. The backend's own fetch is the
|
||||
higher-trust source: its network, its parser, no third-party JavaScript in the path.
|
||||
|
||||
## Consequences
|
||||
|
||||
- `Store.Upsert` decomposes one incoming flat body across two tables and enforces the
|
||||
ownership rule at that seam.
|
||||
- The `updated_at` ordering rule stays on the Bookmark, where Progress lives. Unchanged.
|
||||
- Per-Reader title overrides are deliberately not supported; they would reintroduce the
|
||||
duplication this removes.
|
||||
@@ -0,0 +1,32 @@
|
||||
# The wire format stays flat and deliberately does not mirror the schema
|
||||
|
||||
Status: accepted
|
||||
|
||||
Storage splits a tracked series across two tables (ADR-0003), but `GET /bookmarks` and
|
||||
`PUT /bookmarks/{key}` keep emitting and accepting one **flat** JSON object with `title`,
|
||||
`cover`, `last_chapter` and `latest_chapter` as siblings — exactly the shape they had
|
||||
when there was one table. The server joins on the way out and decomposes on the way in.
|
||||
|
||||
## Why a future reader will find this surprising
|
||||
|
||||
The obvious move after splitting a table is to nest the JSON to match. Don't "fix" this.
|
||||
|
||||
**A nested payload would have broken every installed userscript instantly.** Scripts read
|
||||
`b.title` directly; moving it to `b.series.title` yields `undefined` — no error, just
|
||||
blank rows and a New Chapter signal that silently reports nothing forever. Because
|
||||
Violentmonkey updates roughly once a day per device, the migration relies on old scripts
|
||||
continuing to work during a 14-day grace window. A nested format and that grace window
|
||||
are mutually exclusive.
|
||||
|
||||
**It is also the better contract independently of compatibility.** A client rendering one
|
||||
row needs the title and the reading position together; nesting exports the re-stitching
|
||||
to every browser to mirror a decision about disk layout it should not know about. Keeping
|
||||
them separate lets storage change again later without a client release — which is the
|
||||
whole reason this ADR is worth the paragraph.
|
||||
|
||||
## Consequence
|
||||
|
||||
The flat shape is a contract, not an implementation detail. Changing the storage schema
|
||||
must not change it. It follows the rule already in force for `updated_at`: the server
|
||||
owns the truth and returns the row **as stored**, and clients adopt the response rather
|
||||
than their own payload.
|
||||
@@ -0,0 +1,47 @@
|
||||
# Domain Docs
|
||||
|
||||
How the engineering skills should consume this repo's domain documentation when exploring the
|
||||
codebase. Layout: **single-context** — one `CONTEXT.md` plus `docs/adr/` at the repo root.
|
||||
|
||||
## Before exploring, read these
|
||||
|
||||
- **`CONTEXT.md`** at the repo root — the glossary / ubiquitous language.
|
||||
- **`docs/adr/`** — read ADRs that touch the area you're about to work in.
|
||||
|
||||
If any of these files don't exist, **proceed silently**. Don't flag their absence; don't suggest
|
||||
creating them upfront. The `/domain-modeling` skill (reached via `/grill-with-docs` and
|
||||
`/improve-codebase-architecture`) creates them lazily when terms or decisions actually get resolved.
|
||||
|
||||
Neither exists yet in this repo. The existing `AGENTS.md` / `CLAUDE.md` and `docs/design-system.md`
|
||||
carry the current architecture and design law — read those regardless.
|
||||
|
||||
## File structure
|
||||
|
||||
```
|
||||
/
|
||||
├── CONTEXT.md
|
||||
├── docs/adr/
|
||||
│ ├── 0001-....md
|
||||
│ └── 0002-....md
|
||||
├── backend/
|
||||
└── userscript/
|
||||
```
|
||||
|
||||
If this repo ever splits into genuinely separate contexts, add a root `CONTEXT-MAP.md` pointing at
|
||||
one `CONTEXT.md` per context and update this file.
|
||||
|
||||
## Use the glossary's vocabulary
|
||||
|
||||
When your output names a domain concept (in an issue title, a refactor proposal, a hypothesis, a
|
||||
test name), use the term as defined in `CONTEXT.md`. Don't drift to synonyms the glossary
|
||||
explicitly avoids.
|
||||
|
||||
If the concept you need isn't in the glossary yet, that's a signal — either you're inventing
|
||||
language the project doesn't use (reconsider) or there's a real gap (note it for
|
||||
`/domain-modeling`).
|
||||
|
||||
## Flag ADR conflicts
|
||||
|
||||
If your output contradicts an existing ADR, surface it explicitly rather than silently overriding:
|
||||
|
||||
> _Contradicts ADR-0002 (…) — but worth reopening because…_
|
||||
@@ -0,0 +1,60 @@
|
||||
# Issue tracker: Gitea (`tea` CLI)
|
||||
|
||||
Issues and specs for this repo live as issues on the self-hosted Gitea instance
|
||||
`gitea.violetcrown.my.id` (repo `sulthan/mangaBookmark`). **`gh` does not work here** — use
|
||||
[`tea`](https://gitea.com/gitea/tea) for everything past plain git. Auth lives in `tea login`,
|
||||
not a `GH_TOKEN` env var. `tea` infers the repo from the local clone's `origin`.
|
||||
|
||||
`tea` prints rendered boxes rather than plain text; pass `--output json` (or `-o json`) when a
|
||||
skill needs to parse the result.
|
||||
|
||||
## Conventions
|
||||
|
||||
- **Create an issue**: `tea issue create --title "..." --description "..."` (`--labels`,
|
||||
`--assignees` optional). Multi-line bodies: pass the body through a shell variable or heredoc.
|
||||
- **Read an issue**: `tea issue <number> --comments` (add `-o json` for machine-readable output).
|
||||
- **List issues**: `tea issue list --state open -o json --fields index,title,body,labels,state,author`;
|
||||
filter with `--labels "..."`, `--state open|closed|all`, `--assignee`, `--keyword`.
|
||||
- **Comment**: `tea comment <number> "..."` (alias of `tea comments add`).
|
||||
- **Apply / remove labels**: `tea issue edit <number> --add-labels "..."` / `--remove-labels "..."`.
|
||||
Labels must exist first — see `tea labels list` / `tea labels create --name "..." --color "#rrggbb"`.
|
||||
- **Close**: `tea issue close <number>` (comment separately with `tea comment`; `close` takes no
|
||||
`--comment` flag).
|
||||
|
||||
## Pull requests as a triage surface
|
||||
|
||||
**PRs as a request surface: no.** _(Set to `yes` if this repo treats external PRs as feature
|
||||
requests; `/triage` reads this flag.)_
|
||||
|
||||
When set to `yes`, PRs run through the same labels and states as issues, using the `tea pr`
|
||||
equivalents: `tea pr <number> --comments`, `tea pr list --state open -o json`,
|
||||
`tea pr create --head <branch> --base main --title "..." --description "..."`, `tea comment`,
|
||||
`tea pr close`. Gitea shares one index space across issues and PRs, so a bare `#42` may be either
|
||||
— resolve with `tea pr 42` and fall back to `tea issue 42`.
|
||||
|
||||
## When a skill says "publish to the issue tracker"
|
||||
|
||||
Create a Gitea issue with `tea issue create`.
|
||||
|
||||
## When a skill says "fetch the relevant ticket"
|
||||
|
||||
Run `tea issue <number> --comments`.
|
||||
|
||||
## Wayfinding operations
|
||||
|
||||
Used by `/wayfinder`. The **map** is a single issue; **tickets** are child issues.
|
||||
|
||||
- **Map**: one issue labelled `wayfinder:map` holding the Notes / Decisions-so-far / Fog body.
|
||||
`tea issue create --labels wayfinder:map --title "..." --description "..."`.
|
||||
- **Child ticket**: an issue labelled `wayfinder:<type>` (`research`/`prototype`/`grilling`/`task`)
|
||||
with `Part of #<map>` as the first body line, and a task-list entry in the map body. `tea` has no
|
||||
sub-issue command, so the task list plus the `Part of` line is the canonical link.
|
||||
- **Blocking**: a `Blocked by: #<n>, #<n>` line at the top of the child body. Gitea's native issue
|
||||
dependencies exist in the API but `tea` does not expose them; the body line is the source of
|
||||
truth. A ticket is unblocked when every listed blocker is closed.
|
||||
- **Frontier query**: `tea issue list --state open -o json` scoped to the map's task list; drop any
|
||||
ticket with an open blocker or an assignee; first in map order wins.
|
||||
- **Claim**: `tea issue edit <n> --add-assignees <your-username>` — the session's first write.
|
||||
(`tea` has no `@me` shorthand; use the Gitea username from `tea login list`.)
|
||||
- **Resolve**: `tea comment <n> "<answer>"`, then `tea issue close <n>`, then append a context
|
||||
pointer to the map's Decisions-so-far via `tea issue edit <map> --description "..."`.
|
||||
@@ -0,0 +1,20 @@
|
||||
# Triage Labels
|
||||
|
||||
The skills speak in terms of five canonical triage roles. This file maps those roles to the actual
|
||||
label strings used in this repo's issue tracker (Gitea — see `docs/agents/issue-tracker.md`).
|
||||
|
||||
| Label in mattpocock/skills | Label in our tracker | Meaning |
|
||||
| -------------------------- | -------------------- | ---------------------------------------- |
|
||||
| `needs-triage` | `needs-triage` | Maintainer needs to evaluate this issue |
|
||||
| `needs-info` | `needs-info` | Waiting on reporter for more information |
|
||||
| `ready-for-agent` | `ready-for-agent` | Fully specified, ready for an AFK agent |
|
||||
| `ready-for-human` | `ready-for-human` | Requires human implementation |
|
||||
| `wontfix` | `wontfix` | Will not be actioned |
|
||||
|
||||
When a skill mentions a role (e.g. "apply the AFK-ready triage label"), use the corresponding label
|
||||
string from this table.
|
||||
|
||||
Gitea will not auto-create labels on `tea issue edit --add-labels`; create a missing one first with
|
||||
`tea labels create --name "<label>" --color "#rrggbb"`.
|
||||
|
||||
Edit the right-hand column to match whatever vocabulary you actually use.
|
||||
Reference in New Issue
Block a user