Prove the import against a copy of the real library (#25) (#33)

Closes #25.

Retires the biggest risk in #18 — losing the owner's reading history — on a copy, before production is anywhere near it.

## What was run

A throwaway generator (python3 stdlib `sqlite3`, ~20 lines, **not committed**) read a copy of `bookmarks-20260807-213515.db` and emitted plain SQL: 29 distinct Series first, then 29 Bookmarks referencing them, each `INSERT ... SELECT id FROM owner` so the reader id is resolved rather than hardcoded. The target was a scratch Postgres whose schema and owner Reader were built by the real binary (`go run .` against a throwaway container), not by hand-written DDL. Production was not touched.

## Verified

| check | result |
|---|---|
| Bookmarks total | 29 |
| reading / archived / other | 18 / 11 / 0 |
| Series | 29, equal to the distinct `(site, series_id)` count in the source |
| Readers | 1; Bookmarks not owned by the owner: 0 |
| Field-by-field diff, all 29 rows x 15 columns | 0 differences |
| `GET /bookmarks` over the real read path | 29 rows, values match source |
| `TRUNCATE bookmarks, series;` then re-apply | clean, 29 again |

The spot-check the ticket asked for was widened to a full row-by-row comparison — 29 rows is small enough that sampling was the more expensive option.

## What is committed

`CUTOVER.md` only, plus two cross-links from `REDEPLOY.md`. The generator stays out of the repository: its output is the owner's reading history, and it reads SQLite, which the backend module dropped in ADR-0001. So the runbook specifies the transformation — column mapping, ordering, nullability, quoting, the temp-table ownership trick — rather than shipping a script. `backend/go.mod` gains nothing.

## Review

Two-axis review ran on the diff; six findings applied, all in the runbook:

- Six source columns (`title`, `series_url`, `cover`, `last_chapter`, `last_chapter_url`, `last_chapter_num`) are nullable in SQLite but `NOT NULL` in Postgres and must be coalesced — the opposite of `latest_chapter_num`, the one column where `NULL` is meaningful. The 2026-08-07 export had none; a fresh one is not promised the same.
- `CREATE TEMP TABLE ... ON COMMIT DROP` must sit *inside* the transaction, or psql's autocommit drops it instantly.
- The ownership check now resolves the Reader by Discord id; comparing against `ORDER BY id LIMIT 1` was true by construction and could never fail.
- The spot-check now samples archived and favourite rows explicitly instead of hoping they fall inside `ORDER BY updated_at DESC LIMIT 5`.
- `git pull --ff-only` before `up -d --build`, or a pre-cutover server rebuilds the SQLite image.
- `python3` and `jq` named as prerequisites; column count corrected to sixteen.

Every query in the runbook was executed against the scratch database as written.

`go vet`, `CGO_ENABLED=0 go build ./...` and `go test ./...` all pass — no Go code changed.

Reviewed-on: #33
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
This commit was merged in pull request #33.
This commit is contained in:
2026-08-08 15:15:33 +07:00
committed by sulthan
parent 27cf0955de
commit 2cc1e69f5d
2 changed files with 281 additions and 3 deletions
+277
View File
@@ -0,0 +1,277 @@
# SQLite → Postgres cutover runbook
One-way, one-time. Moves the owner's reading history out of the retired SQLite
volume (`bookmarkmanager_bookmarks-data`, holding `/data/bookmarks.db`) and into
the Postgres schema the migration runner builds. There is no dual-write period:
the old database is read once, at cutover, from a **fresh export** — anything
written to SQLite after the export is lost, so the old API must already be down.
Routine deploys are `REDEPLOY.md`; first-time setup is `DEPLOY.md`. This file is
run once and then only ever read for reference.
Proven end to end on 2026-08-08 against a copy of `bookmarks-20260807-213515.db`
into a scratch Postgres: 29 Bookmarks (18 reading, 11 archived, 7 favourites),
29 Series, all owned by the seeded Reader, and every field of every row matching
the source exactly. Production was not touched.
---
## 0. The generator is throwaway
It is written at cutover, run once, and deleted. It is deliberately **not** in
this repository and never will be:
- Its output is the owner's personal reading history. That does not enter
version control.
- It reads SQLite. The backend module dropped `modernc.org/sqlite` (ADR-0001);
a committed generator would drag the dependency back in through the side door.
So §3 specifies the transformation rather than shipping a script. It is a
twenty-line program against a sixteen-column table (fifteen after `key`, which
is dropped) — writing it from the spec below costs less than maintaining it
would.
Beyond `DEPLOY.md`'s prerequisites (Docker and Compose), this runbook needs
`python3` on the machine running §3 — its stdlib `sqlite3` module is the whole
SQLite dependency — and `jq` for the one read-path check in §5. Neither has to
be the server: §3 only reads the snapshot copy, so it can run on a laptop and
the resulting `import.sql` be copied over.
---
## 1. Stop the old API and take a fresh export
**Order matters.** Export after the API stops, or you migrate a snapshot that is
already stale.
```bash
cd /opt/bookmarkmanager
COMPOSE="docker compose -f docker-compose.yml -f docker-compose.prod.yml"
BACKUP_DIR="$(cd .. && pwd)/bookmarkmanager-backups"; mkdir -p "$BACKUP_DIR"
STAMP=$(date -u +%Y%m%d-%H%M%S)
$COMPOSE stop bookmark-api
# Copy the file straight out of the retired volume. Nothing is writing to it,
# so a plain copy is consistent — no -wal to worry about after a clean stop.
docker run --rm -v bookmarkmanager_bookmarks-data:/from:ro -v "$BACKUP_DIR":/to \
alpine cp /from/bookmarks.db "/to/bookmarks-$STAMP.db"
ls -lh "$BACKUP_DIR/bookmarks-$STAMP.db"
```
Work on a **copy** of that file for the rest of this runbook. The export is the
last line of retreat; nothing below should be able to write to it.
```bash
mkdir -p /tmp/cutover && cp "$BACKUP_DIR/bookmarks-$STAMP.db" /tmp/cutover/snapshot.db
chmod 444 /tmp/cutover/snapshot.db
```
---
## 2. Bring up Postgres with the schema and the owner Reader
The new stack builds its own schema and seeds exactly one Reader from
`OWNER_DISCORD_ID` — do not hand-write either. Pull the Postgres-era commit
first: on a server that has only ever run the SQLite build, `--build` without a
pull silently rebuilds the old image and the checks below fail with
"relation readers does not exist".
```bash
git pull --ff-only
git log --oneline -1
# .env needs the new required vars (DATABASE_URL is built from
# POSTGRES_PASSWORD; TOKEN_KEY, OWNER_DISCORD_ID and the DISCORD_* set are
# required). Compose fails at start for a missing one.
git diff HEAD@{1} HEAD -- .env.example docker-compose.yml docker-compose.prod.yml
$COMPOSE up -d --build
docker logs bookmark-api --tail 20 # -> "listening on :8080"
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks -c '\dt'
# -> bookmarks, readers, schema_migrations, series, sessions
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks \
-c 'select id, discord_id from readers'
# -> exactly one row, and discord_id is the owner's
```
Two rows in `readers`, or zero, means `OWNER_DISCORD_ID` is wrong or the seed
failed. Stop here — the import attaches history to "the oldest reader row", and
that is only unambiguous while there is one.
`bookmarks` and `series` are empty at this point. That is what makes the import
a plain sequence of `INSERT`s with no conflict handling.
---
## 3. Generate the import SQL
Read `/tmp/cutover/snapshot.db` and emit plain SQL on stdout. The old table is
flat and its columns map one-for-one onto the split schema — no transformation
beyond the split itself:
| SQLite `bookmarks` column | lands in | notes |
|---|---|---|
| `site`, `series_id` | both tables | the Series key; the wire `key` column is dropped, it is re-derived as `site:series_id` on read |
| `title`, `series_url`, `cover`, `kind` | `series` | shared facts (ADR-0003) |
| `latest_chapter`, `latest_chapter_num`, `latest_checked_at` | `series` | `latest_chapter_num` is nullable on **both** sides and `NULL` is meaningful — never coerce it to `0` |
| `last_chapter`, `last_chapter_num`, `last_chapter_url` | `bookmarks` | Progress |
| `favorite`, `status`, `updated_at` | `bookmarks` | `favorite` is `0`/`1` in SQLite and a real `boolean` in Postgres — emit `true`/`false` |
| — | `bookmarks.reader_id` | the seeded owner |
**`latest_chapter_num` is the only column where `NULL` survives.** The SQLite
table declares `title`, `series_url`, `cover`, `last_chapter`,
`last_chapter_url` as bare `TEXT` and `last_chapter_num` as bare `REAL` — all
six nullable — while their Postgres targets are `NOT NULL DEFAULT ''` /
`NOT NULL DEFAULT 0`. One `NULL` in any of them aborts the whole import on a
not-null violation. Coalesce them in the `SELECT` (`ifnull(title,'')`,
`ifnull(last_chapter_num,0)`, …) rather than discovering it at §5. The
2026-08-07 export happened to have none; a fresh export is not promised the
same.
Rules the generator must follow:
- **Series first, Bookmarks second.** `bookmarks` has a foreign key onto
`series (site, series_id)`; the reverse order fails on the first row.
- **`SELECT DISTINCT` the Series.** The old key's uniqueness already makes
`(site, series_id)` unique, so this is belt and braces — but if it ever
collapses two rows, the count check in §5 catches it.
- **Never hardcode the reader id.** Emit
`INSERT INTO bookmarks (reader_id, …) SELECT id, … FROM owner`, where `owner`
is a temp table built once at the top:
`CREATE TEMP TABLE owner ON COMMIT DROP AS SELECT id FROM readers ORDER BY id LIMIT 1;`
A literal id is a number nobody verifies; this one cannot be wrong.
- **Wrap the whole file in `BEGIN; … COMMIT;`, temp table included.** Postgres
has transactional DDL and DML: a failure half way leaves an empty database
rather than half a library. The ordering is load-bearing —
`ON COMMIT DROP` outside the transaction means the temp table drops itself
the instant it is created (psql autocommits) and every
`SELECT … FROM owner` then fails.
- **Quote strings by doubling `'`.** Titles contain apostrophes and the URLs
contain `%5C%27` escapes. Emit standard SQL literals only — no `E''` strings,
no backslash escaping (`standard_conforming_strings` is on, so a backslash is
a literal backslash and the URLs survive verbatim).
```bash
python3 gen_import.py /tmp/cutover/snapshot.db > /tmp/cutover/import.sql
wc -l /tmp/cutover/import.sql # -> 2 header + 29 series + 29 bookmarks + framing
```
---
## 4. Review it by eye
29 rows is small enough to actually read, and this is the last point at which a
mistake is free:
```bash
less /tmp/cutover/import.sql
grep -c '^INSERT INTO series' /tmp/cutover/import.sql # -> 29
grep -c '^INSERT INTO bookmarks' /tmp/cutover/import.sql # -> 29
```
Look for: a title whose apostrophe is not doubled, a `favorite` that is still
`0`/`1`, a `latest_chapter_num` that turned into `0`, and any `reader_id`
written as a bare number.
---
## 5. Apply it
```bash
docker cp /tmp/cutover/import.sql "$($COMPOSE ps -q postgres)":/tmp/import.sql
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks -v ON_ERROR_STOP=1 \
-f /tmp/import.sql
```
`ON_ERROR_STOP=1` is not optional: without it `psql` reports the error, keeps
going, and exits `0` on a half-imported database.
Then the checklist. Every number here is asserted, not eyeballed:
```bash
OWNER=$(grep -E '^OWNER_DISCORD_ID=' .env | cut -d= -f2)
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks -x -c "
SELECT (SELECT count(*) FROM bookmarks) AS bookmarks_total,
(SELECT count(*) FROM bookmarks WHERE status='reading') AS reading,
(SELECT count(*) FROM bookmarks WHERE status='archived')AS archived,
(SELECT count(*) FROM series) AS series_total,
(SELECT count(*) FROM readers) AS readers_total,
(SELECT count(*) FROM bookmarks
WHERE reader_id <> (SELECT id FROM readers WHERE discord_id='$OWNER'))
AS not_owned_by_owner;"
```
`not_owned_by_owner` resolves the Reader by **Discord id**, not by
`ORDER BY id LIMIT 1`. The second form is the expression §3 tells the generator
to import with, so comparing against it is true by construction and could never
fail; resolving by Discord id is an independent check that the rows landed on
the identity the owner will actually log in as. If that subquery returns NULL
the whole count comes back `0` for the wrong reason — hence `readers_total`
beside it.
Expected, for the 2026-08-07 export: `29`, `18`, `11`, `29`, `1`, `0`. Against a
different export, the invariants rather than the literals are what hold:
- `bookmarks_total` equals the SQLite row count.
- `reading + archived` equals `bookmarks_total` (nothing was `finished`).
- `series_total` equals `SELECT count(*) FROM (SELECT DISTINCT site, series_id FROM bookmarks)`
in the source.
- `readers_total` is `1` and `not_owned_by_owner` is `0`.
Then spot-check the values themselves against the source — read position,
favourite flag and latest chapter. Take the sample from each bucket explicitly:
`ORDER BY updated_at DESC LIMIT 5` alone returns the most recently *progressed*
rows, which are the ones least likely to be archived.
```bash
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks -c "
SELECT s.title, b.last_chapter, b.last_chapter_num, b.favorite,
s.latest_chapter, b.status
FROM bookmarks b JOIN series s USING (site, series_id)
WHERE b.status='reading' ORDER BY b.updated_at DESC LIMIT 3;"
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks -c "
SELECT s.title, b.last_chapter, b.last_chapter_num, b.favorite,
s.latest_chapter, b.status
FROM bookmarks b JOIN series s USING (site, series_id)
WHERE b.status='archived' ORDER BY b.updated_at DESC LIMIT 2;"
$COMPOSE exec -T postgres psql -U bookmarks -d bookmarks -c "
SELECT s.title, b.last_chapter, b.last_chapter_num, b.favorite,
s.latest_chapter, b.status
FROM bookmarks b JOIN series s USING (site, series_id)
WHERE b.favorite ORDER BY b.updated_at DESC LIMIT 2;"
```
Compare each against the same row in the snapshot — the generator's own source
is the reference, so read it back with the same `python3` you used in §3.
Finally, prove the **read path**, not just the tables — this is the check that
would catch a correct import behind a broken join:
```bash
API=https://bookmark-api.violetcrown.my.id
TOKEN=$(grep -E '^API_TOKEN=' .env | cut -d= -f2) # grace-window credential
curl -s -H "Authorization: Bearer $TOKEN" $API/bookmarks | jq 'length' # -> 29
```
---
## 6. Afterwards
- **Keep the old SQLite volume for a month.** It is already undeclared in
compose, so `docker compose down -v` cannot take it. Remove it by hand once
the Postgres data has been trusted for a while:
`docker volume rm bookmarkmanager_bookmarks-data` (see `REDEPLOY.md` §1).
- **Delete the generator and the working copies:** `rm -rf /tmp/cutover`. The
timestamped export in `$BACKUP_DIR` is the copy that is kept.
- **Take the first Postgres dump immediately** — `REDEPLOY.md` §1. Until that
exists, the only backup of the migrated data is the SQLite file it came from.
If the import is wrong, there is nothing to unpick: drop the rows and start
again from §3 — `TRUNCATE bookmarks, series;` leaves the seeded Reader and the
schema in place.