Cut production over #26

Closed
opened 2026-08-08 06:06:29 +07:00 by sulthan · 4 comments
Owner

Parent

#18

What to build

The one-way door. Export fresh, import, verify, and confirm the whole system works end to end against real data — the owner's library intact, Discord login working, a userscript syncing, and an already-installed script still working on the old token.

There is no dual-write period and no going back except by restoring the retained volume.

Acceptance criteria

  • The export is taken fresh at cutover. The earlier snapshot is not used — it will have drifted.
  • The generated SQL is reviewed before it is applied.
  • Verification is re-run against production: counts, lifecycle buckets, Series count, and ownership all as specified in the import ticket.
  • The owner logs into production with Discord and sees their library.
  • The owner installs a rendered userscript and records a chapter read, end to end, against production.
  • An already-installed userscript still syncs using the retired global token, confirming the grace path works in production rather than only in tests.
  • The poller is confirmed running against Series in production, at unchanged request rate.
  • The old SQLite volume is retained, not deleted. Deleting it is explicitly not part of this ticket.
  • The backup runbook is rewritten for Postgres, the environment contract is updated, and retired variables are removed from documentation.

Blocked by

#24, #25

## Parent #18 ## What to build The one-way door. Export fresh, import, verify, and confirm the whole system works end to end against real data — the owner's library intact, Discord login working, a userscript syncing, and an already-installed script still working on the old token. There is no dual-write period and no going back except by restoring the retained volume. ## Acceptance criteria - [x] The export is taken fresh at cutover. The earlier snapshot is not used — it will have drifted. - [x] The generated SQL is reviewed before it is applied. - [x] Verification is re-run against production: counts, lifecycle buckets, Series count, and ownership all as specified in the import ticket. - [x] The owner logs into production with Discord and sees their library. - [x] The owner installs a rendered userscript and records a chapter read, end to end, against production. - [x] An already-installed userscript still syncs using the retired global token, confirming the grace path works in production rather than only in tests. - [x] The poller is confirmed running against Series in production, at unchanged request rate. - [x] The old SQLite volume is retained, not deleted. Deleting it is explicitly not part of this ticket. - [x] The backup runbook is rewritten for Postgres, the environment contract is updated, and retired variables are removed from documentation. ## Blocked by #24, #25
sulthan added the ready-for-agent label 2026-08-08 06:06:29 +07:00
Author
Owner

Cutover staged and rehearsed against production, but not executed — blocked on the four Discord values in the server's .env.

Branch: feat/cut-production-over (2 commits off main).

Rehearsal against the real host

The runbook was dry-run against a warm copy of the live SQLite volume. Numbers match #25 exactly:

  • 29 bookmarks, 18 reading / 11 archived, all kind=manga
  • 29 distinct (site, series_id) pairs → 29 Series
  • 7 favourites
  • 0 NULLs in the generated SQL; the 3 titles carrying an ASCII apostrophe (I'm Being Misunderstood as a Soccer Genius, The Chaebeol's Youngest Son, The Nebula's Civilization) come out with balanced doubled quotes
  • favorite emitted as true/false, never 0/1; reader_id never a literal
  • Generator output: 61 lines = BEGIN + temp table + 29 series + 29 bookmarks + COMMIT
  • Pre-cutover read-path baseline: GET /bookmarks returns 29

The generator lives only at /tmp/cutover/gen_import.py on the server. Not committed, per #25.

Runbook defects the rehearsal found

CUTOVER.md as merged would have failed mid-cutover. Fixed in this branch:

  1. Wrong volume name. Both runbooks hardcoded bookmarkmanager_bookmarks-data. The compose project is the lowercased directory name, and the checkout is ~/mangaBookmark, so the volume is mangabookmark_bookmarks-data. Now derived exactly — not with docker volume ls --filter name=, which is a substring match that can return several volumes into a -v mount or an rm.
  2. WAL assumed away. §1 claimed a clean stop leaves no -wal. compose stop SIGKILLs after 10s and the live volume currently carries a 1.1 MB -wal; copying bookmarks.db alone would have silently dropped whatever it held. Now asserted, with a fallback.
  3. jq is not installed on the server. §5's read-path check now uses python3, which the runbook already requires.
  4. Stale $VOL. §6 removes the retired volume "once the Postgres data has been trusted for a while" — a month later, in a shell where $VOL is unset. Now self-contained.
  5. Paths disagreed. REDEPLOY.md and DEPLOY.md pointed at /opt/bookmarkmanager, so §6's handoff to REDEPLOY.md §1 sent the operator to a path that does not exist here.
  6. REDEPLOY.md §1 listed the pre-split schema — readers and sessions missing from both the \dt output and the pg_restore --list contents.

Environment contract (last AC)

README.md's config table listed 8 of the backend's 25 variables, omitting the entire DISCORD_* set the backend refuses to start without, plus the poller and userscript-path vars. Completed, with the compose-only variables (POSTGRES_PASSWORD, BOOKMARK_*_HOST, PROXY_NETWORK, TRAEFIK_*) called out as such. Retired variables were already clean: WEB_PASSWORD survives only in ADR-0002 as history, and API_TOKEN is deliberately retained and labelled for the grace window.

Also fixed

Production git pull was broken — the checkout's remote was the HTTPS clone URL with no credential helper, so REDEPLOY.md §2 died on could not read Username. Switched to SSH on Gitea's port 2222 (port 22 is the host's own sshd and rejects every key). Recorded in the troubleshooting table.

Staged on the server, nothing disruptive

  • Code pulled to 2cc1e69; bookmarkmanager-backend:latest rebuilt; postgres:17-alpine pulled — so the downtime window is an import, not a build.
  • .env backed up to .env.pre-cutover.bak; TOKEN_KEY and POSTGRES_PASSWORD generated; API_TOKEN_GRACE_UNTIL=2026-08-22; DISCORD_REDIRECT_URI set.
  • The old API is still serving on the old image. Nothing one-way has happened.

Blocked

OWNER_DISCORD_ID, DISCORD_CLIENT_ID, DISCORD_CLIENT_SECRET, DISCORD_GUILD_ID are empty in ~/mangaBookmark/.env; compose refuses to interpolate. Once they are set, remaining ACs run in order: stop + fresh export, generate, review, up, import, verify, then the owner's Discord login, userscript install, chapter read, and the retired-token sync check.

go test ./... green (uncached).

Cutover staged and rehearsed against production, but **not executed** — blocked on the four Discord values in the server's `.env`. Branch: `feat/cut-production-over` (2 commits off `main`). ## Rehearsal against the real host The runbook was dry-run against a warm copy of the live SQLite volume. Numbers match #25 exactly: - 29 bookmarks, 18 `reading` / 11 `archived`, all `kind=manga` - 29 distinct `(site, series_id)` pairs → 29 Series - 7 favourites - 0 NULLs in the generated SQL; the 3 titles carrying an ASCII apostrophe (`I'm Being Misunderstood as a Soccer Genius`, `The Chaebeol's Youngest Son`, `The Nebula's Civilization`) come out with balanced doubled quotes - `favorite` emitted as `true`/`false`, never `0`/`1`; `reader_id` never a literal - Generator output: 61 lines = `BEGIN` + temp table + 29 series + 29 bookmarks + `COMMIT` - Pre-cutover read-path baseline: `GET /bookmarks` returns 29 The generator lives only at `/tmp/cutover/gen_import.py` on the server. Not committed, per #25. ## Runbook defects the rehearsal found `CUTOVER.md` as merged would have failed mid-cutover. Fixed in this branch: 1. **Wrong volume name.** Both runbooks hardcoded `bookmarkmanager_bookmarks-data`. The compose project is the lowercased *directory* name, and the checkout is `~/mangaBookmark`, so the volume is `mangabookmark_bookmarks-data`. Now derived exactly — not with `docker volume ls --filter name=`, which is a substring match that can return several volumes into a `-v` mount or an `rm`. 2. **WAL assumed away.** §1 claimed a clean stop leaves no `-wal`. `compose stop` SIGKILLs after 10s and the live volume currently carries a 1.1 MB `-wal`; copying `bookmarks.db` alone would have silently dropped whatever it held. Now asserted, with a fallback. 3. **`jq` is not installed on the server.** §5's read-path check now uses `python3`, which the runbook already requires. 4. **Stale `$VOL`.** §6 removes the retired volume "once the Postgres data has been trusted for a while" — a month later, in a shell where `$VOL` is unset. Now self-contained. 5. **Paths disagreed.** `REDEPLOY.md` and `DEPLOY.md` pointed at `/opt/bookmarkmanager`, so §6's handoff to `REDEPLOY.md` §1 sent the operator to a path that does not exist here. 6. **`REDEPLOY.md` §1 listed the pre-split schema** — `readers` and `sessions` missing from both the `\dt` output and the `pg_restore --list` contents. ## Environment contract (last AC) `README.md`'s config table listed 8 of the backend's 25 variables, omitting the entire `DISCORD_*` set the backend refuses to start without, plus the poller and userscript-path vars. Completed, with the compose-only variables (`POSTGRES_PASSWORD`, `BOOKMARK_*_HOST`, `PROXY_NETWORK`, `TRAEFIK_*`) called out as such. Retired variables were already clean: `WEB_PASSWORD` survives only in ADR-0002 as history, and `API_TOKEN` is deliberately retained and labelled for the grace window. ## Also fixed Production `git pull` was broken — the checkout's remote was the HTTPS clone URL with no credential helper, so `REDEPLOY.md` §2 died on `could not read Username`. Switched to SSH on Gitea's port 2222 (port 22 is the host's own sshd and rejects every key). Recorded in the troubleshooting table. ## Staged on the server, nothing disruptive - Code pulled to `2cc1e69`; `bookmarkmanager-backend:latest` rebuilt; `postgres:17-alpine` pulled — so the downtime window is an import, not a build. - `.env` backed up to `.env.pre-cutover.bak`; `TOKEN_KEY` and `POSTGRES_PASSWORD` generated; `API_TOKEN_GRACE_UNTIL=2026-08-22`; `DISCORD_REDIRECT_URI` set. - The old API is still serving on the old image. Nothing one-way has happened. ## Blocked `OWNER_DISCORD_ID`, `DISCORD_CLIENT_ID`, `DISCORD_CLIENT_SECRET`, `DISCORD_GUILD_ID` are empty in `~/mangaBookmark/.env`; compose refuses to interpolate. Once they are set, remaining ACs run in order: stop + fresh export, generate, review, `up`, import, verify, then the owner's Discord login, userscript install, chapter read, and the retired-token sync check. `go test ./...` green (uncached).
Author
Owner

Cutover executed

Production is on Postgres. main at 1b1820d.

§1 — fresh export

Old API stopped first, then exported. The WAL assertion earned its place: the live volume had been carrying a 1.1 MB -wal, and the clean SIGTERM checkpointed it away, leaving bookmarks.db alone at 24 KB — verified before copying rather than assumed.

../mangaBookmark-backups/bookmarks-20260808-090657.db

The earlier snapshot was not reused, and the freshness shows: The-Outcast-Is-Too-Good-at-Martial-Arts reads Chapter 137 in this export where the rehearsal copy had 136.

§2 — schema and owner Reader

Five tables built by the migration runner. Exactly one row in readers, discord_id = the configured owner. bookmarks and series empty before the import.

§3–4 — generated and reviewed

61 lines: BEGIN, temp owner table, 29 series, 29 bookmarks, COMMIT. Read in full before applying. Apostrophes doubled (The Nebula''s Civilization, I''m Being Misunderstood as a Soccer Genius, The Chaebeol''s Youngest Son); Unicode ’ left alone; favorite emitted true/false; every reader_id resolved via SELECT id ... FROM owner, never a literal; the %5C%27 URL escapes survive verbatim.

§5 — applied and verified

bookmarks_total    | 29
reading            | 18
archived           | 11
finished           |  0
series_total       | 29
readers_total      |  1
favorites          |  7
not_owned_by_owner |  0

not_owned_by_owner resolves the Reader by Discord id, independently of the ORDER BY id LIMIT 1 the import used, and readers_total is beside it so a NULL subquery cannot fake a zero.

Rather than sampling, every migrated row was compared against the snapshot it came from — 29 rows × 13 columns, 0 mismatches, including the nullable latest_chapter_num and the updated_at ordering key.

Read path over HTTPS: /healthz ok, /bookmarks 401 unauthenticated, 401 on a wrong credential, 29 rows on the retired global token, CORS preflight 204 echoing https://asurascans.com, web UI 200.

The grace path is audited in production, not just in tests:

auth: retired global token accepted for owner reader 1 (grace until 2026-08-22T00:00:00Z)

§6 — afterwards

  • mangabookmark_bookmarks-data retained, still undeclared in compose.
  • First Postgres dump taken and verified: bookmarks-20260808-091332.dump, pg_restore --list shows all five tables.
  • /tmp/cutover deliberately left in place until the remaining checks pass — it holds the snapshot and the comparison script, which are the reference if anything looks wrong. rm -rf /tmp/cutover afterwards.

A blocker found on the way

The guild-membership check called GET /guilds/{guild}/members/{user} — the Guild resource's Get Guild Member, which requires a Bot token and the application present in the guild. Handed a user Bearer token it answers 401, which discordMember reported as an error, so every sign-in would have rendered "Discord sign-in is unavailable right now". Found before the cutover rather than during it.

guilds.members.read grants Get Current User Guild Member, GET /users/@me/guilds/{guild}/member. Same single-guild question, same privacy property, and it accepts the token we hold. #18 flagged this as verified from Discord's documentation but never from a live flow — it was wrong.

The test stub mirrored the implementation, so the suite was blind to it. It now serves the OAuth path and answers the bot path 401 the way Discord does; reverting the fix fails the test with the exact production symptom.

Still open

  • Owner signs in with Discord and sees the library.
  • Owner installs a rendered userscript and records a chapter read.
  • An already-installed script syncs on the retired token from a real device. The API half is proven above (29 rows, logged); the device half is not.
  • Poller confirmed against Series at unchanged rate. Configuration is logged as interval=10m cooldown=1h batch=14 stagger=20s — identical to the pre-cutover values — and the first wake was due 09:21:21 UTC. Owner is verifying.
## Cutover executed Production is on Postgres. `main` at `1b1820d`. ### §1 — fresh export Old API stopped first, then exported. The WAL assertion earned its place: the live volume had been carrying a 1.1 MB `-wal`, and the clean SIGTERM checkpointed it away, leaving `bookmarks.db` alone at 24 KB — verified before copying rather than assumed. `../mangaBookmark-backups/bookmarks-20260808-090657.db` The earlier snapshot was not reused, and the freshness shows: `The-Outcast-Is-Too-Good-at-Martial-Arts` reads Chapter 137 in this export where the rehearsal copy had 136. ### §2 — schema and owner Reader Five tables built by the migration runner. Exactly one row in `readers`, `discord_id` = the configured owner. `bookmarks` and `series` empty before the import. ### §3–4 — generated and reviewed 61 lines: `BEGIN`, temp `owner` table, 29 series, 29 bookmarks, `COMMIT`. Read in full before applying. Apostrophes doubled (`The Nebula''s Civilization`, `I''m Being Misunderstood as a Soccer Genius`, `The Chaebeol''s Youngest Son`); Unicode `’` left alone; `favorite` emitted `true`/`false`; every `reader_id` resolved via `SELECT id ... FROM owner`, never a literal; the `%5C%27` URL escapes survive verbatim. ### §5 — applied and verified ``` bookmarks_total | 29 reading | 18 archived | 11 finished | 0 series_total | 29 readers_total | 1 favorites | 7 not_owned_by_owner | 0 ``` `not_owned_by_owner` resolves the Reader by Discord id, independently of the `ORDER BY id LIMIT 1` the import used, and `readers_total` is beside it so a NULL subquery cannot fake a zero. Rather than sampling, every migrated row was compared against the snapshot it came from — **29 rows × 13 columns, 0 mismatches**, including the nullable `latest_chapter_num` and the `updated_at` ordering key. Read path over HTTPS: `/healthz` ok, `/bookmarks` 401 unauthenticated, 401 on a wrong credential, **29 rows on the retired global token**, CORS preflight 204 echoing `https://asurascans.com`, web UI 200. The grace path is audited in production, not just in tests: ``` auth: retired global token accepted for owner reader 1 (grace until 2026-08-22T00:00:00Z) ``` ### §6 — afterwards - `mangabookmark_bookmarks-data` **retained**, still undeclared in compose. - First Postgres dump taken and verified: `bookmarks-20260808-091332.dump`, `pg_restore --list` shows all five tables. - `/tmp/cutover` deliberately left in place until the remaining checks pass — it holds the snapshot and the comparison script, which are the reference if anything looks wrong. `rm -rf /tmp/cutover` afterwards. ## A blocker found on the way The guild-membership check called `GET /guilds/{guild}/members/{user}` — the Guild resource's *Get Guild Member*, which requires a Bot token and the application present in the guild. Handed a user Bearer token it answers 401, which `discordMember` reported as an error, so **every sign-in would have rendered "Discord sign-in is unavailable right now"**. Found before the cutover rather than during it. `guilds.members.read` grants *Get Current User Guild Member*, `GET /users/@me/guilds/{guild}/member`. Same single-guild question, same privacy property, and it accepts the token we hold. #18 flagged this as verified from Discord's documentation but never from a live flow — it was wrong. The test stub mirrored the implementation, so the suite was blind to it. It now serves the OAuth path and answers the bot path 401 the way Discord does; reverting the fix fails the test with the exact production symptom. ## Still open - Owner signs in with Discord and sees the library. - Owner installs a rendered userscript and records a chapter read. - An already-installed script syncs on the retired token from a real device. The API half is proven above (29 rows, logged); the device half is not. - Poller confirmed against Series at unchanged rate. Configuration is logged as `interval=10m cooldown=1h batch=14 stagger=20s` — identical to the pre-cutover values — and the first wake was due 09:21:21 UTC. Owner is verifying.
Author
Owner

Owner-verified against production: Discord login works and shows the library, the poller is running normally, and an already-installed userscript is still syncing on the retired global token. Those three are ticked.

The install AC stays open, and it turned up a real bug. Install works on desktop, but on mobile Violentmonkey (Chromium) the Install link does nothing useful — Violentmonkey does not intercept navigation to a .user.js URL there, so the browser just renders the script as text and there is no path to installing it.

Fixed in #35: ?download=1 on the same session-gated install endpoint sets Content-Disposition: attachment, so the file saves and can be added from Violentmonkey's own menu. The setup panel now offers Download links alongside Install. The plain link stays inline deliberately — the updater polls the /u/ path and an attachment disposition there would break auto-update.

Remaining before this closes: install a script on mobile via the new Download link and record a chapter read end to end.

Owner-verified against production: Discord login works and shows the library, the poller is running normally, and an already-installed userscript is still syncing on the retired global token. Those three are ticked. The install AC stays open, and it turned up a real bug. Install works on desktop, but on mobile Violentmonkey (Chromium) the Install link does nothing useful — Violentmonkey does not intercept navigation to a `.user.js` URL there, so the browser just renders the script as text and there is no path to installing it. Fixed in #35: `?download=1` on the same session-gated install endpoint sets `Content-Disposition: attachment`, so the file saves and can be added from Violentmonkey's own menu. The setup panel now offers Download links alongside Install. The plain link stays inline deliberately — the updater polls the `/u/` path and an attachment disposition there would break auto-update. Remaining before this closes: install a script on mobile via the new Download link and record a chapter read end to end.
Author
Owner

I verified that the userscript download is working perfectly fine. and the userscript install is good. i check the userscript install in a desktop browser. The poller working just fine.

I verified that the userscript download is working perfectly fine. and the userscript install is good. i check the userscript install in a desktop browser. The poller working just fine.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: sulthan/mangaBookmark#26