Files
mangaBookmark/backend/internal/latest/sites.go
T
sulthan 08749df050 feat(backend)!: run on Postgres with a migration-owned schema (#28)
Swap modernc.org/sqlite for jackc/pgx/v5 with no observable change:
same endpoints, same wire format, same updated_at ordering rule.

The schema now comes from numbered SQL embedded in the binary and
applied on startup, one transaction each, recorded in
schema_migrations. That replaces two pieces of SQLite-era machinery,
both deleted rather than ported: the column probing (Postgres has ADD
COLUMN IF NOT EXISTS, and there is no legacy database left to probe)
and the Asura key rewrite, which has run clean on every start for
months now that the userscripts strip build hashes before writing. Its
regexp survives as latest.asuraBuildHash, where the poller still needs
it to scope chapter links to a series whose slug carries a rotating
hash.

Types get real: favorite is a boolean, chapter numbers double
precision, timestamps stay unix-ms bigint. SQLite's null-safe IS NOT
becomes IS DISTINCT FROM, which is what implements the rule that only
reading progress reorders a list. Inside COALESCE/NULLIF the status
and kind parameters need an explicit ::text -- there is no target
column to infer from and Postgres refuses to guess.

Tests lose their free t.TempDir() database, so Docker is now a hard
prerequisite for `go test ./...`: internal/pgtest starts one
postgres:17-alpine per test binary and hands each test a database of
its own.

Also lands CONTEXT.md and the four ADRs written while scoping #18.

BREAKING CHANGE: DB_PATH is retired for DATABASE_URL, which is
required and has no default. Compose gains a postgres service on an
internal network with its own volume; POSTGRES_PASSWORD joins .env.
The old bookmarks-data volume is deliberately left undeclared so
`docker compose down -v` cannot take the pre-migration database with
it. main is not deployable until #25 and #26 land.

Closes #20

Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-08 06:52:20 +07:00

149 lines
6.0 KiB
Go

package latest
import (
"net/url"
"regexp"
"strconv"
"strings"
)
// latestChapter is the newest chapter a series page advertises.
type latestChapter struct {
Num float64
Label string
}
// asuraSlugRe pulls the series slug out of a stored series_url.
// Shape verified live 2026-07-26: https://asurascans.com/comics/<slug>, where
// the slug carries a trailing build-hash suffix (e.g. "-f886a8af") that
// rotates on every site redeploy — callers must strip it (asuraBuildHash)
// before using the slug to scope anything.
var asuraSlugRe = regexp.MustCompile(`/comics/([^/?#]+)`)
// asuraBuildHash matches the trailing "-xxxxxxxx" site-wide build ID Asura
// appends to every series slug. It rotates on each site redeploy, so it is
// never part of a stable series_id. Must stay in sync with stripBuildHash in
// userscript/manga-bookmark.user.js.
var asuraBuildHash = regexp.MustCompile(`-[0-9a-f]{8}$`)
// demonicChapterRe matches the pre-redirect anchors demonic series pages link
// through. Both the raw "&" and the HTML-escaped "&amp;" forms occur.
var demonicChapterRe = regexp.MustCompile(`chaptered\.php\?manga=\d+&(?:amp;)?chapter=([0-9.]+)`)
// comixSlugRe pulls the "<id>-<slug>" segment out of a stored series_url.
// Only the id prefix is stable; the slug tail follows the title.
var comixSlugRe = regexp.MustCompile(`/title/([^/?#]+)`)
// kaganeChapterRe matches the chapter numbers in a kagane API response. This
// branch is fed by the browser fetcher, so the body is JSON rather than HTML —
// there are no anchors to scan.
var kaganeChapterRe = regexp.MustCompile(`"chapter_no":"([0-9.]+)"`)
// novelfullSlugRe pulls the series slug out of a stored series_url. novelfull
// series pages are "/<slug>.html"; their chapter anchors are
// "/<slug>/chapter-<n>[-<title-slug>].html". Verified live 2026-08-05.
var novelfullSlugRe = regexp.MustCompile(`^/([^/?#]+)\.html$`)
// lnwSlugRe does the same for lightnovelworld, whose series pages live under
// /novel/<slug>/ while its chapter URLs are flat at the site root:
// "/<slug>-chapter-<n>/", absolute in the page's own anchors. Verified live
// 2026-08-05.
var lnwSlugRe = regexp.MustCompile(`^/novel/([^/?#]+)/?$`)
// latestChapterFrom returns the highest chapter number body advertises for this
// series. ok is false when the body yields nothing usable — an unknown site, an
// empty body, a Cloudflare challenge page, and a site redesign all land here,
// and the caller treats all four identically.
//
// Ported from the userscript's latestChapterFromAnchors (asura L123-133,
// demonic L183-193), including its reason for taking a maximum rather than a
// first or last: neither site lists chapters in a dependable order.
//
// The userscript's asura rule additionally requires the anchor text to match
// /Chapter\s+[\d.]+/i. That check exists only to skip the "First Chapter"
// shortcut, which points at chapter/1 and therefore can never win a maximum, so
// it is redundant here. For asura, scoping the pattern to this series' own slug
// replaces it with a stronger guarantee: a chapter link belonging to some other
// series cannot contribute even if the page starts carrying them. demonic has no
// such guarantee — demonicChapterRe matches any chaptered.php?manga=<id> anchor
// with no per-series scoping, because the stored series_id for demonic is a
// slug, not the numeric id the URL carries, so it cannot easily be scoped.
func latestChapterFrom(site, seriesURL, body string) (latestChapter, bool) {
var re *regexp.Regexp
switch site {
case "asura":
m := asuraSlugRe.FindStringSubmatch(seriesURL)
if m == nil {
return latestChapter{}, false
}
// Stored URLs predating a redeploy may carry a stale build hash;
// chapter hrefs in the fetched body carry the current one. Strip to
// the stable ID and make the hash optional in the pattern, so scoping
// survives rotations.
slug := asuraBuildHash.ReplaceAllString(m[1], "")
// Compiled per call rather than cached: this runs once per fetch, which
// is at most a few times a minute, and the slug varies per series.
re = regexp.MustCompile(`/comics/` + regexp.QuoteMeta(slug) + `(?:-[0-9a-f]{8})?/chapter/([0-9.]+)`)
case "demonic":
re = demonicChapterRe
case "comix":
m := comixSlugRe.FindStringSubmatch(seriesURL)
if m == nil {
return latestChapter{}, false
}
// comix ships an SPA: the served HTML carries a JSON state blob instead
// of chapter anchors, and latestChapterUrl is the only place the newest
// chapter appears. Scoping to this series' id prefix keeps a
// "recommended" strip's entries from winning the maximum.
id := m[1]
if i := strings.Index(id, "-"); i != -1 {
id = id[:i]
}
re = regexp.MustCompile(`"latestChapterUrl":"/title/` + regexp.QuoteMeta(id) + `-[^"]*-chapter-([0-9.]+)"`)
case "kagane":
re = kaganeChapterRe
case "novelfull":
u, err := url.Parse(seriesURL)
if err != nil {
return latestChapter{}, false
}
m := novelfullSlugRe.FindStringSubmatch(u.Path)
if m == nil {
return latestChapter{}, false
}
// Scoped to this series' slug for the same reason asura is: page 1
// carries a "latest chapters" widget and a "you may also like" strip,
// and neither may contribute to the maximum.
re = regexp.MustCompile(`/` + regexp.QuoteMeta(m[1]) + `/chapter-([0-9.]+)`)
case "lightnovelworld":
u, err := url.Parse(seriesURL)
if err != nil {
return latestChapter{}, false
}
m := lnwSlugRe.FindStringSubmatch(u.Path)
if m == nil {
return latestChapter{}, false
}
re = regexp.MustCompile(`lightnovelworld\.net/` + regexp.QuoteMeta(m[1]) + `-chapter-([0-9.]+)/`)
default:
return latestChapter{}, false
}
var best latestChapter
found := false
for _, m := range re.FindAllStringSubmatch(body, -1) {
// [0-9.]+ can swallow a trailing separator, e.g. "chapter/12." in a
// sentence; ParseFloat would reject the whole match.
raw := strings.Trim(m[1], ".")
num, err := strconv.ParseFloat(raw, 64)
if err != nil {
continue
}
if !found || num > best.Num {
best = latestChapter{Num: num, Label: "Chapter " + raw}
found = true
}
}
return best, found
}