206c447f36
Asura series slugs carry a site-wide build hash (-059befe1) that rotates on every redeploy, silently orphaning all asura bookmarks (old-hash URLs 302 to new-hash ones, so detect() yields keys that never match stored rows).
Backend:
- migrateAsuraKeys in OpenStore: one-off idempotent migration rewriting hashed asura keys to the stable hashless ID, merging collisions to newest updated_at; losers deleted before winner rewrite (PK-collision safe) — regression tests included
- poller latestChapterFrom: build hash made optional in chapter-scoping regex so latest_chapter survives redeploys
Userscript:
- stripBuildHash(/-[0-9a-f]{8}$/) applied to seriesId in both asura detect branches; URLs keep full slug (stale hashes 302)
- load-time migration rewrites cached + retry-queued asura keys to the stripped form so a queued PUT cannot resurrect an orphaned row
Docs: AGENTS.md + CLAUDE.md URL-shape notes; design spec at docs/superpowers/specs/2026-07-28-asura-stable-series-id-design.md
Verified: go test -count=1 ./... green (new rotation + collision tests), node --check green, regex verified against live asurascans.com slugs.
Deploy order: backend first (migration runs at OpenStore). Orphan rows created by not-yet-updated userscripts self-heal on next backend restart.
Reviewed-on: #6
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
83 lines
3.3 KiB
Go
83 lines
3.3 KiB
Go
package main
|
|
|
|
import (
|
|
"regexp"
|
|
"strconv"
|
|
"strings"
|
|
)
|
|
|
|
// latestChapter is the newest chapter a series page advertises.
|
|
type latestChapter struct {
|
|
Num float64
|
|
Label string
|
|
}
|
|
|
|
// asuraSlugRe pulls the series slug out of a stored series_url.
|
|
// Shape verified live 2026-07-26: https://asurascans.com/comics/<slug>, where
|
|
// the slug carries a trailing build-hash suffix (e.g. "-f886a8af") that
|
|
// rotates on every site redeploy — callers must strip it (asuraBuildHash)
|
|
// before using the slug to scope anything.
|
|
var asuraSlugRe = regexp.MustCompile(`/comics/([^/?#]+)`)
|
|
|
|
// demonicChapterRe matches the pre-redirect anchors demonic series pages link
|
|
// through. Both the raw "&" and the HTML-escaped "&" forms occur.
|
|
var demonicChapterRe = regexp.MustCompile(`chaptered\.php\?manga=\d+&(?:amp;)?chapter=([0-9.]+)`)
|
|
|
|
// latestChapterFrom returns the highest chapter number body advertises for this
|
|
// series. ok is false when the body yields nothing usable — an unknown site, an
|
|
// empty body, a Cloudflare challenge page, and a site redesign all land here,
|
|
// and the caller treats all four identically.
|
|
//
|
|
// Ported from the userscript's latestChapterFromAnchors (asura L123-133,
|
|
// demonic L183-193), including its reason for taking a maximum rather than a
|
|
// first or last: neither site lists chapters in a dependable order.
|
|
//
|
|
// The userscript's asura rule additionally requires the anchor text to match
|
|
// /Chapter\s+[\d.]+/i. That check exists only to skip the "First Chapter"
|
|
// shortcut, which points at chapter/1 and therefore can never win a maximum, so
|
|
// it is redundant here. For asura, scoping the pattern to this series' own slug
|
|
// replaces it with a stronger guarantee: a chapter link belonging to some other
|
|
// series cannot contribute even if the page starts carrying them. demonic has no
|
|
// such guarantee — demonicChapterRe matches any chaptered.php?manga=<id> anchor
|
|
// with no per-series scoping, because the stored series_id for demonic is a
|
|
// slug, not the numeric id the URL carries, so it cannot easily be scoped.
|
|
func latestChapterFrom(site, seriesURL, body string) (latestChapter, bool) {
|
|
var re *regexp.Regexp
|
|
switch site {
|
|
case "asura":
|
|
m := asuraSlugRe.FindStringSubmatch(seriesURL)
|
|
if m == nil {
|
|
return latestChapter{}, false
|
|
}
|
|
// Stored URLs predating a redeploy may carry a stale build hash;
|
|
// chapter hrefs in the fetched body carry the current one. Strip to
|
|
// the stable ID (same rule as migrateAsuraKeys) and make the hash
|
|
// optional in the pattern, so scoping survives rotations.
|
|
slug := asuraBuildHash.ReplaceAllString(m[1], "")
|
|
// Compiled per call rather than cached: this runs once per fetch, which
|
|
// is at most a few times a minute, and the slug varies per series.
|
|
re = regexp.MustCompile(`/comics/` + regexp.QuoteMeta(slug) + `(?:-[0-9a-f]{8})?/chapter/([0-9.]+)`)
|
|
case "demonic":
|
|
re = demonicChapterRe
|
|
default:
|
|
return latestChapter{}, false
|
|
}
|
|
|
|
var best latestChapter
|
|
found := false
|
|
for _, m := range re.FindAllStringSubmatch(body, -1) {
|
|
// [0-9.]+ can swallow a trailing separator, e.g. "chapter/12." in a
|
|
// sentence; ParseFloat would reject the whole match.
|
|
raw := strings.Trim(m[1], ".")
|
|
num, err := strconv.ParseFloat(raw, 64)
|
|
if err != nil {
|
|
continue
|
|
}
|
|
if !found || num > best.Num {
|
|
best = latestChapter{Num: num, Label: "Chapter " + raw}
|
|
found = true
|
|
}
|
|
}
|
|
return best, found
|
|
}
|