Files
mangaBookmark/backend/internal/latest/sites.go
T
sulthan 180ee78b1f Add comix.to and kagane.to support (#13)
Tracks read progress on comix.to and kagane.to alongside asura and demonic, in both the userscript and the backend.

Implements `docs/superpowers/plans/2026-08-03-comix-kagane-support.md`.

## Userscript

- `comix` adapter — `/title/<id>-<slug>`; only the id prefix is identity (the slug follows the title). No `og:image`, so the cover is matched by `alt`.
- `kagane` adapter — reader URLs are uuids with no chapter number, so it comes out of `og:title`; anchor scanning is structurally impossible, replaced by `latestChapterFromApi` against kagane's same-origin JSON API.
- `seriesId` threaded through `latestChapterFromAnchors` so comix can scope its scan to its own series and a recommendation strip cannot win the maximum.
- `@match` for both hosts, panel chips, v1.6.0.

## Backend

- `latestChapterFrom` cases: comix parses the SSR JSON state blob (`latestChapterUrl`, scoped to the series id); kagane parses API JSON (`chapter_no`).
- Poller allowlist extended; `Poller.BrowserFetch` with `fetcherFor(site)` routes kagane to a browser fetcher. Nil means kagane is not polled at all — never a fallback to the TLS fetcher, which would only ever retrieve a challenge page.
- `BrowserFetcher`: chromedp against a `headless-shell` sidecar. kagane sits behind a Cloudflare JS challenge that no TLS fingerprint clears, and the request is made inside the page rather than by replaying `cf_clearance`.
- `BROWSER_WS_URL` wiring, sidecar in both compose files (no `ports:`, dedicated non-external network), Dockerfile on `golang:1.26-alpine` — chromedp requires go 1.26.
- Web UI `--comix` / `--kagane` tokens in both colour branches.

## Notes for review

- `series_url` is client-supplied and a headless browser is a strong SSRF primitive, so kagane's host is pinned twice: in `fetchableSeriesURL` and again in `kaganeAPIURL`.
- Three chained defects found during verification made the browser path dead under Compose (sidecar flag collision, Chrome's Host-header DNS-rebinding check, the wrong chromedp option). Fixed; the compose comments record the wrong configurations too, so they don't get "simplified" back.
- `ALLOWED_ORIGINS` now includes both new origins. Without it every write from comix/kagane silently fails CORS preflight, parks in the retry queue, and drops at the cap.

## Verification

221 backend tests, 32 userscript tests, static `CGO_ENABLED=0` build, both compose configs.

Two gaps, both real:

1. The userscript on live pages via Violentmonkey needs a human browser profile — not run. Check: comix series page (title/cover, no chapter), comix chapter page (records the number; an *older* chapter must not regress it), comix SPA navigation without reload, kagane series page (og:image cover), kagane reader (number from `og:title`), both chips opening the right sites.
2. The kagane browser path has not completed end-to-end anywhere. Dial/navigate/fetch is confirmed, but Cloudflare 403'd headless-shell's Chrome on every attempt from the dev sandbox, and comix's poll-through-Docker was blocked by that environment's TLS interception. Both environment-dependent rather than branch defects — the first real deploy is the actual verification.

Reviewed-on: #13
Co-authored-by: Sulthan Zaki <sultankiki05@gmail.com>
Co-committed-by: Sulthan Zaki <sultankiki05@gmail.com>
2026-08-03 19:53:45 +07:00

110 lines
4.4 KiB
Go

package latest
import (
"regexp"
"strconv"
"strings"
"mangabm/backend/internal/store"
)
// latestChapter is the newest chapter a series page advertises.
type latestChapter struct {
Num float64
Label string
}
// asuraSlugRe pulls the series slug out of a stored series_url.
// Shape verified live 2026-07-26: https://asurascans.com/comics/<slug>, where
// the slug carries a trailing build-hash suffix (e.g. "-f886a8af") that
// rotates on every site redeploy — callers must strip it (asuraBuildHash)
// before using the slug to scope anything.
var asuraSlugRe = regexp.MustCompile(`/comics/([^/?#]+)`)
// demonicChapterRe matches the pre-redirect anchors demonic series pages link
// through. Both the raw "&" and the HTML-escaped "&amp;" forms occur.
var demonicChapterRe = regexp.MustCompile(`chaptered\.php\?manga=\d+&(?:amp;)?chapter=([0-9.]+)`)
// comixSlugRe pulls the "<id>-<slug>" segment out of a stored series_url.
// Only the id prefix is stable; the slug tail follows the title.
var comixSlugRe = regexp.MustCompile(`/title/([^/?#]+)`)
// kaganeChapterRe matches the chapter numbers in a kagane API response. This
// branch is fed by the browser fetcher, so the body is JSON rather than HTML —
// there are no anchors to scan.
var kaganeChapterRe = regexp.MustCompile(`"chapter_no":"([0-9.]+)"`)
// latestChapterFrom returns the highest chapter number body advertises for this
// series. ok is false when the body yields nothing usable — an unknown site, an
// empty body, a Cloudflare challenge page, and a site redesign all land here,
// and the caller treats all four identically.
//
// Ported from the userscript's latestChapterFromAnchors (asura L123-133,
// demonic L183-193), including its reason for taking a maximum rather than a
// first or last: neither site lists chapters in a dependable order.
//
// The userscript's asura rule additionally requires the anchor text to match
// /Chapter\s+[\d.]+/i. That check exists only to skip the "First Chapter"
// shortcut, which points at chapter/1 and therefore can never win a maximum, so
// it is redundant here. For asura, scoping the pattern to this series' own slug
// replaces it with a stronger guarantee: a chapter link belonging to some other
// series cannot contribute even if the page starts carrying them. demonic has no
// such guarantee — demonicChapterRe matches any chaptered.php?manga=<id> anchor
// with no per-series scoping, because the stored series_id for demonic is a
// slug, not the numeric id the URL carries, so it cannot easily be scoped.
func latestChapterFrom(site, seriesURL, body string) (latestChapter, bool) {
var re *regexp.Regexp
switch site {
case "asura":
m := asuraSlugRe.FindStringSubmatch(seriesURL)
if m == nil {
return latestChapter{}, false
}
// Stored URLs predating a redeploy may carry a stale build hash;
// chapter hrefs in the fetched body carry the current one. Strip to
// the stable ID (same rule as migrateAsuraKeys) and make the hash
// optional in the pattern, so scoping survives rotations.
slug := store.AsuraBuildHash.ReplaceAllString(m[1], "")
// Compiled per call rather than cached: this runs once per fetch, which
// is at most a few times a minute, and the slug varies per series.
re = regexp.MustCompile(`/comics/` + regexp.QuoteMeta(slug) + `(?:-[0-9a-f]{8})?/chapter/([0-9.]+)`)
case "demonic":
re = demonicChapterRe
case "comix":
m := comixSlugRe.FindStringSubmatch(seriesURL)
if m == nil {
return latestChapter{}, false
}
// comix ships an SPA: the served HTML carries a JSON state blob instead
// of chapter anchors, and latestChapterUrl is the only place the newest
// chapter appears. Scoping to this series' id prefix keeps a
// "recommended" strip's entries from winning the maximum.
id := m[1]
if i := strings.Index(id, "-"); i != -1 {
id = id[:i]
}
re = regexp.MustCompile(`"latestChapterUrl":"/title/` + regexp.QuoteMeta(id) + `-[^"]*-chapter-([0-9.]+)"`)
case "kagane":
re = kaganeChapterRe
default:
return latestChapter{}, false
}
var best latestChapter
found := false
for _, m := range re.FindAllStringSubmatch(body, -1) {
// [0-9.]+ can swallow a trailing separator, e.g. "chapter/12." in a
// sentence; ParseFloat would reject the whole match.
raw := strings.Trim(m[1], ".")
num, err := strconv.ParseFloat(raw, 64)
if err != nil {
continue
}
if !found || num > best.Num {
best = latestChapter{Num: num, Label: "Chapter " + raw}
found = true
}
}
return best, found
}