Discord webhook doesnt work #178

Open
opened 2026-09-23 08:51:09 +07:00 by sulthan · 1 comment
Owner

the browser is not reachable, the thing is i dont even get any webhook message from the discord bot. I want the discord bot to be as informatice as they can. Reporting any error event that are happening in the server

the browser is not reachable, the thing is i dont even get any webhook message from the discord bot. I want the discord bot to be as informatice as they can. Reporting any error event that are happening in the server
sulthan added the bugneeds-triage labels 2026-09-23 08:55:48 +07:00
Author
Owner

This was generated by AI during triage.

Triage notes

Split: the "report any error event" half is now #179 (enhancement, needs-triage).
This ticket is the bug — the browser was unreachable and no message arrived.

What the code says should have happened. The path is wired end to end: main
builds a notify.Client whenever DISCORD_WEBHOOK_URL is non-empty and logs
owner notices: enabled / owner notices: disabled (DISCORD_WEBHOOK_URL unset) at
startup, and Poller.ownerNotices judges latest.FaultsFrom beside every pass-log
insert. For an unreachable browser, two of the four conditions apply:

  • The first wake after the browser goes away fails on the first Series, so the pass
    row is stall-shaped (due > 0, checked 0, empty skip) — see the outcomeUnreachable
    return in runLanePass. ConditionStall carries no twelve-hour window, so it
    should fire on that pass, not half a day later.
  • Sibling browser Lanes then skip for RefuseBackoff, and once every browser-backed
    Site's latest pass is SkipSidecarDown/SkipNoFetcher with no sidecar reach inside
    OwnerWindow, ConditionSidecarDown fires once under the empty-Site suppression row.

Repeated silence across a whole outage is therefore not the designed behaviour, which
makes this a real bug rather than the twelve-hour window doing its job.

Candidate causes, ordered.

  1. DISCORD_WEBHOOK_URL is unset on the deployment. #171's own rollout note says the
    feature ships dark — unset is silent by design, and the prod step (create webhook,
    set the variable, redeploy) is easy to have missed.
  2. It is set, and the POST fails — revoked webhook, wrong URL, blocked egress. A
    failed send is logged and the suppression row deliberately left unwritten, so the
    only trace is a owner notice ...: send failed / webhook status NNN log line.
  3. No fault was ever judged. The browser Lanes wake only when 5+ Series are due or one
    has waited 15m; a Lane under both thresholds records SkipAsleep, which
    sidecarDownSince excludes on purpose (an asleep Lane is not evidence about the
    sidecar). With a small kagane/comix queue the Lanes may simply never have probed,
    in which case nothing judged the browser down at all.

Two facts separate them (both from the deployment, not the repo):

  • The startup log line — owner notices: enabled or disabled (DISCORD_WEBHOOK_URL unset).
  • The Lanes landing verdict right now: it shares latest.FaultsFrom with the notifier
    (#173), so if the page names a fault the judgement fired and cause 2 is the live one;
    if it reads healthy, cause 3 is.

Also worth grepping the API logs for owner notice and for
browser unreachable, browser lanes skipping passes.

Two gaps this exposed regardless of which cause wins, for the brief once the cause
is known:

  • A browser loss after at least one successful check in the same pass leaves
    checked > 0, so that pass is neither stall-shaped nor sidecar-down; the first
    evidence waits for the next probe.
  • The notice path cannot report on its own failure — that is #179's territory, and it
    is why this outage was indistinguishable from a quiet server.
> *This was generated by AI during triage.* ## Triage notes Split: the "report any error event" half is now #179 (enhancement, needs-triage). This ticket is the bug — the browser was unreachable and no message arrived. **What the code says should have happened.** The path is wired end to end: `main` builds a `notify.Client` whenever `DISCORD_WEBHOOK_URL` is non-empty and logs `owner notices: enabled` / `owner notices: disabled (DISCORD_WEBHOOK_URL unset)` at startup, and `Poller.ownerNotices` judges `latest.FaultsFrom` beside every pass-log insert. For an unreachable browser, two of the four conditions apply: - The first wake after the browser goes away fails on the first Series, so the pass row is stall-shaped (due > 0, checked 0, empty skip) — see the `outcomeUnreachable` return in `runLanePass`. `ConditionStall` carries **no** twelve-hour window, so it should fire on that pass, not half a day later. - Sibling browser Lanes then skip for `RefuseBackoff`, and once every browser-backed Site's latest pass is `SkipSidecarDown`/`SkipNoFetcher` with no sidecar reach inside `OwnerWindow`, `ConditionSidecarDown` fires once under the empty-Site suppression row. Repeated silence across a whole outage is therefore not the designed behaviour, which makes this a real bug rather than the twelve-hour window doing its job. **Candidate causes, ordered.** 1. `DISCORD_WEBHOOK_URL` is unset on the deployment. #171's own rollout note says the feature ships dark — unset is silent by design, and the prod step (create webhook, set the variable, redeploy) is easy to have missed. 2. It is set, and the POST fails — revoked webhook, wrong URL, blocked egress. A failed send is logged and the suppression row deliberately left unwritten, so the only trace is a `owner notice ...: send failed` / `webhook status NNN` log line. 3. No fault was ever judged. The browser Lanes wake only when 5+ Series are due or one has waited 15m; a Lane under both thresholds records `SkipAsleep`, which `sidecarDownSince` excludes on purpose (an asleep Lane is not evidence about the sidecar). With a small kagane/comix queue the Lanes may simply never have probed, in which case nothing judged the browser down at all. **Two facts separate them** (both from the deployment, not the repo): - The startup log line — `owner notices: enabled` or `disabled (DISCORD_WEBHOOK_URL unset)`. - The Lanes landing verdict right now: it shares `latest.FaultsFrom` with the notifier (#173), so if the page names a fault the judgement fired and cause 2 is the live one; if it reads healthy, cause 3 is. Also worth grepping the API logs for `owner notice` and for `browser unreachable, browser lanes skipping passes`. **Two gaps this exposed regardless of which cause wins**, for the brief once the cause is known: - A browser loss *after* at least one successful check in the same pass leaves checked > 0, so that pass is neither stall-shaped nor sidecar-down; the first evidence waits for the next probe. - The notice path cannot report on its own failure — that is #179's territory, and it is why this outage was indistinguishable from a quiet server.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: sulthan/mangaBookmark#178