Alertmanager Mirror Drift Correction + Sprint 231 Documentation ERRATA (infra/k3s ≠ deployment SSOT)
SummarySprint 231's central conclusion ('alertmanager receiver:null → all 17 rules dropped / 6 days no alerts') was proven wrong by server-side live diagnosis. Root cause: AlgoSu infra/k3s/ is a non-deployed reference mirror; the actual deployment SSOT is aether-gitops (algosu/base/monitoring + overlays/prod). AlgoSu CI only bumps image tags in aether-gitops and does not propagate manifests, and ArgoCD (automated, selfHeal, Synced) only watches aether-gitops. Sprint 231 misread this non-deployed mirror's stale snapshot (receiver:null) as the deployed config. Live alertmanager was already healthy via Sprint 130 B-1 (alertmanager-native discord-default + discord-critical, identity-discord-secret/webhook-url, v0.28.1). Worse, the 'fix' Sprint 231 added to the mirror (monitoring-secrets/discord_webhook + v0.27.0 + webhook_url_file) was non-functional since v0.27.0 does not support file-based discord webhooks. The actual 'log non-collection' cause was Loki OOM (512Mi limit, ~5-day spike), fixed and verified via aether-gitops PR #7 (512Mi→1Gi) (C). This sprint: A-lite aligns the AlgoSu alertmanager mirror to live (v0.28.1 + discord-default/critical + identity-discord-secret/webhook-url + non-deployed banner) and removes the sealed-secrets placeholder; B adds ERRATA to the Sprint 231 ADR/runbook; B+ seeds the loki probe/securityContext live gap. Deployment-neutral (mirror/doc integrity restoration). Critic R1 CLEAN. Lesson: do not assert runtime defects from static manifests (especially a non-deployed mirror) (feedback-source-vs-live-drift).
Goal
- Correct Sprint 231's (full monitoring audit) misjudgment based on server-side live diagnosis.
- Align the AlgoSu
infra/k3s/mirror to the live (aether-gitops) configuration and document the structure that caused the misjudgment (infra/k3s≠ deployment SSOT).
Background
- Sprint 231 saw
receiver: 'null'ininfra/k3s/monitoring/alertmanager.yamland asserted "all 17 alert rules dropped, 6 days of no alerts" as a confirmed defect, then "wired" a Discord receiver. - The user diagnosed live on the server and the conclusion turned out to be wrong.
Root Cause (confirmed via server diagnosis)
- AlgoSu
infra/k3s/is a non-deployed reference mirror. The actual deployment SSOT is aether-gitops (algosu/base/monitoring/+overlays/prod). AlgoSu CI (ci.yml:1073+) only bumps image tags in aether-gitops and does not propagate manifests. ArgoCD (automated, selfHeal, Synced) only watches aether-gitops. - The
receiver: 'null'Sprint 231 read was a stale snapshot of the non-deployed mirror. Live alertmanager was already healthy via Sprint 130 B-1 — alertmanager-nativediscord-default+discord-critical(severity routing, repeat 30m),webhook_url_filereferencingidentity-discord-secret/webhook-url(optional: false), image v0.28.1. - The "fix" Sprint 231 added (
monitoring-secrets/discord_webhook+ keeping v0.27.0 +webhook_url_file) was non-functional — v0.27.0 does not support file-based discord webhooks — pushing the mirror further from live. - The actual "log non-collection" cause was Loki OOM, not the pipeline (limit 512Mi, restartCount 17, last OOMKilled 2026-05-18, ~5-day spike exceeding the limit).
Key Decisions
- Mirror role = demote to reference + fix only what's broken (full mirror sync isn't worth the burden for a solo/Free Tier project). A non-deployed banner blocks future misreads.
- A-lite: replace AlgoSu
alertmanager.yamlwith the live manifest (v0.28.1 + discord-default/critical + identity-discord-secret/webhook-url + optional:false), remove the sealed-secretsALERTMANAGER_DISCORD_WEBHOOKplaceholder (reuse the existing secret). Deployment-neutral (ArgoCD does not reference infra/k3s). - B: the Sprint 231 ADR/runbook are historical records, so keep the body and add an ERRATA block at the top to correct the facts.
- C (done on server): Loki OOM hardening via aether-gitops PR #7 (
5b07bf6), 512Mi→1Gi / 128Mi→256Mi req, verified (loki Running, no OOM recurrence, ample node headroom).
Work Summary (start e5f1046 → squash cc92924, PR #401)
e080ff1fix(infra): align alertmanager.yaml to live (v0.28.1, discord-default/critical, identity-discord-secret/webhook-url optional:false, title/message templates) + non-deployed banner header + remove sealed-secrets placeholder.6c4a55fdocs: runbook §0 ERRATA + §1.2/§4-A/B/C factual corrections (mirror ≠ deployment source, alerting healthy in live, Loki OOM as the real cause, loki hardening gap D1 seed) + ERRATA block atop ADR sprint-231 KR+EN.
Verification
- alertmanager.yaml valid YAML (3 docs, route discord-default + critical route, secret identity-discord-secret/webhook-url optional:false, image v0.28.1) ·
check-grafana-metrics.mjsexit 0. - ADR gates (index 170, adr-en, links 0, doc-refs 0) + adr-conversion OK.
- Critic (Codex gpt-5.5,
codex review --base e5f1046): R1 CLEAN ("changes mainly align the reference Alertmanager manifest and documentation with the stated live configuration. I did not find any introduced functional issue"). - CI #401:
Secret & Env Scanfailed once (gitleaks download 504 flake) → re-run passed → autoMerge SQUASH. All gates (esp.Quality — monitoringBLOCKING) green.