Monitoring 1033 Error Resolution
Date
ImpactMedium
Decisions
D1: Zero-downtime Deployment with blog Deployment preStop Hook
- Context: Cloudflare Error 1033 recurred on blog.algo-su.com. Sprint 69 resolved the same error due to missing cloudflared
--metricsflag, but this time the cloudflared pod itself was normal (Running, 0 restarts,--metrics 0.0.0.0:2000applied). Root cause was: during ArgoCD rolling update in a replicas=1 environment, old pod readiness probe failure → Service endpoint gap → cloudflared receiving upstream 502 → Cloudflare 1033 propagation. - Choice: Add
preStop: exec: command: ["sh", "-c", "sleep 5"]lifecycle hook to blog Deployment spec. Old pod continues serving traffic for 5 seconds after receiving SIGTERM, eliminating the endpoint gap until the new pod reaches Ready state. - Alternatives: (a) Scale to replicas=2 — OCI ARM resource limits (24GB memory, 6 services + monitoring running) make maintaining 2 replicas costly, rejected. (b) Apply only
maxSurge=1, maxUnavailable=0strategy — already default but does not prevent endpoint gap on readiness failure, rejected. (c) Recreate strategy — allows intentional downtime, rejected. - Code Paths:
infra/k3s/blog.yaml(or corresponding aether-gitops manifest)
Patterns
P1: Zero-downtime Rolling Update Pattern for replicas=1 Services
- Where: blog Deployment
spec.template.spec.containers[].lifecycle.preStop - When to Reuse: When traffic interruption during rolling update needs to be prevented in a Deployment with replicas=1. Core mechanism: (1) Start new pod first with
maxSurge=1, maxUnavailable=0, (2) ApplypreStop: sleep Nto old pod (N = new pod readiness time + buffer) to maintain traffic serving for a period after SIGTERM.terminationGracePeriodSecondsmust be greater than or equal to the preStop sleep time.
Metrics
- Commits: 2 (46e4525, 6b2063e)
- Files changed: 3 (+98/-0)
- Service impact: blog.algo-su.com — Error 1033 prevented during rolling update, monitoring.algo-su.com — Grafana access restored