The Notification Risk a Health Check Missed
The false notifications caused by stale snapshot comparisons after recovery and the freshness guard that stopped them
After Hangangjari launched, two push notifications went out just after runtime checks indicated recovery from an operational incident.
The Hangangjari API was responding again. The healthcheck had returned to normal. K3s pods were up, and workers were running. Those checks covered runtime recovery, but not the quality of the first events produced afterward.
Right after recovery, the parking worker directly compared the last snapshot before the incident with the first snapshot after recovery and interpreted the difference as a new change. As a result, two parking-threshold notifications were sent to real users.
APNs delivery, user-setting matching, and suppression rules all followed their configured policies. The false sends originated in change detection, which treated a stale snapshot as the immediately previous state.
Events Left Behind by the Incident
The incident began when the operating node ran low on disk space.
Hangangjari’s backend runs on K3s on a small server at home. The API, workers, Postgres, and Redis all run on the same node. As its disk nearly filled, K3s reported DiskPressure.
Backup data filled the same filesystem used by the environment running the pods. Kubernetes reports the DiskPressure node condition when disk space or inodes reach an eviction threshold. Pod eviction and restrictions on scheduling new pods can then affect the API, DB, Redis, and workers together.
The first recovery steps were to free disk space, restore K3s state, and check the health check. Afterward, host-disk free space and DiskPressure were added to monitoring alerts, and the backup policy was changed to retain only recent backups on that disk.
Those runtime checks did not examine which snapshot the restarted worker would use as its comparison baseline.
While the worker was stopped, the last snapshot still remained in Postgres. When the worker ran again, a new snapshot arrived too. Each was a valid parking state. So the existing logic compared them.
The existing logic compared the two values without checking the collection gap.
sequenceDiagram autonumber participant Worker as Parking worker participant DB as Postgres participant Fact as Fact writer participant Push as Push pipeline participant User as User Worker->>DB: Store pre-incident snapshot (AVAILABLE 169) Note over Worker,DB: Parking worker stopped during K3s DiskPressure Note over Worker,DB: Parking changes in this gap were not collected Worker->>DB: Store post-recovery snapshot (FULL 0) Worker->>DB: Fetch previous snapshot for comparison DB-->>Worker: Return pre-incident snapshot Worker->>Fact: Compare previous and current values Note right of Fact: Age of previous value was not checked Fact->>Push: Create parking_space_threshold fact Push->>User: Send 2 push notifications
This comparison is used in a continuous collection flow. If remaining spaces were 40 a moment ago and 20 now, the user may want to know about that change.
In this case, the comparison spanned a long incident gap. The code treated an old value as the immediate predecessor and processed the first post-recovery state as a new change.
Event Quality Beyond the Healthcheck
For this service, a normal health check means a request reached the API process and received a response. DB and Redis readiness is checked with readiness probes. Neither check validates the time gap between worker snapshots or the freshness of a new push fact.
The read path still needs the latest parking state after recovery. The cache should also be updated, because the post-recovery value is newer than the pre-incident value.
Push candidates come from differences between two states. If the two times are too far apart, even a fresh current value cannot be described as “just changed.”
Read responses and change detection therefore use different post-recovery conditions.
| Treatment | Post-recovery choice |
|---|---|
| Store parking state | Continue |
| Update Redis cache | Continue |
| App and widget response | Use latest state |
| Create push fact | Stop if previous snapshot is old |
| Enqueue outbox item | Stop if the fact is already expired |
The post-recovery snapshot is used for read responses and cache, but it does not create a push fact when compared with a stale previous value.
Stale Previous Values Excluded from Comparison
A freshness guard now runs before parking snapshots are compared.
Hangangjari parking status is collected at a short interval. The existing stale window is sufficiently longer than that interval, so the same threshold now separates ordinary collection delays from recovery gaps.
If the difference between observed_at on the previous snapshot and current snapshot exceeds the stale window, the current snapshot is stored but no push fact is created.
current.observed_at - previous.observed_at > stale window
When this condition triggers, the system records previous_snapshot_stale and stops.
Tests define the boundary behavior. Differences inside the stale window are treated as normal collection delay and can create facts. Differences beyond the window are treated as recovery gaps and do not create facts. The incident pattern of directly comparing an old snapshot with the post-recovery snapshot now stops on this path.
The current post-recovery parking state remains in DB. Cache is updated. The app and widgets read the latest value. The freshness guard affects only fact creation when the previous value is stale.
The current state remains available on the read path, while changes whose occurrence time is unknown are excluded from push.
One More Filter Before the Outbox
Push is not sent immediately. The system creates a fact, checks user settings, puts it in the outbox, and a dispatcher sends it to APNs. Time passes between those stages. If workers are backed up or processing a backlog after restart, an event that was valid when the fact was created may already be old before entering the outbox.
So parking facts now use the same stale window for expires_at. The candidate builder checks expiration before creating an outbox entry.
flowchart TB
Current["Collect new snapshot"] --> Store["Store in DB / update Redis cache"]
Current --> Age{"Does previous snapshot exceed<br/>the stale window?"}
Age -->|Yes| StopA["Stop fact creation"]
Age -->|No| Fact["Create push fact"]
Fact --> Expired{"Has fact expired?"}
Expired -->|Yes| StopB["Stop before enqueueing outbox"]
Expired -->|No| Outbox["Enqueue outbox item"]
The first guard prevents comparing current values with old previous values. The second guard prevents already-late facts from entering the send queue. They look similar, but they block different failure points.
Reason codes distinguish a stale previous snapshot, a fact that expired before the outbox, and suppression caused by user settings.
Notifications Intentionally Left Unsent
During the incident, a parking lot may have changed from available to full and remained full after recovery. The two snapshots do not reveal when that change happened, so they cannot support the claim that it “just became full.”
No notification is created from that stale pair of snapshots.
The app screen still shows the latest state and its refresh time. A push notification, however, pulls attention outside the app and can affect movement decisions. The system withholds a notification when the change time is unknown.
Only changes confirmed by consecutive post-recovery snapshots become new notification candidates.
Recovery Verification Through Push Delivery
Before this, recovery checks were mostly runtime-state checks.
- Does the API respond?
- Are pods up again?
- Is the worker running again?
- Are DB and Redis connected?
- Does deployed state match the desired revision?
Recovery checks now add:
- Are newly created post-recovery events safe to send to users?
After DiskPressure, recovery verification now covers pods and health checks, the worker’s first collection time, source freshness, suppression reasons, and outbox growth. Recovery was complete only after confirming that no more parking-threshold outbox entries of the same type appeared.
The DiskPressure and eviction conditions were checked against Kubernetes’ Node-pressure Eviction documentation. The scope of health checks and readiness was checked against the Kubernetes probe documentation and the service configuration.
Comments
No comments yet. Be the first to leave one.
Pending review