The Notification Risk a Health Check Missed

The false notifications caused by stale snapshot comparisons after recovery and the freshness guard that stopped them

After Hangangjari launched, two push notifications went out just after runtime checks indicated recovery from an operational incident.

The Hangangjari API was responding again. The healthcheck had returned to normal. K3s pods were up, and workers were running. Those checks covered runtime recovery, but not the quality of the first events produced afterward.

Right after recovery, the parking worker directly compared the last snapshot before the incident with the first snapshot after recovery and interpreted the difference as a new change. As a result, two parking-threshold notifications were sent to real users.

APNs delivery, user-setting matching, and suppression rules all followed their configured policies. The false sends originated in change detection, which treated a stale snapshot as the immediately previous state.

Events Left Behind by the Incident

The incident began when the operating node ran low on disk space.

Hangangjari’s backend runs on K3s on a small server at home. The API, workers, Postgres, and Redis all run on the same node. As its disk nearly filled, K3s reported DiskPressure.

Backup data filled the same filesystem used by the environment running the pods. Kubernetes reports the DiskPressure node condition when disk space or inodes reach an eviction threshold. Pod eviction and restrictions on scheduling new pods can then affect the API, DB, Redis, and workers together.

The first recovery steps were to free disk space, restore K3s state, and check the health check. Afterward, host-disk free space and DiskPressure were added to monitoring alerts, and the backup policy was changed to retain only recent backups on that disk.

Those runtime checks did not examine which snapshot the restarted worker would use as its comparison baseline.

While the worker was stopped, the last snapshot still remained in Postgres. When the worker ran again, a new snapshot arrived too. Each was a valid parking state. So the existing logic compared them.

The existing logic compared the two values without checking the collection gap.

sequenceDiagram
  autonumber
  participant Worker as Parking worker
  participant DB as Postgres
  participant Fact as Fact writer
  participant Push as Push pipeline
  participant User as User

  Worker->>DB: Store pre-incident snapshot (AVAILABLE 169)
  Note over Worker,DB: Parking worker stopped during K3s DiskPressure
  Note over Worker,DB: Parking changes in this gap were not collected
  Worker->>DB: Store post-recovery snapshot (FULL 0)
  Worker->>DB: Fetch previous snapshot for comparison
  DB-->>Worker: Return pre-incident snapshot
  Worker->>Fact: Compare previous and current values
  Note right of Fact: Age of previous value was not checked
  Fact->>Push: Create parking_space_threshold fact
  Push->>User: Send 2 push notifications

This comparison is used in a continuous collection flow. If remaining spaces were 40 a moment ago and 20 now, the user may want to know about that change.

In this case, the comparison spanned a long incident gap. The code treated an old value as the immediate predecessor and processed the first post-recovery state as a new change.

Event Quality Beyond the Healthcheck

For this service, a normal health check means a request reached the API process and received a response. DB and Redis readiness is checked with readiness probes. Neither check validates the time gap between worker snapshots or the freshness of a new push fact.

The read path still needs the latest parking state after recovery. The cache should also be updated, because the post-recovery value is newer than the pre-incident value.

Push candidates come from differences between two states. If the two times are too far apart, even a fresh current value cannot be described as “just changed.”

Read responses and change detection therefore use different post-recovery conditions.

TreatmentPost-recovery choice
Store parking stateContinue
Update Redis cacheContinue
App and widget responseUse latest state
Create push factStop if previous snapshot is old
Enqueue outbox itemStop if the fact is already expired

The post-recovery snapshot is used for read responses and cache, but it does not create a push fact when compared with a stale previous value.

Stale Previous Values Excluded from Comparison

A freshness guard now runs before parking snapshots are compared.

Hangangjari parking status is collected at a short interval. The existing stale window is sufficiently longer than that interval, so the same threshold now separates ordinary collection delays from recovery gaps.

If the difference between observed_at on the previous snapshot and current snapshot exceeds the stale window, the current snapshot is stored but no push fact is created.

current.observed_at - previous.observed_at > stale window

When this condition triggers, the system records previous_snapshot_stale and stops.

Tests define the boundary behavior. Differences inside the stale window are treated as normal collection delay and can create facts. Differences beyond the window are treated as recovery gaps and do not create facts. The incident pattern of directly comparing an old snapshot with the post-recovery snapshot now stops on this path.

The current post-recovery parking state remains in DB. Cache is updated. The app and widgets read the latest value. The freshness guard affects only fact creation when the previous value is stale.

The current state remains available on the read path, while changes whose occurrence time is unknown are excluded from push.

One More Filter Before the Outbox

Push is not sent immediately. The system creates a fact, checks user settings, puts it in the outbox, and a dispatcher sends it to APNs. Time passes between those stages. If workers are backed up or processing a backlog after restart, an event that was valid when the fact was created may already be old before entering the outbox.

So parking facts now use the same stale window for expires_at. The candidate builder checks expiration before creating an outbox entry.

flowchart TB
  Current["Collect new snapshot"] --> Store["Store in DB / update Redis cache"]
  Current --> Age{"Does previous snapshot exceed<br/>the stale window?"}
  Age -->|Yes| StopA["Stop fact creation"]
  Age -->|No| Fact["Create push fact"]
  Fact --> Expired{"Has fact expired?"}
  Expired -->|Yes| StopB["Stop before enqueueing outbox"]
  Expired -->|No| Outbox["Enqueue outbox item"]

The first guard prevents comparing current values with old previous values. The second guard prevents already-late facts from entering the send queue. They look similar, but they block different failure points.

Reason codes distinguish a stale previous snapshot, a fact that expired before the outbox, and suppression caused by user settings.

Notifications Intentionally Left Unsent

During the incident, a parking lot may have changed from available to full and remained full after recovery. The two snapshots do not reveal when that change happened, so they cannot support the claim that it “just became full.”

No notification is created from that stale pair of snapshots.

The app screen still shows the latest state and its refresh time. A push notification, however, pulls attention outside the app and can affect movement decisions. The system withholds a notification when the change time is unknown.

Only changes confirmed by consecutive post-recovery snapshots become new notification candidates.

Recovery Verification Through Push Delivery

Before this, recovery checks were mostly runtime-state checks.

  • Does the API respond?
  • Are pods up again?
  • Is the worker running again?
  • Are DB and Redis connected?
  • Does deployed state match the desired revision?

Recovery checks now add:

  • Are newly created post-recovery events safe to send to users?

After DiskPressure, recovery verification now covers pods and health checks, the worker’s first collection time, source freshness, suppression reasons, and outbox growth. Recovery was complete only after confirming that no more parking-threshold outbox entries of the same type appeared.

The DiskPressure and eviction conditions were checked against Kubernetes’ Node-pressure Eviction documentation. The scope of health checks and readiness was checked against the Kubernetes probe documentation and the service configuration.

Comments

Comments

    Image preview