An Investigation Order for Stale Data
An operations sequence for tracing stale data through the request, API, worker, and source
Hangangjari’s backend runs on a K3s cluster on a small server at home. When a user sees stale parking information, the investigation checks request arrival, API health, and the worker’s last collection in order.
The user-facing API reaches K3s ingress through Cloudflare edge and Tunnel. The tunnel hides the origin server and keeps the user API path separate from the operator access path.
Regardless of server location, users see API response time and data freshness. The operating procedure checks request paths, collection state, and recoverability on the small server.
Whether the Request Reaches the API
The first check is the route a user request takes before reaching the API. A blocked route, a single failing feature inside the API, and stale data are different problems.
flowchart LR Client["iOS app / widget"] --> Edge["Cloudflare edge"] Edge --> Tunnel["Cloudflare Tunnel"] Tunnel --> Ingress["K3s ingress"] Ingress --> API["FastAPI service"] API --> Redis["Redis"] API --> PG["Postgres"]
The brand site, support page, and privacy page are separated from the API. Static web pages and app APIs have different failure modes, so their status is checked by runtime.
flowchart TB Site["hangangjari.app<br/>brand / support / privacy"] --> Pages["Static site"] APIName["API domain"] --> Edge["Cloudflare edge"] Edge --> Tunnel["Cloudflare Tunnel"] Tunnel --> Runtime["K3s cluster"]
The tunnel routes traffic from the edge network to the service while keeping the origin hidden. User requests and operator access use separate paths, allowing incident response to classify and recover each path independently.
Investigation Order for a Stale Value
The K3s inventory follows the order used to investigate stale values: API, workers, stateful stores, ingress, metrics, and logs.
| Component | Role |
|---|---|
| FastAPI API | App API read by the iOS app and widgets |
| Parking worker | Collects parking status |
| Outing worker | Collects events, notices, facilities, and realtime context |
| Forecast worker | Builds forecasts and warms home summary |
| Push worker | Builds candidates, checks send rules, and dispatches the outbox |
| Postgres/PostGIS | Reference data, history, and audit records |
| Redis | Cache, latest status, and supporting indexes |
| K3s ingress | User API routing |
| Prometheus/Grafana/logs | Metrics, logs, and dashboards |
flowchart TB Desired["Server state store<br/>desired state"] --> GitOps["GitOps sync"] GitOps --> API["API deployment"] GitOps --> Workers["Worker deployment"] GitOps --> State["Postgres / Redis"] GitOps --> Ingress["Ingress"] API --> Signals["Metrics / logs"] Workers --> Signals State --> Signals
The operating inventory includes deployment paths, backups, metrics and logs, and recovery procedures.
Dashboards That Narrow the Cause
The dashboard gathers signals that narrow the causes to inspect after a user report.
flowchart TD Checks["Things to check"] --> APIQ["API<br/>status / latency / error rate"] Checks --> WorkerQ["Workers<br/>last success / rows / freshness"] Checks --> DataQ["Data<br/>cache hit/miss / forecast runs / outbox backlog"] Checks --> InfraQ["K3s<br/>ingress / GitOps state / backup state"] APIQ --> Triage["See which part is shaking"] WorkerQ --> Triage DataQ --> Triage InfraQ --> Triage
Initial triage uses these signals:
- Is the API alive?
- Are protected routes being accepted and rejected correctly?
- Is latency spiking on a specific endpoint?
- Are workers continuously collecting fresh data?
- Is the last success time by source too old?
- Is Redis cache abnormally empty?
- Is the forecast worker creating recent runs?
- Is the push outbox building up?
- Are DB backup and restore drills healthy?
The API and workers expose these internal metrics for a collector to pull periodically.
Component State After a User Report
The first step after a report is to narrow the failure location through request-path and component state.
flowchart TD
Alert["User report or smoke failure"] --> Public{"User API healthy?"}
Public -->|No| Runtime{"Ingress and K3s healthy?"}
Runtime -->|Yes| Edge["Check edge and tunnel"]
Runtime -->|No| Pods["Check pods and GitOps state"]
Public -->|Yes| Feature{"Only one API feature failing?"}
Feature -->|Yes| Fresh{"Freshness degraded?"}
Fresh -->|Yes| Worker["Check workers and ingestion runs"]
Fresh -->|No| Cache["Check cache/display response and DB query"]
Feature -->|No| Client["Check client cache or widget snapshot"]
This order separates an API outage, source failure, stale cache, and a client reading an old snapshot.
Data Classified by Recovery Priority
Backup status includes both file creation and restore-drill results.
DB backup state and restore viability are part of the operating checks. Much public data can be collected again, while user settings, push subscriptions, app events, and forecast/backtest history require preserved backups.
Recovery procedures classify data as rebuildable or preserved. An empty Redis cache follows a rebuild path; damaged Postgres reference data follows a backup-restore path verified through restore drills.
From a Stale Value to Its Cause
When stale data appears, the investigation starts with the request path and API health, then moves through the feature worker and source freshness to cache and client snapshots. Postgres data leads to backup and restore-drill checks; Redis query values lead to the rebuild path.
Comments
No comments yet. Be the first to leave one.
Pending review