An Investigation Order for Stale Data

An operations sequence for tracing stale data through the request, API, worker, and source

Hangangjari’s backend runs on a K3s cluster on a small server at home. When a user sees stale parking information, the investigation checks request arrival, API health, and the worker’s last collection in order.

The user-facing API reaches K3s ingress through Cloudflare edge and Tunnel. The tunnel hides the origin server and keeps the user API path separate from the operator access path.

Regardless of server location, users see API response time and data freshness. The operating procedure checks request paths, collection state, and recoverability on the small server.

Whether the Request Reaches the API

The first check is the route a user request takes before reaching the API. A blocked route, a single failing feature inside the API, and stale data are different problems.

flowchart LR
  Client["iOS app / widget"] --> Edge["Cloudflare edge"]
  Edge --> Tunnel["Cloudflare Tunnel"]
  Tunnel --> Ingress["K3s ingress"]
  Ingress --> API["FastAPI service"]
  API --> Redis["Redis"]
  API --> PG["Postgres"]

The brand site, support page, and privacy page are separated from the API. Static web pages and app APIs have different failure modes, so their status is checked by runtime.

flowchart TB
  Site["hangangjari.app<br/>brand / support / privacy"] --> Pages["Static site"]
  APIName["API domain"] --> Edge["Cloudflare edge"]
  Edge --> Tunnel["Cloudflare Tunnel"]
  Tunnel --> Runtime["K3s cluster"]

The tunnel routes traffic from the edge network to the service while keeping the origin hidden. User requests and operator access use separate paths, allowing incident response to classify and recover each path independently.

Investigation Order for a Stale Value

The K3s inventory follows the order used to investigate stale values: API, workers, stateful stores, ingress, metrics, and logs.

ComponentRole
FastAPI APIApp API read by the iOS app and widgets
Parking workerCollects parking status
Outing workerCollects events, notices, facilities, and realtime context
Forecast workerBuilds forecasts and warms home summary
Push workerBuilds candidates, checks send rules, and dispatches the outbox
Postgres/PostGISReference data, history, and audit records
RedisCache, latest status, and supporting indexes
K3s ingressUser API routing
Prometheus/Grafana/logsMetrics, logs, and dashboards
flowchart TB
  Desired["Server state store<br/>desired state"] --> GitOps["GitOps sync"]
  GitOps --> API["API deployment"]
  GitOps --> Workers["Worker deployment"]
  GitOps --> State["Postgres / Redis"]
  GitOps --> Ingress["Ingress"]
  API --> Signals["Metrics / logs"]
  Workers --> Signals
  State --> Signals

The operating inventory includes deployment paths, backups, metrics and logs, and recovery procedures.

Dashboards That Narrow the Cause

The dashboard gathers signals that narrow the causes to inspect after a user report.

flowchart TD
  Checks["Things to check"] --> APIQ["API<br/>status / latency / error rate"]
  Checks --> WorkerQ["Workers<br/>last success / rows / freshness"]
  Checks --> DataQ["Data<br/>cache hit/miss / forecast runs / outbox backlog"]
  Checks --> InfraQ["K3s<br/>ingress / GitOps state / backup state"]
  APIQ --> Triage["See which part is shaking"]
  WorkerQ --> Triage
  DataQ --> Triage
  InfraQ --> Triage

Initial triage uses these signals:

  • Is the API alive?
  • Are protected routes being accepted and rejected correctly?
  • Is latency spiking on a specific endpoint?
  • Are workers continuously collecting fresh data?
  • Is the last success time by source too old?
  • Is Redis cache abnormally empty?
  • Is the forecast worker creating recent runs?
  • Is the push outbox building up?
  • Are DB backup and restore drills healthy?

The API and workers expose these internal metrics for a collector to pull periodically.

Component State After a User Report

The first step after a report is to narrow the failure location through request-path and component state.

flowchart TD
  Alert["User report or smoke failure"] --> Public{"User API healthy?"}
  Public -->|No| Runtime{"Ingress and K3s healthy?"}
  Runtime -->|Yes| Edge["Check edge and tunnel"]
  Runtime -->|No| Pods["Check pods and GitOps state"]
  Public -->|Yes| Feature{"Only one API feature failing?"}
  Feature -->|Yes| Fresh{"Freshness degraded?"}
  Fresh -->|Yes| Worker["Check workers and ingestion runs"]
  Fresh -->|No| Cache["Check cache/display response and DB query"]
  Feature -->|No| Client["Check client cache or widget snapshot"]

This order separates an API outage, source failure, stale cache, and a client reading an old snapshot.

Data Classified by Recovery Priority

Backup status includes both file creation and restore-drill results.

DB backup state and restore viability are part of the operating checks. Much public data can be collected again, while user settings, push subscriptions, app events, and forecast/backtest history require preserved backups.

Recovery procedures classify data as rebuildable or preserved. An empty Redis cache follows a rebuild path; damaged Postgres reference data follows a backup-restore path verified through restore drills.

From a Stale Value to Its Cause

When stale data appears, the investigation starts with the request path and API health, then moves through the feature worker and source freshness to cache and client snapshots. Postgres data leads to backup and restore-drill checks; Redis query values lead to the rebuild path.

Comments

Comments

    Image preview