Incident Triage and Recovery Criteria for a Mini PC Server

Operations notes on tracing user requests and verifying mini PC backups through restoration

Hangangjari’s backend runs on a mini PC at home. The App Store app and widgets depend on its API, ingestion workers, and notification path. The incident check order runs from user requests to source data so stale parking information and delayed notifications can be traced to a segment.

Operations Starting from User Requests

Early checks focused on API and worker processes, DB and Redis connections, and app responses. After public release, incidents are narrowed through the user request path, specific APIs, worker ingestion, source freshness, and cache in that order.

flowchart LR
  Client["iOS app / widget"] --> Public["User request path<br/>edge / ingress"]
  Public --> API["API service"]
  API --> Data["Postgres / Redis"]
  Workers["Workers<br/>ingestion / forecast / push"] --> Data
  Data --> API
  Runtime["GitOps state"] --> API
  Runtime --> Workers
  Signals["Metrics / logs / dashboards"] --> API
  Signals --> Workers

If a segment stops, the app fails requests or displays stale information. Each segment needs an observable state and last-success time regardless of host size.

Separate User and Management Paths

User APIs and operator access use different authentication and exposure policies. The app API is published on the user request path, while operator access remains behind a separate protection policy. Incident triage also checks the user request path, running services, and operator control path independently.

Server State Kept in Git

Deployable images and Kubernetes resource state are stored in Git, and the GitOps controller reconciles the runtime toward them.

flowchart LR
  AppRepo["App repository"] --> CI["CI<br/>tests / build"]
  CI --> Image["Versioned image"]
  CI --> Desired["Desired deploy state"]
  Desired --> Sync["GitOps sync"]
  Sync --> Runtime["API / workers / data services"]
  Runtime --> Smoke["Post-deploy smoke check"]

Incident review uses these records:

  • Deployment images remain traceable to commits.
  • Runtime drift from Git state is visible.
  • Deployment state and post-deploy results are recorded together.
  • Direct server changes receive a separate operating record.

Git state cannot verify DB migrations, worker rollout order, operating configuration, or external-source failures. Those are checked in post-deploy verification and operational dashboards.

Dashboards Built from the Investigation Order

The dashboard places these signals in incident-triage order:

  • Is the user API alive?
  • Is the API alive but only one screen response failing?
  • Are workers still collecting source data?
  • Has the last success time exceeded the stale threshold?
  • Did the cache hit/failure ratio change abnormally?
  • Is the forecast run recent?
  • Is the push outbox building up?
  • Were backups created, and was restore verified?

Overall API state, screen-specific responses, workers, source freshness, and cache fallback appear in that order. Each graph narrows the next component or log to inspect.

The Path to Follow During an Incident

Incident triage follows the request path to isolate the failing segment.

flowchart TD
  Symptom["User report or smoke failure"] --> Public{"User API healthy?"}
  Public -->|No| Path{"Edge/ingress path issue?"}
  Path -->|Yes| Network["Check user request path"]
  Path -->|No| Runtime["Check runtime"]
  Public -->|Yes| Feature{"Only one screen failing?"}
  Feature -->|Yes| Data{"Source freshness or cache issue?"}
  Data -->|Yes| Worker["Check workers / ingestion / cache"]
  Data -->|No| API["Check API response / DB query"]
  Feature -->|No| Client["Check app cache / widget snapshot"]

The dashboard checks edge, API, workers, and source data in order. This separates a blocked request path from stale source data and identifies the component to recover before a code change.

Backups Verified Through Restore

Backups are verified by restoring and reading data in a separate environment. Data is classified to set recovery priority.

DataCharacter
Values that can be collected again from public sourcesRebuildable
User notification settings and push subscriptionsUser data that must be recovered
App events and audit recordsEvidence for operations review
Forecast/backtest historyEvidence for forecast review
CacheRebuildable

User notification settings and push subscriptions are restored first. Cache and public data that can be collected again are rebuilt, while audit and forecast history remain preserved for incident and forecast analysis.

Recovery Completion Criteria

Recovery checks include source freshness, cache fallback, push backlog, and restored data as well as API and worker processes. Because this path runs on one mini PC, host resources and application state appear in the same check order.

Comments

Comments

    Image preview