Incident Triage and Recovery Criteria for a Mini PC Server
Operations notes on tracing user requests and verifying mini PC backups through restoration
Hangangjari’s backend runs on a mini PC at home. The App Store app and widgets depend on its API, ingestion workers, and notification path. The incident check order runs from user requests to source data so stale parking information and delayed notifications can be traced to a segment.
Operations Starting from User Requests
Early checks focused on API and worker processes, DB and Redis connections, and app responses. After public release, incidents are narrowed through the user request path, specific APIs, worker ingestion, source freshness, and cache in that order.
flowchart LR Client["iOS app / widget"] --> Public["User request path<br/>edge / ingress"] Public --> API["API service"] API --> Data["Postgres / Redis"] Workers["Workers<br/>ingestion / forecast / push"] --> Data Data --> API Runtime["GitOps state"] --> API Runtime --> Workers Signals["Metrics / logs / dashboards"] --> API Signals --> Workers
If a segment stops, the app fails requests or displays stale information. Each segment needs an observable state and last-success time regardless of host size.
Separate User and Management Paths
User APIs and operator access use different authentication and exposure policies. The app API is published on the user request path, while operator access remains behind a separate protection policy. Incident triage also checks the user request path, running services, and operator control path independently.
Server State Kept in Git
Deployable images and Kubernetes resource state are stored in Git, and the GitOps controller reconciles the runtime toward them.
flowchart LR AppRepo["App repository"] --> CI["CI<br/>tests / build"] CI --> Image["Versioned image"] CI --> Desired["Desired deploy state"] Desired --> Sync["GitOps sync"] Sync --> Runtime["API / workers / data services"] Runtime --> Smoke["Post-deploy smoke check"]
Incident review uses these records:
- Deployment images remain traceable to commits.
- Runtime drift from Git state is visible.
- Deployment state and post-deploy results are recorded together.
- Direct server changes receive a separate operating record.
Git state cannot verify DB migrations, worker rollout order, operating configuration, or external-source failures. Those are checked in post-deploy verification and operational dashboards.
Dashboards Built from the Investigation Order
The dashboard places these signals in incident-triage order:
- Is the user API alive?
- Is the API alive but only one screen response failing?
- Are workers still collecting source data?
- Has the last success time exceeded the stale threshold?
- Did the cache hit/failure ratio change abnormally?
- Is the forecast run recent?
- Is the push outbox building up?
- Were backups created, and was restore verified?
Overall API state, screen-specific responses, workers, source freshness, and cache fallback appear in that order. Each graph narrows the next component or log to inspect.
The Path to Follow During an Incident
Incident triage follows the request path to isolate the failing segment.
flowchart TD
Symptom["User report or smoke failure"] --> Public{"User API healthy?"}
Public -->|No| Path{"Edge/ingress path issue?"}
Path -->|Yes| Network["Check user request path"]
Path -->|No| Runtime["Check runtime"]
Public -->|Yes| Feature{"Only one screen failing?"}
Feature -->|Yes| Data{"Source freshness or cache issue?"}
Data -->|Yes| Worker["Check workers / ingestion / cache"]
Data -->|No| API["Check API response / DB query"]
Feature -->|No| Client["Check app cache / widget snapshot"]
The dashboard checks edge, API, workers, and source data in order. This separates a blocked request path from stale source data and identifies the component to recover before a code change.
Backups Verified Through Restore
Backups are verified by restoring and reading data in a separate environment. Data is classified to set recovery priority.
| Data | Character |
|---|---|
| Values that can be collected again from public sources | Rebuildable |
| User notification settings and push subscriptions | User data that must be recovered |
| App events and audit records | Evidence for operations review |
| Forecast/backtest history | Evidence for forecast review |
| Cache | Rebuildable |
User notification settings and push subscriptions are restored first. Cache and public data that can be collected again are rebuilt, while audit and forecast history remain preserved for incident and forecast analysis.
Recovery Completion Criteria
Recovery checks include source freshness, cache fallback, push backlog, and restored data as well as API and worker processes. Because this path runs on one mini PC, host resources and application state appear in the same check order.
Comments
No comments yet. Be the first to leave one.
Pending review