Public Data Prepared Outside the App
Why public-data collection moved to workers while the API serves prepared values and freshness
As the iOS app and widgets split into several screens and touchpoints, the backend took on the job of preparing external data in advance so the app could read it immediately.
Early on, I thought the app could call public-data APIs directly. That was simpler for a small first screen: call the parking API and display its response.
The server then had less to do, and failures stayed close to the screen code. Hangangjari, however, combined sources with different update rates and failure modes. Handling all of them in the app exposed slow sources and empty responses directly to users.
As features grew, the limits of direct app calls became clear.
Each external source had different field names, update methods, and status expressions. Some data changed often, while some changed only once a day. When an external source was slow or failed, the app still needed to carry the last verified value, updated time, and source status together.
Prediction was better suited to the server as well. Hourly parking changes and park congestion changes needed accumulated data and background jobs. Notifications were similar. To notify users of status changes or congestion signals when they did not open the app, the server had to handle subscriptions and APNs delivery.
For that reason, the backend used FastAPI, and I separated request-serving APIs from workers that prepare values in the background.
flowchart LR
Sources["Official public sources<br/>Parking · Events · Facilities · Real-time signals"] --> Workers["Workers<br/>Collect · Normalize · Validate"]
Workers --> PG[("PostgreSQL / PostGIS<br/>Reference data · History · Location")]
Workers --> Redis[("Redis<br/>Current state · Forecast summaries")]
PG --> Forecast["Forecasts<br/>Hourly changes"]
Forecast --> Redis
PG --> Notify["Notification candidates<br/>Condition changes"]
Redis --> Notify
Notify --> APNs["APNs"]
PG --> API["FastAPI<br/>Prepared responses"]
Redis --> API
API --> Client["iOS app · Widgets"]
APNs --> Client
In this setup, the API focuses on reading as much as possible. I did not put external scraping, new prediction generation, or notification-candidate calculation inside the user request path. If external sources are slow or fail when a user opens the app, that slowness and failure would appear directly on the screen.
The current backend is roughly split into these roles.
- API: responses the app and widgets read immediately
- Parking worker: parking-lot list and current status collection
- Outing worker: event, facility, notice, and real-time city-data collection
- Forecast worker: parking and park congestion change calculation
- Notification worker: selecting notification candidates and sending them through APNs
Data is stored in PostgreSQL/PostGIS, while frequently read state and forecast responses are cached in Redis. Postgres keeps reference data and history, and Redis keeps short-lived responses that can be read quickly.
During an incident, I checked the value read by the API, the worker’s last run, and the external source state in order.
A Read-Only API
Early on, the API and batch jobs were closer together. Their different execution schedules led me to separate collection, prediction, notification, and API response, then split metrics and health checks by component.
When a public-data source fails, the app may show an empty or stale value. The API returns the last successful value, updated time, and source state together so the screen can explain the current collection state.
Prediction followed the same logic. A forecast was a supporting signal for understanding hourly changes, with its own cache lifetime and update path. The API kept forecast values separate from current observations.
Work Kept Out of the Request Path
For a small app, the API could just fetch the needed data, transform it, and return it during the request. But Hangangjari receives requests at the moment users decide to move. If an external source is slow or fails at that moment, the app slows down too and the widget looks empty.
I removed the following jobs from the API request path.
- Calling external sources directly.
- Slow normalization and coordinate correction.
- Rebuilding predictions.
- Calculating notification candidates.
- Repeatedly retrying failed sources.
Workers handle these jobs periodically, and the API reads already-prepared values quickly. The API answers user requests; workers prepare values behind the scenes. They slow down for different reasons and require different checks.
After this split, failure diagnosis also changed. I could distinguish whether the API was slow, whether a worker failed to prepare a value, or whether a source was delayed.
How Stale Values Appear
A last successful value can remain useful after a source failure. Hangangjari returns it with its updated time instead of presenting it as current.
For example, a parking count from five minutes ago can be useful. A value from forty minutes ago requires caution. Event information may be fine after a few hours, but real-time parking information is different. Each data type has its own acceptable freshness.
So the response carried the value together with the updated time, the last success time, and whether the source was currently failing.
Forecasts Kept Separate from Current Values
Parking forecasts cannot include every event, weather change, control, or unexpected incident. So the screen displays forecast availability separately from current parking status.
Prediction sits one level below the current value. Remaining spaces and updated time come first, and prediction helps users understand hourly change. When deciding whether to go right now, the current value is central. When asking whether it might be better to go later, prediction helps.
Workers collect external data in advance and store source state with updated times. When a user opens the app or widget, the API reads and returns those prepared values.
Comments
No comments yet. Be the first to leave one.
Pending review