Separating No Data from Collection Failure
Distinguishing empty public-data responses from parser failures and managing each source policy
A public-data parser reads JSON or HTML and determines whether the response is truly empty, its shape has changed, a park name uses an unfamiliar form, or the last successful fetch is too old. If the parser misses a response-shape change, or schema drift, it can store an empty response and a parsing failure as the same result.
A zero-row day and a parsing failure can produce the same empty-looking screen. Some days genuinely have no event rows; on others, an official page has changed and the parser can no longer read it.
If both are stored as the same value, the app displays a collection failure as zero events.
Hangangjari’s parser evaluates the official source’s properties, update interval, schema drift, park mapping, source URL, and whether a value can appear on screen.
Which Sources Can Reach the Screen
Each public-data source receives a grade that determines how its values can reach the screen.
| Grade | Meaning | How the app uses it |
|---|---|---|
| Official API | Documented public API | Can be a primary source |
| Official web/AJAX | JSON/HTML publicly called by an official website | Primary or supporting source |
| Official HTML/RSS | HTML/RSS from an official page | Used after parser stability checks |
| Curated static | Human-reviewed static data | Facility, mapping, and correction data |
| Discovery only | Supporting signal for finding missing data | Not directly exposed in user responses |
| Do not use | Login, personalization, payment, vehicle number, captcha, private API | Not used |
This distinction determines the wording shown to users. The app can say “confirmed from an official source” only when an official source is the baseline. A Discovery only source can reveal gaps but never supplies a value directly to a user response.
The source set is limited to sources whose role can be explained. Information that can change where a user goes, especially parking and control information, is displayed with its source and observation time.
Collection Schedules Outside the Code
Hangangjari keeps collection intervals in a schedule catalog. It groups jobs by source ID and execution style in one place.
The current categories are:
| Category | Example | Execution style |
|---|---|---|
| Parking master | Parking reference-data sync | Kubernetes CronJob |
| Parking status | Real-time parking status | worker scheduler |
| Outing facility | Convenience and park facilities | Kubernetes CronJob |
| Outing event | Event information | worker scheduler |
| Outing notice | Notices and controls | worker scheduler |
| Realtime context | Seoul real-time city data | worker scheduler |
| Transit dataset | Transit and access helper data | CronJob |
| Forecast generation | Forecast creation | worker scheduler |
| Forecast backtest | Forecast validation | CronJob |
The system tracks the execution location together with the failed source, stale state, and parser-specific row-count drift.
Sources that need short, repeated checks run inside a worker scheduler. Master, facility, and validation jobs that run daily or every few hours use Kubernetes CronJob. Their cadence and failure impact determine the execution model.
From Source Data to Screen Data
flowchart LR Catalog["Collection schedule catalog"] --> Fetch["Fetch via source adapter"] Fetch --> Parse["Parser<br/>schema_hash<br/>parser version"] Parse --> Normalize["Domain normalization"] Normalize --> Validate["Validation<br/>required fields<br/>park mapping<br/>time"] Validate --> Upsert["Postgres upsert"] Upsert --> Runs["ingestion_runs<br/>status<br/>row count<br/>error"] Upsert --> ReadModel["API screen-ready response"] Runs --> Health["Source status"]
Raw source data goes through five steps before becoming a screen value.
- The catalog provides each source’s contract and execution interval.
- The adapter fetches official or public data.
- The parser turns raw data into candidate values the app can handle.
- The normalizer aligns parks, time, status, and URLs to the app model.
- The validator and repository write data to the DB and leave an ingestion run.
These stages separate collection failure from zero rows. If the app cannot show a current value, the response retains an unavailable or stale state.
Park Names and Times for Display
External data rarely arrives in the shape the app wants. Hangangjari normalizes these fields separately.
| Item | Reason |
|---|---|
| Park mapping | Official filter names, park names, coordinates, and keywords can differ |
| Time | Start, end, registered, modified, and collected times must be separated |
| Status | Scheduled, ongoing, ended, canceled, and unknown are mapped into app states |
| Freshness | Old successful data must not look current |
| Source URL | Users should be able to verify the source |
| Raw payload | Needed for debugging, but not returned directly in app responses |
Park mapping needed particular care. Hangang parks look familiar, but each source describes them slightly differently. Contexts such as “Jamwon,” “Banpo,” and “Banpo/Jamwon” differ.
The place model decides whether to treat them as one place or separate places. String similarity and coordinate distance provide candidates, while reviewed mappings prevent obvious places from being split or distinct places from being merged. These rules now live in the app’s place model instead of individual parsers.
Collection Results as Data
The system stores collection-run details alongside fetched results so operators have evidence to inspect when something goes wrong.
erDiagram
DATA_SOURCES ||--o{ INGESTION_RUNS : reports
DATA_SOURCES ||--o{ OUTING_SIGNALS : publishes
DATA_SOURCES ||--o{ OUTING_FACILITIES : provides
PARKS ||--o{ OUTING_FACILITIES : contains
OUTING_SIGNALS ||--o{ OUTING_SIGNAL_PARK_LINKS : maps
PARKS ||--o{ OUTING_SIGNAL_PARK_LINKS : receives
ingestion_runs is the window into where collection failed.
- When did it succeed?
- How many rows were read?
- Did the response shape hash change?
- Did the status distribution suddenly change?
- Did row-count drift appear?
- What error caused failure?
These values distinguish “the app is slow” from “the source response changed.” They also make it possible to check whether fixing a parser actually improved collection.
Translation Separate from Collection Failure
Displaying Korean-first official data on localized screens requires conditions for translating event names, notices, and facility names and for preserving the original.
Hangangjari separates fetched source text from display copy. The translation cache is a value for screen rendering; it does not replace the original source. Translation failure also must not become collection failure.
Collection Records Behind Source Grades
Even after selecting an official source, zero rows and collection failure can collapse into the same empty screen if collection results are not recorded.
Collection success, row count, schema hash, and freshness are stored as inspectable data. Raw payloads are not returned directly to users, and park mapping plus localized display are separated from source preservation.
Source grades, failure records, and zero-row rules make each screen value traceable to its collection result.
Comments
No comments yet. Be the first to leave one.
Pending review