Separating No Data from Collection Failure

Distinguishing empty public-data responses from parser failures and managing each source policy

A public-data parser reads JSON or HTML and determines whether the response is truly empty, its shape has changed, a park name uses an unfamiliar form, or the last successful fetch is too old. If the parser misses a response-shape change, or schema drift, it can store an empty response and a parsing failure as the same result.

A zero-row day and a parsing failure can produce the same empty-looking screen. Some days genuinely have no event rows; on others, an official page has changed and the parser can no longer read it.

If both are stored as the same value, the app displays a collection failure as zero events.

Hangangjari’s parser evaluates the official source’s properties, update interval, schema drift, park mapping, source URL, and whether a value can appear on screen.

Which Sources Can Reach the Screen

Each public-data source receives a grade that determines how its values can reach the screen.

GradeMeaningHow the app uses it
Official APIDocumented public APICan be a primary source
Official web/AJAXJSON/HTML publicly called by an official websitePrimary or supporting source
Official HTML/RSSHTML/RSS from an official pageUsed after parser stability checks
Curated staticHuman-reviewed static dataFacility, mapping, and correction data
Discovery onlySupporting signal for finding missing dataNot directly exposed in user responses
Do not useLogin, personalization, payment, vehicle number, captcha, private APINot used

This distinction determines the wording shown to users. The app can say “confirmed from an official source” only when an official source is the baseline. A Discovery only source can reveal gaps but never supplies a value directly to a user response.

The source set is limited to sources whose role can be explained. Information that can change where a user goes, especially parking and control information, is displayed with its source and observation time.

Collection Schedules Outside the Code

Hangangjari keeps collection intervals in a schedule catalog. It groups jobs by source ID and execution style in one place.

The current categories are:

CategoryExampleExecution style
Parking masterParking reference-data syncKubernetes CronJob
Parking statusReal-time parking statusworker scheduler
Outing facilityConvenience and park facilitiesKubernetes CronJob
Outing eventEvent informationworker scheduler
Outing noticeNotices and controlsworker scheduler
Realtime contextSeoul real-time city dataworker scheduler
Transit datasetTransit and access helper dataCronJob
Forecast generationForecast creationworker scheduler
Forecast backtestForecast validationCronJob

The system tracks the execution location together with the failed source, stale state, and parser-specific row-count drift.

Sources that need short, repeated checks run inside a worker scheduler. Master, facility, and validation jobs that run daily or every few hours use Kubernetes CronJob. Their cadence and failure impact determine the execution model.

From Source Data to Screen Data

flowchart LR
  Catalog["Collection schedule catalog"] --> Fetch["Fetch via source adapter"]
  Fetch --> Parse["Parser<br/>schema_hash<br/>parser version"]
  Parse --> Normalize["Domain normalization"]
  Normalize --> Validate["Validation<br/>required fields<br/>park mapping<br/>time"]
  Validate --> Upsert["Postgres upsert"]
  Upsert --> Runs["ingestion_runs<br/>status<br/>row count<br/>error"]
  Upsert --> ReadModel["API screen-ready response"]
  Runs --> Health["Source status"]

Raw source data goes through five steps before becoming a screen value.

  1. The catalog provides each source’s contract and execution interval.
  2. The adapter fetches official or public data.
  3. The parser turns raw data into candidate values the app can handle.
  4. The normalizer aligns parks, time, status, and URLs to the app model.
  5. The validator and repository write data to the DB and leave an ingestion run.

These stages separate collection failure from zero rows. If the app cannot show a current value, the response retains an unavailable or stale state.

Park Names and Times for Display

External data rarely arrives in the shape the app wants. Hangangjari normalizes these fields separately.

ItemReason
Park mappingOfficial filter names, park names, coordinates, and keywords can differ
TimeStart, end, registered, modified, and collected times must be separated
StatusScheduled, ongoing, ended, canceled, and unknown are mapped into app states
FreshnessOld successful data must not look current
Source URLUsers should be able to verify the source
Raw payloadNeeded for debugging, but not returned directly in app responses

Park mapping needed particular care. Hangang parks look familiar, but each source describes them slightly differently. Contexts such as “Jamwon,” “Banpo,” and “Banpo/Jamwon” differ.

The place model decides whether to treat them as one place or separate places. String similarity and coordinate distance provide candidates, while reviewed mappings prevent obvious places from being split or distinct places from being merged. These rules now live in the app’s place model instead of individual parsers.

Collection Results as Data

The system stores collection-run details alongside fetched results so operators have evidence to inspect when something goes wrong.

erDiagram
  DATA_SOURCES ||--o{ INGESTION_RUNS : reports
  DATA_SOURCES ||--o{ OUTING_SIGNALS : publishes
  DATA_SOURCES ||--o{ OUTING_FACILITIES : provides
  PARKS ||--o{ OUTING_FACILITIES : contains
  OUTING_SIGNALS ||--o{ OUTING_SIGNAL_PARK_LINKS : maps
  PARKS ||--o{ OUTING_SIGNAL_PARK_LINKS : receives

ingestion_runs is the window into where collection failed.

  • When did it succeed?
  • How many rows were read?
  • Did the response shape hash change?
  • Did the status distribution suddenly change?
  • Did row-count drift appear?
  • What error caused failure?

These values distinguish “the app is slow” from “the source response changed.” They also make it possible to check whether fixing a parser actually improved collection.

Translation Separate from Collection Failure

Displaying Korean-first official data on localized screens requires conditions for translating event names, notices, and facility names and for preserving the original.

Hangangjari separates fetched source text from display copy. The translation cache is a value for screen rendering; it does not replace the original source. Translation failure also must not become collection failure.

Collection Records Behind Source Grades

Even after selecting an official source, zero rows and collection failure can collapse into the same empty screen if collection results are not recorded.

Collection success, row count, schema hash, and freshness are stored as inspectable data. Raw payloads are not returned directly to users, and park mapping plus localized display are separated from source preservation.

Source grades, failure records, and zero-row rules make each screen value traceable to its collection result.

Comments

Comments

    Image preview