Source Catalog and Ingestion Runs

Recording source schedules and ingestion outcomes to distinguish no data from collection failure

Public-data collection records each source’s schedule, screen state on failure, row-count changes, and last successful run.

Hangangjari defines each job and its schedule in a collection plan, or source catalog, and stores each execution result in an ingestion run. Linking them identifies when each source failed.

Collection Targets and Cadence in One Place

The collection plan keeps each source ID, job ID, owner, role, interval, and stale threshold in one place.

flowchart LR
  Catalog["Collection plan"] --> Scheduler["Worker scheduler"]
  Catalog --> Cron["Kubernetes CronJob"]
  Scheduler --> Jobs["Parking / outing / forecast jobs"]
  Cron --> Batch["Reference-data and backtest jobs"]

The list includes sources of these shapes.

CategoryExample
parking masterSync parking-lot reference data
parking statusPoll realtime parking values
outing facilityCollect facility data
outing eventCollect event data
outing noticeCollect notices and controls
realtime contextCollect crowding, weather, and traffic context
forecastBuild forecasts and backtests

Each job records its execution location together with whether it supports a visible value, affects freshness, or appears directly on the user screen.

This catalog covers only jobs that read external sources and change screen-data freshness. Notification candidate creation and outbox dispatch belong to the delivery ledger, so the push worker is excluded.

Recording the Final Ingestion State

An ingestion run records how a collection job finished.

flowchart LR
  Fetch["Fetch source"] --> Parse["Parse"]
  Parse --> Normalize["Normalize"]
  Normalize --> Upsert["Upsert domain rows"]
  Upsert --> Run["Ingestion-run record"]
  Run --> Health["Source health"]
  Health --> API["API freshness/source state"]

When Hangangjari’s outing ingestion succeeds, it records:

  • Source ID.
  • Start and end time.
  • Success or failure.
  • Number of rows read.
  • Status distribution.
  • Schema hash.
  • Failure error summary.

Operators use these values to identify parser failures. The API also uses them to produce source-check results such as fresh, stale, and unavailable.

Run records are stored with result rows so past source failures and parser changes remain inspectable.

“No Data” Versus “Collection Failed”

Showing every collection failure as “none” makes an actual zero-row result indistinguishable from an incident.

Zero events is different from a failed event source. No facility is different from a facility parser producing empty rows. Stale realtime context is different from a disabled source.

The source catalog and ingestion runs preserve these state distinctions.

StateMeaningSignal to users
freshRecent successful data existsUsable as reference
partialOnly some sources are availableLimited reference
staleLast success is oldOn-site difference possible
unavailableCollection failed or source is unavailableDistinct from none

Row Counts and Schema Changes

Public-data sources can change without notice. Field names or HTML structure can change, and rows for a specific park can disappear.

Collection results therefore include schema hash and row count. These values distinguish “there are fewer events today” from “the parser only read half of them.”

Status distribution gives the same kind of signal. If every facility value suddenly becomes unknown, or if crowding distribution is abnormally skewed, the source or parser may be the problem.

Screen Guidance from Ingestion Results

Source-check results feed both operations views and freshness copy in the app and widgets.

The Hangangjari API classifies sources as fresh, stale, or unavailable based on the last success time and stale threshold per source. This information explains sources and refresh state on screen, and it becomes warning text and unavailable reasons in widgets.

The dashboard’s stale state becomes “information may not be current” in the app and warning text in widgets.

erDiagram
  DATA_SOURCES ||--o{ INGESTION_RUNS : records
  DATA_SOURCES ||--o{ OUTING_SIGNALS : publishes
  DATA_SOURCES ||--o{ OUTING_FACILITIES : supplies
  INGESTION_RUNS }o--|| SOURCE_HEALTH : derives

Ingestion Records Used by the API

The source catalog defines schedules and stale thresholds. Each ingestion run retains status, row count, status distribution, and schema hash. The API reads the latest run to distinguish zero rows from collection failure and reflects the same state in operational alerts, app freshness copy, and widget warnings.

Comments

Comments

    Image preview