Confidence and Backtesting for Parking Forecasts

Showing forecast evidence by separating parking baselines, ranges, confidence, and backtests

Hangangjari’s parking forecast is a reference value for estimating parking risk at arrival time. The response separates current and forecast values and presents the expected remaining spaces with a forecast range, confidence, and source freshness. Backtests compare past forecasts with observed values.

Four Inputs to a Forecast

The parking forecast combines four kinds of evidence.

EvidenceMeaning
Recent statusLatest confirmed remaining spaces and capacity
Recent movementRate of change between recent confirmed values
Historical baselineHistorical mean and variance for the same horizon and time range
Recent-window statisticsMoving average, volatility, and trend correction for the recent window

When a current value exists, recent movement is weighted strongly. The farther the horizon, the more weight shifts to the historical baseline. When no current value exists, the historical baseline or recent-window statistics are used instead. If there is no evidence, the forecast stays empty.

Weights and thresholds can change with operational results. The API also records the evidence used and the reasons its strength was reduced.

flowchart LR
  Snapshots["Recent status"] --> Trend["Change per minute"]
  History["Historical baseline"] --> Estimate["p50 estimate"]
  Features["Recent-window statistics"] --> Estimate
  Trend --> Estimate
  Estimate --> Quantiles["p10 / p50 / p90"]
  Quantiles --> Risk["Risk level"]
  Quantiles --> Status["Expected crowding"]

Forecast Range and Risk

Forecast values use a p10, p50, and p90 range instead of a single estimate.

  • p50: representative expected remaining spaces.
  • p10: low remaining spaces under a conservative view.
  • p90: high remaining spaces under an optimistic view.

Risk uses all three values while placing greater weight on the conservative p10 range.

Confidence from Evidence Strength

Confidence is an evidence-strength value calculated from horizon, source freshness, and sample counts. The screen copy and operational review use the same value.

The current calculation applies these adjustments:

  • Confidence decreases as the horizon gets farther away.
  • Confidence decreases when the source is stale.
  • Sufficient historical samples raise confidence slightly.
  • Sufficient recent-window samples raise confidence slightly.

Reason codes record why confidence fell, including a stale source, a distant horizon, or use of the historical baseline.

flowchart TD
  Horizon["Forecast horizon"] --> Confidence
  Freshness["Source freshness"] --> Confidence
  Samples["Historical sample count"] --> Confidence
  Features["Recent-window sample count"] --> Confidence
  Confidence --> Reason["Reason codes"]

Forecasts Stored as One Bundle

Each forecast generation is stored as a run. A run includes model version, creation time, and forecast horizon. On success it records the number of stored rows and the observed time used as the base. On failure it closes the run as failed.

sequenceDiagram
  autonumber
  participant Job as Forecast job
  participant Repo as Forecast repository
  participant Gen as Forecast generator

  Job->>Repo: Create new forecast-run record
  Gen->>Repo: Read lots/history/baseline/statistics
  Gen->>Gen: Calculate values by horizon
  Gen->>Repo: Store forecast rows
  Gen->>Repo: Close run with base observed time

A run identifies when the forecast visible through the API was calculated and keeps results separated by model version.

Forecast Copy Informed by Backtests

Hangangjari backtests attach labels from confirmed values near each forecast’s target arrival time, then calculate metrics by horizon.

The values are roughly:

  • Sample count.
  • Average error similar to MAE.
  • Ratio of actual values inside the p10-p90 range.
  • Ratio of samples with stale sources.
  • Gate pass/fail.
flowchart LR
  ForecastRows["Past forecast rows"] --> Labels["Attach actual values near arrival time"]
  Labels --> Score["Compare forecast and actual value"]
  Score --> Metrics["Horizon/model metrics"]
  Metrics --> Gate["Quality gate"]

Backtest results inform copy and confidence by horizon. Horizons with larger errors or weaker p10-p90 coverage receive lower confidence and corresponding reason codes. Metrics are compared by model version after formula changes.

Forecast Evidence Included in the Response

The parking forecast response includes p10, p50, p90, risk level, confidence, and freshness. Reason codes support screen copy and debugging, while runs and model versions keep backtest results separate from earlier calculations.

Comments

Comments

    Image preview