Team Ai
Datasetpublic

Saxon0520/LiveHouse-TS-test

LiveHouse-TS A prospective benchmark for univariate time-series forecasting. Real streams. Predictions frozen in advance. Results as the future unfolds. Explore the leaderboard ↗   ·   Submit a model   ·   Public data   ·   Evaluation   ·   Cite Forecast first. Evaluate later. LiveHouse-TS asks a simple question: how well does a model predict data that has not arrived yet? Models receive… See the full description on the dataset page: https://huggingface.co/datasets/Saxon0520/LiveHouse-TS-test.

sourceHugging Faceapache-2.0updated 2h agoView on Hugging Face
0likes3.3kdownloads
Dataset Card

<a id="livehouse-ts"></a>

<p align="center"> <img src="https://huggingface.co/spaces/Saxon0520/LiveHouse-TS-test/resolve/main/assets/readme-hero.svg" alt="LiveHouse-TS — Forecast before the future arrives." width="100%"> </p>

<h1 align="center">LiveHouse-TS</h1>

<p align="center"> A prospective benchmark for univariate time-series forecasting.<br> <strong>Real streams. Predictions frozen in advance. Results as the future unfolds.</strong> </p>

<p align="center"> <a href="https://saxon0520-livehouse-ts-test.static.hf.space/index.html#leaderboard"><strong>Explore the leaderboard ↗</strong></a> &nbsp; · &nbsp; <a href="#join-the-benchmark">Submit a model</a> &nbsp; · &nbsp; <a href="https://huggingface.co/datasets/Saxon0520/LiveHouse-TS-test">Public data</a> &nbsp; · &nbsp; <a href="#evaluation-protocol">Evaluation</a> &nbsp; · &nbsp; <a href="#citation">Cite</a> </p>

<p align="center"> <a href="#quick-start"><img src="https://img.shields.io/badge/Python-3.10%2B-6366f1?style=flat-square" alt="Python 3.10 or newer"></a> <a href="#evaluation-protocol"><img src="https://img.shields.io/badge/Evaluation-Prospective-0f766e?style=flat-square" alt="Prospective evaluation"></a> <a href="#paper-model-roster"><img src="https://img.shields.io/badge/Models-API%20%2B%20Baselines-7c3aed?style=flat-square" alt="Hosted APIs and statistical baselines"></a> <a href="https://github.com/ATMSaxon/LiveHouse-TS-test/blob/main/LICENSE.txt"><img src="https://img.shields.io/badge/License-Apache%202.0-475569?style=flat-square" alt="Apache 2.0 license"></a> </p>


Forecast first. Evaluate later.

LiveHouse-TS asks a simple question: how well does a model predict data that has not arrived yet? Models receive only observations available at the forecast cutoff. Predictions are frozen before the target window begins, then scored once the complete target is available.

ObservePredictCompare
Collect public streams across weather, finance, mobility and more.Give every model the same historical context and freeze its forecast.Explore point and probabilistic errors, model comparisons and archived results.

Development preview. The links above point to the personal test environment. Not every source or model is ready for evaluation; incomplete data and invalid forecasts are excluded. See model coverage and data availability for current limitations.

Explore the benchmark

<a href="https://saxon0520-livehouse-ts-test.static.hf.space/index.html#leaderboard"> <img src="https://huggingface.co/spaces/Saxon0520/LiveHouse-TS-test/resolve/main/assets/readme-benchmark.jpg" alt="Actual LiveHouse-TS test website: model error heatmap across datasets grouped by domain" width="100%"> </a>

Actual test-site screenshot captured on October 7, 2026. This heatmap is an exploratory error comparison, not an official ranking; values change as new results arrive. The header waveform is decorative, not benchmark data.

Compare datasets by domain, inspect head-to-head results, follow historical ranks, and overlay forecasts with observed outcomes.

**Open the explorer →** &nbsp; · &nbsp; Browse the archive &nbsp; · &nbsp; Inspect forecasts

Join the benchmark

Wrap → Validate → Submit. Expose your model through the example API, check its response contract, then submit its endpoint and model metadata. Install the SDK with the quick start below.

<a id="submit-your-own-model"></a>

Participants host their model behind one small HTTP API. Start from `examples/model_api`, replace forecast_one, and expose:

text
GET  /health
POST /forecast

The request contains opaque series IDs, historical timestamps and values, frequency, horizon, and requested quantiles. It never contains future targets, private metrics, dataset-internal IDs, or another model's predictions. Remote endpoints must use stable HTTPS.

Each input must have one output with prediction_length finite numeric values in mean and in each requested quantile (0.1 through 0.9 in steps of 0.1). Quantiles must be nondecreasing at each time step. Point-only models may repeat their point forecast at every quantile, as the demo does; this is a degenerate distribution, not a calibrated uncertainty estimate. Admission and live evaluation apply the same checks before storing predictions.

The example echoes protocol_version and model at the top level and series_id on each output. Existing value-only responses still work; if these fields are returned, they must match the request. Optional output timestamps must have explicit timezones and follow the context's last timestamp on the requested grid, including any forecast gap. Shifted timestamps, mismatched identities and duplicate quantile levels are rejected, not repaired. The TSFM adapter also checks returned metadata.item_id and timestamps when present.

Validate an endpoint before submission:

bash
livehouse-ts-validate organization/model https://forecast.example.org/forecast

The command exits zero with status: ok, or exits nonzero with a safe failure code (for example forecast_misaligned, unexpected_series, or crossed_quantiles). It does not print provider response bodies. CI exercises the example over real HTTP through admission, prediction freezing, scoring and public export, using synthetic test data only.

Submit the following through the GitHub community-model form:

  1. 1.model ID and display name;
  2. 2.model card URL;
  3. 3.public endpoint-code URL;
  4. 4.stable HTTPS /forecast URL;
  5. 5.version and organization.

After endpoint validation, maintainers record an admission timestamp. A model is evaluated only on tasks created after admission; no historical backfill is used.

Quick start

bash
git clone https://github.com/ATMSaxon/LiveHouse-TS-test.git
cd LiveHouse-TS-test
python -m venv .venv
source .venv/bin/activate
python -m pip install -e '.[dev,baselines,api]'
pytest -q

Create a schema-v2 database or add the v2 audit tables to an existing v1 database:

python
from livehouse_ts.data_schema import connect

database = connect("livehouse.sqlite")

Implement DataSource.fetch() to return a normalized DataBatch, then call collect(source, SQLiteRepository(database)). A data adapter is responsible for recording event time, public availability time, and ingestion time separately.

Run one cycle against the fourteen currently enabled data streams:

bash
livehouse-ts-cycle private/livehouse.sqlite public

This collects past observations, resolves due tasks, issues dataset-specific forecasts, and writes public CSV/JSON artifacts. Seasonal Naive is always included. Add HTTPS models through LIVEHOUSE_MODELS_JSON:

json
[{"model_id":"organization/model","endpoint_url":"https://forecast.example.org/forecast"}]

Evaluation protocol

For each task, LiveHouse-TS:

  1. 1.selects context whose available_time is no later than the task cutoff;
  2. 2.calls every admitted model with the same context and horizon;
  3. 3.freezes predictions before target_start;
  4. 4.waits until the target window ends and every target value is available;
  5. 5.computes metrics with a recorded metric version;
  6. 6.publishes only public-format-v1 aggregate results and bounded visual examples.

Old pre-schema results are intentionally excluded from the new leaderboard.

Implementation details, task windows and source availability are documented in the technical reference.

Metrics and ranking

Metric v2 computes MSE, RMSE, MAE and quantile-approximate CRPS using the context population standard deviation only (unit scale for a near-constant context). MAPE uses the median forecast in original units and is omitted near zero targets. CRPS uses the nine equally spaced quantiles 0.1–0.9 from paper Appendix E.

Official ranking follows Appendix E: compare MSE and CRPS separately on shared releases, average the two win/tie/loss outcomes, then macro-average by dataset and eligible opponent. Each pair needs 30 shared releases, five datasets and seven days of target-time span. A model needs three eligible opponents and membership in the largest connected comparison component to receive an official rank. Ties share rank; missing comparisons are not ties. Otherwise status is provisional, awaiting forecasts/scores, or failed.

The lower-is-better normalized composite score remains a diagnostic, not the official ranking key; zero-error baseline denominators are omitted from this diagnostic only. Diagnostic Elo starts at 1500, K=16 per shared-task pair, with simultaneous updates and deterministic target-time/task-ID ordering. It is not the official rank and the paper does not specify these exact Elo parameters. Average rank uses tied shared-task composite ranks, macro-averaged by dataset. RTG is 100*(baseline MSE-model MSE)/(baseline MSE+model MSE), averaged by dataset (both zero gives zero). Stability is the sample standard deviation of daily dataset-balanced normalized composite scores. Improvement is Kendall tau-a of normalized squared errors for repeated predictions of the same target and truth revision, averaged by target and dataset; fewer than two predictions gives no value. Coverage uses only reference tasks after the model's admission.

These implementation choices are explicit; this repository does not yet claim full numerical reproduction of every paper table or the original experiment.

Core workflow

text
public sources -> normalized observations -> frozen forecast
               -> future observations arrive -> metrics -> public leaderboard

The public SDK has five modules:

  • —data.py: streaming-data adapters and SQLite ingestion;
  • —data_schema.py: canonical objects and SQLite schema v2;
  • —models.py: the model protocol, hosted adapters, and four statistical baselines;
  • —metrics.py: versioned point and probabilistic metrics;
  • —eval.py: forecast freezing, scoring, and public export.

Runtime data, model weights, raw responses, frozen forecasts, and result files are not stored in this Git repository.

LiveHouse series

This repository is intentionally scoped to univariate time series. Future spatio-temporal graph or multivariate benchmarks will use separate LiveHouse projects while sharing versioned concepts such as datasets, entities, variables, tasks, models, and releases.

Update log

October 10, 2026 — focused benchmark explorer

  • —Put Elo first and Archive last; removed the result-date card and seven redundant or unwanted plots.
  • —Kept twelve complementary charts; paired win/tie/loss counts now appear in the comparison summary.
  • —Replaced the multicolored Elo bars with a sorted, single-color dot plot and aligned numeric ratings.
  • —Kept the scoring, official ranks, archive data and forecasting protocol unchanged.

October 7, 2026 — README refresh

  • —Added a branded overview, an actual benchmark screenshot and direct participation links.
  • —Moved operational reference material behind the introductory workflow; evaluation behavior is unchanged.

SQLite schema v2 — audit storage and source completion

  • —Checked model-response identity, optional forecast timestamps and duplicate quantile levels; added real HTTP endpoint-to-scoring integration coverage.
  • —Allowed previously captured GBFS periods to settle through a new fetch outage; counted only newly stored observations.
  • —Used frozen inputs consistently for the Improvement diagnostic as well as scoring and forecast examples.
  • —Added private exact-response archives and immutable canonical task inputs; migrated v1 without deleting results.
  • —Implemented all 17 source adapters; enabled hourly NWS aggregation and GBFS snapshot accumulation.
  • —Added quarter-hour GBFS-only collection without increasing hosted model inference frequency.
  • —Kept NCEI, NDBC and GDELT disabled pending usable target data; no substitution or gap filling.

0.2.0 — evaluation v2 / metrics v2 (SQLite schema remains v1)

  • —Added Binance hourly/second/monthly, CoinGecko and World Bank streams and exact calendar task grids.
  • —Backfilled due second-level target windows without backdating data availability.
  • —Added isolated hosted-model diagnostics without changing benchmark admission.
  • —Fixed per-dataset issue clocks, reserved inference time, and bounded concurrent model calls.
  • —Added transport-only data retries, safe failure reasons, and partial-run Actions failures.
  • —Added dataset-specific task windows with explicit history warm-up status.
  • —Added NASA POWER UTC parsing and complete-hour USGS earthquake counts.
  • —Kept NCEI daily mean temperature disabled while its target field is absent.
  • —Enforced inference completion deadlines and gap-aware timestamp alignment.
  • —Separated new results from legacy scores without deleting historical data.
  • —Added nine-quantile context-normalized scoring and paper eligibility rules.
  • —Corrected RTG, temporal stability, fixed-target Improvement, and tied ranks.
  • —Kept unscored models visible and labelled data provenance.
  • —Added USGS discharge, NOAA water level and Wikimedia pageviews adapters.

0.1.0 — schema v1

  • —Reduced the public project to five core modules and one HTTP model example.
  • —Added canonical SQLite storage with separate event, availability and ingest times.
  • —Added the livehouse-ts-v1 model and metric protocols.
  • —Added two real Open-Meteo streams and the hourly single-writer operator.
  • —Added normalized ranking, win rate, Elo, RTG, stability and availability.
  • —Started a new leaderboard containing only schema-v1 releases.

Future schema, protocol, metric, data-source, or ranking changes receive an explicit version entry here.

Citation

bibtex
@misc{livehouse_ts,
  title  = {LiveHouse-TS: A Live Benchmark for Time-Series Forecasting},
  author = {ThinkCat Lab},
  year   = {2026},
  url    = {https://github.com/ATMSaxon/LiveHouse-TS-test}
}

Apache-2.0 licensed. LiveHouse-TS builds on the prospective evaluation direction of GIFT-Eval while maintaining its own live-data schema and protocol.


Technical reference

Models · Data streams · Architecture · Evaluation flow · Database fields · Maintainer checklist

Paper model roster

Self-hosting status: the eight foundation models are not yet self-hosted. The adapters below are optional hosted-provider integrations, not completed local deployments. GPU/server configuration and actual weight loading remain deferred. Do not enable paper to claim self-hosted results.

LIVEHOUSE_MODEL_SET=baselines enables Seasonal Naive, Moving Average (24), ARIMA(1,1,1), and ETS (additive trend, no seasonality). LIVEHOUSE_MODEL_SET=paper additionally enables all eight foundation models:

LIVEHOUSE_MODEL_SET=tsfm selects the seven TSFM.ai adapters without requiring the independent TabPFN client. LIVEHOUSE_TSFM_MODELS_JSON can restrict this to an explicitly validated list of IDs; an empty setting means the complete seven. In the personal sandbox, the 2026-10-04 synthetic admission check passed Chronos-2, TiRex, Toto-1.0, Moirai-2.0, Chronos-Bolt and Sundial. TimesFM-2.5 returned crossed quantiles and is not admitted by that check; predictions are not repaired or substituted. This is contract validation, not a forecasting-quality result. The isolated 2026-10-05 recheck found six adjacent-quantile crossings (maximum 0.1061) in the raw six-step synthetic response. TabPFN is absent from the TSFM.ai catalog checked that day and still requires its own TABPFN_TOKEN.

ModelProvider / original registry identifier
Chronos-2TSFM.ai: amazon/chronos-2
TiRexTSFM.ai: NX-AI/TiRex-1.1-gifteval
TimesFM-2.5TSFM.ai: google/timesfm-2.5-200m-pytorch
Toto-1.0TSFM.ai: Datadog/Toto-Open-Base-1.0
Moirai-2.0TSFM.ai: Salesforce/moirai-2.0-R-small
Chronos-BoltTSFM.ai: amazon/chronos-bolt-base
SundialTSFM.ai: thuml/sundial-base-128m
TabPFN-TSPrior Labs client: priorlabs/tabpfn-ts

Install .[baselines] for statistical inference and additionally .[tabpfn] for the paper roster. Core imports remain dependency-light. The TabPFN package versions match the previous repository's client environment; its hosted checkpoint is provider-controlled, so this does not establish exact reproduction of the paper's historical weights. The other hosted model IDs are also not immutable weight hashes.

In Actions, set secrets TSFM_API_KEY and TABPFN_TOKEN, then manually run Validate paper models with paper. This makes real provider calls on a synthetic series, uses inference quota, and does not publish benchmark results. Only after it passes, set repository variable LIVEHOUSE_MODEL_SET=paper to enable the hourly operator. minimal is the default and keeps the reference baseline plus explicitly configured HTTPS endpoints. Missing credentials fail before an operator run can change HF data.

These are prospective runs, not imported paper scores. Statistical models use a deterministic 200-sample centered residual bootstrap; failed fits are recorded as failures, never silently replaced by a different model. TiRex asks for at least 24 future steps and retains the requested prefix, matching the old registry's hosted-horizon workaround. The persistence demo is not a paper model.

Evaluation implementation and data streams

Evaluation v2 checks the clock both before and after inference, rejects late responses, verifies regular context/target grids, and forecasts across the gap between the last available context and the future target before slicing out the scored horizon. Context must have both availability and ingestion times no later than cutoff. The CLI does not accept backdated issue times. The first successful scoring of a run is immutable; later source revisions do not replace it.

The operator now takes a fresh cutoff for each dataset after earlier work has finished. All models on that task still receive the same context, horizon and deadline. It chooses the first native-grid timestamp strictly after a reserved inference budget: max(model timeout) * ceil(model count / workers) + 30 seconds, with at most four inference workers and a 60-second planning allowance for local models without a declared timeout. This allowance is not a forced kill for arbitrary local code: actual late results are still rejected. Forecasting covers the resulting gap before slicing the target horizon. SQLite reads/writes stay on the evaluator thread; workers perform only model calls and validation. Queued calls check the deadline again before contacting a model.

An irregular context grid is reported once under status.json.skipped_tasks before dispatch, rather than recorded as a failure of every model. No values are filled, shifted or discarded to make the grid pass. Data GETs retry up to three times (1/2-second waits) for truncated connections, transport timeouts and 408/502/503/504 responses. Authentication errors, rate limits and malformed JSON are not retried. Model inference calls are not automatically retried. NOAA water levels are fetched with explicit UTC begin/end timestamps covering the last 72 hours, so overlap is requested again and missing original records can be appended when the provider makes them available. This does not interpolate gaps or rewrite existing first-ingestion times.

Failures carry safe reason codes such as deadline_before_inference, deadline_after_inference, context_irregular, invalid_horizon, missing_quantiles, crossed_quantiles, http_503 or incomplete_response. Unknown errors expose only their exception class, never response bodies or tokens. A partial cycle preserves and publishes its valid results, then fails the Actions job with a summary of affected tasks. A green workflow therefore no longer silently covers forecast/collection failures. History warm-up remains a normal pending state. Existing failed runs and scores are not rewritten.

Data coverage

Enabled streamValue typeTask frequencyContext / forecast steps
Open-Meteo Shanghai temperatureModel estimate, not station truth1h336 / 24
Open-Meteo Shanghai PM2.5Model estimate, not station truth1h336 / 24
USGS Potomac dischargeGauge measurement5min (current API; paper lists 15min)288 / 72
NOAA San Francisco water levelStation measurement6min240 / 60
Wikimedia Time series pageviewsReported count1d30 / 7
NASA POWER Shanghai temperatureGridded meteorological model estimate1h (explicit UTC)336 / 24
USGS global earthquake catalogue countsReported event count, complete hours1h168 / 24
Binance BTC/USDT hourly closeCompleted candle closing price1h96 / 24
Binance BTC/USDT one-second closeCompleted candle closing price1s900 / 60
Binance BTC/USDT monthly closeCompleted candle closing price1mo60 / 12
CoinGecko Bitcoin priceHourly market price1h96 / 24
World Bank China GDPReported annual statistic, current USD1y40 / 5
NWS KSFO temperatureLast valid measurement in each completed UTC hour1h96 / 24
Citi Bike W 52 St & 11 AveLast observed report in each completed UTC quarter-hour15min96 / 4

Availability is conservatively recorded at first ingestion, not inferred from event time. Irregular or missing grids are rejected, not silently filled. These fourteen streams do not yet reproduce the paper's full 17-dataset inventory. Window lengths live in data.py and travel with each dataset's metadata. The operator requires the full configured context; shorter histories appear in status.json.pending_context and do not create undersized tasks. Unregistered custom adapters retain the 168/6 default (minimum 24 points). USGS water uses 288/72 to preserve the paper's 24-hour context and six-hour horizon at its current five-minute cadence. Full evaluation remains hourly; a GBFS-only collector runs between those evaluations to obtain quarter-hour snapshots. Exported release rows record actual horizon and frequency. Earlier six-step tasks keep their original definitions and results; new tasks have window lengths in their identity, so changing a window never rewrites an existing task. Aggregates currently include both old and new task windows.

NASA is requested in UTC, not the API's default local solar time, and its fill value is excluded. Its release delay is handled as a forecast gap, not hidden. The earthquake weekly feed usually provides 167 complete hours on first fetch; the 168-step context warms up as subsequent hourly collections accumulate. Zero counts are emitted only inside complete feed coverage, and partial edge hours are excluded. Feed truncation/duplicate event IDs are rejected.

All 17 source adapters are implemented. Three are not enabled for evaluation: NCEI USW00014732 still has no TAVG (0 of 85 returned records on 2026-10-05), NDBC 46013 lacks a continuous 10-minute wave-height grid, and GDELT has returned HTTP 429. NCEI is not replaced by TMAX/TMIN or another station. NDBC keeps actual timestamps and missing values absent, rather than filling its missing wave heights. GDELT requests four adjacent 48-hour timelines with no smoothing and six-second spacing, covering eight days while preserving 15-minute resolution: the provider changes to hourly resolution at 72 hours, so a 7-day query cannot be called 15-minute data. API gaps and missing values are never filled. The operator's check_sources action probes all 17 without writing to HF.

NWS follows pagination until its 96-hour context is covered, selects the last finite measurement per completed UTC hour, rejects MADIS X/Q/B quality flags, and preserves empty hours as gaps. This is an explicit hourly aggregation, not a claim that every native station report is hourly. Raw responses preserve the original measurement times. The 2026-10-05 probe found an entirely missing temperature hour at 2026-10-03 20:00 UTC (the only report had a null temperature). Such a gap prevents task creation until it is repaired by the provider or leaves the 96-hour context.

GBFS fixes station 66db237e-0aca-11e7-82f6-3863bb44ef7c (W 52 St & 11 Ave). Raw station reports keep their last_reported times in the snapshot variable. Only completed quarter-hours are aggregated into the univariate target, using the last actually observed report; no empty bucket is filled. Stale feeds and inactive stations are rejected. The initial 96-period context takes at least 24 hours to accumulate. GitHub schedules are best-effort: missing intervals remain gaps until a complete context is available, not manufactured history. Snapshot-only runs preserve the last full evaluation's health status; a successful collector cannot clear outstanding model failures. An upstream fetch failure still closes previously captured, completed GBFS buckets, so available truth can be scored without another model call. The failure remains reported; missing buckets are not filled. Operator observation counts report newly inserted rows, not repeated polls of existing revisions.

The Binance one-second adapter fetches the latest 1,000 candles for context and explicitly backfills each due 60-second target window on the next operator run. Backfilled observations become available at ingestion, not retrospectively at the candle timestamp. This is second-resolution data with hourly task issue, not the paper's every-second evaluation schedule. All Binance adapters use the original candle-open labels and exclude candles whose close time has not passed. CoinGecko's off-grid latest sample is excluded without rounding timestamps or filling gaps. World Bank checks the country, indicator and page completeness, skips missing values and uncompleted years, and labels each annual value at January 1.

Month/year task grids use real calendar boundaries, including leap years, in context checks, forecast gaps, target construction and scoring. TSFM requests translate 1mo/1y to month-start/year-start aliases MS/YS. Monthly and annual tasks will remain pending until the complete future target is published; historical observations are context, not retroactive live scores.

Storage and deployment

System architecture

The deployment has three boundaries: the evaluator owns private state and credentials; model services receive historical inputs only; the browser reads public exports only. No API keys belong in the Space or public Dataset.

mermaid
flowchart LR
    Sources[Public data APIs] -->|observations| Runner[Single evaluator worker]
    Private[(Private HF SQLite)] -->|restore pinned revision| Runner
    Runner -->|historical context only| Models[Model endpoints]
    Models -->|mean and quantiles| Runner
    Runner -->|commit database first| Private
    Runner -->|derived CSV and JSON| Public[Public HF Dataset]
    Public -->|read-only| Space[HF Space leaderboard]
    Public -->|pinned revisions| Archive[Daily archive]

The operator restores a pinned database into an ephemeral working directory, collects data, resolves old predictions, creates new tasks, and uploads the private database before publishing public artifacts. HF parent-commit checks reject conflicting writes. This is not a cross-repository transaction: private upload may succeed while public upload fails. source_revision.json identifies the private commit and code revision supporting each public export. The next operator resumes from the latest private state; it does not delete or roll back successful private commits. Only one evaluator should be scheduled.

Evaluation flow

mermaid
flowchart TD
    Restore[Restore private SQLite] --> Collect[Collect and normalize streams]
    Collect --> Resolve[Score due runs only when the full target grid exists]
    Resolve --> Context{Full causal context available?}
    Context -->|No| Wait[Record pending context]
    Context -->|Yes| Task[Create future task with dataset-specific window]
    Task --> Admission{Model admitted and deadline not reached?}
    Admission -->|No| Failure[Record failed model-task pair]
    Admission -->|Yes| Infer[Request gap plus target horizon]
    Infer --> Validate{Valid output and returned before target start?}
    Validate -->|No| Failure
    Validate -->|Yes| Freeze[Store target predictions and completion time]
    Freeze --> Publish[Upload private state then public exports]
    Wait --> Publish
    Failure --> Publish
    Publish --> Next[Next scheduled cycle resolves pending targets]

Forecast input contains no future labels. Context is selected using both available_time and ingested_at at the task cutoff. The model predicts the unobserved gap plus target horizon; gap steps are discarded before storage. Failures affect one model-task pair, not other models. A task-level status is only a summary: use forecast_runs and scores to determine each model's state.

SQLite data dictionary (schema v2)

The v1→v2 migration only adds raw-response archives and frozen task inputs. Existing observations, predictions and scores are preserved; evaluation and metric definitions remain v2. The browser/export JSON contract remains v1.

Types below are SQLite types. PK means primary key; FK means foreign key. Unless marked nullable, fields are declared NOT NULL (SQLite TEXT primary keys are an exception in the DDL; the repository always supplies their IDs). All event/cutoff times written by the repository are UTC ISO-8601 with six fractional digits. JSON fields hold objects, not additional relational tables.

TableFieldTypeMeaning / constraints
schema_metadataversionINTEGERPK; applied migration version, currently 2
schema_metadataapplied_atTEXTSchema initialization time
sourcessource_idTEXTPK; provider/source identifier
sourcesnameTEXTHuman-readable source name
sourcesurlTEXTProvider endpoint, no credentials
sourceslicenseTEXTAttribution/license description; may be unverified provider terms
sourcesfetchintervalsecondsINTEGERNullable; intended polling interval, not an enforced scheduler
datasetsdataset_idTEXTPK; benchmark stream identifier
datasetssource_idTEXTFK → sources.source_id
datasetsnameTEXTDisplay name
datasetsdomainTEXTWeather, hydrology, ocean, etc.; descriptive category
datasetsfrequencyTEXTTask grid step, e.g. 1h, 5min, 1d; aggregation is recorded separately
datasetstimezoneTEXTDefault UTC
datasetsbenchmark_familyTEXTDefault time_series; future series extension point
datasetsmodalityTEXTDefault univariate; does not implement multivariate evaluation
datasetsmetadata_jsonTEXTDefault {}; contextlength, predictionlength, valuekind, availabilitybasis and source details
entitiesdataset_idTEXTComposite PK and FK → datasets
entitiesentity_idTEXTComposite PK; station, location or series ID within dataset
entitiesnameTEXTEntity display name
entitieskindTEXTDefault series; e.g. location
entitieslatitudeREALNullable; latitude
entitieslongitudeREALNullable; longitude
entitiesmetadata_jsonTEXTDefault {}; entity metadata
variablesdataset_idTEXTComposite PK and FK → datasets
variablesvariable_idTEXTComposite PK; variable within dataset
variablesnameTEXTVariable display name
variablesunitTEXTDefault empty; measurement unit
variablesroleTEXTCHECK target or covariate; current tasks use target only
variablesvalue_typeTEXTDefault float
variablesaggregationTEXTDefault instantaneous; descriptive aggregation metadata
variablesmetadata_jsonTEXTDefault {}; variable metadata
observationsdataset_idTEXTComposite PK; part of both entity/variable FKs
observationsentity_idTEXTComposite PK; FK with dataset_id → entities
observationsvariable_idTEXTComposite PK; FK with dataset_id → variables
observationsevent_timeTEXTComposite PK; when the event/sample occurred
observationsavailable_timeTEXTWhen this value was first known available; collectors conservatively use ingestion time
observationsingested_atTEXTAcquisition time; cannot precede available_time
observationsvalueREALFinite numeric value, validated by repository
observationsrevision_idTEXTComposite PK; default v1; distinguishes changed values
observationsquality_flagTEXTDefault empty; source quality annotation
observationsraw_hashTEXTSHA-256 key into raw_archives for v2 collections; legacy/custom records may only have a hash or empty value
raw_archivesraw_hashTEXTPK; SHA-256 of the uncompressed canonical response-bundle JSON
raw_archivesdataset_idTEXTFK → datasets; source stream
raw_archivesfetched_atTEXTReceipt time of the last response in this collection
raw_archivesresponses_zlibBLOBZlib-compressed JSON list of URL, receipt timestamp and base64-encoded exact response bytes; private only
modelsmodel_idTEXTPK; stable submitted model/version identity
modelsdisplay_nameTEXTDisplay name
modelsversionTEXTDeclared implementation version; not necessarily a weight hash
modelsendpoint_urlTEXTNullable; server-side endpoint, excluded from public model export
modelscode_urlTEXTNullable; public implementation link
modelsadmitted_atTEXTNullable; first admission cutoff; preserved on metadata update
modelsenabledINTEGERDefault 1; CHECK 0 or 1
forecast_taskstask_idTEXTPK; includes evaluation version, stream, target start and window lengths
forecast_tasksdataset_idTEXTFK → datasets; part of entity/variable FKs
forecast_tasksentity_idTEXTFK with dataset_id → entities
forecast_tasksvariable_idTEXTFK with dataset_id → variables
forecast_taskscutoff_timeTEXTLatest permitted context availability/ingestion time
forecast_taskscontext_startTEXTInclusive context start
forecast_taskscontext_endTEXTInclusive context end, before target_start
forecast_taskstarget_startTEXTFirst scored timestamp; forecast must finish strictly before it
forecast_taskstarget_endTEXTLast scored timestamp, inclusive
forecast_taskshorizonINTEGERPositive number of target steps
forecast_tasksfrequencyTEXTFrozen task frequency
forecast_tasksstatusTEXTpending, forecasted, scored or failed; coarse task summary
forecast_taskscreated_atTEXTTask creation cutoff
forecast_runsrun_idTEXTPK; model execution identifier
forecast_runstask_idTEXTFK → forecasttasks; UNIQUE together with modelid
forecast_runsmodel_idTEXTFK → models; one attempt per model-task pair
forecast_runscreated_atTEXTCompletion time for successful inference; attempt time for failures
forecast_runsprotocol_versionTEXTHistorical column name; v2 stores evaluator version, not HTTP protocol version
forecast_runsstatusTEXTCHECK ok or failed; ok means frozen, not necessarily scored
forecast_runserrorTEXTDefault empty; safe failure reason code (legacy rows have exception class), never credentials or provider response
forecast_inputstask_idTEXTPK and FK → forecast_tasks; one shared canonical input per task
forecast_inputsinput_hashTEXTSHA-256 of input_json
forecast_inputsinput_jsonTEXTExact timestamps, values, opaque series ID, frequency, requested quantiles, full prediction length and discarded gap
forecast_inputscreated_atTEXTTime the input snapshot was first frozen, before inference
predictionsrun_idTEXTComposite PK and FK → forecast_runs
predictionsstepINTEGERComposite PK; zero-based target step, CHECK >= 0
predictionsstatisticTEXTComposite PK; mean or q0.1 … q0.9
predictionsvalueREALFinite forecast value in source units
scoresrun_idTEXTComposite PK and FK → forecast_runs
scoresmetric_nameTEXTComposite PK; MSE, RMSE, MAE, MAPE, CRPS
scoresmetric_versionTEXTComposite PK; isolates metric definitions
scoresvalueREALMetric value; undefined MAPE is omitted, not zero
scoresscored_atTEXTTruth-availability cutoff used for scoring
releasesrelease_idTEXTPK; reserved, not populated by current operator
releasesschema_versionINTEGERReserved release schema version
releasescreated_atTEXTReserved release creation time
releasespublic_revisionTEXTNullable; reserved public commit reference

observations uses the five-column primary key (dataset_id, entity_id, variable_id, event_time, revision_id). Repeated identical revisions do not overwrite first-ingestion timestamps. Two secondary indexes cover series/event time and available_time. Foreign keys are enabled per connection and SQLite uses WAL. New operator collections store original HTTP response bytes privately, including every page of a paginated retrieval. Model-neutral canonical inputs are frozen before inference and reused for scoring normalization and public forecast examples. Adapter-specific wire encoding remains defined by the pinned code revision. Legacy tasks without an explicit snapshot retain causal reconstruction from their original observations; raw bodies that were never saved are not reconstructed or claimed retroactively.

Public releases.csv is a derived join of tasks, runs and scores; it is not an export of the reserved releases table. models.csv exposes admitted model metadata but never endpoint credentials. history.json contains daily aggregate snapshots and pinned public commit IDs. source_revision.json maps each export to its private SQLite and code commits; database_schema_version is separate from the public schema_version. Private database downloads require HF authorization; the browser never receives that token.

Joining the evaluation: participant and maintainer checklist

  1. 1.Host the model behind the example's HTTPS POST /forecast contract. Install .[api] and run uvicorn examples.model_api.app:app --host 0.0.0.0 --port 8000 for local development; add HTTPS at your deployment boundary.
  2. 2.Return one output per input, with finite mean and all nine requested quantile arrays of the requested length. A point model must declare and repeat its point predictions across quantiles, not fabricate uncertainty. Never query future observations inside the model service.
  3. 3.Run livehouse-ts-validate organization/model-version https://host/forecast. This is a small contract check, not a performance or latency certification. Also test the largest dataset horizon plus publication gap you will serve.
  4. 4.Submit the community-model form with model ID, version/weight revision, model card, code URL, endpoint, owner and declared output type. Use a new model ID when weights change; metadata upsert does not version old runs.
  5. 5.A maintainer validates the service and adds it to LIVEHOUSE_MODELS_JSON in the evaluator's repository secrets. The operator records admission on its next run; forecasts are issued only for newly created, eligible tasks. The form is not an automatic registration or approval service.
  6. 6.Wait for targets to arrive. The model first shows awaitingforecast or awaitingscores, then provisional. Official ranks require the evidence thresholds documented above; no pre-admission backfill is performed.

For hosted TSFM services, keep TSFM_API_KEY server-side. A successful provider login or model listing does not establish inference availability: validate the actual model response, all requested quantiles and deadlines before enabling recurring evaluation. Provider quotas and billing are separate from HF storage.

The private Hugging Face Dataset contains the canonical SQLite database, normalized observations, tasks, and frozen forecasts. The public Dataset contains stable CSV, JSON, and metrics-only release exports. The Hugging Face Space is a read-only presentation layer with no token or private-data access. A single evaluator writer updates Saxon0520/LiveHouse-TS-test-private-data first, then publishes derived artifacts to the public Saxon0520/LiveHouse-TS-test Dataset using optimistic commit checks.

GitHub Actions provides an hourly single-writer operator and a manual HF resource/Space deployment workflow. Both use the repository secret HF_TOKEN. Optional external models are stored in LIVEHOUSE_MODELS_JSON. The Static Space itself receives no secret.

Run Verify Hugging Face backup in GitHub Actions to check recovery without changing either Dataset. It restores the private SQLite revision referenced by the public export, checks database integrity and foreign keys, and reproduces the four public CSV tables and leaderboard JSON (except its generation time). The run reports counts only; it does not upload the private database. The public Dataset exposes leaderboard, datasets, models, and releases as separate Dataset Viewer configurations.

Interactive result explorer

The benchmark page follows the subset/explorer approach of TabArena: category buttons select a domain, a dataset selector narrows to an individual stream, and model checkboxes control all analysis panels and the detailed table. The server exports dataset_leaderboards alongside domain aggregates; dataset selection never reuses global scores. Official eligibility thresholds remain unchanged, so a single-dataset view has diagnostics but no official rank.

Twelve interactive charts cover Elo ratings first, the accuracy/coverage frontier, cross-domain and dataset heatmaps, error distributions, daily trends, paired-task scatter plots, per-dataset paired advantages, pairwise win heatmaps, performance profiles, archived Elo/rank/error history and multi-model forecast overlays. Paired win/tie/loss counts appear in the comparison summary rather than a separate chart. Archive is the final page section. Domain cells open the dataset explorer; paired scatter points narrow the current dataset. Plotly 4.1.1 is loaded from its official CDN; charts support hover, zoom, legend toggles and PNG export. Tables and data links still work if the chart library is unavailable. Only nonempty selections are plotted. Elo uses an ordered dot plot with a visible numeric axis and aligned ratings; the gray row guides are not confidence intervals. Its displayed axis spans at least 100 Elo points so tiny differences are not stretched across the chart.

analysis.json contains the same already-public scored records as releases.csv plus schema, metric/evaluation versions and export timestamp. It contains no private contexts, future targets, endpoint secrets or error bodies. Its timestamp must match the selected leaderboard; an older archive without this file shows an unavailable diagnostic state rather than borrowing current releases.

Pair comparisons use only matching dataset/release/metric-version keys. The paired advantage is 100*(error_B-error_A)/(error_A+error_B) (zero when both are zero); positive means A is better. The paired scatter's logarithmic mode uses log10(1+error) and raw-value tick/hover labels, retaining exact zero errors. Performance profiles use the intersection of tasks for the selected models with scores, averaged within dataset then equally across datasets. Profiles show the share within a factor of the best error; if the best is zero, only zero-error models count at finite factors. These selected-cohort views do not replace or recompute the official rank in the leaderboard.

forecast_examples.json publishes at most one resolved task for each of 20 datasets, with at most 120 context and 120 target points per example. Only current-version, scored forecasts are included, with the mean and 10/50/90% quantiles. Models must share the exact task and the same observed truth values to appear together; different truth revisions are not overlaid. Raw response archives, unscored predictions and future labels remain private. The gallery is checked against the selected publication timestamp and metric version; older snapshots without a gallery retain their original single-example view.

Error distributions weight each release equally and are exploratory; daily trends average within dataset/day before averaging equally across available datasets. Baseline normalization matches dataset, release and metric version; missing/zero-error references yield missing values, never zero scores. Pairwise heatmaps compare MSE and CRPS outcomes separately on shared releases and average equally by dataset. They show shared counts but do not impose official eligibility gates, so they must not be read as official ranks or significance tests. Historical rank lines are explicitly global. No latency, cost or confidence-interval chart is fabricated where that evidence is not recorded.

The Domain selector recomputes scores, Elo, coverage and failures using only that domain's tasks, then updates the table, comparison charts and dataset list. It does not merely filter globally scored rows. The five-dataset official-rank threshold is unchanged, so small domain slices can stay provisional. Old snapshots without domain exports offer only All domains. Historical rank lines remain global; forecast examples follow dataset and model selections. Display sorting does not change stored official ranks. The Date selector and Latest/Archived status identify the selected snapshot; publication timestamps remain in the downloadable exports rather than a separate result-date card.

Custom website domain (deferred)

Domain binding requires the owner's chosen hostname and DNS access. HF's native custom-domain feature currently requires PRO or Team/Enterprise and a public or protected Space. Register the hostname in Space settings, add the CNAME specified by HF (documented target hf.space), and verify the status becomes ready and HTTPS works before changing public links. Do not buy a domain or upgrade a plan without owner approval. See HF custom domain documentation. The existing HF address remains the working entry point until binding is done.

Historical ranks and Backtesting Archive

The Space includes Elo and composite-score comparison charts, daily rank history, and a version selector that loads the table, datasets, and forecast example from the same pinned public revision. Elo is displayed separately; current primary rank uses eligible pairwise wins. Archived v1 results retain their former composite-score ordering and must not be compared as one ranking regime with evaluation v2.

The existing hourly operator adds one snapshot per UTC day to history.json, using the previously published HF commit. The first captured version for a day is retained unchanged. Each archive entry links to that commit's leaderboard, release scores, forecast example, and private-source revision reference; it does not publish the private database. These are archived prospective results, not retrospective inference or imported paper scores. Archiving starts when this feature is deployed; outages leave gaps instead of fabricated history. The chart shows the latest 30 archived days; older versions remain selectable and downloadable. HF commit history must be retained for archive links to work.

<p align="right"><a href="#livehouse-ts">Back to top ↑</a></p>