Team Ai
Datasetpublic

latency-sensitive-bench/benchmark-datasets

LAGEN benchmark-datasets Training data, episode-level evaluations and paper experiment manifests. Experiment Entry Sim2Real calibration 30-case held-out calibration Visual history Visual history Latency in prompt Latency in prompt Latency transfer Latency transfer Task transfer Task transfer VLA fine-tuning scope VLA fine-tuning scope Mean vs. profile training Mean vs. profile training Observation stride Observation stride Context window Context window… See the full description on the dataset page: https://huggingface.co/datasets/latency-sensitive-bench/benchmark-datasets.

sourceHugging Faceotherupdated 1h agoView on Hugging Face
0likes2.7kdownloads
Dataset Card

LAGEN benchmark-datasets

Training data, episode-level evaluations and paper experiment manifests.

ExperimentEntry
Sim2Real calibration30-case held-out calibration
Visual historyVisual history
Latency in promptLatency in prompt
Latency transferLatency transfer
Task transferTask transfer
VLA fine-tuning scopeVLA fine-tuning scope
Mean vs. profile trainingMean vs. profile training
Observation strideObservation stride
Context windowContext window

Benchmark release inventory

Artifact retention

HAIC and Extreme Parkour releases are retired. Model bundles retain the published evaluation checkpoint, or the latest checkpoint when no evaluation selection exists. Optimizer and trainer recovery state are not release assets. Identical dataset copies use the canonical task paths. Original configurations and experiment evidence remain source records.

Figure 2 latency degradation

Latency-degradation catalogue preserves all 15 curves, exact evaluation records, fixed checkpoint sources and provenance gaps. Six MIKASA checkpoints are reused, and all six confirmed game checkpoints are published in benchmark-models with unchanged SHA256 and size.

latency-sensitive-bench/benchmark-datasets · Team Ai