datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
inference-datagroot_n1.7_inference_on_diff_data
GROOT Inference Analysis Log
Evaluation records for a GR00T policy trained on the task
"pick octopus and place inside brown basket", run on a Unitree G1 at 20 Hz
with the ego_view stereo camera.
Six training checkpoints (50, 100, 150, 200, 250, 300 demonstration episodes) were each
evaluated on 50 inference episodes. Every episode is recorded here with video,
per-tick state/action logs and run metadata.
Success rate
Checkpoint (training episodes)
Success… See the full description on the dataset page: https://huggingface.co/datasets/rahul-ai-01/groot_n1.7_inference_on_diff_data.inference-dataset
VLM-Gym Inference Dataset
This dataset contains pre-defined test episodes and initial states for evaluating Vision-Language Models (VLMs) on the VLM-Gym benchmark.
Dataset Structure
inference-dataset/
├── test_set_easy/ # Easy difficulty test episodes (JSONL)
├── test_set_hard/ # Hard difficulty test episodes (JSONL)
├── initial_states_easy/ # Initial environment states for easy episodes (JSON)
├── initial_states_hard/ # Initial environment… See the full description on the dataset page: https://huggingface.co/datasets/VisGym/inference-dataset.Inference_Step1XEditdata_inference_pythia_6_9bp2-etf-active-inference-resultsinference-scratchcode-world-model-inference-examples-40
Inference examples
This directory contains 40 numbered, independent inference examples.
Every example uses only its public number; source case names and internal paths
are intentionally omitted.
Each numbered directory contains:
first_frame.png: exact 1536x864 generated RGB first frame used by inference.
prompts/*.txt: the exact rolling long-inference prompts used for the result.
condition/*.npz: ordered lossless condition chunks.
metadata.json: frame count, FPS, prompt windows… See the full description on the dataset page: https://huggingface.co/datasets/NTU-yiwen/code-world-model-inference-examples-40.hf-inference-providers-datamodel-inference-activationstamil-inference-resultspeculators-ci-datasets
speculator-tutorial
Raw vs. on-policy regenerated conversation data for training speculative-decoding
drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside
so you can see exactly what regeneration changes and why it matters.
Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B.
Why regenerate at all?
A speculative-decoding drafter is trained to predict what the verifier would say next.
If you train it… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/speculators-ci-datasets.vllm-inference-benchmarksqwen3.8-max-glm5.2-kimi-k3-distillation
Multi-Teacher Distillation Dataset (57,937 traces)
A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains.
Teachers
Teacher
Provider
Traces
Qwen3.8-Max-Preview
Alibaba Cloud Model Studio
48,283
GLM-5.2
Z.AI Coding Plan
5,307
Kimi Code K3
Moonshot AI (Kimi)
4,347… See the full description on the dataset page: https://huggingface.co/datasets/inferenceport-ai/qwen3.8-max-glm5.2-kimi-k3-distillation.cci-multilingualrules-model-inference-activationssovereign-shadow-inference-bench
Sovereign Shadow Inference Bench
A public, versioned evidence surface for independent Hugging Face shadow inference beside Sovereign's primary OpenRouter/Revolver route.
What this dataset proves
The seed record in data/shadow_receipts.jsonl was produced by one real Hugging Face Inference Providers request. It records provider/model identity, request bounds, latency, hashes, literal-match outcome, source revision, and an immutable receipt hash.
What it… See the full description on the dataset page: https://huggingface.co/datasets/Thorsu/sovereign-shadow-inference-bench.Qwen-SFT-Inference-Outputseval_pi0_inference_only_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 70,
"total_frames": 40583,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:70"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tersooawai/eval_pi0_inference_only_dataset.amortized-inference-training
Amortized Inference Training Trajectories
Generated trajectories are organized as source_dataset/geometry_id/material_id/trajectory.
Each geometry has index.jsonl; the global decisions live under index/.
Only records with status == "accepted" and trainable == true passed the recorded legacy
topology gate. rejected arrays are intentionally absent. Historical trajectories for which the
topology gate was never run are retained as unreviewed and must not be used for training until… See the full description on the dataset page: https://huggingface.co/datasets/BDXXN/amortized-inference-training.figma-Playground-Inference-for-PRO-s-WebsiteRoboTryOn_inferenceinference-benchmarkerInference_OfflineRLmodel-inference-responsesVILA-inference-demosbeatrix-captured-interactive-inferences
Beatrix captured interactive inferences
Byte-by-byte internals of real inference runs on mini-beatrix-2.5s,
the 237.1M full-splat byte model
(model). Nothing
here is simulated or sampled from a proxy: each capture is one greedy
generation, re-run through a single instrumented forward pass that walks
the blocks by hand and records what every layer did at every byte.
Each prompt is captured twice over the same byte sequence — once on
the bare core and once with the library's top… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/beatrix-captured-interactive-inferences.rollout_colorlogo_inference_merged_v1
SO-101 Colorlogo Inference Rollouts (Merged v1)
This repository contains raw inference rollouts collected with an SO-101 follower arm.
Dataset summary
Format: LeRobot Dataset v3.0
Episodes: 11
Frames: 41,675
Frame rate: 30 FPS
Cameras: observation.images.front and observation.images.handeye
Tasks: red block (8 episodes), green block (1 episode), blue block (2 episodes)
Robot type: so_follower
The original local datasets were merged without changing observations… See the full description on the dataset page: https://huggingface.co/datasets/ted88168/rollout_colorlogo_inference_merged_v1.cci-multilingualrules-model-inference-responsesown-your-inference-data
Own your inference — benchmark data
Raw benchmark artifacts behind the "own your inference" talk: what a self-hosted GLM-5.2 endpoint can and cannot do, measured on three deployments.
Companion code repository: Gladiator07/own-your-inference. It holds the runners that produced this data and a script that rebuilds every table from these artifacts on a laptop.
The three deployments
Experiment
Hardware
Weights
KV cache
h200-fp8-kv
8x NVIDIA H200
GLM-5.2… See the full description on the dataset page: https://huggingface.co/datasets/Gladiator/own-your-inference-data.so101-real-inference-c2-h1-20261004
SO-101 real inference: C2 and H1
31 attended real inference episodes from October 4, 2026, evaluated with the same fine-tuned SmolVLA checkpoint 020000. This is not the pretrained-reference policy and not RL evaluation.
Condition
Runs
Success
Failure
Unresolved cutoff
iPhone clips
C2: blue, 2 RPM, normal direction
15
2
1
12
14
H1: blue, 3 RPM, normal direction
16
0
1
15
16
C2 episodes 8/9 succeeded; episode5 failed and has no iPhone clip. Its robot cameras… See the full description on the dataset page: https://huggingface.co/datasets/ewinchell07/so101-real-inference-c2-h1-20261004.
