inference
inference-datagroot_n1.7_inference_on_diff_data
GROOT Inference Analysis Log
Evaluation records for a GR00T policy trained on the task
"pick octopus and place inside brown basket", run on a Unitree G1 at 20 Hz
with the ego_view stereo camera.
Six training checkpoints (50, 100, 150, 200, 250, 300 demonstration episodes) were each
evaluated on 50 inference episodes. Every episode is recorded here with video,
per-tick state/action logs and run metadata.
Success rate
Checkpoint (training episodes)
Success… See the full description on the dataset page: https://huggingface.co/datasets/rahul-ai-01/groot_n1.7_inference_on_diff_data.inference-dataset
VLM-Gym Inference Dataset
This dataset contains pre-defined test episodes and initial states for evaluating Vision-Language Models (VLMs) on the VLM-Gym benchmark.
Dataset Structure
inference-dataset/
├── test_set_easy/ # Easy difficulty test episodes (JSONL)
├── test_set_hard/ # Hard difficulty test episodes (JSONL)
├── initial_states_easy/ # Initial environment states for easy episodes (JSON)
├── initial_states_hard/ # Initial environment… See the full description on the dataset page: https://huggingface.co/datasets/VisGym/inference-dataset.Inference_Step1XEditdata_inference_pythia_6_9bp2-etf-active-inference-results
