Team Ai
Datasetpublic

nosuke113/libero-plus-evaluation

LIBERO-plus Evaluation Dataset (180 held-out tasks) This dataset contains the evaluation tasks used for validating the pi0.5 LIBERO-plus LoRA model on the LIBERO-plus benchmark. It consists of a meticulously designed set of 180 tasks, distributed along 4 perturbation axes to rigorously evaluate the robustness and generalization capabilities of Vision-Language-Action (VLA) models: Base — Standard BDDL scenes. Camera (0, 0, 100, 0, 0), fixed object poses, no noise. View — Camera… See the full description on the dataset page: https://huggingface.co/datasets/nosuke113/libero-plus-evaluation.

sourceHugging Faceapache-2.0updated 16d agoView on Hugging Face
0likes63downloads
Dataset Card

LIBERO-plus Evaluation Dataset (180 held-out tasks)

This dataset contains the evaluation tasks used for validating the pi0.5 LIBERO-plus LoRA model on the LIBERO-plus benchmark.

It consists of a meticulously designed set of 180 tasks, distributed along 4 perturbation axes to rigorously evaluate the robustness and generalization capabilities of Vision-Language-Action (VLA) models:

  1. 1.Base — Standard BDDL scenes. Camera (0, 0, 100, 0, 0), fixed object poses, no noise.
  2. 2.View — Camera orbit angle (0-356 deg), zoom distance (100-191), and gaze tilt are varied.
  3. 3.Noise — Image corruptions: Gaussian noise, motion blur, fog, glass distortion (level 8-48).
  4. 4.Moved — Object initial positions shifted by several cm from default.

Data Format

The dataset provides a val_tasks.csv file that lists the specific suite and task names to be evaluated using the libero benchmark code.

ColumnDescription
partEvaluation section (general or axis)
axisPerturbation axis category
kindSpecific perturbation kind
suiteLIBERO suite name
task_idFull task string
base_taskOriginal base task name
paramPerturbation parameters

Usage

You can load this list of tasks and feed it into the standard LIBERO evaluation pipeline to benchmark your model.