nosuke113/libero-plus-evaluation
LIBERO-plus Evaluation Dataset (180 held-out tasks) This dataset contains the evaluation tasks used for validating the pi0.5 LIBERO-plus LoRA model on the LIBERO-plus benchmark. It consists of a meticulously designed set of 180 tasks, distributed along 4 perturbation axes to rigorously evaluate the robustness and generalization capabilities of Vision-Language-Action (VLA) models: Base — Standard BDDL scenes. Camera (0, 0, 100, 0, 0), fixed object poses, no noise. View — Camera… See the full description on the dataset page: https://huggingface.co/datasets/nosuke113/libero-plus-evaluation.
LIBERO-plus Evaluation Dataset (180 held-out tasks)
This dataset contains the evaluation tasks used for validating the pi0.5 LIBERO-plus LoRA model on the LIBERO-plus benchmark.
It consists of a meticulously designed set of 180 tasks, distributed along 4 perturbation axes to rigorously evaluate the robustness and generalization capabilities of Vision-Language-Action (VLA) models:
- Base — Standard BDDL scenes. Camera (0, 0, 100, 0, 0), fixed object poses, no noise.
- View — Camera orbit angle (0-356 deg), zoom distance (100-191), and gaze tilt are varied.
- Noise — Image corruptions: Gaussian noise, motion blur, fog, glass distortion (level 8-48).
- Moved — Object initial positions shifted by several cm from default.
Data Format
The dataset provides a val_tasks.csv file that lists the specific suite and task names to be evaluated using the libero benchmark code.
Usage
You can load this list of tasks and feed it into the standard LIBERO evaluation pipeline to benchmark your model.
