datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gorilla-openfunctions-v1Luciole-PostTraining-Dataset-1.1
Table of Contents
Dataset Description
Curation Rationale
Bias, Risks, and Limitations
Data Subsets
Sample Metadata
Downloading the Data
Available Configurations
Loading Examples
Accessing Data Through the Directory Hierarchy
Details on Data Sources
Citation
Acknowledgements
Contact
Dataset Description
The Luciole-PostTraining-Dataset-1.1 is a curated collection of open, instruction-style text data designed for language model post training. It includes a mixture of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-PostTraining-Dataset-1.1.Challenge-phase1-dataset
Post-training for Robotics Foundation Models — Challenge Phase 1 Dataset
This is the public Phase 1 dataset for the RSS 2026 Workshop & Challenge on Post-training for Robotics Foundation Models.
Workshop website: https://posttraining-for-robotics.github.io/
The dataset contains real-robot bimanual manipulation trajectories on three benchmark tasks, collected on a bimanual YAM follower teleoperated by a GELLO leader arm. Every frame is timestamp-aligned across joint state, action… See the full description on the dataset page: https://huggingface.co/datasets/Posttraining-RFM-RSS2026/Challenge-phase1-dataset.Salesforce-xlam-function-calling-60kLLaVa-CC3M-PostTraining-clip-vit-base-patch16Challenge-phase2-rollouts-dataset
Challenge phase2 rollouts dataset
Policy rollout episodes collected for the RSS2026 Challenge (phase 2), in LeRobot v2.1 format.
Sub-datasets: 23 (one folder each)
Total episodes: 808 | Total frames: 4,289,763
Format: LeRobot codebase_version v2.1
Tasks: insert-mouse-battery, seal-water-bottle-cap, tower-of-hanoi-game
Policies: generalist / specialist (per team)
Each frame carries observation.commander_state distinguishing human teleoperation from autopilot policy inference… See the full description on the dataset page: https://huggingface.co/datasets/Posttraining-RFM-RSS2026/Challenge-phase2-rollouts-dataset.posttraining-eval-resultsAutoIF-instruct-61k-with-funcsAutoIF-instruct-61kpost-training-v1Organic-Chemistry-VLM-PostTraining
Dataset Card for "Chemistry_text_to_image"
More Information needed
lucky-initialization-posttraining-100m-v1OpenOrcaNousResearch-hermes-function-calling-v1stingning-ultrachatlucky-initialization-posttraining-100m-v3post-training-trackio-datasetpost-training-benchmarks-viewerfunction-calling-1.0gorilla-apibenchpost-training-ruler-readings
Readings from a preference-pair post-training study — mostly what did not move
One sentence: a table of every measurement we took while trying to teach a model to write structural judgements from a corpus of real incident write-ups — including the eight cells where the knob turned out to be flat, with the interval attached.
Most published evaluation artifacts show what worked. This one is mostly the opposite: each row is a knob we turned, the reading we got, and whether it… See the full description on the dataset page: https://huggingface.co/datasets/laa1991/post-training-ruler-readings.glm47-pie-cpp-posttraining-data
GLM-4.7-Flash PIE C++ Post-Training Data
The exact prepared dataset used for the GLM-4.7-Flash C++ performance
post-training runs.
Splits
File
Rows
Purpose
sft/train.jsonl
7,864
Supervised fine-tuning
grpo/train.jsonl
7,887
GRPO prompt and reward evaluation
eval/validation.jsonl
1,259
Full held-out evaluation
eval/validation_mini126.jsonl
126
Fast evaluation
eval/validation_mini4.jsonl
4
Smoke evaluation
tasks.tar.gz
9,146 task JSONs
Reward… See the full description on the dataset page: https://huggingface.co/datasets/TokenBender/glm47-pie-cpp-posttraining-data.glaiveai-glaive-function-calling-v2video-commons-posttraining-manifest
Video Commons Post-training Manifest
Scene-level train, validation, and sealed-holdout metadata for a controlled
video adaptation experiment. The repository publishes source URLs, per-item
licenses, media hashes, actions, and dimensions; it does not redistribute
the source videos.
Each upstream asset remains governed by its own license recorded in
data/manifest.parquet. There is intentionally no blanket dataset license
that overrides those upstream terms.
Use the downloader in… See the full description on the dataset page: https://huggingface.co/datasets/hang010412/video-commons-posttraining-manifest.Post-Training_Answer_Style_Alignmentlilm1-230m-posttraining
LiLM1-230M post-training data
This dataset contains the selected post-training data for LiLM1-230M.
Method
The records combine general assistant text with structured tool-use examples.
The configurations preserve the binding stage, the ratio study, and the
selected 4:8 continuation.
Configurations
Configuration
Content
binding-repair
Tool binding data
ratio-study
Three training splits used for ratio selection
ratio-evaluation
Shared… See the full description on the dataset page: https://huggingface.co/datasets/glouriousgautam/lilm1-230m-posttraining.allenai-WildChat-1M-gpt4-endatabricks-dolly-15kgpt4-self-instructmeta-math-MetaMathQA
