Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Post-training-Data-Flywheel /gorilla-openfunctions-v1text10K<n<100K0 likes2k downloads2y agoHugging Face02OpenLLM-France /Luciole-PostTraining-Dataset-1.1 Table of Contents Dataset Description Curation Rationale Bias, Risks, and Limitations Data Subsets Sample Metadata Downloading the Data Available Configurations Loading Examples Accessing Data Through the Directory Hierarchy Details on Data Sources Citation Acknowledgements Contact Dataset Description The Luciole-PostTraining-Dataset-1.1 is a curated collection of open, instruction-style text data designed for language model post training. It includes a mixture of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-PostTraining-Dataset-1.1.text1M<n<10M6 likes1.2k downloads9d agoHugging Face03Posttraining-RFM-RSS2026 /Challenge-phase1-dataset Post-training for Robotics Foundation Models — Challenge Phase 1 Dataset This is the public Phase 1 dataset for the RSS 2026 Workshop & Challenge on Post-training for Robotics Foundation Models. Workshop website: https://posttraining-for-robotics.github.io/ The dataset contains real-robot bimanual manipulation trajectories on three benchmark tasks, collected on a bimanual YAM follower teleoperated by a GELLO leader arm. Every frame is timestamp-aligned across joint state, action… See the full description on the dataset page: https://huggingface.co/datasets/Posttraining-RFM-RSS2026/Challenge-phase1-dataset.video1K<n<10K2 likes693 downloads4mo agoHugging Face04Post-training-Data-Flywheel /Salesforce-xlam-function-calling-60ktext10K<n<100K0 likes598 downloads2y agoHugging Face05Martingkc /LLaVa-CC3M-PostTraining-clip-vit-base-patch16text100K<n<1M0 likes236 downloads6mo agoHugging Face06Posttraining-RFM-RSS2026 /Challenge-phase2-rollouts-dataset Challenge phase2 rollouts dataset Policy rollout episodes collected for the RSS2026 Challenge (phase 2), in LeRobot v2.1 format. Sub-datasets: 23 (one folder each) Total episodes: 808 | Total frames: 4,289,763 Format: LeRobot codebase_version v2.1 Tasks: insert-mouse-battery, seal-water-bottle-cap, tower-of-hanoi-game Policies: generalist / specialist (per team) Each frame carries observation.commander_state distinguishing human teleoperation from autopilot policy inference… See the full description on the dataset page: https://huggingface.co/datasets/Posttraining-RFM-RSS2026/Challenge-phase2-rollouts-dataset.video1K<n<10K0 likes200 downloads3mo agoHugging Face07benchpress /posttraining-eval-results0 likes191 downloads5mo agoHugging Face08Post-training-Data-Flywheel /AutoIF-instruct-61k-with-funcstext10K<n<100K8 likes189 downloads2y agoHugging Face09Post-training-Data-Flywheel /AutoIF-instruct-61ktext10K<n<100K18 likes162 downloads2y agoHugging Face10AIGym /post-training-v1text100K<n<1M0 likes119 downloads1y agoHugging Face111-800-SHARED-TASKS /Organic-Chemistry-VLM-PostTraining Dataset Card for "Chemistry_text_to_image" More Information needed image100K<n<1M6 likes75 downloads2y agoHugging Face12lilywchen /lucky-initialization-posttraining-100m-v10 likes73 downloads2mo agoHugging Face13Post-training-Data-Flywheel /OpenOrcatext1M<n<10M0 likes71 downloads2y agoHugging Face14Post-training-Data-Flywheel /NousResearch-hermes-function-calling-v1text1K<n<10K0 likes64 downloads2y agoHugging Face15Post-training-Data-Flywheel /stingning-ultrachattext1M<n<10M0 likes61 downloads2y agoHugging Face16lilywchen /lucky-initialization-posttraining-100m-v3tabularn<1K0 likes58 downloads2mo agoHugging Face17ToluClassics /post-training-trackio-datasettabularn<1K0 likes57 downloads2mo agoHugging Face18HuggingFaceTB /post-training-benchmarks-viewertabularn<1K3 likes47 downloads1y agoHugging Face19Post-training-Data-Flywheel /function-calling-1.00 likes44 downloads2y agoHugging Face20Post-training-Data-Flywheel /gorilla-apibenchtext10K<n<100K0 likes43 downloads2y agoHugging Face21laa1991 /post-training-ruler-readings Readings from a preference-pair post-training study — mostly what did not move One sentence: a table of every measurement we took while trying to teach a model to write structural judgements from a corpus of real incident write-ups — including the eight cells where the knob turned out to be flat, with the interval attached. Most published evaluation artifacts show what worked. This one is mostly the opposite: each row is a knob we turned, the reading we got, and whether it… See the full description on the dataset page: https://huggingface.co/datasets/laa1991/post-training-ruler-readings.texttext-generationn<1K0 likes43 downloads3d agoHugging Face22TokenBender /glm47-pie-cpp-posttraining-data GLM-4.7-Flash PIE C++ Post-Training Data The exact prepared dataset used for the GLM-4.7-Flash C++ performance post-training runs. Splits File Rows Purpose sft/train.jsonl 7,864 Supervised fine-tuning grpo/train.jsonl 7,887 GRPO prompt and reward evaluation eval/validation.jsonl 1,259 Full held-out evaluation eval/validation_mini126.jsonl 126 Fast evaluation eval/validation_mini4.jsonl 4 Smoke evaluation tasks.tar.gz 9,146 task JSONs Reward… See the full description on the dataset page: https://huggingface.co/datasets/TokenBender/glm47-pie-cpp-posttraining-data.texttext-generation10K<n<100K0 likes42 downloads3mo agoHugging Face23Post-training-Data-Flywheel /glaiveai-glaive-function-calling-v2text10K<n<100K1 likes41 downloads2y agoHugging Face24hang010412 /video-commons-posttraining-manifest Video Commons Post-training Manifest Scene-level train, validation, and sealed-holdout metadata for a controlled video adaptation experiment. The repository publishes source URLs, per-item licenses, media hashes, actions, and dimensions; it does not redistribute the source videos. Each upstream asset remains governed by its own license recorded in data/manifest.parquet. There is intentionally no blanket dataset license that overrides those upstream terms. Use the downloader in… See the full description on the dataset page: https://huggingface.co/datasets/hang010412/video-commons-posttraining-manifest.textn<1K0 likes41 downloads12d agoHugging Face25Tuyentd /Post-Training_Answer_Style_Alignmentaudion<1K0 likes38 downloads1y agoHugging Face26glouriousgautam /lilm1-230m-posttraining LiLM1-230M post-training data This dataset contains the selected post-training data for LiLM1-230M. Method The records combine general assistant text with structured tool-use examples. The configurations preserve the binding stage, the ratio study, and the selected 4:8 continuation. Configurations Configuration Content binding-repair Tool binding data ratio-study Three training splits used for ratio selection ratio-evaluation Shared… See the full description on the dataset page: https://huggingface.co/datasets/glouriousgautam/lilm1-230m-posttraining.tabulartext-generation10K<n<100K0 likes34 downloads1mo agoHugging Face27Post-training-Data-Flywheel /allenai-WildChat-1M-gpt4-entext100K<n<1M0 likes29 downloads2y agoHugging Face28Post-training-Data-Flywheel /databricks-dolly-15ktext10K<n<100K0 likes28 downloads2y agoHugging Face29Post-training-Data-Flywheel /gpt4-self-instructtext10K<n<100K0 likes27 downloads2y agoHugging Face30Post-training-Data-Flywheel /meta-math-MetaMathQAtext100K<n<1M0 likes23 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.