datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
doc-formats-jsonl-1
[doc] formats - jsonl - 1
This dataset contains one jsonl file at the root.
sounio-code-examples
Sounio Curated Code Examples
Curated compile-clean .sio examples for training and evaluating code models on
Sounio, a self-hosted systems and scientific programming language for epistemic
computing, uncertainty propagation, and algebraic effects.
This directory is the Cx-1 expansion lane for
chiuratto-AIgourakis/sounio-code-examples.
Current batch
Examples: 5,000
Metadata files: 5,000
Compiler gate: bin/souc check pass rate 5,000/5,000
Utility layer: 5,000… See the full description on the dataset page: https://huggingface.co/datasets/chiuratto-AIgourakis/sounio-code-examples.lisbet-exampleschat_formatted_examplesdoc-formats-csv-1
[doc] formats - csv - 1
This dataset contains one csv file at the root:
data.csv
kind,sound
dog,woof
cat,meow
pokemon,pika
human,hello
The YAML section of the README does not contain anything related to loading the data (only the size category metadata):
---
size_categories:
- n<1K
---
doc-formats-parquet-1code-world-model-inference-examples-40
Inference examples
This directory contains 40 numbered, independent inference examples.
Every example uses only its public number; source case names and internal paths
are intentionally omitted.
Each numbered directory contains:
first_frame.png: exact 1536x864 generated RGB first frame used by inference.
prompts/*.txt: the exact rolling long-inference prompts used for the result.
condition/*.npz: ordered lossless condition chunks.
metadata.json: frame count, FPS, prompt windows… See the full description on the dataset page: https://huggingface.co/datasets/NTU-yiwen/code-world-model-inference-examples-40.example_datasetsThis is a collection of example datasets from various sources that the datasets module in pgmpy supports. All credits
for creating the datasets go to the original authors. The following datasets are included (along with their LICENCE).
The licences are included in the respective dataset folders as well.
example-causal-datasets: CC0 1.0 Universal. Last synced on
2026-02-05.
nslm: CC0 1.0 Universal. Last downloaded on 2026-03-04.
Tuebingen-pair-wise-dataset: Last downloaded on 2026-03-02.… See the full description on the dataset page: https://huggingface.co/datasets/pgmpy/example_datasets.example_short_storiesEmbodiedSWE-Task-Examples
EmbodiedSWE Simulation Task Examples
打开任务 MP4 浏览页
28 个官网场景 · 26 段官网 MP4 · 10 个官方数据来源各 3 条完整 episode,共 30 条多相机示教。shoe_knot 与 dice 暂无官方展示视频,保留任务目录和评分标准。
序号规则: 本页 C01–C28 是按官网目录顺序生成的辅助编号,官方任务以 suite.scene 标识。LeRobot task_index 和 episode_index 都只在各自 source_repo 内有效,绝不能把各库的 task_index=0 当成同一任务。
相机: 按官方实际提供的 front/front_wide/high/side/corner/wrist 展示,单臂 wrist 不伪装成 left/right。官网演示只展示其原有视角。
评分: 来自当前公开代码的 BaseGrader 与具体 RUBRIC,progress 是归一化权重和。live=当前值;once=历史最好;final=交付末端才评估。历史数据生成使用的… See the full description on the dataset page: https://huggingface.co/datasets/Travor278/EmbodiedSWE-Task-Examples.unsplash-examplesdocvqa_1200_examplesrepo-to-space-example-videos
Gradio Space Example Inputs — Videos
A small, curated, freely-licensed pool of videos used as gr.Examples for
Gradio Spaces that wrap video-input generation models (image-to-video,
video-to-video, motion controls, etc.). Sister dataset for images:
linoyts/repo-to-space-example-inputs.
When a Space takes video input, the agent building the Space picks 2–3 clips
whose caption + categories match the model's task, downloads them via
hf_hub_download, runs any model-specific… See the full description on the dataset page: https://huggingface.co/datasets/linoyts/repo-to-space-example-videos.multitask_german_examples_32kmath_examplesdoc-audio-6
[doc] audio dataset 6
This dataset contains 4 audio files in the /train directory, with a CSV metadata file providing another data column.
rli-example-deliverables
rli-example-deliverables
Example AI deliverables on the RLI public set
Dataset Structure
This dataset contains project folders organized by task ID (public_001 through public_010).
Each project folder contains:
human_deliverable/ - Reference outputs created by human experts
project/ - Project specifications and inputs
brief.md - Task description and requirements
inputs/ - Input files provided for the task
Usage
from huggingface_hub import… See the full description on the dataset page: https://huggingface.co/datasets/cais/rli-example-deliverables.DynamicVLA-Task-Examples
DynamicVLA Simulation Task Examples
打开 MP4 任务浏览页
DOM has 3 task families, not 3 task_index values: 140,932 structured instruction IDs and 207,306 episodes. This subset selects 6 demonstrations per family (18 total), preserving actual task_index/episode_index and structured instructions. All 3 published cameras are single-arm cameras: opposite, wrist, side, not top/left/right arms.
所有示例是完整 H.264 MP4。单臂相机保留官方名称:opst_cam / wrist_cam / side_cam。每条有同步合并视频、各路原视频与官方动作/状态 Parquet。… See the full description on the dataset page: https://huggingface.co/datasets/Travor278/DynamicVLA-Task-Examples.multimodal-example
Multimodal Example Dataset
Small example dataset for testing multimodal (vision-language) fine-tuning with ms-swift.
Structure
├── train.jsonl # 10 training samples
├── test.jsonl # 2 validation samples
├── images/ # All referenced images (400x300 JPEG)
│ ├── dog_portrait.jpg
│ ├── forest_river.jpg
│ ├── laptop_desk.jpg
│ ├── mountain_lake.jpg
│ ├── ocean_rocks.jpg
│ ├── coffee_cup.jpg
│ ├── bookshelf.jpg
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/f13rnd/multimodal-example.kabr-worked-examples
Dataset Card for KABR Worked Examples
This dataset is comprised of manually annotated bounding box detections, mini-scenes, behavior annotations, and associated telemetry
for three drone video sessions that were used for kabr-tools case studies. Drone video was collected at Mpala Research Centre in January 2023; please see the full video dataset for more information on original video context.
Dataset Details
Annotations were created to evaluate the kabr-tools… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/kabr-worked-examples.kiwi_example_traj
KIWI Example Trajectories
Processed 6-DoF trajectories from bimanual recordings: left and right hand-mounted camera poses and a head-mounted marker pose. Tracks within an episode share a metric coordinate frame and clock.
Trajectories for scene_001 / episode_001 and scene_002 / episode_001 are stored as Parquet in the example split. Camera recordings with audio removed are available under raw/ for scene_001 / episode_001 and scene_002 / episode_001.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/azur0licht/kiwi_example_traj.ReflexBench-Task-Examples
ReflexBench Simulation Task Examples
打开 MP4 任务浏览页
Six official task_index values 0–5. Single Franka arm; only fixed_cam and wrist_cam exist. Do not relabel the wrist as left/right. Success rules describe evaluation, not an unprovided per-demo score.
所有示例是完整 H.264 MP4。单臂相机保留官方名称:fixed_cam / wrist_cam。每条有同步合并视频、各路原视频与官方动作/状态 Parquet。
评估标准根据官方代码和文档整理;发布数据未保存逐条示教 success 标签,故 actual_episode_success 为 null。本文不编造 0–100 得分。
官方 task_index
Task / instruction scenario
分类
直接浏览
0… See the full description on the dataset page: https://huggingface.co/datasets/Travor278/ReflexBench-Task-Examples.Taiwanese-Minnan-Example-Sentences
Taiwanese Minnan Example Sentences
The dataset consists of a collection of example sentences designed to aid in recognizing Taiwanese Minnan (Taiwanese Hokkien) for automatic speech recognition (ASR) tasks. This dataset is sourced from the Ministry of Education in Taiwan and aims to provide valuable linguistic resources for researchers and developers working on speech recognition systems.
Dataset Features
Source: Ministry of Education, Taiwan (Sutian Resource Center)
Text:… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Example-Sentences.doc-image-6
[doc] image dataset 6
This dataset contains 4 jpeg files in the train/images/ subdirectory, along with a train/metadata.csv file that provides the data for other columns. The metadata file contains relative paths to the images.
z-image-examples
Z-Image Turbo Portrait Dataset
This dataset contains 126 portrait prompts and their corresponding image outputs, demonstrating the capabilities of the Z-Image Turbo text-to-image model.
Model Information
Model Name: Z-Image Turbo
Hugging Face Repository: Tongyi-MAI/Z-Image-Turbo
Dataset Contents
prompts.jsonl: A JSONL file containing the 126 text prompts used for generation. Each entry includes a unique ID and the prompt text.
outputs/: Directory containing… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/z-image-examples.hydro_cali_agent_examplesycophancy_examples
Sycophancy Examples
Two sycophancy evaluation datasets from Kei et al., "Reward hacking can generalise across settings".
Original source: GeodesicResearch/Obfuscation_Generalization
Files
File
Examples
Description
sycophancy_opinion_political.jsonl
5,000
Political opinion questions with persona-aligned "sycophantic" answers
sycophancy_fact.jsonl
401
Factual questions where the persona holds a misconception; sycophantic answer agrees with the misconception… See the full description on the dataset page: https://huggingface.co/datasets/camgeodesic/sycophancy_examples.example-space-to-dataset-imageDemo to save data from a Space to a Dataset. Goal is to provide reusable snippets of code.
Documentation: https://huggingface.co/docs/huggingface_hub/main/en/guides/upload#scheduled-uploads
Space: https://huggingface.co/spaces/Wauplin/space_to_dataset_saver/
JSON dataset: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-json
Image dataset: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-image
Image (zipped) dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-image.example-vlm-sft-dataset
Sample VLM Reference Dataset
Dataset Description
This is a reference dataset demonstrating the proper VLM SFT (Supervised Fine-Tuning) format for vision-language model training. It contains 10 minimal conversational examples that show the exact structure and formatting required for VLM training pipelines.
⚠️ Important: This dataset is for format reference only - not intended for actual model training. Use this as a template to understand the required data structure for… See the full description on the dataset page: https://huggingface.co/datasets/alay2shah/example-vlm-sft-dataset.diffuman4d_example
Diffuman4D Example Test Data
Project Page | Paper | Code | Model
This repo provides several example scenes for testing Diffuman4D. Note that these scenes are not part of the official DNA-Rendering dataset. For more testing (and training) data, see this repo.
Usage
See the GitHub repo for detailed usage.
Cite
@inproceedings{jin2025diffuman4d,
title={Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models}… See the full description on the dataset page: https://huggingface.co/datasets/krahets/diffuman4d_example.
