datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
doc-formats-jsonl-1
[doc] formats - jsonl - 1
This dataset contains one jsonl file at the root.
sounio-code-examples
Sounio Curated Code Examples
Curated compile-clean .sio examples for training and evaluating code models on
Sounio, a self-hosted systems and scientific programming language for epistemic
computing, uncertainty propagation, and algebraic effects.
This directory is the Cx-1 expansion lane for
chiuratto-AIgourakis/sounio-code-examples.
Current batch
Examples: 5,000
Metadata files: 5,000
Compiler gate: bin/souc check pass rate 5,000/5,000
Utility layer: 5,000… See the full description on the dataset page: https://huggingface.co/datasets/chiuratto-AIgourakis/sounio-code-examples.chat_formatted_examplesGEN3C-Testing-Example
GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control
CVPR 2025 (Highlight)
Xuanchi Ren*,
Tianchang Shen*
Jiahui Huang,
Huan Ling,
Yifan Lu,
Merlin Nimier-David,
Thomas Müller,
Alexander Keller,
Sanja Fidler,
Jun Gao
* indicates equal contribution
Paper, Project Page
Abstract: We present GEN3C, a generative video model with precise Camera Control and
temporal 3D Consistency. Prior video models already generate realistic videos,
but they tend to leverage… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/GEN3C-Testing-Example.multimodal-example
Multimodal Example Dataset
Small example dataset for testing multimodal (vision-language) fine-tuning with ms-swift.
Structure
├── train.jsonl # 10 training samples
├── test.jsonl # 2 validation samples
├── images/ # All referenced images (400x300 JPEG)
│ ├── dog_portrait.jpg
│ ├── forest_river.jpg
│ ├── laptop_desk.jpg
│ ├── mountain_lake.jpg
│ ├── ocean_rocks.jpg
│ ├── coffee_cup.jpg
│ ├── bookshelf.jpg
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/f13rnd/multimodal-example.example-space-to-dataset-jsonDemo to save data from a Space to a Dataset. Goal is to provide reusable snippets of code.
Documentation: https://huggingface.co/docs/huggingface_hub/main/en/guides/upload#scheduled-uploads
Space: https://huggingface.co/spaces/Wauplin/space_to_dataset_saver/
JSON dataset: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-json
Image dataset: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-image
Image (zipped) dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-json.hydro_cali_agent_examplesycophancy_examples
Sycophancy Examples
Two sycophancy evaluation datasets from Kei et al., "Reward hacking can generalise across settings".
Original source: GeodesicResearch/Obfuscation_Generalization
Files
File
Examples
Description
sycophancy_opinion_political.jsonl
5,000
Political opinion questions with persona-aligned "sycophantic" answers
sycophancy_fact.jsonl
401
Factual questions where the persona holds a misconception; sycophantic answer agrees with the misconception… See the full description on the dataset page: https://huggingface.co/datasets/camgeodesic/sycophancy_examples.OpenSeek-Synthetic-Reasoning-Data-Examples
OpenSeek-Reasoning-Data
OpenSeek [Github|Blog]
Recent reseach has demonstrated that the reasoning ability of LLMs originates from the pre-training stage, activated by RL training. Massive raw corpus containing complex human reasoning process, but lack of generalized and effective synthesis method to extract these reasoning process.
News
🔥🔥🔥[2025/02/25] We publish some math, code, and general knowledge domain reasoning data synthesized from the current pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/OpenSeek-Synthetic-Reasoning-Data-Examples.flutter-full-examples-v1
Flutter Codegen: Full Examples
Synthetic dataset of complete Flutter/Dart widgets, each paired with the goal
that describes them and (optionally) starting code. Unlike flutter-codegen-diff-steps,
there's no step history or diff structure here -- each row is a single, standalone
goal -> complete file example.
This is the whole-code counterpart to flutter-diff-steps-v1, intended for
training/evaluating a baseline that generates the entire file in one shot, to
compare against the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-full-examples-v1.chinese_law_examples
1000 examples of law items
law_item.jsonl contains 1000 samples of current and effective Chinese laws. e.g.
{"title": "《中华人民共和国劳动合同法(2012修正)》",
"classification": "类别 : 劳动合同营商环境优化 ",
"num": "第十九条",
"contents": "第十九条【试用期】劳动合同期限三个月以上不满一年的,试用期不得超过一个月;劳动合同期限一年以上不满三年的,试用期不得超过二个月;三年以上固定期限和无固定期限的劳动合同,试用期不得超过六个月。同一用人单位与同一劳动者只能约定一次试用期。以完成一定工作任务为期限的劳动合同或者劳动合同期限不满三个月的,不得约定试用期。试用期包含在劳动合同期限内。劳动合同仅约定试用期的,试用期不成立,该期限为劳动合同期限。"}
Using BGE Embedding to compute similarity between query and… See the full description on the dataset page: https://huggingface.co/datasets/pandalla/chinese_law_examples.agent-traces-examplegeorgian-booking-safety-examples
Synthetic Georgian booking-safety examples
Twelve hand-annotated design examples. Not an ASR/NLU benchmark, a training
corpus, or a measurement of any deployed product.
This package makes the OMO AI educational examples available as a
small, flat JSONL table with a complete, lossless copy of each original case.
The source material and this packaging were prepared with AI assistance.
All utterances, identities, times, state and destination responses are fictional.
There are no… See the full description on the dataset page: https://huggingface.co/datasets/vajelski/georgian-booking-safety-examples.spider-rollouts-web-search-qwen2.5-7b-gaia-32-examplesopengloss-v1.3-query-examples-flat
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Query Examples v1.3 (Flattened)
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples-flat.llm-behavioral-drift-examples
LLM-Behavioral-Drift-Examples
Training examples for SFT used to induce behavioral drift.
Dataset Description
Various training examples in an AI office assistant setting (email, calendar, docs, and misc. info.) skewed in particular ways (aggression, irrelevancy/tangential information, and excessive verbosity).
Examples produced by Gemini 2.5 Flash.
Example Usage
aggressive_dataset = load_dataset(
f"{username}/{repo_name}",
data_files="aggressive.jsonl"… See the full description on the dataset page: https://huggingface.co/datasets/6S-bobby/llm-behavioral-drift-examples.worked-examples-metamath-v0
Full MetaMathQA worked-examples pack
Source: meta-math/MetaMathQA (all 395k, all types).
Train: 355,688 instances (90% of families)
Holdout: 39,155 (eval/holdout_bare.jsonl)
Docs: arms/<arm>/docs.jsonl.gz (gunzip to use)
Tokens: tokenized/<arm>/shard-00000.npy (dolma2, EOS 100257)
Arm stats
{
"fade_shuffled": {
"n_docs": 1873620,
"n_tokens": 453279629
}
}
cursor-traces-exampleThis dataset was generated using teich by TeichAI
My Agent Traces
This directory contains raw agent trace files generated by teich.
JSONL files: 9
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/cursor-traces-example.agenda-parser-models-example-agent-traces
Agenda Parser — fine-tuned agent models
Three Gemma 4 models fine-tuned to drive the Agenda Parser's ReAct agent: at each step
the model emits a single JSON action {"thought","tool","args"} over two toolkits —
meeting-agenda packets and Michigan local-government law (Open Meetings Act, FOIA,
the Michigan Compiled Laws via Cornell LII). This card doubles as the project write-up; the
dataset itself (bottom) is a gallery of example traces from the three models.
tier
base… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-models-example-agent-traces.LORE-examples
LORE Examples
A small set of matched multimodal examples from LORE, for the
MIMIC model — enough to try inference,
embedding, and generation across DNA, RNA, and protein modalities without wiring
up your own data.
Each example is a single biological entity (a transcript and/or its protein) with
several co-observed modalities. Rows are drawn from the held-out (validation) split
of MIMIC's training data, so they are in-distribution and length-bounded to the
model's context… See the full description on the dataset page: https://huggingface.co/datasets/polymathic-ai/LORE-examples.literary-genre-examples
Literary Genre Dataset
This dataset contains a curated list of 86 fiction and nonfiction genres, each accompanied by a representative example paragraph. The example texts illustrate the typical tone, writing style, and content characteristics for each genre.
Genres Covered: 86 total, spanning popular and niche categories in both fiction and nonfiction.
Genre Types: Marked as either Fiction or Nonfiction.
Example Paragraphs: Each genre includes a sample paragraph written to capture… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/literary-genre-examples.Chinese-DeepSeek-V3.2-Exp-chat-example
deepseek/deepseek-v3.2-exp (6.6K) 中文数据集样本
一、前言
本报告基于 deepseek/deepseek-v3.2-exp 模型(官方 API,8K 上下文窗口)进行数据集评测与可视化展示。测试数据集共包含 6,655 轮对话,语言覆盖以中文为主,辅以部分混合语种及非中文输入。本次报告旨在总结模型的对话特征、输入输出长度分布及上下文预算消耗情况,并为后续应用和优化提供参考。
二、数据与方法
数据来源:用户构建的 6,655 轮真实中文对话样本。
估算方法:
中文字符近似为 1 Token;
英文 4 字符 ≈ 1 Token;
用于规模与上下文预算对比,而非精确 Token 计数。
统计维度:
平均 Prompt/Output 长度(字符与估算 Token);
总 Token 占上下文窗口比例;
语言分布(Prompt 语言类型);
对话长度分布(用户提问、助手回答、总对话长度)。
三、总体结果
1. 样本概况… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-DeepSeek-V3.2-Exp-chat-example.repro-fuse-full-spectrum-unlearnable-examples-via-spectral-equalization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
evaluation_exampleschinese_verdict_examples
verdicts examples
verdicts_200.jsonl contains 200 examples of verdicts from Chinese Judgements Online, we process the datasets for semantic retrieval
using BGE to compute similarity between query and verdict
from FlagEmbedding import FlagModel
from datasets import load_dataset
dataset = load_dataset("FarReelAILab/verdicts")
model = FlagModel('BAAI/bge-large-zh-v1.5',
query_instruction_for_retrieval="为这个句子生成表示以用于检索相关文章:",
use_fp16=True)… See the full description on the dataset page: https://huggingface.co/datasets/pandalla/chinese_verdict_examples.Sam_Altman_OpenAI_Podcast_XScraper_Example
🔍 X-Twitter Scraper: Real-Time Tweet Search & Scrape Tool
Search and scrape X-Twitter for posts by keyword, account, or trending topics.A simple, no-code tool to pull real-time, relevant content in LLM-ready JSON format — perfect for agents, RAG systems, or content workflows.
👉 Start Searching & Scraping on Hugging Face
✨ Features
⚡ Real-Time FetchStream the latest tweets as they’re posted — no delay.
🎯 Flexible SearchSearch by keywords, #hashtags, $cashtags… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/Sam_Altman_OpenAI_Podcast_XScraper_Example.dictation-cleanup-examples
Dictation cleanup examples
A sample of the hand-written cases behind
SpeakoFlow Mini, published so the
conventions the model follows are inspectable rather than described.
Seven cases in each of fifteen categories, spread across short, medium and long transcripts.
Every case was written by hand. None of it is captured speech.
This is not a benchmark
Read that before using it for anything.
These cases are drawn from the training pool, not from the held-out set the… See the full description on the dataset page: https://huggingface.co/datasets/SpeakoFlow/dictation-cleanup-examples.CoT-O1-examplesExample-Dataset
Example-Dataset
Example-Dataset — это русскоязычный легковесный диалоговый датасет, предназначенный для обучения и дообучения (Fine-Tuning / SFT) чат-ботов и диалоговых ИИ-ассистентов.
⚠️ Статус: Датасет находится в активной разработке и будет регулярно обновляться и дополняться новыми примерами диалогов и сценариями общения.
📊 Текущая статистика и характеристики
Количество диалогов на текущий момент: 81
Формат структуры: Multi-turn (многошаговые диалоги:… See the full description on the dataset page: https://huggingface.co/datasets/longtimedevs/Example-Dataset.code-examples
MongoDB Code Examples
This dataset contains code examples of using MongoDB technologies. These code examples come from the
MongoDB documentation and developer blog.
The dataset is updated regularly to stay relatively up-to-date with the latest published content.
Schema
The dataset includes the code example text and useful metadata for working with the code examples. Every code example in the dataset includes the following:
export interface CodeExampleDatasetEntry… See the full description on the dataset page: https://huggingface.co/datasets/mongodb-eai/code-examples.
