datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hub-stats
Changelog
NEW Changes March 11th 2026
Added new split: arxiv_papers, sourced from the Hugging Face /api/papers endpoint
papers continues to point to daily_papers.parquet, which is the Daily Papers feed
NEW Changes July 25th
added baseModels field to models which shows the models that the user tagged as base models for that model
Example:
{
"models": [
{
"_id": "687de260234339fed21e768a",
"id": "Qwen/Qwen3-235B-A22B-Instruct-2507"
}
],
"relation":… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/hub-stats.CFAD
CFAD
Benchmark-ready packaging of the CFAD (Chinese Fake Audio Detection) clean test
set (arXiv 2207.12308), for speech anti-spoofing and
synthetic / deepfake voice detection on Mandarin Chinese speech.
Overview
CFAD is a large-scale Chinese fake-audio detection corpus. This repo packages the clean
version's two test partitions:
test_seen — spoof systems and real corpora also present in the train/dev splits.
test_unseen — spoof systems and real corpora held out… See the full description on the dataset page: https://huggingface.co/datasets/SpeechAntiSpoofingBenchmarks/CFAD.SWE-chat
SWE-chat: Coding Agent Interactions From Real Users in the Wild
📄 Paper: arxiv.org/abs/2604.20779
🌐 Website: swe-chat.com
[!NOTE]
This is a copy of SALT-NLP/SWE-chat with a traces config added as the default, so the Hub's dataset viewer renders sessions as agent traces. The original files are unchanged; see Agent Traces for how traces/ was built.
Dataset Summary
SWE-chat captures real-world AI coding sessions from developers using AI coding assistants… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/SWE-chat.CFAD
CFAD
Benchmark-ready packaging of the CFAD (Chinese Fake Audio Detection) clean test
set (arXiv 2207.12308), for speech anti-spoofing and
synthetic / deepfake voice detection on Mandarin Chinese speech.
Overview
CFAD is a large-scale Chinese fake-audio detection corpus. This repo packages the clean
version's two test partitions:
test_seen — spoof systems and real corpora also present in the train/dev splits.
test_unseen — spoof systems and real corpora held out… See the full description on the dataset page: https://huggingface.co/datasets/pupengleileileilei/CFAD.gr00t-x-embodiment-sim-gr1-pouring-v3
GR00T X-Embodiment Sim: GR1 Pouring (LeRobot v3.0 conversion)
Format test: one subset of nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim (gr1_full_upper_body.Pouring) converted from LeRobot v2.0 to v3.0, to preview how NVIDIA's GR00T datasets render on the Hub.
Source: NVIDIA, CC-BY-4.0. All data is NVIDIA's; only the file layout changed.
Robot: Fourier GR-1 (GR1FixedLowerBody), 1,000 episodes, 267,780 frames at 20 fps, one 256×256 front_view camera.
Conversion: v2.0 → v2.1… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/gr00t-x-embodiment-sim-gr1-pouring-v3.Fable-5-tracesA simple dataset of the raw Fable 5 Claude session logs we could get our hands on before it was taken away (no clue if it's coming back).
The raw trace files live in sessions/*.jsonl. Cache files, paste-cache files, shell history, and merged COT training exports are intentionally omitted so Hugging Face Datasets can load the repo through the agent-traces path.
A pretty viewer for dataset:… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/Fable-5-traces.hf-coding-tools-traces
HuggingFace AI Coding Tools — Agent Traces
This dataset rehydrates the benchmark results from
davidkling/hf-coding-tools-dashboard
into the JSONL session format consumed by the
Hugging Face Agent Trace Viewer.
What's inside
32 sessions, one per (tool, model, effort, thinking) configuration
9,130 query → response turns total (≈18,260 events)
Tools covered: claude_code, codex, copilot, cursor
Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/hf-coding-tools-traces.react-code-instructions
React Code Instructions
Popular Queries
Number of instructions by Model
Unnested Messages
Instructions Added Per Day
Dataset of Claude Artifact esque React Apps generated by Llama 3.1 70B, Llama 3.1 405B, and Deepseek Chat V3.
Examples
Virtual Fitness Trainer Website
LinkedIn Clone
iPhone Calculator
Chipotle Waitlist
Apple Store
glue
Dataset Card for GLUE
Dataset Summary
GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems.
Supported Tasks and Leaderboards
The leaderboard for the GLUE benchmark can be found at this address. It comprises the following tasks:
ax
A manually-curated evaluation dataset for fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/glue.crema-d
CREMA-D
Viewer-compatible mirror of the audio portion of CREMA-D (Crowd-sourced Emotional Multimodal Actors Dataset).
This repository republishes the audio files from myleslinder/crema-d as Hub-native Parquet shards so the Hugging Face viewer, search, and filtering work without a custom loading script.
What is included
7,442 WAV clips
91 actors
12 fixed sentence prompts
6 emotion labels: anger, disgust, fear, happy, neutral, sad
4 intensity codes: LO, MD, HI, XX… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/crema-d.factory-traces
simready-usd-web-viewers
SimReady assets in browser USD viewers
Six real SimReady OpenUSD packages from the Hub, loaded in five browser USD libraries and a pre-converted GLB baseline. Each cell is what the library drew.
Asset
Reference
three.js 0.174
three.js 0.186
three.js 0.186 + crawler
Needle
tinyusdz
cinevva usdjs
GLB (pre-converted)
LG laptopusdc · 11 MB
❌ zip error
✅ renders1.3 s · 279 MB
✅ renders1.3 s · 378 MB
✅ renders1.2 s · 1.3 GB
✅ renders1.3 s · 446 MB
✅ renders2.2 s · 327 MB
✅… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/simready-usd-web-viewers.agent-sessions-list
Agent Traces
agent-sessions-list is a small index of real agent session trace files from Claude Code, Codex, Hermes Agent, Factory/Droid, and Pi.
How to find your Agent Traces
Agent
Typical local session directory
Claude Code
~/.claude/projects
Codex
~/.codex/sessions
Codex archive
~/.codex/archived_sessions
Hermes Agent
~/.hermes/state.db; export with hermes sessions export <output>.jsonl
Factory/Droid
~/.factory/sessions
Pi… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/agent-sessions-list.mmu_cfa_cfa4
mmu_cfa_cfa4 HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_cfa_cfa4.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be installed… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_cfa_cfa4.VisionFoundry-10K
VisionFoundry-10K
VisionFoundry-10K is a synthetic visual question answering (VQA) dataset with 10,000 image-question-answer triples spanning 10 vision-centric tasks. The data is produced by the VisionFoundry pipeline: an LLM generates task-aware questions, answers, and detailed text-to-image prompts; a text-to-image model synthesizes images; and a strong multimodal verifier filters samples for alignment.
VisionFoundry: Teaching VLMs Visual Perception with… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/VisionFoundry-10K.medicaid-provider-spending
Medicaid Provider Spending
This dataset contains provider-level Medicaid spending data aggregated from outpatient and professional claims with valid HCPCS codes, covering January 2018 through December 2024. It provides insights into how Medicaid dollars are distributed across providers and procedures nationwide.
Provider details (name, address, taxonomy) are sourced from the NPPES NPI Registry (February 2026 dissemination).
Data Description
Attribute
Value… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/medicaid-provider-spending.mmu_cfa_cfa3
mmu_cfa_cfa3 HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_cfa_cfa3.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be installed… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_cfa_cfa3.en-cfa
CFA exam questions
🌐 The Fin AI
Formerly TheFinAI/flare-cfa (the old name redirects here).
Legacy FLARE/PIXIU-era task (loaded by the PIXIU harness). The authors have not yet specified a license for this dataset; treat it as research-only until clarified, and respect the terms of the original sources.
Task
multiple-choice QA (CFA exam)
Language
en
License
other (unspecified)
Quick Start
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/en-cfa.hermes-agent-trace-samples-2026-06-05
Hermes Agent Raw Session Samples
Five public-safe raw Hermes Agent session exports generated on 2026-06-05 with the Hermes CLI using Hugging Face Inference Providers.
Each file in sessions/ is the exact single-session output from:
hermes sessions export sessions/<session_id>.jsonl --session-id <session_id>
No derived tables, flattened rows, SQLite database, or formatted JSON copies are included.
mmu_cfa_snii
mmu_cfa_snii HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_cfa_snii.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be installed… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_cfa_snii.pi-diff-review
Coding agent session traces for badlogicgames/pi-diff-review
This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each sessions/*.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/pi-diff-review.pi-mono-fresh
pi-mono-fresh
cfahlgren1/pi-mono-fresh is a straight mirror of the JSONL files from badlogicgames/pi-mono.
What is included
627 .jsonl files mirrored from the source dataset.
manifest.jsonl, if present in the source dataset.
No schema changes, filtering, or content transformations.
Provenance
Source dataset: badlogicgames/pi-mono
Source snapshot mirrored: dac2a1d3ba12dda597b973a791a77618ccb5f413
Mirror created: 2026-04-06
Mirrored by: cfahlgren1… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/pi-mono-fresh.finmmeval-cfa-cpa
Financial Exam MCQ Training Dataset
A bilingual training dataset of financial and accounting multiple-choice questions in English and Chinese, formatted for instruction tuning and answer selection tasks.
Dataset Structure
Format: Multiple-choice questions
Language: English and Chinese
Domain: Accounting, finance, auditing, taxation, and financial regulations
Size: 596 examples
Files:
train-00000-of-00001-en.parquet
train-00000-of-00001-cn.parquet… See the full description on the dataset page: https://huggingface.co/datasets/Tomas08119993/finmmeval-cfa-cpa.codex-sessions
Codex Sessions
Archive of raw OpenAI Codex CLI session files, plus a derived one-row-per-session view.
Files
rollout-*.jsonl: untouched raw Codex session files stored for fidelity
sessions.jsonl: derived file where 1 row = 1 session
Derived Sessions Format
sessions.jsonl contains one JSON object per session with this shape:
{
"session_id": "019d2fac-0b38-70f0-baff-a394265d8291",
"file_name":… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/codex-sessions.model-toolcall-research
Model Toolcall Research
This dataset stores newline-delimited agent traces from bounded research runs on model repository tool-schema support.
The Dataset Viewer is configured to index only .jsonl files:
toolcall_traces loads trace files under traces/**/*.jsonl.
research_session loads top-level provenance/session traces from *.jsonl.
The archive/ directory preserves the earlier .trace.json uploads for reference, but those files are newline-delimited JSON streams rather than… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/model-toolcall-research.mmu_cfa_seccsn
mmu_cfa_seccsn HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_cfa_seccsn.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_cfa_seccsn.fin-cfa-graphgen
Fin-CFA-GraphGen: 785K Knowledge-Guided Financial QA Examples
Fin-CFA-GraphGen is a large-scale English dataset for financial instruction tuning, financial question answering, and domain-specific language-model post-training. It contains 785,149 synthetic question–answer examples generated from CFA curriculum and exam-preparation books with the GraphGen knowledge-driven data-generation method.
The dataset and its role in the post-training pipeline are described in Data-Centric… See the full description on the dataset page: https://huggingface.co/datasets/whoisjiji/fin-cfa-graphgen.web-fetch-harness-traces
Native web fetch harness traces
Separate, lightly sanitized native JSONL traces comparing URL-fetch behavior in Claude Code and Codex CLI against:
https://huggingface.co/datasets/nyu-mll/glue
Captured on 2026-09-15. No shell HTTP client, browser automation, or MCP fetcher was used.
Files
data/claude-code.jsonl: Claude Code's stream-json events.
data/codex.jsonl: Codex's persisted native rollout JSONL.
Both traces are exposed together in the default subset and… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/web-fetch-harness-traces.ICAI_SM_CFA
ICAI Study Material
A personal collection of ICAI study material (MTP, RTP, practice sets and subject modules), stored here so it can be read folder by folder without OTP or password logins.
Read it on the website: https://bramhalab.github.io/icai-library-for-we/
Credits
The Institute of Chartered Accountants of India (ICAI) for all the study material. The content belongs to ICAI.
Hugging Face for free storage.
Disclaimer
This is an unofficial… See the full description on the dataset page: https://huggingface.co/datasets/Brand1809/ICAI_SM_CFA.cfaibed
