Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cfahlgren1 /hub-stats Changelog NEW Changes March 11th 2026 Added new split: arxiv_papers, sourced from the Hugging Face /api/papers endpoint papers continues to point to daily_papers.parquet, which is the Daily Papers feed NEW Changes July 25th added baseModels field to models which shows the models that the user tagged as base models for that model Example: { "models": [ { "_id": "687de260234339fed21e768a", "id": "Qwen/Qwen3-235B-A22B-Instruct-2507" } ], "relation":… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/hub-stats.tabular1M<n<10M81 likes13k downloads8h agoHugging Face02SpeechAntiSpoofingBenchmarks /CFAD CFAD Benchmark-ready packaging of the CFAD (Chinese Fake Audio Detection) clean test set (arXiv 2207.12308), for speech anti-spoofing and synthetic / deepfake voice detection on Mandarin Chinese speech. Overview CFAD is a large-scale Chinese fake-audio detection corpus. This repo packages the clean version's two test partitions: test_seen — spoof systems and real corpora also present in the train/dev splits. test_unseen — spoof systems and real corpora held out… See the full description on the dataset page: https://huggingface.co/datasets/SpeechAntiSpoofingBenchmarks/CFAD.audioaudio-classification10K<n<100K0 likes1.2k downloads4mo agoHugging Face03cfahlgren1 /SWE-chat SWE-chat: Coding Agent Interactions From Real Users in the Wild 📄 Paper: arxiv.org/abs/2604.20779 🌐 Website: swe-chat.com [!NOTE] This is a copy of SALT-NLP/SWE-chat with a traces config added as the default, so the Hub's dataset viewer renders sessions as agent traces. The original files are unchanged; see Agent Traces for how traces/ was built. Dataset Summary SWE-chat captures real-world AI coding sessions from developers using AI coding assistants… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/SWE-chat.text-generation1M<n<10M0 likes1.2k downloads15d agoHugging Face04pupengleileileilei /CFAD CFAD Benchmark-ready packaging of the CFAD (Chinese Fake Audio Detection) clean test set (arXiv 2207.12308), for speech anti-spoofing and synthetic / deepfake voice detection on Mandarin Chinese speech. Overview CFAD is a large-scale Chinese fake-audio detection corpus. This repo packages the clean version's two test partitions: test_seen — spoof systems and real corpora also present in the train/dev splits. test_unseen — spoof systems and real corpora held out… See the full description on the dataset page: https://huggingface.co/datasets/pupengleileileilei/CFAD.audioaudio-classification10K<n<100K0 likes756 downloads23d agoHugging Face05cfahlgren1 /gr00t-x-embodiment-sim-gr1-pouring-v3 GR00T X-Embodiment Sim: GR1 Pouring (LeRobot v3.0 conversion) Format test: one subset of nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim (gr1_full_upper_body.Pouring) converted from LeRobot v2.0 to v3.0, to preview how NVIDIA's GR00T datasets render on the Hub. Source: NVIDIA, CC-BY-4.0. All data is NVIDIA's; only the file layout changed. Robot: Fourier GR-1 (GR1FixedLowerBody), 1,000 episodes, 267,780 frames at 20 fps, one 256×256 front_view camera. Conversion: v2.0 → v2.1… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/gr00t-x-embodiment-sim-gr1-pouring-v3.tabularrobotics100K<n<1M1 likes658 downloads8d agoHugging Face06cfahlgren1 /Fable-5-tracesA simple dataset of the raw Fable 5 Claude session logs we could get our hands on before it was taken away (no clue if it's coming back). The raw trace files live in sessions/*.jsonl. Cache files, paste-cache files, shell history, and merged COT training exports are intentionally omitted so Hugging Face Datasets can load the repo through the agent-traces path. A pretty viewer for dataset:… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/Fable-5-traces.tabularn<1K19 likes613 downloads4mo agoHugging Face07cfahlgren1 /hf-coding-tools-traces HuggingFace AI Coding Tools — Agent Traces This dataset rehydrates the benchmark results from davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 32 sessions, one per (tool, model, effort, thinking) configuration 9,130 query → response turns total (≈18,260 events) Tools covered: claude_code, codex, copilot, cursor Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/hf-coding-tools-traces.tabularn<1K0 likes436 downloads5mo agoHugging Face08cfahlgren1 /react-code-instructions React Code Instructions Popular Queries Number of instructions by Model Unnested Messages Instructions Added Per Day Dataset of Claude Artifact esque React Apps generated by Llama 3.1 70B, Llama 3.1 405B, and Deepseek Chat V3. Examples Virtual Fitness Trainer Website LinkedIn Clone iPhone Calculator Chipotle Waitlist Apple Store text10K<n<100K158 likes330 downloads2y agoHugging Face09cfahlgren1 /glue Dataset Card for GLUE Dataset Summary GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems. Supported Tasks and Leaderboards The leaderboard for the GLUE benchmark can be found at this address. It comprises the following tasks: ax A manually-curated evaluation dataset for fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/glue.tabulartext-classification1M<n<10M0 likes296 downloads24d agoHugging Face10cfahlgren1 /crema-d CREMA-D Viewer-compatible mirror of the audio portion of CREMA-D (Crowd-sourced Emotional Multimodal Actors Dataset). This repository republishes the audio files from myleslinder/crema-d as Hub-native Parquet shards so the Hugging Face viewer, search, and filtering work without a custom loading script. What is included 7,442 WAV clips 91 actors 12 fixed sentence prompts 6 emotion labels: anger, disgust, fear, happy, neutral, sad 4 intensity codes: LO, MD, HI, XX… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/crema-d.audio1K<n<10K0 likes285 downloads6mo agoHugging Face11cfahlgren1 /factory-traces tabularn<1K0 likes283 downloads4mo agoHugging Face12cfahlgren1 /simready-usd-web-viewers SimReady assets in browser USD viewers Six real SimReady OpenUSD packages from the Hub, loaded in five browser USD libraries and a pre-converted GLB baseline. Each cell is what the library drew. Asset Reference three.js 0.174 three.js 0.186 three.js 0.186 + crawler Needle tinyusdz cinevva usdjs GLB (pre-converted) LG laptopusdc · 11 MB ❌ zip error ✅ renders1.3 s · 279 MB ✅ renders1.3 s · 378 MB ✅ renders1.2 s · 1.3 GB ✅ renders1.3 s · 446 MB ✅ renders2.2 s · 327 MB ✅… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/simready-usd-web-viewers.imagen<1K1 likes278 downloads1d agoHugging Face13cfahlgren1 /agent-sessions-list Agent Traces agent-sessions-list is a small index of real agent session trace files from Claude Code, Codex, Hermes Agent, Factory/Droid, and Pi. How to find your Agent Traces Agent Typical local session directory Claude Code ~/.claude/projects Codex ~/.codex/sessions Codex archive ~/.codex/archived_sessions Hermes Agent ~/.hermes/state.db; export with hermes sessions export <output>.jsonl Factory/Droid ~/.factory/sessions Pi… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/agent-sessions-list.tabularn<1K0 likes232 downloads4mo agoHugging Face14UniverseTBD /mmu_cfa_cfa4 mmu_cfa_cfa4 HATS Catalog Collection This is the collection of HATS catalogs representing mmu_cfa_cfa4. This dataset is part of the Multimodal Universe, a large-scale collection of multimodal astronomical data. For full details, see the paper: The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data. Access the catalog We recommend the use of the LSDB Python framework to access HATS catalogs. LSDB can be installed… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_cfa_cfa4.tabularn<1K0 likes207 downloads5mo agoHugging Face15cfahlgren1 /VisionFoundry-10K VisionFoundry-10K VisionFoundry-10K is a synthetic visual question answering (VQA) dataset with 10,000 image-question-answer triples spanning 10 vision-centric tasks. The data is produced by the VisionFoundry pipeline: an LLM generates task-aware questions, answers, and detailed text-to-image prompts; a text-to-image model synthesizes images; and a strong multimodal verifier filters samples for alignment. VisionFoundry: Teaching VLMs Visual Perception with… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/VisionFoundry-10K.imagevisual-question-answering10K<n<100K0 likes174 downloads25d agoHugging Face16cfahlgren1 /medicaid-provider-spending Medicaid Provider Spending This dataset contains provider-level Medicaid spending data aggregated from outpatient and professional claims with valid HCPCS codes, covering January 2018 through December 2024. It provides insights into how Medicaid dollars are distributed across providers and procedures nationwide. Provider details (name, address, taxonomy) are sourced from the NPPES NPI Registry (February 2026 dissemination). Data Description Attribute Value… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/medicaid-provider-spending.tabular100M<n<1B3 likes163 downloads8mo agoHugging Face17UniverseTBD /mmu_cfa_cfa3 mmu_cfa_cfa3 HATS Catalog Collection This is the collection of HATS catalogs representing mmu_cfa_cfa3. This dataset is part of the Multimodal Universe, a large-scale collection of multimodal astronomical data. For full details, see the paper: The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data. Access the catalog We recommend the use of the LSDB Python framework to access HATS catalogs. LSDB can be installed… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_cfa_cfa3.tabularn<1K0 likes161 downloads5mo agoHugging Face18TheFinAI /en-cfagated CFA exam questions 🌐 The Fin AI Formerly TheFinAI/flare-cfa (the old name redirects here). Legacy FLARE/PIXIU-era task (loaded by the PIXIU harness). The authors have not yet specified a license for this dataset; treat it as research-only until clarified, and respect the terms of the original sources. Task multiple-choice QA (CFA exam) Language en License other (unspecified) Quick Start from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/en-cfa.textquestion-answering1K<n<10K2 likes160 downloads3d agoHugging Face19cfahlgren1 /hermes-agent-trace-samples-2026-06-05 Hermes Agent Raw Session Samples Five public-safe raw Hermes Agent session exports generated on 2026-06-05 with the Hermes CLI using Hugging Face Inference Providers. Each file in sessions/ is the exact single-session output from: hermes sessions export sessions/<session_id>.jsonl --session-id <session_id> No derived tables, flattened rows, SQLite database, or formatted JSON copies are included. tabularn<1K0 likes160 downloads4mo agoHugging Face20UniverseTBD /mmu_cfa_snii mmu_cfa_snii HATS Catalog Collection This is the collection of HATS catalogs representing mmu_cfa_snii. This dataset is part of the Multimodal Universe, a large-scale collection of multimodal astronomical data. For full details, see the paper: The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data. Access the catalog We recommend the use of the LSDB Python framework to access HATS catalogs. LSDB can be installed… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_cfa_snii.tabularn<1K0 likes150 downloads5mo agoHugging Face21cfahlgren1 /pi-diff-review Coding agent session traces for badlogicgames/pi-diff-review This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each sessions/*.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/pi-diff-review.tabulartext-generationn<1K0 likes148 downloads6mo agoHugging Face22cfahlgren1 /pi-mono-fresh pi-mono-fresh cfahlgren1/pi-mono-fresh is a straight mirror of the JSONL files from badlogicgames/pi-mono. What is included 627 .jsonl files mirrored from the source dataset. manifest.jsonl, if present in the source dataset. No schema changes, filtering, or content transformations. Provenance Source dataset: badlogicgames/pi-mono Source snapshot mirrored: dac2a1d3ba12dda597b973a791a77618ccb5f413 Mirror created: 2026-04-06 Mirrored by: cfahlgren1… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/pi-mono-fresh.text0 likes136 downloads6mo agoHugging Face23Tomas08119993 /finmmeval-cfa-cpa Financial Exam MCQ Training Dataset A bilingual training dataset of financial and accounting multiple-choice questions in English and Chinese, formatted for instruction tuning and answer selection tasks. Dataset Structure Format: Multiple-choice questions Language: English and Chinese Domain: Accounting, finance, auditing, taxation, and financial regulations Size: 596 examples Files: train-00000-of-00001-en.parquet train-00000-of-00001-cn.parquet… See the full description on the dataset page: https://huggingface.co/datasets/Tomas08119993/finmmeval-cfa-cpa.textquestion-answeringn<1K1 likes133 downloads6mo agoHugging Face24cfahlgren1 /codex-sessions Codex Sessions Archive of raw OpenAI Codex CLI session files, plus a derived one-row-per-session view. Files rollout-*.jsonl: untouched raw Codex session files stored for fidelity sessions.jsonl: derived file where 1 row = 1 session Derived Sessions Format sessions.jsonl contains one JSON object per session with this shape: { "session_id": "019d2fac-0b38-70f0-baff-a394265d8291", "file_name":… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/codex-sessions.textn<1K0 likes129 downloads7mo agoHugging Face25cfahlgren1 /model-toolcall-research Model Toolcall Research This dataset stores newline-delimited agent traces from bounded research runs on model repository tool-schema support. The Dataset Viewer is configured to index only .jsonl files: toolcall_traces loads trace files under traces/**/*.jsonl. research_session loads top-level provenance/session traces from *.jsonl. The archive/ directory preserves the earlier .trace.json uploads for reference, but those files are newline-delimited JSON streams rather than… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/model-toolcall-research.tabularn<1K0 likes118 downloads5mo agoHugging Face26UniverseTBD /mmu_cfa_seccsn mmu_cfa_seccsn HATS Catalog Collection This is the collection of HATS catalogs representing mmu_cfa_seccsn. This dataset is part of the Multimodal Universe, a large-scale collection of multimodal astronomical data. For full details, see the paper: The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data. Access the catalog We recommend the use of the LSDB Python framework to access HATS catalogs. LSDB can be… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_cfa_seccsn.tabularn<1K0 likes117 downloads5mo agoHugging Face27whoisjiji /fin-cfa-graphgen Fin-CFA-GraphGen: 785K Knowledge-Guided Financial QA Examples Fin-CFA-GraphGen is a large-scale English dataset for financial instruction tuning, financial question answering, and domain-specific language-model post-training. It contains 785,149 synthetic question–answer examples generated from CFA curriculum and exam-preparation books with the GraphGen knowledge-driven data-generation method. The dataset and its role in the post-training pipeline are described in Data-Centric… See the full description on the dataset page: https://huggingface.co/datasets/whoisjiji/fin-cfa-graphgen.tabulartext-generation100K<n<1M0 likes117 downloads1mo agoHugging Face28cfahlgren1 /web-fetch-harness-traces Native web fetch harness traces Separate, lightly sanitized native JSONL traces comparing URL-fetch behavior in Claude Code and Codex CLI against: https://huggingface.co/datasets/nyu-mll/glue Captured on 2026-09-15. No shell HTTP client, browser automation, or MCP fetcher was used. Files data/claude-code.jsonl: Claude Code's stream-json events. data/codex.jsonl: Codex's persisted native rollout JSONL. Both traces are exposed together in the default subset and… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/web-fetch-harness-traces.tabularn<1K0 likes105 downloads26d agoHugging Face29Brand1809 /ICAI_SM_CFA ICAI Study Material A personal collection of ICAI study material (MTP, RTP, practice sets and subject modules), stored here so it can be read folder by folder without OTP or password logins. Read it on the website: https://bramhalab.github.io/icai-library-for-we/ Credits The Institute of Chartered Accountants of India (ICAI) for all the study material. The content belongs to ICAI. Hugging Face for free storage. Disclaimer This is an unofficial… See the full description on the dataset page: https://huggingface.co/datasets/Brand1809/ICAI_SM_CFA.document1K<n<10K0 likes98 downloads2d agoHugging Face30250class /cfaibeddocumentn<1K0 likes85 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.