Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01m-a-p /FineFineWeb-sample FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.tabulartext-classification100M<n<1B4 likes34k downloads2y agoHugging Face02agentlans /common-crawl-sample Common Crawl sample A small unofficial random subset of the famous Common Crawl dataset. 60 random segment WET files were downloaded from Common Crawl on 2024-05-12. Lines between 500 and 5000 characters long (inclusive) were kept. Only unique texts were kept. No other filtering. Languages Each text was assigned to one of the language codes using the GCLD3 Python package. The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.texttext-generation1M<n<10M8 likes5.1k downloads2y agoHugging Face03EleutherAI /rpj-v2-sampleThis is a mirror of the sample-10B subset of RedPajama-Data-V2 which we have re-uploaded in order to resolve issues with the original download script. Getting Started RedPajama-V2 is an open dataset for training large language models. The dataset includes over 100B text documents coming from 84 CommonCrawl snapshots and processed using the CCNet pipeline. Out of these, there are 30B documents in the corpus that additionally come with quality signals. In addition, we also provide the… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/rpj-v2-sample.texttext-generation1M<n<10M2 likes4.3k downloads2y agoHugging Face04samforce /structural-cognition-papers Structural Cognition Papers (English) English-language research papers on Structural Cognition Theory by Lin Xiaohei (林小黑, Lin Xiaohei,非林业/林小) — the first discoverer of structural cognition. 锚点: 公众号「今晚狗蛋看局 | 环缝互换 | 对立面互换 | 危机 | Huanfeng Interchange | Loop-Gap Interchange」| Gitee论文库 | Zenodo DOI全集 | GitHub Pages品牌页 Overview A unified structural framework for cognition, physics, AI, and social systems. Four axioms (canonical): 结构先于语义 / 耦合即认知 / 观察者自指 / 退相干离散台阶 +… See the full description on the dataset page: https://huggingface.co/datasets/samforce/structural-cognition-papers.documenttext-generationn<1K0 likes3.4k downloads4d agoHugging Face05ksolovev /fine-news-sample Fine-News Sample Fine-News Sample contains 1,000,000 rows sampled from the Fine-News corpus. The sample covers all 117 capture months and 388 language-and-script labels in that corpus. Each selected row preserves its article text, source metadata, and sampling weight. The full corpus derives from INFINI-NEWS, which extracts article text from Common Crawl News web archives. At a glance Measure Value Rows 1,000,000 Distinct document IDs 1,000,000 Sum… See the full description on the dataset page: https://huggingface.co/datasets/ksolovev/fine-news-sample.texttext-generation1M<n<10M0 likes2.7k downloads2d agoHugging Face06SamuelChien821 /devopsbench-100 DevOpsBench-100 DevOpsBench-100 is a synthetic long-horizon software-engineering / SRE agent benchmark: 100 tasks over one executable world ("NovaCart", a mid-size e-commerce SaaS) with 72 SQLite tables, 1451 seeded rows, a 38-file monorepo with 417 commits, and 97 MCP tools spanning a first-party engineering stack (tickets, PRs, CI, deployments, canaries, migrations, feature flags, metrics, alerts, incidents, chat, knowledge base) plus deliberately disagreeing vendor-shaped… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/devopsbench-100.texttext-generationn<1K0 likes2.5k downloads1mo agoHugging Face07ai4bharat /samanantar Dataset Card for Samanantar Dataset Summary Samanantar is the largest publicly available parallel corpora collection for Indic language: Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, Telugu. The corpus has 49.6M sentence pairs between English to Indian Languages. Supported Tasks and Leaderboards [More Information Needed] Languages Samanantar contains parallel sentences between English (en) and 11 Indic… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/samanantar.texttext-generation10M<n<100M46 likes2.4k downloads2y agoHugging Face08sammshen /lmcache-agentic-traces LMCache Agentic Dataset Collection A curated dataset collection of 787 multi-turn agentic LLM sessions (24,881 total LLM iterations) designed for benchmarking stateful LLM serving systems. Every session exhibits at least 5 turns with prefix growth and builds to at least 10K tokens of context — making it ideal for evaluating tiered KV Cache solutions like LMCache. Motivation Modern LLM agents (coding assistants, research agents, tool-calling systems) make dozens of… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/lmcache-agentic-traces.tabulartext-generation10K<n<100K17 likes2.1k downloads4mo agoHugging Face09SamuelChien821 /salesbench-100 SalesBench-100 SalesBench-100 is a synthetic long-horizon sales-agent benchmark with 100 original workflows across Salesforce, HubSpot, Gong, and a seeded evidence room. Each task begins with a high-level employee request and has its own authored causal rule and provider transition. Identity, operating facts, authority, governed policy, live-system indexes, and exceptions are separated so no mounted business asset publishes a selected option or precomputed change. Every task has… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/salesbench-100.documenttext-generationn<1K1 likes1.9k downloads1mo agoHugging Face10olm /olm-CC-MAIN-2022-49-sampling-ratio-olm-0.15114822547 Dataset Card for OLM November/December 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the November/December 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabulartext-generation10M<n<100M3 likes1.7k downloads4y agoHugging Face11SamuelChien821 /hubbench HubBench 1.4.0 One Blobfish-authored, oracle-proven benchmark family per Harbor Hub professional-domain cluster. Every task is an employee decision worked over a dependent chain of evidence — never a lookup — against mock stateful tools over an isolated SQLite world. The agent reaches the world only through its public surfaces (MCP over streamable HTTP, a terminal tool CLI, a REST API, and a web console); a deterministic verifier (HubScore) grades the finished world from… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/hubbench.documentquestion-answeringn<1K0 likes1.4k downloads1mo agoHugging Face12asenion-ai /sampled-local-resumes sampled-local-resumes This dataset contains synthetic resume data sampled from local folders (20% sample from each folder). License This dataset is released under the Apache License 2.0. Please see the LICENSE and NOTICE files for details. Attribution Copyright 2025 Fairly AI Inc. dba Asenion This dataset includes data released by Fairly AI Inc. dba Asenion under the Apache License, Version 2.0. You may obtain a copy of the License at:… See the full description on the dataset page: https://huggingface.co/datasets/asenion-ai/sampled-local-resumes.text-generation1K<n<10K0 likes1.3k downloads1y agoHugging Face13bertin-project /mc4-es-sampled50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "https://commoncrawl.org", and processed by AllenAI.texttext-generation1M<n<10M2 likes1.3k downloads4y agoHugging Face14TheFinAI /dolma3_300B_samplegated Dolma 3 — 300B-token sample 🌐 The Fin AI Pretraining / reference corpus released by The Fin AI. Source: Dolma 3 mix (AllenAI) — https://huggingface.co/allenai. Source A ~300B-token sample of AllenAI's Dolma 3 mix; Dolma is released under ODC-BY 1.0. Structure Rows: 187,823,645 Columns: source, date, text, token_count, category Quick Start from datasets import load_dataset ds = load_dataset("TheFinAI/dolma3_300B_sample"… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/dolma3_300B_sample.tabulartext-generation100M<n<1B0 likes1.1k downloads3d agoHugging Face15DynaMath /DynaMath_Sample Dataset Card for DynaMath [💻 Github] [🌐 Homepage][📖 Preprint Paper] Dataset Details 🔈 Notice DynaMath is a dynamic benchmark with 501 seed question generators. This dataset is only a sample of 10 variants generated by DynaMath. We encourage you to use the dataset generator on our github site to generate random datasets to test. 🌟 About DynaMath The rapid advancements in Vision-Language Models (VLMs) have shown significant potential in tackling… See the full description on the dataset page: https://huggingface.co/datasets/DynaMath/DynaMath_Sample.imagemultiple-choice1K<n<10K9 likes726 downloads2y agoHugging Face16samuki-hf /thinking-rollouts thinking-rollouts Unconstrained rollouts from thinking (chain-of-thought) models on DS-1000 and LiveCodeBench, CoT saved verbatim alongside the final answer. Format per genlm/rollouts issue #5; schema is a superset of temperature-sweep-data. Hive-partitioned Parquet, thinking_mode folded into the model tag: rollouts/domain=<dataset>/model=<tag>/temp=<temp>/data.parquet (tags like qwen3-8b-think, qwen3-1.7b-nothink). 100 samples/instance. Columns: model, thinking_mode, temp… See the full description on the dataset page: https://huggingface.co/datasets/samuki-hf/thinking-rollouts.tabulartext-generation10M<n<100M2 likes663 downloads2mo agoHugging Face17Sambarboi /climbmix-tokenized-20480-diloco ClimbMix, retokenized and shuffled for three-worker DiLoCo This is a document-preserving, three-way split of NVIDIA's Nemotron-ClimbMix, retokenized with a 20,480-entry byte-level BPE tokenizer. Each document ends in <|endoftext|>. The Arrow IPC streams use transparent Zstandard buffer compression. A deterministic whole-shard holdout is shared by every worker for validation and is excluded from training. Training part Documents Tokens Files Compressed size 000 15,709… See the full description on the dataset page: https://huggingface.co/datasets/Sambarboi/climbmix-tokenized-20480-diloco.text-generation1M<n<10M0 likes649 downloads3mo agoHugging Face18samuki-hf /tool-use Tool-use rollouts (Qwen3, think/nothink) Tool-augmented code-generation rollouts: Qwen3-8B and Qwen3-14B, each in thinking and non-thinking mode, on DS-1000, LiveCodeBench (Python) and Multilingual-LCB (OCaml). During generation the model can call a run_code tool (up to 3 rounds) that executes its candidate in a sandbox (pinned DS-1000 env / LCB public tests / OCaml compile+publics) and returns real output. Design: 100 samples per instance at temperature 0.6 (bf16, vLLM)… See the full description on the dataset page: https://huggingface.co/datasets/samuki-hf/tool-use.tabulartext-generation1M<n<10M1 likes598 downloads2mo agoHugging Face19sarvamai /samvaad-hi-v1100k high-quality conversations in English, Hindi, and Hinglish curated exclusively with an Indic context. texttext-generation100K<n<1M69 likes591 downloads2y agoHugging Face20Samarth0710 /reviewarena ReviewArena ReviewArena accompanies the NeurIPS Evaluations & Datasets submission ReviewArena: A Large-Scale Cross-Conference Dataset and Benchmark for LLM Peer Review. This release is a large, multi-conference corpus of peer-reviewed papers + their reviews + author rebuttals + acceptance decisions, harvested from OpenReview and aligned with OCR'd full-text markdown of each paper PDF where available. 51,529 papers 196,099 reviews 558,785 OCR'd PDF pages (markdown inlined… See the full description on the dataset page: https://huggingface.co/datasets/Samarth0710/reviewarena.tabulartext-generation10K<n<100K3 likes576 downloads3mo agoHugging Face21FredyRivera-dev /LLaDA-Sample-10BT Dataset: LLaDA-Sample-10BTBase: HuggingFaceFW/fineweb (subset sample-10BT)Purpose: Training LLaDA (Large Language Diffusion Models) Preprocessing Tokenizer: GSAI-ML/LLaDA-8B-Instruct Chunking: Up to 4,096 tokens per chunk (1% of chunks randomly sized between 1–4,096 tokens) Noisy masking: Applied with noise factor ε = 1×10⁻³ Fields per chunk (PyTorch tensors): input_ids noisy_input_ids mask t (time scalar) Statistics Total chunks: ~2,520,000… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/LLaDA-Sample-10BT.text-generation1B<n<10B2 likes570 downloads4mo agoHugging Face22MiniLLM /pile-diff_samp-qwen_1.8B-qwen_104M-r0.5This repository contains the refined pre-training corpus from the paper MiniPLM: Knowledge Distillation for Pre-Training Language Models. Code: https://github.com/thu-coai/MiniPLM text-generation0 likes429 downloads2y agoHugging Face23JamesConley /fineweb-sample-22.95B-512 FineWeb-Sample-22.95B-512 Dataset Description This dataset contains approximately 22.95 billion tokens (22,948,244,480 tokens) sampled from the FineWeb-Edu dataset. Each text sample is capped at a maximum of 512 tokens. Dataset Statistics Total Tokens: ~22.95B (22,948,244,480) Max Tokens per Sample: 512 Max Characters per Sample: 5,120 (10 chars/token estimate) Source Dataset: FineWeb-Edu 350BT Random Seed: 42 Dataset Structure The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/JamesConley/fineweb-sample-22.95B-512.texttext-generation10M<n<100M0 likes419 downloads11mo agoHugging Face24xu-song /cc100-samplesThe cc100-samples is a subset which contains first 10,000 lines of cc100. Languages To load a language which isn't part of the config, all you need to do is specify the language code in the config. You can find the valid languages in Homepage section of Dataset Description: https://data.statmt.org/cc-100/ E.g. dataset = load_dataset("cc100-samples", lang="en") VALID_CODES = [ "am", "ar", "as", "az", "be", "bg", "bn", "bn_rom", "br", "bs", "ca", "cs", "cy", "da", "de", "el"… See the full description on the dataset page: https://huggingface.co/datasets/xu-song/cc100-samples.texttext-generation1M<n<10M6 likes409 downloads2y agoHugging Face25FredyRivera-dev /LLaDA-Sample-ES Dataset: LLaDA-Sample-ES Base: crscardellino/spanish_billion_words Purpose: Training LLaDA (Large Language Diffusion Models) Preprocessing Tokenizer: GSAI-ML/LLaDA-8B-Instruct Chunking: Up to 4,096 tokens per chunk (1% of chunks randomly sized between 1–4,096 tokens) Noisy masking: Applied with noise factor ε = 1×10⁻³ Fields per chunk (PyTorch tensors): input_ids noisy_input_ids mask t (time scalar) Statistics Total chunks: ~ 652,089 Shards: 65… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/LLaDA-Sample-ES.text-generation100M<n<1B1 likes406 downloads4mo agoHugging Face26SamuelChien821 /counselbench-100 CounselBench-100 CounselBench-100 v3.2.5 is a synthetic legal-work benchmark with 100 authored matters across ten practice workflows. Every task has a natural employee request, a 97-asset evidence room, twelve portfolio decisions, 5–9 supported actions, 3–7 evidence holds, and a distinct deep multi-provider MCP trajectory. The answer is not preclassified in the evidence. Each portfolio item requires an immutable identity join, an operative-authority and revision lookup, a… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/counselbench-100.documentquestion-answeringn<1K0 likes397 downloads1mo agoHugging Face27BEE-spoke-data /TxT360-5M-sample-en BEE-spoke-data/TxT360-5M-sample-en english only sample from LLM360/TxT360: min length 256 GPT-4 tokens max length 24576 GPT-4 tokens GPT-4 tiktoken token count: token_count count 5.000000e+06 mean 1.003614e+03 std 1.424231e+03 min 2.570000e+02 25% 4.020000e+02 50% 6.220000e+02 75% 1.050000e+03 max 2.457400e+04 Total count: 5018.07 M tokens texttext-generation10M<n<100M3 likes389 downloads10mo agoHugging Face28voidful /agent-sft-stitch-zh-tts-taste-codec-chat-sample Gemma 4 E2B Taste-S multi-turn codec SFT This dataset contains 37,362 complete Traditional Chinese agent dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers 229,434 synthesized speech segments, approximately 520.5 hours of audio before codec extraction. Every assistant speech segment is represented without Gemma native audio tags: <SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY> The first assistant output starts immediately with <SAY>. [SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.tabulartext-generation10K<n<100K0 likes380 downloads3mo agoHugging Face29samirun974 /eurobaskettexttext-generation1K<n<10K0 likes350 downloads2y agoHugging Face30samaya-ai /FrontierFinancegated FrontierFinance: A benchmark for measuring the frontier intelligence of finance AI agents. arXiv report | website | grading code 1. Overview Investors use AI agents across their entire workflow — idea screening and discovery, company and market research, financial data collection and modeling, portfolio tracking, and catalyst monitoring. Measuring how well an AI system performs across this range is both important and hard: a benchmark must be broad enough to span… See the full description on the dataset page: https://huggingface.co/datasets/samaya-ai/FrontierFinance.texttext-generationn<1K12 likes350 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.