Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ryanmarten /OpenThoughts-1k-sample [!NOTE] We have released a paper for OpenThoughts! See our paper here. Open-Thoughts-1k-sample This is a 1k sample of the OpenThoughts-114k dataset. Open synthetic reasoning dataset with high-quality examples covering math, science, code, and puzzles! Inspect the content with rich formatting with Curator Viewer. Available Subsets default subset containing ready-to-train data used to finetune the OpenThinker-7B and OpenThinker-32B models: ds =… See the full description on the dataset page: https://huggingface.co/datasets/ryanmarten/OpenThoughts-1k-sample.text1K<n<10K63 likes1.1m downloads1y agoHugging Face02m-a-p /FineFineWeb-sample FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.tabulartext-classification100M<n<1B4 likes34k downloads2y agoHugging Face03axolotl-ai-co /evolkit-logprobs-pipeline-75k-v2-sampletextn<1K1 likes15k downloads2y agoHugging Face04OriginFlow-AI /origindata-preview-samplegated OriginData (preview sample) OriginData is the world's first large-scale, real-world dataset combining hand pose and force annotations. Spanning 28 domains and 1,170 real-world tasks, it captures how human hands interact with the physical world, providing force, pose, and semantic annotations supported by high-precision multimodal calibration for embodied AI. This preview contains 100.62 hours across 7,717 episodes, delivered in LeRobot v3.0 format with stereo RGB video, hand… See the full description on the dataset page: https://huggingface.co/datasets/OriginFlow-AI/origindata-preview-sample.tabularrobotics10M<n<100M15 likes15k downloads1d agoHugging Face05Samuelsantos777 /psg-audio-v3-unofficial-mirror PSG-Audio v3 — Unofficial Complete Mirror Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset. This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community. Overview PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.textaudio-classificationn<1K1 likes11k downloads3mo agoHugging Face06moonshine-ai /audio_samples_1kaudio0 likes9k downloads7mo agoHugging Face07Research-EAI /essential-web-1t-sample-fdc-partitioned 🌐 Essential-Web: FDC Level-2 Partitioned Dataset 📋 Dataset Description This dataset contains a 1 trillion token sample from Essential-Web, partitioned by Free Decimal Correspondence (FDC) level-2 categories. Essential-Web is a 24-trillion-token web dataset with extensive document-level metadata designed to enable rapid dataset curation through SQL-like filtering. 🔍 Free Decimal Correspondence (FDC) The FDC taxonomy is an open classification system… See the full description on the dataset page: https://huggingface.co/datasets/Research-EAI/essential-web-1t-sample-fdc-partitioned.text100M<n<1B5 likes8k downloads1y agoHugging Face08sam-paech /wildchat_creative_writing_annotated_10ktabular10K<n<100K2 likes7.2k downloads9mo agoHugging Face09knkarthick /samsum Dataset Card for SAMSum Corpus Dataset Description Links Homepage: hhttps://arxiv.org/abs/1911.12237v2 Repository: https://arxiv.org/abs/1911.12237v2 Paper: https://arxiv.org/abs/1911.12237v2 Point of Contact: https://huggingface.co/knkarthick Dataset Summary The SAMSum dataset contains about 16k messenger-like conversations with summaries. Conversations were created and written down by linguists fluent in English. Linguists were asked to… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/samsum.textsummarization10K<n<100K44 likes5.5k downloads1y agoHugging Face10agentlans /common-crawl-sample Common Crawl sample A small unofficial random subset of the famous Common Crawl dataset. 60 random segment WET files were downloaded from Common Crawl on 2024-05-12. Lines between 500 and 5000 characters long (inclusive) were kept. Only unique texts were kept. No other filtering. Languages Each text was assigned to one of the language codes using the GCLD3 Python package. The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.texttext-generation1M<n<10M8 likes5.1k downloads2y agoHugging Face11Samsoup /cosmos_qatext10K<n<100K0 likes5k downloads3y agoHugging Face12olm /olm-CC-MAIN-2022-21-sampling-ratio-0.14775510204 Dataset Card for OLM May 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the May 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes4.7k downloads4y agoHugging Face13RMT-team /babilong-1k-samples BABILong (1000 samples) : a long-context needle-in-a-haystack benchmark for LLMs Preprint is on arXiv and code for LLM evaluation is available on GitHub. BABILong Leaderboard with top-performing long-context models. bAbI + Books = BABILong BABILong is a novel generative benchmark for evaluating the performance of NLP models in processing arbitrarily long documents with distributed facts. It contains 9 configs, corresponding to different sequence lengths in tokens: 0k… See the full description on the dataset page: https://huggingface.co/datasets/RMT-team/babilong-1k-samples.text10K<n<100K4 likes4.6k downloads2y agoHugging Face14EleutherAI /rpj-v2-sampleThis is a mirror of the sample-10B subset of RedPajama-Data-V2 which we have re-uploaded in order to resolve issues with the original download script. Getting Started RedPajama-V2 is an open dataset for training large language models. The dataset includes over 100B text documents coming from 84 CommonCrawl snapshots and processed using the CCNet pipeline. Out of these, there are 30B documents in the corpus that additionally come with quality signals. In addition, we also provide the… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/rpj-v2-sample.texttext-generation1M<n<10M2 likes4.3k downloads2y agoHugging Face15stanford-cs336 /owt-sampleThese files were created with the following script: from datasets import load_dataset from tqdm import tqdm import io dataset = load_dataset("Skylion007/openwebtext")['train'] split_dataset = dataset.train_test_split(train_size=2400000, test_size=60000, seed=0) with io.open('data/owt_train.txt','w') as fopen: listout = [] for data in tqdm(split_dataset['train']): listout.append(data['text']+'<|endoftext|>') if len(listout) > 1000: _ =… See the full description on the dataset page: https://huggingface.co/datasets/stanford-cs336/owt-sample.text10M<n<100M7 likes3.4k downloads3y agoHugging Face16hngl /swebench-verified-sample-100-Qwen3-30B-evaltext10K<n<100K0 likes3.1k downloads11mo agoHugging Face17samuelstevens /BirdSet BirdSet (Mirror) This dataset repository is a convenience mirror of BirdSet. Attribution Please credit the original BirdSet authors and resources: Original dataset card: https://huggingface.co/datasets/DBD-research-group/BirdSet Original project repository: https://github.com/DBD-research-group/BirdSet Paper: https://arxiv.org/abs/2403.10380 Citation If you use this dataset, please cite the original BirdSet paper:… See the full description on the dataset page: https://huggingface.co/datasets/samuelstevens/BirdSet.audioaudio-classification1M<n<10M0 likes2.9k downloads8mo agoHugging Face18ksolovev /fine-news-sample Fine-News Sample Fine-News Sample contains 1,000,000 rows sampled from the Fine-News corpus. The sample covers all 117 capture months and 388 language-and-script labels in that corpus. Each selected row preserves its article text, source metadata, and sampling weight. The full corpus derives from INFINI-NEWS, which extracts article text from Common Crawl News web archives. At a glance Measure Value Rows 1,000,000 Distinct document IDs 1,000,000 Sum… See the full description on the dataset page: https://huggingface.co/datasets/ksolovev/fine-news-sample.texttext-generation1M<n<10M0 likes2.7k downloads2d agoHugging Face19olm /olm-CC-MAIN-2017-22-sampling-ratio-0.16178770949 Dataset Card for OLM May 2017 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the May 2017 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M0 likes2.6k downloads4y agoHugging Face20samatv256 /jev-decisions-v1 Jev Decisions v1 12M canonical agent-decision records for tool selection, routing, value prediction, completion, and local agent control. Jev Decisions v1 is a derived, decision-oriented corpus built from public agent trajectory datasets. It canonicalizes heterogeneous trajectories into a shared learning interface: state + available candidate decisions -> target / outcome / eligibility mini-Jev is a related open decision-model project. Its earlier v1 baseline was not trained on… See the full description on the dataset page: https://huggingface.co/datasets/samatv256/jev-decisions-v1.textreinforcement-learning10M<n<100M12 likes2.6k downloads12d agoHugging Face21olm /olm-CC-MAIN-2022-27-sampling-ratio-0.16142697881 Dataset Card for OLM June/July 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the June/July 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes2.6k downloads4y agoHugging Face22SamuelChien821 /devopsbench-100 DevOpsBench-100 DevOpsBench-100 is a synthetic long-horizon software-engineering / SRE agent benchmark: 100 tasks over one executable world ("NovaCart", a mid-size e-commerce SaaS) with 72 SQLite tables, 1451 seeded rows, a 38-file monorepo with 417 commits, and 97 MCP tools spanning a first-party engineering stack (tickets, PRs, CI, deployments, canaries, migrations, feature flags, metrics, alerts, incidents, chat, knowledge base) plus deliberately disagreeing vendor-shaped… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/devopsbench-100.texttext-generationn<1K0 likes2.5k downloads1mo agoHugging Face23ai4bharat /samanantar Dataset Card for Samanantar Dataset Summary Samanantar is the largest publicly available parallel corpora collection for Indic language: Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, Telugu. The corpus has 49.6M sentence pairs between English to Indian Languages. Supported Tasks and Leaderboards [More Information Needed] Languages Samanantar contains parallel sentences between English (en) and 11 Indic… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/samanantar.texttext-generation10M<n<100M46 likes2.4k downloads2y agoHugging Face24sam-paech /livecodebench-code_generation_litetext1K<n<10K0 likes2.2k downloads1y agoHugging Face25sammshen /lmcache-agentic-traces LMCache Agentic Dataset Collection A curated dataset collection of 787 multi-turn agentic LLM sessions (24,881 total LLM iterations) designed for benchmarking stateful LLM serving systems. Every session exhibits at least 5 turns with prefix growth and builds to at least 10K tokens of context — making it ideal for evaluating tiered KV Cache solutions like LMCache. Motivation Modern LLM agents (coding assistants, research agents, tool-calling systems) make dozens of… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/lmcache-agentic-traces.tabulartext-generation10K<n<100K17 likes2.1k downloads4mo agoHugging Face26olm /olm-CC-MAIN-2022-33-sampling-ratio-0.20 Dataset Card for OLM August 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 20% of the August 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes2.1k downloads4y agoHugging Face27stablellama /Krea-2-Raw_samplesThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw. NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model. Raw is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on Krea 2 Raw Quality testing Data source The images were created in ComfyUI with the bf16 version of… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples.tabulartext-to-image1K<n<10K0 likes2k downloads1mo agoHugging Face28bezzam /vibevoice_samplesSource: https://github.com/vibevoice-community/VibeVoice/tree/main/demo audion<1K0 likes2k downloads2mo agoHugging Face29nanotron /minipile_100_samplestextn<1K2 likes1.9k downloads2y agoHugging Face30SamuelChien821 /salesbench-100 SalesBench-100 SalesBench-100 is a synthetic long-horizon sales-agent benchmark with 100 original workflows across Salesforce, HubSpot, Gong, and a seeded evidence room. Each task begins with a high-level employee request and has its own authored causal rule and provider transition. Identity, operating facts, authority, governed policy, live-system indexes, and exceptions are separated so no mounted business asset publishes a selected option or precomputed change. Every task has… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/salesbench-100.documenttext-generationn<1K1 likes1.9k downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.