Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01m-a-p /FineFineWeb-sample FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.tabulartext-classification100M<n<1B4 likes34k downloads2y agoHugging Face02ksolovev /fine-news-sample Fine-News Sample Fine-News Sample contains 1,000,000 rows sampled from the Fine-News corpus. The sample covers all 117 capture months and 388 language-and-script labels in that corpus. Each selected row preserves its article text, source metadata, and sampling weight. The full corpus derives from INFINI-NEWS, which extracts article text from Common Crawl News web archives. At a glance Measure Value Rows 1,000,000 Distinct document IDs 1,000,000 Sum… See the full description on the dataset page: https://huggingface.co/datasets/ksolovev/fine-news-sample.texttext-generation1M<n<10M0 likes2.7k downloads2d agoHugging Face03TheFinAI /dolma3_300B_samplegated Dolma 3 — 300B-token sample 🌐 The Fin AI Pretraining / reference corpus released by The Fin AI. Source: Dolma 3 mix (AllenAI) — https://huggingface.co/allenai. Source A ~300B-token sample of AllenAI's Dolma 3 mix; Dolma is released under ODC-BY 1.0. Structure Rows: 187,823,645 Columns: source, date, text, token_count, category Quick Start from datasets import load_dataset ds = load_dataset("TheFinAI/dolma3_300B_sample"… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/dolma3_300B_sample.tabulartext-generation100M<n<1B0 likes1.1k downloads3d agoHugging Face04voidful /agent-sft-stitch-zh-tts-taste-codec-chat-sample Gemma 4 E2B Taste-S multi-turn codec SFT This dataset contains 37,362 complete Traditional Chinese agent dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers 229,434 synthesized speech segments, approximately 520.5 hours of audio before codec extraction. Every assistant speech segment is represented without Gemma native audio tags: <SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY> The first assistant output starts immediately with <SAY>. [SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.tabulartext-generation10K<n<100K0 likes380 downloads3mo agoHugging Face05Mayank6255 /fineweb_2_samples_hq fineweb_2_samples_hq FINEWEB2-HQ dataset Dataset Structure This dataset contains 5 JSONL files with a total size of 26415.31 MB. Files: ukr_Cyrl_sample_001.jsonl: 6163.40 MB ron_Latn_sample_001.jsonl: 3739.69 MB kor_Hang_sample_001.jsonl: 4120.89 MB hin_Deva_sample_001.jsonl: 6681.96 MB heb_Hebr_sample_001.jsonl: 5709.37 MB Usage from datasets import load_dataset dataset = load_dataset("path/to/this/dataset") Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.tabulartext-generation10M<n<100M0 likes321 downloads1y agoHugging Face06superviselab /multimodal-video-annotation-samples Video Annotation Samples – SuperviseLab SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories. Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.tabularvideo-classificationn<1K1 likes173 downloads6mo agoHugging Face07WhissleAI /egocentric-activity-sample Egocentric Activity Sample Dataset A small-scale egocentric (first-person) video dataset with Ego4D-style annotations, designed for quick prototyping and experimentation with egocentric video understanding tasks. Dataset Summary Metric Value Video clips 19 Total duration ~9.5 minutes Resolution 960x540 (540p) FPS 30 Narrations 99 NLQ queries 57 Moment annotations 19 FHO actions 57 Total size ~54 MB Activities Covered… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/egocentric-activity-sample.tabularvideo-classificationn<1K0 likes143 downloads6mo agoHugging Face08aklein4 /fineweb-edu-sample-10BT-shuffled 📚 FineWeb-Edu (Shuffled) The samples in HuggingFaceFW/fineweb-edu don't appear to be fully shuffled, leading to oscillating loss curves. This dataset contains a shuffled version of the sample-10BT sample from HuggingFaceFW/fineweb-edu. Shuffling was performed using the following script: import datasets data = datasets.load_dataset( "HuggingFaceFW/fineweb-edu", "sample-10BT", split="train", streaming=False, ) data_shuffled = data.shuffle(seed=42)… See the full description on the dataset page: https://huggingface.co/datasets/aklein4/fineweb-edu-sample-10BT-shuffled.tabulartext-generation1M<n<10M1 likes131 downloads1y agoHugging Face09akumalondon /Rail_Freight_Logistics_Company_Email_Archive_Sample Ukrainian Rail-Freight Correspondence Corpus (Sample) Real operational correspondence from a working freight forwarding business, and the documents attached to it — consignment notes, service acts, invoices, wagon manifests. Not scraped, not synthetic, and never published anywhere before. This is a de-identified sample released for evaluation. It is drawn from a larger private archive; see Full archive below. Published by Akuma London · akumalondon.com Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.tabulartext-generation1K<n<10K0 likes108 downloads28d agoHugging Face10mashu-data /reddit-comments-sample Reddit Comment Trees Sample — Initial snapshot Initial sample: 43,913 posts and 59,874 comments across three communities. This is a selected, structurally checked snapshot, with incomplete subreddit coverage. It is not a complete three-month archive. Overview Posts and associated comments from r/LocalLLaMA, r/wallstreetbets, and r/SkincareAddiction. The requested post window is June 1–August 31, 2026 in Asia/Shanghai, with UTC bounds 2026-05-31 16:00:00… See the full description on the dataset page: https://huggingface.co/datasets/mashu-data/reddit-comments-sample.tabulartext-generation100K<n<1M0 likes101 downloads28d agoHugging Face11TheFinAI /dolma3_300B_sample_shuffledgated dolma3_300B_sample_shuffled 🌐 The Fin AI Source: Dolma 3 mix (AllenAI) — https://huggingface.co/allenai. A ~300B-token sample of AllenAI's Dolma 3 mix; Dolma is released under ODC-BY 1.0. License: ODC BY. Global row-level shuffle of TheFinAI/dolma3_300B_sample. Source data uses per-row Bernoulli sampling (p ≈ 0.0506) from allenai/dolma3_mix-6T-1025-7B to produce ~300B cl100k tokens preserving the original Dolma3 mix ratios. However the source parquets cluster records by… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/dolma3_300B_sample_shuffled.tabulartext-generation100M<n<1B0 likes98 downloads3d agoHugging Face12agentlans /m-a-p-FineFineWeb-sample Unofficial m-a-p/FineFineWeb Sample This dataset is a processed, lightweight sample of the original m-a-p/FineFineWeb, a comprehensive corpus designed for fine-grained domain web text studies. Sampling Methodology To create this subset, the following processing steps were taken: Selection: 100 random .jsonl files were chosen from the original dataset. Extraction: 10,000 rows were downloaded per selected file. Processing: The extracted rows were combined and… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/m-a-p-FineFineWeb-sample.tabulartext-classification1M<n<10M0 likes96 downloads4mo agoHugging Face13agentlans /thomas-yanxin-MT-SFT-ShareGPT-sample MT-SFT-ShareGPT Sample Dataset This dataset provides a sample of the thomas-yanxin/MT-SFT-ShareGPT dataset with English and Chinese subsets. Dataset Contents train.jsonl: Contains 1/10 of the original data, shuffled EN.jsonl: English conversations from train.jsonl ZH.jsonl: Chinese conversations from train.jsonl Each row represents a conversation with an optional system message, followed by human and GPT turns. Columns from the original dataset are preserved, with… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/thomas-yanxin-MT-SFT-ShareGPT-sample.tabulartext-generation1M<n<10M0 likes89 downloads10mo agoHugging Face14OwnedByDanes /Usenet-Corpus-1980-2013-Threaded-Samples Usenet Corpus 1980–2013 — Threaded (Samples) A small, browsable showcase sample of the Usenet Corpus 1980–2013 — Threaded dataset: Usenet posts reconstructed into conversations via thread_id, thread_position, and thread_depth. This repo is a free preview; the full, commercially-licensed corpus (405.6M posts, 190.8M threads, 102.5B tokens) is at: Full threaded dataset (gated): https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded Cleaned (unthreaded)… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded-Samples.tabulartext-generation10K<n<100K0 likes83 downloads29d agoHugging Face15BoomQ /fineweb-edu-2016-qwen2-sample FineWeb-Edu 2016 / Qwen2 Consistency sample — not the completed year. Documents: 900. Actual recounted Qwen2 tokens: 937,977. Source: HuggingFaceFW/fineweb-edu, revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9. The input inventory covers 9 crawl directories. date is the integer crawl year 2016, not an article publication date. Original text is preserved without cleaning, normalization, deduplication, truncation, or added formatting. Source token counts are not used. Optional… See the full description on the dataset page: https://huggingface.co/datasets/BoomQ/fineweb-edu-2016-qwen2-sample.tabulartext-generationn<1K0 likes76 downloads1mo agoHugging Face16agentlans /HuggingFaceFW-finewiki-sample HuggingFaceFW/finewiki sample A uniformly randomized subset of HuggingFaceFW/finewiki, created to provide a smaller and more manageable dataset for analysis, fine-tuning, and benchmarking. Overview This sample includes Wikipedia articles from languages with more than one million pages. Sampling is performed uniformly at random instead of alphabetically to ensure unbiased representation. Language Inclusion Criteria Languages were selected based on page count and… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/HuggingFaceFW-finewiki-sample.tabulartext-generation100K<n<1M0 likes72 downloads11mo agoHugging Face17Tinuade /common-crawl-docx-sample Common Crawl DOCX Sample A sample of normalized text extracted from DOCX records in Common Crawl. Source Common Crawl release: CC-MAIN-YYYY-NN Source index: Common Crawl URL Index Pipeline: marin-community/marin Pipeline revision: REPLACE_WITH_GIT_SHA Records were selected using declared DOCX MIME type, detected DOCX MIME type, or a .docx URL suffix. Only successful, non-truncated index records were eligible. Processing The pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/Tinuade/common-crawl-docx-sample.tabulartext-generation1K<n<10K0 likes71 downloads24d agoHugging Face18dzcorpora /algerian-darija-customer-service-samplegated Algerian Darija customer messages — stratified sample 500 spontaneous Algerian Darija messages, written by real customers, drawn from a first-party corpus of 869,166 customer messages. Every message here is unique after normalization, de-identified, and typed by a human — nothing elicited, translated, scraped or generated. Algerian Darija (ISO 639-3 arq) is spoken by around 45 million people and is one of the worst-covered varieties in current language models. For scale: PADIC… See the full description on the dataset page: https://huggingface.co/datasets/dzcorpora/algerian-darija-customer-service-sample.tabulartext-generation1K<n<10K2 likes69 downloads26d agoHugging Face19Taxonomy-Aligned-Conversational-Tutor /TACTBench-Samples TACTBench Demonstration Samples This repository contains five full-context demonstration examples from TACTBench. It does not contain the TACT training set or the remaining hidden TACTBench evaluation set. The samples use the same full-history representation as the benchmark evaluation and illustrate direct correction, error explanation, guided revision, clarification checking, affective feedback, and retry elicitation. Data data/demo.jsonl: five complete… See the full description on the dataset page: https://huggingface.co/datasets/Taxonomy-Aligned-Conversational-Tutor/TACTBench-Samples.tabulartext-generationn<1K0 likes64 downloads17d agoHugging Face20finaleads /french-corpus-llm-sample French Corpus LLM — Sample 500 (v1.4.0) FINALEADS LLC builds compliance-ready training datasets for French regulated industries. We turn 2.66 billion tokens of French finance, regulatory, and economic open data into audit-trailed, pseudonymized, AI Act Article 10-documented shares — so foundation model and regtech teams can ship into European enterprises without a data-lineage gap. This is a public sample of 500 stratified documents drawn from the French Premium Web Corpus v1.4.0… See the full description on the dataset page: https://huggingface.co/datasets/finaleads/french-corpus-llm-sample.documenttext-generationn<1K1 likes59 downloads5mo agoHugging Face21killdevil111 /DataShield-Sample-Risk DataShield This dataset releases sample-level risk scores for DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment, accepted to the EMNLP Main Conference. For the method, code, and complete documentation, see the DataShield GitHub repository. Dataset configurations Configuration Source dataset Rows dolly15k databricks/databricks-dolly-15k 15,011 alpaca52k tatsu-lab/alpaca 51,974 from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/killdevil111/DataShield-Sample-Risk.tabulartext-generation10K<n<100K0 likes52 downloads2mo agoHugging Face22wealthschema /household-samples WealthSchema Synthetic Household Samples 8 synthetic U.S. households, one per life stage plus one high-net-worth: a small free preview of what a complete, internally consistent household financial profile looks like. Each record covers the people, income, assets, debts, insurance, taxes, goals and a monthly trajectory. No real person is behind any of it. Built for teams that need realistic households to design, demo or test financial software: planning tools, robo-advisors… See the full description on the dataset page: https://huggingface.co/datasets/wealthschema/household-samples.tabulartabular-classificationn<1K0 likes50 downloads16d agoHugging Face23dFusionAILabs /sample-fusion-intelligence-traces Sample Fusion Intelligence Traces Structured AI reasoning traces from dFusion's Fusion Intelligence system. Each record captures a complete agentic workflow: a real user query on a domain-specific topic, the full message chain including system prompts, tool calls, search results, intermediate reasoning steps, and a final synthesized answer — along with human feedback. These are not synthetic benchmarks. They are traces from real queries submitted by real users on live financial… See the full description on the dataset page: https://huggingface.co/datasets/dFusionAILabs/sample-fusion-intelligence-traces.tabularquestion-answeringn<1K0 likes48 downloads7mo agoHugging Face24debolut /amazon-reviews-2023-all-beauty-sample Amazon Reviews 2023 – All_Beauty (Sampled) This dataset is a sampled subset of the McAuley-Lab/Amazon-Reviews-2023 All_Beauty category, prepared for the YZM2022 Data Mining homework (Assoc. Prof. Dr. Arzu Kakisim). Sampling strategy Source: full All_Beauty reviews (701K) and metadata (112K items). 3-core filtering (each user and item has at least 3 interactions, iterated to convergence). Cap to the most recent 60 000 interactions, re-applied 3-core. Metadata restricted… See the full description on the dataset page: https://huggingface.co/datasets/debolut/amazon-reviews-2023-all-beauty-sample.tabulartext-classification10K<n<100K0 likes44 downloads5mo agoHugging Face25agentlans /ClimbMix-sample Unofficial NVIDIA Nemotron-ClimbMix (Subsampled) This dataset is a curated, subsampled version of OptimalScale/ClimbMix, which itself is a detokenized version of NVIDIA's official pretraining dataset, nvidia/Nemotron-ClimbMix. It is designed for researchers and developers looking for a smaller, well-shuffled slice of the Nemotron pretraining data for quick experimentation, testing, or ablation studies. Processing Method To create this streamlined version, the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/ClimbMix-sample.tabulartext-generation1M<n<10M0 likes44 downloads4mo agoHugging Face26CL-From-Nothing /rose_code_samples rose_code samples (pass@8 rollouts) vLLM pass@8 samples on the CL-From-Nothing/rose_code train split (23,688 codeforces stdin/stdout problems), scored by the deepcoder verifier (reward=1.0 iff all test cases pass). Qwen3-1.7B/ — student model rollouts. 23,688 questions × 8 samples = 189,504 lines. Qwen3-4B-Thinking-2507/ — teacher model rollouts. Sampling: temperature 0.7, top_p 0.9, max_tokens 16384, 8 samples/question (pass@8). Each cluster file holds a contiguous… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rose_code_samples.tabulartext-generation100K<n<1M0 likes39 downloads4mo agoHugging Face27dlab-spp /reflection-sample-2k SPP Reflection 2k Sample A 2,000-row sample (seed 42) of dlab-spp/reflection-10m, in the identical format, for quick inspection of the data from Synthetic Persona Pretraining (SPP): Alignment from Token Zero. 📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero 📦 Full dataset: dlab-spp/reflection-10m (~10M documents). Each row pairs a pretraining document with a synthetic, value-laden reflection (first- and third-person) grounded in a value constitution.… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-sample-2k.tabulartext-generation1K<n<10K0 likes38 downloads2mo agoHugging Face28Jackrong /DeepSeek-v3.1-reasoner-Distilled-math-samples DeepSeek-V3.1 Distillation with NVIDIA Nemotron-Post-Training-Dataset-v2 (Math Subset) The release of DeepSeek-V3.1 has attracted wide attention in the AI community. Its significant improvements in reasoning ability provide a new opportunity to explore optimization of domain-specific models. To investigate the potential of this model in complex mathematical reasoning tasks, I selected the math subset from NVIDIA’s newly released Nemotron-Post-Training-Dataset-v2 as seed problems and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-v3.1-reasoner-Distilled-math-samples.tabularquestion-answeringn<1K1 likes34 downloads1y agoHugging Face29YMinglai /GSM-DC-Dataset-Sample GSM-DC Test Dataset This dataset contains the test set for GSM-DC (Grade School Math with Distractor Chains), a synthetic math reasoning dataset with controlled complexity. Dataset Details Total Problems: 6300 Operation Counts (OP): 16-22 (out-of-distribution test set) Problem Types: Graph-based mathematical reasoning problems Noise Levels: Light, Medium, Hard (distractor difficulty) Dataset Structure Each problem in all_problems.json contains: problem_text:… See the full description on the dataset page: https://huggingface.co/datasets/YMinglai/GSM-DC-Dataset-Sample.tabularquestion-answering1K<n<10K0 likes33 downloads9mo agoHugging Face30alirezaaminzadeh /meetscribe-meeting-samples MeetScribe Meeting Samples Synthetic bilingual (EN/FA) enterprise meeting transcripts with labeled action items. File Language Domain operations_review_en EN Production / maintenance operations_review_en.json EN JSON ASR (Whisper format) safety_board_fa FA HSE safety board procurement_sync_en EN Procurement / RFQ maintenance_planning_fa FA Maintenance planning Usage python scripts/build_dataset.py Generates meetings.jsonl with extracted… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/meetscribe-meeting-samples.tabularsummarizationn<1K0 likes28 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.