Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Mayank6255 /fineweb_2_samples_hq fineweb_2_samples_hq FINEWEB2-HQ dataset Dataset Structure This dataset contains 5 JSONL files with a total size of 26415.31 MB. Files: ukr_Cyrl_sample_001.jsonl: 6163.40 MB ron_Latn_sample_001.jsonl: 3739.69 MB kor_Hang_sample_001.jsonl: 4120.89 MB hin_Deva_sample_001.jsonl: 6681.96 MB heb_Hebr_sample_001.jsonl: 5709.37 MB Usage from datasets import load_dataset dataset = load_dataset("path/to/this/dataset") Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.tabulartext-generation10M<n<100M0 likes321 downloads1y agoHugging Face02superviselab /multimodal-video-annotation-samples Video Annotation Samples – SuperviseLab SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories. Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.tabularvideo-classificationn<1K1 likes173 downloads6mo agoHugging Face03cfahlgren1 /hermes-agent-trace-samples-2026-06-05 Hermes Agent Raw Session Samples Five public-safe raw Hermes Agent session exports generated on 2026-06-05 with the Hermes CLI using Hugging Face Inference Providers. Each file in sessions/ is the exact single-session output from: hermes sessions export sessions/<session_id>.jsonl --session-id <session_id> No derived tables, flattened rows, SQLite database, or formatted JSON copies are included. tabularn<1K0 likes160 downloads4mo agoHugging Face04fluid-concepts /multimodal-peer-collaboration-samplesgated Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges. ▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.audion<1K1 likes148 downloads22d agoHugging Face05alirezaaminzadeh /docflow-invoice-samples-fa DocFlow Invoice Samples — Persian & Bilingual Synthetic invoice dataset for evaluating DocFlow AI field extraction pipelines. Published by Aria AI Engineering Team. Dataset Summary Property Value Samples 50 (synthetic, OCR-friendly) Languages Persian (FA), English (EN) Formats PNG images + JSON annotations Use case Invoice OCR benchmarking, AP automation R&D Synthetic Yes — no real PII Fields Annotated vendor_name… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/docflow-invoice-samples-fa.imageimage-to-textn<1K0 likes129 downloads3mo agoHugging Face06OwnedByDanes /Usenet-Corpus-1980-2013-Threaded-Samples Usenet Corpus 1980–2013 — Threaded (Samples) A small, browsable showcase sample of the Usenet Corpus 1980–2013 — Threaded dataset: Usenet posts reconstructed into conversations via thread_id, thread_position, and thread_depth. This repo is a free preview; the full, commercially-licensed corpus (405.6M posts, 190.8M threads, 102.5B tokens) is at: Full threaded dataset (gated): https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded Cleaned (unthreaded)… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded-Samples.tabulartext-generation10K<n<100K0 likes83 downloads29d agoHugging Face07Taxonomy-Aligned-Conversational-Tutor /TACTBench-Samples TACTBench Demonstration Samples This repository contains five full-context demonstration examples from TACTBench. It does not contain the TACT training set or the remaining hidden TACTBench evaluation set. The samples use the same full-history representation as the benchmark evaluation and illustrate direct correction, error explanation, guided revision, clarification checking, affective feedback, and retry elicitation. Data data/demo.jsonl: five complete… See the full description on the dataset page: https://huggingface.co/datasets/Taxonomy-Aligned-Conversational-Tutor/TACTBench-Samples.tabulartext-generationn<1K0 likes64 downloads17d agoHugging Face08ssergaroo /english-classics-parallel-samples Booklern English classics: parallel samples Paragraph-aligned opening passages of public-domain English classics with a translation into Arabic, Spanish, Japanese, Brazilian Portuguese, Russian, Turkish, Chinese, published by Booklern, a bilingual book reader for learning English through real books. Each book is read on Booklern with a sentence-by-sentence translation under the English, read-aloud audio, a dictionary and vocabulary tools; the rows here are the same opening… See the full description on the dataset page: https://huggingface.co/datasets/ssergaroo/english-classics-parallel-samples.tabulartranslation1K<n<10K0 likes52 downloads2d agoHugging Face09wealthschema /household-samples WealthSchema Synthetic Household Samples 8 synthetic U.S. households, one per life stage plus one high-net-worth: a small free preview of what a complete, internally consistent household financial profile looks like. Each record covers the people, income, assets, debts, insurance, taxes, goals and a monthly trajectory. No real person is behind any of it. Built for teams that need realistic households to design, demo or test financial software: planning tools, robo-advisors… See the full description on the dataset page: https://huggingface.co/datasets/wealthschema/household-samples.tabulartabular-classificationn<1K0 likes50 downloads17d agoHugging Face10alirezaaminzadeh /talentmatch-resume-samples TalentMatch Resume Samples Synthetic enterprise resumes and job descriptions with expert HR rankings for benchmark evaluation. Contents screenings.jsonl — model vs expert ranks per JD/resume pair manifest.json — corpus metadata benchmark_report.json — reproducible metrics Usage import json with open("screenings.jsonl") as f: for line in f: print(json.loads(line)) Built by Aria AI. tabularn<1K0 likes40 downloads2mo agoHugging Face11CL-From-Nothing /rose_code_samples rose_code samples (pass@8 rollouts) vLLM pass@8 samples on the CL-From-Nothing/rose_code train split (23,688 codeforces stdin/stdout problems), scored by the deepcoder verifier (reward=1.0 iff all test cases pass). Qwen3-1.7B/ — student model rollouts. 23,688 questions × 8 samples = 189,504 lines. Qwen3-4B-Thinking-2507/ — teacher model rollouts. Sampling: temperature 0.7, top_p 0.9, max_tokens 16384, 8 samples/question (pass@8). Each cluster file holds a contiguous… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rose_code_samples.tabulartext-generation100K<n<1M0 likes39 downloads4mo agoHugging Face12Jackrong /DeepSeek-v3.1-reasoner-Distilled-math-samples DeepSeek-V3.1 Distillation with NVIDIA Nemotron-Post-Training-Dataset-v2 (Math Subset) The release of DeepSeek-V3.1 has attracted wide attention in the AI community. Its significant improvements in reasoning ability provide a new opportunity to explore optimization of domain-specific models. To investigate the potential of this model in complex mathematical reasoning tasks, I selected the math subset from NVIDIA’s newly released Nemotron-Post-Training-Dataset-v2 as seed problems and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-v3.1-reasoner-Distilled-math-samples.tabularquestion-answeringn<1K1 likes34 downloads1y agoHugging Face13alirezaaminzadeh /contractguard-clause-samples ContractGuard Clause Samples — Bilingual Contract Clause Corpus Synthetic labeled clause corpus for training and evaluating the ContractGuard clause classifier and risk engine. Published by Aria AI Engineering Team. Dataset Summary Property Value Clauses 875 (635 train / 240 test) Languages English (500) + Persian (375) Categories 12 (payment, delivery, warranty, confidentiality, termination, liability, penalty, IP, dispute resolution, force majeure… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/contractguard-clause-samples.tabulartext-classificationn<1K0 likes32 downloads2mo agoHugging Face14alirezaaminzadeh /meetscribe-meeting-samples MeetScribe Meeting Samples Synthetic bilingual (EN/FA) enterprise meeting transcripts with labeled action items. File Language Domain operations_review_en EN Production / maintenance operations_review_en.json EN JSON ASR (Whisper format) safety_board_fa FA HSE safety board procurement_sync_en EN Procurement / RFQ maintenance_planning_fa FA Maintenance planning Usage python scripts/build_dataset.py Generates meetings.jsonl with extracted… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/meetscribe-meeting-samples.tabularsummarizationn<1K0 likes28 downloads2mo agoHugging Face15Lambent /100k-finewebedu-samples2048tTotal tokens in matching entries: 140_782_625 Average tokens per entry: 1407.83 tabular100K<n<1M0 likes23 downloads2y agoHugging Face16Lambent /100k-finewebedu-samples4096tTotal tokens in matching entries: 275_639_417 Average tokens per entry: 2756.39 tabular100K<n<1M0 likes23 downloads2y agoHugging Face17skrishna /filtered_toxic_samplestabular1K<n<10K0 likes17 downloads3y agoHugging Face18alirezaaminzadeh /enterprise-rag-samples OrgMind Enterprise Policy Samples Synthetic organizational policy documents and QA benchmark pairs for OrgMind RAG Studio. Contents File Description chunks.jsonl Semantic chunks with page/paragraph citation metadata qa_pairs.jsonl Curated questions with expected document + keywords benchmark_report.json Reproducible retrieval metrics eval_results.json Benchmark summary (no per-row details) manifest.json Corpus statistics… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/enterprise-rag-samples.tabularn<1K0 likes14 downloads2mo agoHugging Face19Lambent /20k-finewebedu-samples-512tTotal tokens in matching entries: 7595043 Average tokens per entry: 379.75 tabular10K<n<100K0 likes9 downloads2y agoHugging Face20Lambent /1M-finewebedu-samples1024tTotal tokens in matching entries: 732318957 Average tokens per entry: 732.32 tabular1M<n<10M0 likes9 downloads2y agoHugging Face21Lambent /100k-finewebedu-samples-8kttabular100K<n<1M1 likes9 downloads2y agoHugging Face22ermandmand /humanoid-balance-control-samplestabularn<1K0 likes9 downloads10mo agoHugging Face23psdn-ai /code-workflow-samplesgated Code Workflow Samples This sample shows paired developer workflow examples for reviewing prompt, code, test, error, and output structure before scoping a larger code dataset. What This Shows Input-output pairs from practical coding workflows Metadata for task type, files, outputs, and review context A compact view of schema consistency for code-centric examples Dataset Specifications Field Value Modality Code I/O pairs Domain… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/code-workflow-samples.tabulartext-generationn<1K1 likes9 downloads3mo agoHugging Face24skrishna /challenging_toxic_samplestabular1K<n<10K0 likes8 downloads3y agoHugging Face25Lambent /100k-finewebedu-samples256tTotal tokens in matching entries: 19672962 Average tokens per entry: 196.73 tabular100K<n<1M0 likes8 downloads2y agoHugging Face26Lambent /20k-finewebedu-samples-8kttabular10K<n<100K0 likes8 downloads2y agoHugging Face27Lambent /1M-finewebedu-samples4096tTotal tokens in matching entries: 2_698_501_256 tabular100K<n<1M0 likes6 downloads2y agoHugging Face28gohsyi /samples_gsm8k_cot_2024-12-04T19-03-11.038885.jsonltabular1K<n<10K0 likes6 downloads2y agoHugging Face29alirezaaminzadeh /mailroom-doc-samples MailRoom Document Corpus Synthetic enterprise mailroom documents for classification training and evaluation. Contents File Records Description documents.jsonl 48 12 base samples × 4 variants (augmented) manifest.json — Dataset metadata benchmark_report.json — Classifier evaluation metrics Categories leave_request — HR leave/time-off forms correspondence — Official letters and memos invoice — Tax invoices and financial documents… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/mailroom-doc-samples.tabularn<1K0 likes6 downloads2mo agoHugging Face30junyoung01 /samplesimage1K<n<10K0 likes6 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.