Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01m-a-p /FineFineWeb-sample FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.tabulartext-classification100M<n<1B4 likes34k downloads2y agoHugging Face02sail /regmix-data-sample RegMix Data Sample Dataset Description The RegMix Data Sample is a curated dataset derived from the Pile-Uncopyrighted, specifically designed for the RegMix paper (https://huggingface.co/papers/2407.01492). This dataset aims to facilitate the automatic identification of high-performing data mixtures for language model pre-training by formulating it as a regression task. Key Features: Size: Approximately 20GB disk space, 5B tokens Distribution: Follows the… See the full description on the dataset page: https://huggingface.co/datasets/sail/regmix-data-sample.text100K<n<1M2 likes764 downloads2y agoHugging Face03BAAI /RoboBrain-X0-Sample-DataThe dataset is currently being uploaded. Please wait a moment. tabularn<1K2 likes453 downloads1y agoHugging Face04io-intelligence /FoldingTShirt_DualArxR5a_Samples FoldingTShirt_DualArxR5a_Samples 100 real-robot teleoperation episodes for “Fold the T-shirt on the table.” on a DualArxR5a dual-arm robot. Format: raw MCAP (ROS 2 / rosbag2). Source Collected with TeleXperience, IO-AI’s product for real-robot teleoperation and data collection. An operator drives the robot; TeleXperience writes time-aligned RGB, joint commands, joint states, gripper targets, and end-effector poses to MCAP. Product page:… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/FoldingTShirt_DualArxR5a_Samples.textroboticsn<1K0 likes423 downloads2mo agoHugging Face05Icey444 /tsv_sampleThis folder is the canonical export for the sampled evaluation TSVs. Files: HRBench4K.tsv — 300 rows HRBench8K.tsv — 300 rows MathVision_MINI.tsv — 300 rows MathVista_MINI.tsv — 300 rows MMBench_en_dev.tsv — 300 rows MME_RealWorld_Lite.tsv — 300 rows MMMU_val.tsv — 300 rows MMStar.tsv — 300 rows MMVet.tsv — 218 rows POPE.tsv — 300 rows RealworldQA.tsv — 300 rows SEED_Bench.tsv — 300 rows VStarBench.tsv — 191 rows Notes: The nine VLMEvalKit-backed TSVs were regenerated from official source… See the full description on the dataset page: https://huggingface.co/datasets/Icey444/tsv_sample.tabularn<1K0 likes398 downloads3mo agoHugging Face06Mayank6255 /fineweb_2_samples_hq fineweb_2_samples_hq FINEWEB2-HQ dataset Dataset Structure This dataset contains 5 JSONL files with a total size of 26415.31 MB. Files: ukr_Cyrl_sample_001.jsonl: 6163.40 MB ron_Latn_sample_001.jsonl: 3739.69 MB kor_Hang_sample_001.jsonl: 4120.89 MB hin_Deva_sample_001.jsonl: 6681.96 MB heb_Hebr_sample_001.jsonl: 5709.37 MB Usage from datasets import load_dataset dataset = load_dataset("path/to/this/dataset") Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.tabulartext-generation10M<n<100M0 likes321 downloads1y agoHugging Face07MedAgentGym /SampledTrajstext10K<n<100K4 likes288 downloads1y agoHugging Face08LianeMarilin /enterprise-agent-aa-samples Dataset Card Dataset Description Enterprise Agent AA Samples contains three executable enterprise-agent scenarios grounded in frozen public data from UCI, the City of Chicago, and SEC Company Facts. The package combines bilingual task briefs, deterministic stateful environments, normalized tool-use trajectories, source-derived reference outputs, and binary rubric checks. Task: enterprise tool-use and agent-trajectory evaluation Languages: English and Chinese… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/enterprise-agent-aa-samples.texttext-generationn<1K1 likes243 downloads1mo agoHugging Face09HCAI-Lab-GT /dolma3-6t-sample-10000-docs-finance-and-business HCAI-Lab/dolma3-6t-sample-10000-docs-finance-and-business Filename-derived finance_and_business slice of HCAI-Lab/dolma3-6t-sample-10000-docs, pinned to revision 561e73c7e0ad35c04f386bae1e3dd39dfb6755e7. Extraction rule The corpus contains every source .jsonl.zst file whose filename contains the literal segment -finance_and_business-. Source paths and compressed file contents are preserved byte-for-byte. This is a coarse WebOrganizer finance_and_business category… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs-finance-and-business.texttext-generation100K<n<1M0 likes233 downloads2mo agoHugging Face10AxiomSetLabs /stem-scientific-code-sample AxiomSet Labs STEM Scientific-Code Sample A 30-task sample of STEM reasoning and scientific-code problems across five domains. Domains Biology: 6 tasks Chemistry: 6 tasks Materials Science: 6 tasks Mathematics: 6 tasks Physics: 6 tasks Each task contains two subproblems and one main problem, including prompts, scientific background, testing templates, and reference solutions. Files data/sample.jsonl — one task per line; used by the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/AxiomSetLabs/stem-scientific-code-sample.texttext-generationn<1K0 likes174 downloads2mo agoHugging Face11superviselab /multimodal-video-annotation-samples Video Annotation Samples – SuperviseLab SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories. Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.tabularvideo-classificationn<1K1 likes173 downloads6mo agoHugging Face12SHPDRG /zhongyi-ancient-books-sample-10 未分卷中医药古籍样本数据集 本数据集收录 10 种中医药古籍数字化资料,每种古籍保留一个独立目录,包含 Markdown 识别文本、PDF 原件和分卷表。数据来源于“未分卷中医药古籍 / Part1 / 未分卷ZP”内部整理目录,本次发布为其中 10 个书目的开源样本。 数据不是 JSON-only。仓库中既有结构化清单 JSONL,也有古籍全文 Markdown、PDF 原件和 Excel 分卷表。ModelScope 的 configs 指向 metadata/books_manifest.jsonl,用于数据预览和 SDK 加载;完整古籍文件在 data/books/ 下按书目目录保存。 如需了解更多中医药古籍数字化数据、定制语料整理、OCR 处理、知识库建设或批量授权合作,可发送邮件至 zhouhaoran@shujuyoupu.com。 数据集简介 数据类型:中医药古籍数字化文本、PDF 原件、分卷元数据。 书目数量:10 种。 文件数量:30 个核心文件,另含目录占位文件。… See the full description on the dataset page: https://huggingface.co/datasets/SHPDRG/zhongyi-ancient-books-sample-10.documentn<1K0 likes165 downloads2d agoHugging Face13Dynamicresponselabs /JASON-High-Stakes-AI-Evaluation-Samples J.A.S.O.N. Evaluation Sample Previews V01-V29 Dynamic Response Labs develops specialized data and evaluation resources for high-stakes AI. This public preview introduces the breadth of the J.A.S.O.N. Framework through 29 domain volumes spanning financial stress, operational disruption, coercion and exploitation, cyber incidents, healthcare finance, automated systems, and other consequential contexts. The collection contains 31 compact preview records. It is designed to help… See the full description on the dataset page: https://huggingface.co/datasets/Dynamicresponselabs/JASON-High-Stakes-AI-Evaluation-Samples.texttext-generationn<1K0 likes163 downloads13d agoHugging Face14cfahlgren1 /hermes-agent-trace-samples-2026-06-05 Hermes Agent Raw Session Samples Five public-safe raw Hermes Agent session exports generated on 2026-06-05 with the Hermes CLI using Hugging Face Inference Providers. Each file in sessions/ is the exact single-session output from: hermes sessions export sessions/<session_id>.jsonl --session-id <session_id> No derived tables, flattened rows, SQLite database, or formatted JSON copies are included. tabularn<1K0 likes160 downloads4mo agoHugging Face15nvidia /Arena-DROID-Camera-Sensitivity-Workflow-Sample Arena DROID Camera Sensitivity Workflow Sample Dataset Description Arena-DROID-Camera-Sensitivity-Workflow-Sample is a compact set of episode-level results generated by an Isaac Lab-Arena simulation experiment. It lets users run the documented camera sensitivity analysis without first executing the policy-evaluation sweep. The experiment evaluates an OpenPI pi05 policy on a DROID Rubik's-cube pick-and-place task while independently varying the wrist-camera… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-DROID-Camera-Sensitivity-Workflow-Sample.tabularroboticsn<1K1 likes156 downloads1mo agoHugging Face16yoheikobashi /fineweb-edu-100BT-samples-not-in-10BTtext100K<n<1M0 likes155 downloads2y agoHugging Face17fluid-concepts /multimodal-peer-collaboration-samplesgated Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges. ▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.audion<1K1 likes148 downloads22d agoHugging Face18syntheticcfo /synthetic-cfo-sample synthetic cfo public sample: a labelled SAP ECC fiscal year One complete synthetic company for one fiscal year on SAP ECC table shapes, generated from accounting rules alone, with fraud planted and recorded in a ground-truth answer key at the moment it was planted. No real company or personal data at any stage. There is no original: nothing was sampled, masked or anonymised. Engine version 1.11.2. 18 tables, 11,737 rows, 200 labelled fraud records across 18 distinct schemes.… See the full description on the dataset page: https://huggingface.co/datasets/syntheticcfo/synthetic-cfo-sample.tabulartabular-classification10K<n<100K0 likes144 downloads16d agoHugging Face19Desperado1103 /Pandent_Sample PanDent Sample Release This repository provides a 300-case sample of PanDent, including 240 cases from the public-source portion of the dataset and 60 de-identified in-house cases. The sample is released to facilitate inspection of image quality, structured clinical annotations, annotation format, and structure--language correspondence. The complete PanDent dataset contains 9,524 dental panoramic radiographs, including 9,019 public-source cases and 505 in-house cases.… See the full description on the dataset page: https://huggingface.co/datasets/Desperado1103/Pandent_Sample.imageimage-to-textn<1K0 likes133 downloads17d agoHugging Face20alirezaaminzadeh /docflow-invoice-samples-fa DocFlow Invoice Samples — Persian & Bilingual Synthetic invoice dataset for evaluating DocFlow AI field extraction pipelines. Published by Aria AI Engineering Team. Dataset Summary Property Value Samples 50 (synthetic, OCR-friendly) Languages Persian (FA), English (EN) Formats PNG images + JSON annotations Use case Invoice OCR benchmarking, AP automation R&D Synthetic Yes — no real PII Fields Annotated vendor_name… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/docflow-invoice-samples-fa.imageimage-to-textn<1K0 likes129 downloads2mo agoHugging Face21tommypraxis /egocentric-samplesgated Praxis · Egocentric evaluation samples Real-world hand, object and tool interaction across six camera views. Explore task recordings with MCAP sensor data, MP4 previews, camera intrinsics and extrinsics, and linked metadata. Contact Tommy · tommy@praxisrobotics.io for evaluation access and tailored data requirements. At a glance Included in this release Coverage Samples 735 Mapped activity duration 24.16 hours Camera views 6 per sample Primary… See the full description on the dataset page: https://huggingface.co/datasets/tommypraxis/egocentric-samples.textn<1K0 likes129 downloads10d agoHugging Face22bingwow /tripsapien-ai-itinerary-validation-samples TripSapien public data CC BY 4.0 sample data for AI itinerary validation: pasted travel plans, expected validation categories, comparison tables, and prompts that show where TripSapien fits in the AI-travel workflow. TripSapien: https://www.tripsapien.com Canonical methodology: https://www.tripsapien.com/research/ai-itinerary-validation Lineage: TripSapien was previously Tripnostic and ValidaTrip, and originally TripPaste. Why this exists AI travel planners write… See the full description on the dataset page: https://huggingface.co/datasets/bingwow/tripsapien-ai-itinerary-validation-samples.textn<1K0 likes128 downloads3mo agoHugging Face23otellm /gsma-sample-data Telecom Benchmark Suite This repository contains a lightweight benchmarking framework for evaluating Large Language Models (LLMs) on various telecom-related tasks. It is organized into two main folders: data/ — contains the datasets in .json format scripts/ — contains task-specific scripts for prompting and evaluation Each dataset in data/ has a corresponding Python script in scripts/ that defines: A system prompt (guidance text for the LLM) A prompt-building function for… See the full description on the dataset page: https://huggingface.co/datasets/otellm/gsma-sample-data.text1K<n<10K2 likes123 downloads11mo agoHugging Face24pinzhenchen /wmt26-mist-sample Update Log 22 June 2026 (latest) - we updated our data mix because some BELEBELE samples did not have the context. If you downloaded data before 22 June, please download the new version. 16 June 2026 - first version Summary The wmt26-mist-sample is a multilingual mix provided by the WMT26 MIST shared task organizers as a starting point for fine-tuning multilingual LLMs. It contains three types of tasks, to cover same-language and cross-lingual comprehension and… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/wmt26-mist-sample.textquestion-answering10K<n<100K1 likes122 downloads4mo agoHugging Face25SueMintony /CLAMP-Sampled-Continuations-and-Demos CLAMP Sampled Continuations and VLABench Demos This dataset accompanies HLR/CLAMP, the Constrained Language-Action Model Planner. It contains 20 closed-loop VLABench case studies: 10 successful and 10 unsuccessful executions. Each case includes the input image and mask, instruction, prompt and entity metadata, CLAMP and ground-truth plans, evaluation output, execution video, and a frame manifest. Layout demo_dataset/ ├── manifest.json ├── README.md ├──… See the full description on the dataset page: https://huggingface.co/datasets/SueMintony/CLAMP-Sampled-Continuations-and-Demos.imagen<1K0 likes121 downloads2mo agoHugging Face26malaysia-ai /starcoderdata-sampletabular100K<n<1M0 likes118 downloads3y agoHugging Face27emolero /gsma_sample GSMA Open-Telco Sample Dataset Sample data from the GSMA Open-Telco LLM Benchmarks—the first dedicated evaluation framework for assessing LLM performance on telecommunications-specific tasks. Subsets Subset Samples Task telemath 100 Telecom-specific mathematical reasoning (signal processing, link budgets, throughput modeling) teleqna 1,000 Multiple-choice Q&A on telecom standards and domain knowledge telelogs 100 Root cause analysis for 5G network… See the full description on the dataset page: https://huggingface.co/datasets/emolero/gsma_sample.text1K<n<10K1 likes115 downloads10mo agoHugging Face28orai-nlp /MMLU_HT_eu_sample MMLU Human Translated Sample for Basque A subset of 270 samples manually translated to Basque from the MMLU dataset (Hendrycks et al., 2020). The corresponding 250 English samples are also provided. The MMLU dataset is a massive multitask test consisting of multiple-choice questions from various branches of knowledge. The test spans subjects in the humanities, social sciences, hard sciences, and other areas that are important for some people to learn. Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/orai-nlp/MMLU_HT_eu_sample.textn<1K0 likes113 downloads2y agoHugging Face29ssuresh /nemo-stage1-50M-samples NeMo Stage1 Pretraining Dataset - 50M Samples This dataset contains 50 million text samples for NeMo model pretraining (Stage 1). The dataset is organized in chunks for efficient loading and processing. Dataset Details Total Samples: ~50,000,000 Format: JSONL (JSON Lines) Structure: Each sample contains {"id": number, "text": "content"} Chunks: 47 files (chunk_000.jsonl to chunk_046.jsonl) Samples per chunk: ~1,000,000 Language: English Task: Text generation pretraining… See the full description on the dataset page: https://huggingface.co/datasets/ssuresh/nemo-stage1-50M-samples.texttext-generation10M<n<100M0 likes113 downloads1y agoHugging Face30cottonwood-development /synthetic-crm-sample CRM 700 — Free Sample (70 records across 3 tables) This is a free 70-record sample of the full 700-record commercial dataset. Records are split across three relational tables: customers, products, and orders. Foreign-key relationships are intact — every order references a valid customer and a valid product. What's in this sample 10 customer records in customers.jsonl 16 product records in products.jsonl 44 order records in orders.jsonl Referential integrity:… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-crm-sample.tabularothern<1K0 likes110 downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.