Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01genrobot2025 /10Kh-RealOmin-OpenDatagated Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Update Notes:Stage 3 data upload completed. 13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use 3,000+ contributors and 10,000+ real household scenarios with exceptional diversity… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.videoroboticsn>1T277 likes210k downloads6mo agoHugging Face02Open-Bee /Honey-Data-15M Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality. Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Honey-Data-15M.imageimage-text-to-text10M<n<100M120 likes40k downloads7mo agoHugging Face03JoTalbot /ua-open-data Україна: дзеркало відкритих даних (data.gov.ua) Автоматичне дзеркало публічних наборів data.gov.ua, яке підтримує пайплайн JoTalbot/ukraine. Набори Набір Файлів Джерело Єдиний державний реєстр юридичних осіб, фізичних осіб-підприємців та громадських формувань 6 — Реєстр декларацій родинних зв’язків та доброчесності 14 — Державний судновий реєстр України 9 — Публічні закупівлі на сайті Prozorro 1 — Інформація щодо стану розгляду справ 5 —… See the full description on the dataset page: https://huggingface.co/datasets/JoTalbot/ua-open-data.2 likes30k downloads1m agoHugging Face04opendatalab /OmniDocBench OmniDocBench English | 简体中文 OmniDocBench is an evaluation dataset for diverse document parsing in real-world scenarios, with the following characteristics: Diverse Document Types: The evaluation set contains 1651 PDF pages, covering 10 document types, 5 layout types and 5 language types. Coverage includes academic literature, research and financial reports, newspapers, textbooks, exam papers, magazines, handwritten notes, historical documents, and more. Rich Annotations:… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/OmniDocBench.image1K<n<10K113 likes27k downloads4mo agoHugging Face05google-research-datasets /nq_open Dataset Card for nq_open Dataset Summary The NQ-Open task, introduced by Lee et.al. 2019, is an open domain question answering benchmark that is derived from Natural Questions. The goal is to predict an English answer string for an input English question. All questions can be answered using the contents of English Wikipedia. Supported Tasks and Leaderboards Open Domain Question-Answering, EfficientQA Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/nq_open.textquestion-answering10K<n<100K36 likes25k downloads3y agoHugging Face06OpenGalaxea /Galaxea-Open-World-Datasetgated Galaxea Open-World Dataset Key Features 500+ hours of real-world mobile manipulation data. All data collected using one uniform robotic embodiment (R1-Lite) for consistency. Fine-grained subtask language annotations (bilingual Chinese/English). Covers residential, kitchen, retail, and officesettings. Dataset in LeRobot v2.1 format. Dataset Structure The dataset is organized as 227 task-level tar.gz archives under the lerobot/ directory. Each… See the full description on the dataset page: https://huggingface.co/datasets/OpenGalaxea/Galaxea-Open-World-Dataset.videon>1T54 likes25k downloads6mo agoHugging Face07open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.tabulartext-retrieval1M<n<10M14 likes24k downloads2h agoHugging Face08TacVerse /opendataLanguage: English (current) · 中文 Representative frames from TacVerse's bimanual demonstrations. Collected with XTac-UMI-G1 grippers, released as LeRobot datasets. TacVerse Open Data Collection of 122 LeRobot v3.0 task datasets — 17,690 episodes, 370.2 hours, 40.0M frames, ~145 GB. Each subfolder is a standalone LeRobot dataset (meta/info.json, data/, videos/). Collection timestamps have been removed from titles and metadata. Every frame carries six synchronized video… See the full description on the dataset page: https://huggingface.co/datasets/TacVerse/opendata.tabularrobotics10M<n<100M7 likes18k downloads23d agoHugging Face09AnnaZhang /waymo_open_dataset_v_1_4_35 likes12k downloads1y agoHugging Face10griffinlabs /Galaxea-Open-World-Dataset-LeRobot-v3.0Galaxea Open-World Dataset taken from OpenGalaxea/Galaxea-Open-World-Dataset, converted to LeRobot Datasets v3.0 format using lerobot.datasets.v30.convert_dataset_v21_to_v30. Missing subsets The subset Boil_The_Water_20250714_006 is missing due to the original files having some episodes at 62 fps, which causes the conversion script to crash with an error. The subset Put_The_Items_Into_The_Storage_Box_20250929_002_007 is missing due to it having 7 DoF arms rather than 6 DoF.… See the full description on the dataset page: https://huggingface.co/datasets/griffinlabs/Galaxea-Open-World-Dataset-LeRobot-v3.0.roboticsn>1T2 likes6.3k downloads5mo agoHugging Face11opendatalab /Sci-Base Sci-Base: The Largest AI-Ready Scientific Foundation Dataset 🌌 The Sciverse Data Foundation Sciverse is a comprehensive, multi-layered scientific data foundation designed to provide the ultimate data infrastructure for the AI for Science (AI4S) community. As scientific research becomes increasingly data-driven, Sciverse supplies the essential, high-quality data resources required to build robust scientific knowledge systems and accelerate research. Sciverse… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/Sci-Base.text1M<n<10M40 likes6.1k downloads5mo agoHugging Face12Gramscii-IT /european-open-data-catalogue Open Data catalogue This repository publishes independently versioned metadata and licensed source snapshots: A discovery catalogue with 69000 dataset entries, covering the providers listed in the discovery table below. 3 independently pinned availability indexes with 911,795 joint combinations across 35 datasets, built from complete source responses within the explicitly declared scope. Licensed Cruscotto source snapshots, stored separately from the metadata, preserve the… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/european-open-data-catalogue.text100K<n<1M1 likes5.8k downloads13h agoHugging Face13opendatalab /AICC🔧 🔧 Our New-Gen Html Parser MinerU-HTML Now Realease! AICC: AI-ready Common Crawl Dataset Paper | Project page News [2025-12-24] 🔥 CC-MinerU-Code Updated! We have updated our specialized high-quality code dataset CC-MinerU-Code, containing 4.58M samples, also extracted from the full Common Crawl corpus. Download: CC-MinerU-Code Each record includes language, code_language, and Markdown-formatted content with fenced code blocks. Here is a sample: {… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/AICC.texttext-generation1B<n<10B115 likes5.5k downloads10mo agoHugging Face14open-reaction-database /ord-data ord-data Getting the Data The datasets live under data/ and are stored with Git LFS. LFS reads are redirected to the Hugging Face mirror via .lfsconfig, so dataset objects are fetched from Hugging Face's CDN rather than from GitHub's shared (and limited) LFS bandwidth. This is automatic — you do not need to configure anything. Option 1: Clone the repository git clone https://github.com/open-reaction-database/ord-data.git With Git LFS installed… See the full description on the dataset page: https://huggingface.co/datasets/open-reaction-database/ord-data.text1M<n<10M8 likes4.8k downloads1mo agoHugging Face15mlfoundations /open_lm_test_data_v20 likes4.6k downloads3y agoHugging Face16OpenDataArena /Spark-234K Spark-234K: Skeleton-Guided Scientific Reasoning from Large-Scale Literature 🎉 Accepted to EMNLP 2026 Findings! Spark-234K is a scientific reasoning dataset containing 234K question-answer pairs synthesized from frontier scientific literature. Instead of directly generating QA pairs from full papers, SPARK first distills each paper into a compact reasoning skeleton—preserving its central claim, supporting evidence, quantitative relations, assumptions, and boundary… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/Spark-234K.texttext-generation100K<n<1M122 likes4.6k downloads1mo agoHugging Face17opendatalab /SlimPajama-Meta-rater Annotated SlimPajama Dataset Dataset Description This dataset contains the first fully annotated SlimPajama dataset with comprehensive quality metrics for data-centric large language model research. The dataset includes approximately 580 billion tokens from the training set of the original SlimPajama dataset, annotated across 25 different quality dimensions. Note: This dataset contains only the training set portion of the original SlimPajama dataset, which is why the… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater.tabulartext-generation10M<n<100M7 likes4.4k downloads1y agoHugging Face18OpenDataArena /OpenDataArena-scored-data-2603 OpenDataArena-scored-data-2603 This repository provides a scored SFT dataset collection currently featuring 63 high-quality instruction-following datasets with nearly 25 million samples. The core value lies in its 30-dimensional scoring: every sample has been evaluated on metrics such as IFD, PPL, Deita_Quality, and 27 others, enabling fine-grained data selection for filtering, curriculum learning, and mixture optimization. Key features: 30 metrics per sample — From lexical… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/OpenDataArena-scored-data-2603.text10M<n<100M9 likes4.1k downloads5mo agoHugging Face19Artyom2003 /open-webui-data1 likes4.1k downloads7h agoHugging Face20ShawnChamberlain /open-economic-quant-research-data Open Economic & Quant Research Data Versioned research content for CasualLab, Macroeconomics, Mortgage Rate Lock-In and Housing Market Dynamics, Tariff Incidence, and Order Flow to Price Impact, including project code, publishable data, fixtures, reports, tests, and reproducibility documentation. Repository structure CasualLab/: causal inference and policy-simulation research content. Macroeconomics/: vintage-aware forecasting and public-source adapter research… See the full description on the dataset page: https://huggingface.co/datasets/ShawnChamberlain/open-economic-quant-research-data.documenttabular-classificationn<1K0 likes3.9k downloads2mo agoHugging Face21data-is-better-together /open-image-preferences-v1 Open Image Preferences Prompt: Anime-style concept art of a Mayan Quetzalcoatl biomutant, dystopian world, vibrant colors, 4K. Image 1 Image 2 Prompt: 8-bit pixel art of a blue knight, green car, and glacier landscape in Norway, fantasy style, colorful and detailed. Image 1… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/open-image-preferences-v1.imagetext-to-image1K<n<10K31 likes3.8k downloads2y agoHugging Face22Goku-OpenLab /open-models-prompt-datasets 🖼️ Open Models Prompt Dataset 🖼️ The ultimate open models image prompt dataset (10GB+). 5400+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for various open-source AI image models and the resulting generated images. The entire dataset exceeds 10GB and contains 5400+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/open-models-prompt-datasets.image1K<n<10K2 likes3.5k downloads3mo agoHugging Face23open-world-agents /example_datasetDataset preview available at: https://huggingface.co/spaces/open-world-agents/visualize_dataset videon<1K0 likes3.3k downloads1y agoHugging Face24open-travel /japan-travel-mcp-data Japan Travel MCP — Data The runtime data for the japan-travel-mcp Model Context Protocol server. Comprehensive Japanese travel data for AI agents, built from public official sources, covering all 47 prefectures and 1,938 local government entities. Code lives on GitHub: github.com/ookami0210/japan-travel-mcp Data lives here. The npm package downloads this dataset on first run. Why this dataset exists Japan's tourism information — created to reach the world — is… See the full description on the dataset page: https://huggingface.co/datasets/open-travel/japan-travel-mcp-data.text-retrieval100K<n<1M0 likes3.1k downloads2h agoHugging Face25Open-Bee /Bee-Training-Data-Stage2 Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality. Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Bee-Training-Data-Stage2.imageimage-to-text10M<n<100M6 likes2.7k downloads7mo agoHugging Face26opendatalab /ChartVerse-SFT-1.8MChartVerse-SFT-1800K is an extended large-scale chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the opendatalab/ChartVerse project. For more details about our method, datasets, and full model series, please visit our Project Page. This dataset contains all verified correct samples without failure rate filtering. Unlike SFT-600K which excludes easy samples (r=0), SFT-1800K includes the complete set of truth-anchored QA pairs for maximum coverage and scale.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/ChartVerse-SFT-1.8M.imagevisual-question-answering1M<n<10M139 likes2.7k downloads8mo agoHugging Face27snehasis19 /opendatalab-experimental-nmr-peaks OpenDataLab Experimental NMR Peaks Dataset Dataset Description This dataset contains experimental NMR (Nuclear Magnetic Resonance) peak sequences extracted from the OpenDataLab experimental spectra database. The dataset includes both H-NMR and C-NMR peak sequences for chemical compounds, along with their SMILES representations and molecular formulas. Dataset Summary Total Samples: 533,595 compounds Batches: 333 batch files Data Source: Experimental spectra… See the full description on the dataset page: https://huggingface.co/datasets/snehasis19/opendatalab-experimental-nmr-peaks.textother100K<n<1M0 likes2.6k downloads8mo agoHugging Face28OpenDataArena /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M65 likes2.5k downloads8mo agoHugging Face29leeoxiang /open-audio-data1 likes2.3k downloads1mo agoHugging Face30hac541309 /open-lid-datasetThis dataset is built from the open source data accompanying "An Open Dataset and Model for Language Identification" (Burchell et al., 2023) The repository containing the actual data can be found here : https://github.com/laurieburchell/open-lid-dataset. The license for this recreation itself follows the original upstream dataset as GPLv3+. However, individual datasets within it follow each of their own licenses. The "src" column lists the sources. "lang" column lists the language code in… See the full description on the dataset page: https://huggingface.co/datasets/hac541309/open-lid-dataset.text100M<n<1B4 likes2.2k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.