Team Ai
20 results

open-data

genrobot2025 /10Kh-RealOmin-OpenDatagated Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Update Notes:Stage 3 data upload completed. 13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use 3,000+ contributors and 10,000+ real household scenarios with exceptional diversity… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.videoroboticsn>1T277 likes210k downloads6mo agoHugging FaceOpen-Bee /Honey-Data-15M Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality. Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Honey-Data-15M.imageimage-text-to-text10M<n<100M120 likes40k downloads7mo agoHugging FaceJoTalbot /ua-open-data Україна: дзеркало відкритих даних (data.gov.ua) Автоматичне дзеркало публічних наборів data.gov.ua, яке підтримує пайплайн JoTalbot/ukraine. Набори Набір Файлів Джерело Єдиний державний реєстр юридичних осіб, фізичних осіб-підприємців та громадських формувань 6 — Реєстр декларацій родинних зв’язків та доброчесності 14 — Державний судновий реєстр України 9 — Публічні закупівлі на сайті Prozorro 1 — Інформація щодо стану розгляду справ 5 —… See the full description on the dataset page: https://huggingface.co/datasets/JoTalbot/ua-open-data.textn<1K2 likes30k downloads2h agoHugging Faceopendatalab /OmniDocBench OmniDocBench English | 简体中文 OmniDocBench is an evaluation dataset for diverse document parsing in real-world scenarios, with the following characteristics: Diverse Document Types: The evaluation set contains 1651 PDF pages, covering 10 document types, 5 layout types and 5 language types. Coverage includes academic literature, research and financial reports, newspapers, textbooks, exam papers, magazines, handwritten notes, historical documents, and more. Rich Annotations:… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/OmniDocBench.image1K<n<10K113 likes27k downloads4mo agoHugging Facegoogle-research-datasets /nq_open Dataset Card for nq_open Dataset Summary The NQ-Open task, introduced by Lee et.al. 2019, is an open domain question answering benchmark that is derived from Natural Questions. The goal is to predict an English answer string for an input English question. All questions can be answered using the contents of English Wikipedia. Supported Tasks and Leaderboards Open Domain Question-Answering, EfficientQA Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/nq_open.textquestion-answering10K<n<100K36 likes25k downloads3y agoHugging FaceOpenGalaxea /Galaxea-Open-World-Datasetgated Galaxea Open-World Dataset Key Features 500+ hours of real-world mobile manipulation data. All data collected using one uniform robotic embodiment (R1-Lite) for consistency. Fine-grained subtask language annotations (bilingual Chinese/English). Covers residential, kitchen, retail, and officesettings. Dataset in LeRobot v2.1 format. Dataset Structure The dataset is organized as 227 task-level tar.gz archives under the lerobot/ directory. Each… See the full description on the dataset page: https://huggingface.co/datasets/OpenGalaxea/Galaxea-Open-World-Dataset.videon>1T54 likes25k downloads6mo agoHugging Face