Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jat-project /jat-dataset JAT Dataset Dataset Description The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent. Paper: https://huggingface.co/papers/2402.09844 Usage >>> from datasets import load_dataset >>> dataset =… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.imagereinforcement-learning100M<n<1B78 likes202k downloads3y agoHugging Face02AyoubChLin /Company-document-dataset-v2 Company Documents v2 Generation complete: all 13 document types have completed export and upload checkpoints. Synthetic, born-digital business documents rendered from four open sample databases, with exact gold labels: 353,580 PDFs (404,514 pages) of 13 document types in English and French, issued by 60 synthetic companies, each with its own letterhead, numbering and wording. Successor of CompanyDocuments (2,677 PDFs, 4 types). Dataset overview property… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/Company-document-dataset-v2.documentdocument-question-answering100K<n<1M2 likes12k downloads3d agoHugging Face03d0rj /LLaVA-OneVision-Data-ru LLaVA-OneVision-Data-ru Translated lmms-lab/LLaVA-OneVision-Data dataset into Russian language using Google translate. Almost all datasets have been translated, except for the following: ["tallyqa(cauldron,llava_format)", "clevr(cauldron,llava_format)", "VisualWebInstruct(filtered)", "figureqa(cauldron,llava_format)", "magpie_pro(l3_80b_mt)", "magpie_pro(qwen2_72b_st)", "rendered_text(cauldron)", "ureader_ie"] Usage import datasets data =… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/LLaVA-OneVision-Data-ru.imagetext-generation1M<n<10M4 likes2k downloads2y agoHugging Face04doctor-ghezelbaash /dr-saeid-ghezelbaash-entity-data Dr. Saeed Ghezelbash Public Knowledge Graph Dr. Saeed Ghezelbash / دکتر سعید قزلباش is a physician practising aesthetic medicine in Kermanshah, Iran, with Iran Medical Council registration 167430. This physician-authored distribution connects his clinical assessment, botulinum toxin and filler guidance, revision and second-opinion approach, educational media, research authorship and stable professional identity. The current files reproduce the verified final website distribution… See the full description on the dataset page: https://huggingface.co/datasets/doctor-ghezelbaash/dr-saeid-ghezelbaash-entity-data.imagequestion-answering10K<n<100K1 likes1.9k downloads2h agoHugging Face05ysn-rfd /text-dataset-tiny-code-script-py-format USED of tahamajs/medicine_ds_persian for .parquet file USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file USED of Abirate/english_quotes for .jsonl file NEW FILES (05/12/2025) NEW FILES (12/26/2025) NEW FILES (02/15/2026) texttext-generation10K<n<100K3 likes1.7k downloads4mo agoHugging Face06openbmb /RLHF-V-Dataset Dataset Card for RLHF-V-Dataset Project Page | Paper | GitHub Updates [2024.05.28] 📃 Our RLAIF-V paper is accesible at arxiv now! [2024.05.20] 🎉 We release a new feedback dataset, RLAIF-V-Dataset, which is a large-scale diverse-task multimodal feedback dataset constructed using open-source models. You can download the corresponding dataset and models (7B, 12B) now! [2024.04.11] 🔥 Our data is used in MiniCPM-V 2.0, an end-side multimodal large language model that… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/RLHF-V-Dataset.imagetext-generation1K<n<10K79 likes996 downloads2y agoHugging Face07ken-sungmin /propagator-multimodal-pretraining-data Propagator Multimodal Pretraining Data This public dataset contains tokenized multimodal pretraining data prepared for the Propagator model family. It combines language, image-grounded, and speech/audio-token examples into a single training format. This is not a raw text or image browsing dataset. The examples have already been converted into compact binary token frames for model training, with a manifest that records the source groups and file layout. Source Code… See the full description on the dataset page: https://huggingface.co/datasets/ken-sungmin/propagator-multimodal-pretraining-data.texttext-generation0 likes881 downloads4mo agoHugging Face08HabibaAbderrahim /Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset Description This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations. It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.imagetranslationn<1K0 likes795 downloads1y agoHugging Face09mlech26l /liquidrandom-data liquidrandom-data Diverse seed data for ML/LLM training data generation pipelines. Used by the liquidrandom Python package. Dataset Summary This dataset contains 520,080 seed data samples across 24 categories, generated using a hierarchical taxonomy tree approach with LLM-based quality validation and fuzzy deduplication. Data is stored as Parquet with zstd compression. Categories Category Samples File Coding Tasks 30,069… See the full description on the dataset page: https://huggingface.co/datasets/mlech26l/liquidrandom-data.tabulartext-generation100K<n<1M0 likes630 downloads3mo agoHugging Face10Z02Z /steam-games-dataset Overview Information of more than 120,000 games published on Steam. Maintained by Fronkon Games. This dataset has been created with this code (MIT) and use the API provided by Steam, the largest gaming platform on PC. Data is also collected from Steam Spy. Only published games, no DLCs, episodes, music, videos, etc. Here is a simple example of how to parse json information: # Simple parse of the 'games.json' file. import os import json dataset = {} if… See the full description on the dataset page: https://huggingface.co/datasets/Z02Z/steam-games-dataset.imagetext-generation100K<n<1M2 likes386 downloads8mo agoHugging Face11mazafard /portuguese-ocr-datasettask_categories: image-to-text task_ids: optical-character-recognition text-recognition Portuguese OCR Dataset A comprehensive dataset for Portuguese OCR (Optical Character Recognition) generated from classic Portuguese literature with diverse fonts and visual styles. Dataset Description This dataset contains 20000 text images for OCR training, created from Portuguese books from Project Gutenberg. Each image contains a complete Portuguese sentence with proper… See the full description on the dataset page: https://huggingface.co/datasets/mazafard/portuguese-ocr-dataset.imageimage-to-textn<1K2 likes320 downloads1y agoHugging Face12ilsp /panellinies-exams-dataset Dataset Card for Pan-Ex Dataset Pan-Ex is a comprehensive dataset derived from publicly available exam questions and official solutions used for student admission to Higher Education Institutions in Greece via the Panhellenic Examinations (Πανελλαδικές Εξετάσεις). The dataset includes questions with the following features: Subjects: Greek Language, Ancient Greek, History, Latin, Biology, Physics, Chemistry, Computer Science, Economics, and Mathematics. Educational Levels:… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/panellinies-exams-dataset.imagequestion-answering1K<n<10K0 likes315 downloads5d agoHugging Face13ismatsamadov /azerbaijan-court-data Azerbaijan Court System Dataset The most comprehensive open dataset of Azerbaijan's judicial system — 1.64 million structured records and 1.54 million court decision PDFs (~160 GB) covering court decisions, active cases, scheduled hearings, court registries, judges, lawyers, and mediator organizations. Built for AI engineers, legal tech startups, and researchers who need real-world legal data at scale. Quick Start Load with Hugging Face datasets from datasets… See the full description on the dataset page: https://huggingface.co/datasets/ismatsamadov/azerbaijan-court-data.imagetext-classification1M<n<10M2 likes270 downloads6mo agoHugging Face14mayiwen /PaperAudit_Dataset PaperAudit Origin Data This directory contains the original paper data downloaded and preprocessed for the PaperAudit project. The data includes papers from top-tier machine learning conferences with their parsed content, metadata, synthetic error annotations, and review information. PaperAudit Dataset Overview This repository is part of the full PaperAudit Dataset, which includes: PaperAudit_Dataset/ ├── PaperAudit_Origin_Data/ # Original paper data (raw + preprocessed)… See the full description on the dataset page: https://huggingface.co/datasets/mayiwen/PaperAudit_Dataset.imagetext-classification1K<n<10K0 likes266 downloads9mo agoHugging Face15Yfeng0216 /PriOrGen_DATA PriorGen_Data PriorGen 使用的数据独立存放于此目录,代码项目位于 /mnt/datadisk0/PriorGen。数据按用途组织: PriorGen_Data/ ├── features/uni/ # UNI WSI patch features (.pt) ├── features/plip/ # PLIP WSI patch features (.pt) ├── annotations/organs/ # TCGA organ report annotations and images ├── splits/ # train/val/test CSV and label mappings ├── knowledge_banks/ # organ knowledge banks (.pt/.json) ├── anchors/ # PIB anchor features (.pt) ├── feature_labels.csv # 每个 UNI/PLIP… See the full description on the dataset page: https://huggingface.co/datasets/Yfeng0216/PriOrGen_DATA.documentimage-classification10K<n<100K0 likes260 downloads2d agoHugging Face16tracki /companies-dataset Tracki - Synthetic Companies Dataset 10,196 synthetic companies x 18 columns. Every one of the 15 content fields was written by a language model; nothing is rule-generated. Built for the Tracki final project (RUNI - Intro to Data Science): describe a company, and Tracki returns the 3 most similar companies (embeddings). This would be used for subsribing to their social media and websites. Every row is fictional. Where the generator's name prior collided with a real trademark it… See the full description on the dataset page: https://huggingface.co/datasets/tracki/companies-dataset.imagesentence-similarity10K<n<100K1 likes251 downloads1mo agoHugging Face17Abd0r /nanog-cancer-data NanoG - Cancer Foundation-Model Training Data Multimodal cancer corpus for NanoG1 (generative multimodal pretraining). Literature, structured biology, imaging, and grounded <simulate> traces. Hub: Abd0r/nanog-cancer-dataAuthor: Syed Abdur Rehman Ali (@Abd0r) · 17 · independent Train exclusion (hard): NCI-60 is out of training. Skip records whose source / path / text refer to NCI-60. Prefer NCI-ALMANAC, TCGA-sim, Polymathic, PMC/PubMed, TCGA omics, imaging. How… See the full description on the dataset page: https://huggingface.co/datasets/Abd0r/nanog-cancer-data.imagetext-generation10M<n<100M1 likes219 downloads3mo agoHugging Face18pkchwy /letterboxd-all-movie-data Letterboxd Film Dataset This dataset contains a comprehensive collection of 847,209 films from the Letterboxd platform, including movie information, user reviews, and ratings. Dataset Summary Total Films: 847,209 File Size: ~1.12 GB (1,120,572,122 bytes) Format: JSONL (JSON Lines) Language: Primarily English, with some multilingual content Data Structure Each line contains a JSON object with the following fields: { "url":… See the full description on the dataset page: https://huggingface.co/datasets/pkchwy/letterboxd-all-movie-data.imagetext-classification100K<n<1M7 likes214 downloads1y agoHugging Face19community-datasets /hebrew_this_world Dataset Card for HebrewSentiment Dataset Summary HebrewThisWorld is a data set consists of 2028 issues of the newspaper 'This World' edited by Uri Avnery and were published between 1950 and 1989. Released under the AGPLv3 license. Data Annotation: Supported Tasks and Leaderboards Language modeling Languages Hebrew Dataset Structure csv file with "," delimeter Data Instances Sample: { "issue_num": 637, "page_count": 16… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/hebrew_this_world.imagetext-generation1K<n<10K1 likes207 downloads2y agoHugging Face20tellang /yeji-datasets ██████╗ █████╗ ██╗ ██╗ ██████╗ █████╗ ████████╗ █████╗ ██╔══██╗██╔══██╗██║ ██║ ██╔══██╗██╔══██╗╚══██╔══╝██╔══██╗ ██████╔╝███████║██║ █╗ ██║ ██║ ██║███████║ ██║ ███████║ ██╔══██╗██╔══██║██║███╗██║ ██║ ██║██╔══██║ ██║ ██╔══██║ ██║ ██║██║ ██║╚███╔███╔╝ ██████╔╝██║ ██║ ██║ ██║ ██║ ╚═╝ ╚═╝╚═╝ ╚═╝ ╚══╝╚══╝ ╚═════╝ ╚═╝ ╚═╝ ╚═╝ ╚═╝ ╚═╝ ⚡ ORIGINAL DATA ARCHIVE ⚡ > ACCESS GRANTED: YEJI DATASETS > TYPE: Raw Data Collection… See the full description on the dataset page: https://huggingface.co/datasets/tellang/yeji-datasets.imagetext-generationn<1K0 likes190 downloads9mo agoHugging Face21BEE-spoke-data /upvoteweb-posts upvoteweb: posts Posts in upvoteweb. configs [!IMPORTANT]There are several configs representing different permutations of this dataset. Load the relevant config for the task you are interested in. Overview of configs: default: largely unfiltered/unprocessed original data eduscored: the "eduscore" predicted on the text column with huggingface's trained classifier en-clean: filter language for en and language_score for > 0.6. Run clean-text on the text col, preserving… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/upvoteweb-posts.imagetext-generation10M<n<100M1 likes181 downloads10mo agoHugging Face22noel7Y /data-agent-benchmarks LongHorizon Full Data-Agent Benchmarks Companion data artifacts for five complete evaluation tracks: DataSciBench full55 / 167 metric entries DABStep full450 DABStep-Research full100 DSBench Modeling full74 LongDS full68 / 2,225 turns The companion GitHub repository contains processed manifests, evaluation code, historical API ReAct baseline code, download/preparation tools, and the frozen source lock. artifact_manifest.json records every uploaded object's size, SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.imagequestion-answering1 likes178 downloads1mo agoHugging Face23nubuwwat /khatme-nubuwwat-ocr-dataset Khatme-Nubuwwat Urdu OCR Corpus This is a structure-aware, fully OCR'd text dataset of around 215 Urdu Khatme Nubuwat books/volumes (approximately 86,557 pages of text) The majority of the books were in Urdu Nastaliq font, with Arabic Naskh and English text present minimally as well. The text dataset is paried with source-page scans. The books are composed of Nastaliq prose with heavy references to Quran and Hadith. Effort was made to ensure that the OCR pipeline transcribed the… See the full description on the dataset page: https://huggingface.co/datasets/nubuwwat/khatme-nubuwwat-ocr-dataset.imagetext-generation10K<n<100K2 likes157 downloads2mo agoHugging Face24lmms-lab /LLaVA-OneVision-Mid-Data Dataset Card for LLaVA-OneVision Due to unknow reasons, we are unable to process dataset with large amount into required HF format. So we directly upload the json files and image folders (compressed into tar.gz files). You can use the following link to directly download and decompress them. https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Mid-Data/tree/main/evol_instruct We provide the whole details of LLaVA-OneVision Dataset. In this dataset, we include the data splits… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Mid-Data.imagetext-generation100K<n<1M21 likes142 downloads2y agoHugging Face25Conditiondark /engsvg-dataset-v2 EngSVG: engineering drawings with solver-verified ground truth 140,000 tasks over 10,000 procedurally generated 2D engineering designs, each drawing verified by a solver before it was allowed into the dataset. What makes it unusual Every design is certified before use. Trusses are solved with a linear axial finite-element method and admitted only when the equilibrium residual falls below 1e-5; plates pass manufacturing geometry checks. 10,000 of 10,000 designs… See the full description on the dataset page: https://huggingface.co/datasets/Conditiondark/engsvg-dataset-v2.imagetext-generation100K<n<1M0 likes114 downloads16d agoHugging Face26InfiX-ai /InfiGUIAgent-DataThis repository contains trajectory data related to reasoning that was used in the second stage of training in InfiGUIAgent. For more information, please refer to our repo. imagetext-generation1K<n<10K6 likes111 downloads2y agoHugging Face27SoyVitou /62k-images-khmer-printed-dataset 62k Khmer-English Printed Dataset This repository contains a dataset of Khmer and English printed text images for training, validation, and testing. The dataset is stored in parquet format and managed using Git Large File Storage (LFS). Installation Prerequisites Before cloning this repository, make sure you have Git LFS installed: Install Git LFS Linux/macOS:curl -s https://packagecloud.io/install/repositories/github/git-lfs/script.deb.sh | sudo… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/62k-images-khmer-printed-dataset.imagetext-generation10K<n<100K2 likes107 downloads2y agoHugging Face28gllllll /glsl-opengl-educational-dataset GLSL/OpenGL & WebGPU Universal Educational Dataset for AI Training Curated, statically validated, and educational dataset of GLSL, WGSL, and HLSL shaders, OpenGL/WebGL programs, and real-time graphics pipelines. Key Features Multi-Stage Shaders: Vertex, Fragment, Compute, Geometry, Tessellation. Static Validation: Validated against Khronos glslangValidator. Universal Shading Targets: Multi-target transpilation (WGSL, HLSL, MSL). Rich Annotations: Includes… See the full description on the dataset page: https://huggingface.co/datasets/gllllll/glsl-opengl-educational-dataset.imagetext-generationn<1K1 likes103 downloads2mo agoHugging Face29True2456 /iq-terrain-vlm-dataset IQ Terrain VLM Dataset A high-fidelity, mathematically pristine Vision-Language Model (VLM) dataset designed specifically to teach models the procedural graphics and raymarching techniques of Inigo Quilez. Dataset Summary Most coding datasets rely on broadly scraped, often buggy code from GitHub or StackOverflow. This dataset takes a highly targeted approach: Mathematical Ground Truth: All GLSL code and mathematical concepts are sourced directly from Inigo… See the full description on the dataset page: https://huggingface.co/datasets/True2456/iq-terrain-vlm-dataset.imagetext-generation1K<n<10K0 likes95 downloads2mo agoHugging Face30dezoe /pcb-defect-multi-modal-dataset 印制电路板焊接及装联缺陷多模态数据集 面向 PCB 焊接与装联缺陷识别、资料辅助根因分析和维修建议生成的中文多模态研究数据集。数据集将光学、扫描声学显微(SAM)和 X 射线等缺陷图像与来源标签、中文任务文本,以及可用的检测测量和试验工况关联,支持视觉语言模型的数据准备、指令微调实验和检索增强问答。 版本 v3 包含 5,000 条图像/缺陷区域样本、15,000 条任务指令,覆盖焊点连接质量、焊料分布与形态、焊接层结构及元件装联位置等方面的 14 类目标缺陷: 目标缺陷 含义与标注范围 虚焊 焊接界面未形成可靠的冶金结合,可能表现为接触不稳或间歇导通。 冷焊 热过程不足导致焊料未充分熔融或未形成合格连接,结合来源的热过程与外观信息标注。 少锡 焊料量或有效焊脚不足,焊接部位的填充或覆盖不充分。 多锡 焊料过量,形成过大的焊脚或明显堆积。 连锡 焊料连接本应相互隔离的端子或焊盘,形成焊料桥接。 开焊 应有焊接连接的部位缺少完整连接,或连接部位发生分离。 润湿不良… See the full description on the dataset page: https://huggingface.co/datasets/dezoe/pcb-defect-multi-modal-dataset.imageimage-to-text1K<n<10K1 likes95 downloads6d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.