datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jat-dataset
JAT Dataset
Dataset Description
The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent.
Paper: https://huggingface.co/papers/2402.09844
Usage
>>> from datasets import load_dataset
>>> dataset =… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.Company-document-dataset-v2
Company Documents v2
Generation complete: all 13 document types have completed export and upload checkpoints.
Synthetic, born-digital business documents rendered from four open sample databases, with exact gold
labels: 353,580 PDFs (404,514 pages) of 13 document types in
English and French, issued by 60 synthetic companies,
each with its own letterhead, numbering and wording. Successor of
CompanyDocuments (2,677 PDFs, 4 types).
Dataset overview
property… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/Company-document-dataset-v2.LLaVA-OneVision-Data-ru
LLaVA-OneVision-Data-ru
Translated lmms-lab/LLaVA-OneVision-Data dataset into Russian language using Google translate.
Almost all datasets have been translated, except for the following:
["tallyqa(cauldron,llava_format)", "clevr(cauldron,llava_format)", "VisualWebInstruct(filtered)", "figureqa(cauldron,llava_format)", "magpie_pro(l3_80b_mt)", "magpie_pro(qwen2_72b_st)", "rendered_text(cauldron)", "ureader_ie"]
Usage
import datasets
data =… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/LLaVA-OneVision-Data-ru.dr-saeid-ghezelbaash-entity-data
Dr. Saeed Ghezelbash Public Knowledge Graph
Dr. Saeed Ghezelbash / دکتر سعید قزلباش is a physician practising aesthetic medicine in Kermanshah, Iran, with Iran Medical Council registration 167430. This physician-authored distribution connects his clinical assessment, botulinum toxin and filler guidance, revision and second-opinion approach, educational media, research authorship and stable professional identity.
The current files reproduce the verified final website distribution… See the full description on the dataset page: https://huggingface.co/datasets/doctor-ghezelbaash/dr-saeid-ghezelbaash-entity-data.text-dataset-tiny-code-script-py-format
USED of tahamajs/medicine_ds_persian for .parquet file
USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file
USED of Abirate/english_quotes for .jsonl file
NEW FILES (05/12/2025)
NEW FILES (12/26/2025)
NEW FILES (02/15/2026)
RLHF-V-Dataset
Dataset Card for RLHF-V-Dataset
Project Page | Paper | GitHub
Updates
[2024.05.28] 📃 Our RLAIF-V paper is accesible at arxiv now!
[2024.05.20] 🎉 We release a new feedback dataset, RLAIF-V-Dataset, which is a large-scale diverse-task multimodal feedback dataset constructed using open-source models. You can download the corresponding dataset and models (7B, 12B) now!
[2024.04.11] 🔥 Our data is used in MiniCPM-V 2.0, an end-side multimodal large language model that… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/RLHF-V-Dataset.propagator-multimodal-pretraining-data
Propagator Multimodal Pretraining Data
This public dataset contains tokenized multimodal pretraining data prepared for the Propagator model family. It combines language, image-grounded, and speech/audio-token examples into a single training format.
This is not a raw text or image browsing dataset. The examples have already been converted into compact binary token frames for model training, with a manifest that records the source groups and file layout.
Source Code… See the full description on the dataset page: https://huggingface.co/datasets/ken-sungmin/propagator-multimodal-pretraining-data.Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset
Description
This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations.
It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.liquidrandom-data
liquidrandom-data
Diverse seed data for ML/LLM training data generation pipelines.
Used by the liquidrandom Python package.
Dataset Summary
This dataset contains 520,080 seed data samples across 24 categories,
generated using a hierarchical taxonomy tree approach with LLM-based quality validation
and fuzzy deduplication. Data is stored as Parquet with zstd compression.
Categories
Category
Samples
File
Coding Tasks
30,069… See the full description on the dataset page: https://huggingface.co/datasets/mlech26l/liquidrandom-data.steam-games-dataset
Overview
Information of more than 120,000 games published on Steam. Maintained by Fronkon Games.
This dataset has been created with this code (MIT) and use the API provided by Steam, the largest gaming platform on PC. Data is also collected from Steam Spy.
Only published games, no DLCs, episodes, music, videos, etc.
Here is a simple example of how to parse json information:
# Simple parse of the 'games.json' file.
import os
import json
dataset = {}
if… See the full description on the dataset page: https://huggingface.co/datasets/Z02Z/steam-games-dataset.portuguese-ocr-datasettask_categories:
image-to-text
task_ids:
optical-character-recognition
text-recognition
Portuguese OCR Dataset
A comprehensive dataset for Portuguese OCR (Optical Character Recognition) generated from classic Portuguese literature with diverse fonts and visual styles.
Dataset Description
This dataset contains 20000 text images for OCR training, created from Portuguese books from Project Gutenberg. Each image contains a complete Portuguese sentence with proper… See the full description on the dataset page: https://huggingface.co/datasets/mazafard/portuguese-ocr-dataset.panellinies-exams-dataset
Dataset Card for Pan-Ex Dataset
Pan-Ex is a comprehensive dataset derived from publicly available exam questions and official solutions used for student admission to Higher Education Institutions in Greece via the Panhellenic Examinations (Πανελλαδικές Εξετάσεις).
The dataset includes questions with the following features:
Subjects: Greek Language, Ancient Greek, History, Latin, Biology, Physics, Chemistry, Computer Science, Economics, and Mathematics.
Educational Levels:… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/panellinies-exams-dataset.azerbaijan-court-data
Azerbaijan Court System Dataset
The most comprehensive open dataset of Azerbaijan's judicial system — 1.64 million structured records and 1.54 million court decision PDFs (~160 GB) covering court decisions, active cases, scheduled hearings, court registries, judges, lawyers, and mediator organizations.
Built for AI engineers, legal tech startups, and researchers who need real-world legal data at scale.
Quick Start
Load with Hugging Face datasets
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/ismatsamadov/azerbaijan-court-data.PaperAudit_Dataset
PaperAudit Origin Data
This directory contains the original paper data downloaded and preprocessed for the PaperAudit project. The data includes papers from top-tier machine learning conferences with their parsed content, metadata, synthetic error annotations, and review information.
PaperAudit Dataset Overview
This repository is part of the full PaperAudit Dataset, which includes:
PaperAudit_Dataset/
├── PaperAudit_Origin_Data/ # Original paper data (raw + preprocessed)… See the full description on the dataset page: https://huggingface.co/datasets/mayiwen/PaperAudit_Dataset.PriOrGen_DATA
PriorGen_Data
PriorGen 使用的数据独立存放于此目录,代码项目位于 /mnt/datadisk0/PriorGen。数据按用途组织:
PriorGen_Data/
├── features/uni/ # UNI WSI patch features (.pt)
├── features/plip/ # PLIP WSI patch features (.pt)
├── annotations/organs/ # TCGA organ report annotations and images
├── splits/ # train/val/test CSV and label mappings
├── knowledge_banks/ # organ knowledge banks (.pt/.json)
├── anchors/ # PIB anchor features (.pt)
├── feature_labels.csv # 每个 UNI/PLIP… See the full description on the dataset page: https://huggingface.co/datasets/Yfeng0216/PriOrGen_DATA.companies-dataset
Tracki - Synthetic Companies Dataset
10,196 synthetic companies x 18 columns. Every one of the 15 content
fields was written by a language model; nothing is rule-generated. Built for the Tracki final
project (RUNI - Intro to Data Science): describe a company, and Tracki returns the
3 most similar companies (embeddings). This would be used for subsribing to their social media and websites.
Every row is fictional. Where the generator's name prior collided with a real trademark it… See the full description on the dataset page: https://huggingface.co/datasets/tracki/companies-dataset.nanog-cancer-data
NanoG - Cancer Foundation-Model Training Data
Multimodal cancer corpus for NanoG1 (generative multimodal pretraining). Literature, structured biology, imaging, and grounded <simulate> traces.
Hub: Abd0r/nanog-cancer-dataAuthor: Syed Abdur Rehman Ali (@Abd0r) · 17 · independent
Train exclusion (hard): NCI-60 is out of training. Skip records whose source / path / text refer to NCI-60. Prefer NCI-ALMANAC, TCGA-sim, Polymathic, PMC/PubMed, TCGA omics, imaging.
How… See the full description on the dataset page: https://huggingface.co/datasets/Abd0r/nanog-cancer-data.letterboxd-all-movie-data
Letterboxd Film Dataset
This dataset contains a comprehensive collection of 847,209 films from the Letterboxd platform, including movie information, user reviews, and ratings.
Dataset Summary
Total Films: 847,209
File Size: ~1.12 GB (1,120,572,122 bytes)
Format: JSONL (JSON Lines)
Language: Primarily English, with some multilingual content
Data Structure
Each line contains a JSON object with the following fields:
{
"url":… See the full description on the dataset page: https://huggingface.co/datasets/pkchwy/letterboxd-all-movie-data.hebrew_this_world
Dataset Card for HebrewSentiment
Dataset Summary
HebrewThisWorld is a data set consists of 2028 issues of the newspaper 'This World' edited by Uri Avnery and were published between 1950 and 1989. Released under the AGPLv3 license.
Data Annotation:
Supported Tasks and Leaderboards
Language modeling
Languages
Hebrew
Dataset Structure
csv file with "," delimeter
Data Instances
Sample:
{
"issue_num": 637,
"page_count": 16… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/hebrew_this_world.yeji-datasets
██████╗ █████╗ ██╗ ██╗ ██████╗ █████╗ ████████╗ █████╗
██╔══██╗██╔══██╗██║ ██║ ██╔══██╗██╔══██╗╚══██╔══╝██╔══██╗
██████╔╝███████║██║ █╗ ██║ ██║ ██║███████║ ██║ ███████║
██╔══██╗██╔══██║██║███╗██║ ██║ ██║██╔══██║ ██║ ██╔══██║
██║ ██║██║ ██║╚███╔███╔╝ ██████╔╝██║ ██║ ██║ ██║ ██║
╚═╝ ╚═╝╚═╝ ╚═╝ ╚══╝╚══╝ ╚═════╝ ╚═╝ ╚═╝ ╚═╝ ╚═╝ ╚═╝
⚡ ORIGINAL DATA ARCHIVE ⚡
> ACCESS GRANTED: YEJI DATASETS
> TYPE: Raw Data Collection… See the full description on the dataset page: https://huggingface.co/datasets/tellang/yeji-datasets.upvoteweb-posts
upvoteweb: posts
Posts in upvoteweb.
configs
[!IMPORTANT]There are several configs representing different permutations of this dataset. Load the relevant config for the task you are interested in.
Overview of configs:
default: largely unfiltered/unprocessed original data
eduscored: the "eduscore" predicted on the text column with huggingface's trained classifier
en-clean: filter language for en and language_score for > 0.6. Run clean-text on the text col, preserving… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/upvoteweb-posts.data-agent-benchmarks
LongHorizon Full Data-Agent Benchmarks
Companion data artifacts for five complete evaluation tracks:
DataSciBench full55 / 167 metric entries
DABStep full450
DABStep-Research full100
DSBench Modeling full74
LongDS full68 / 2,225 turns
The companion GitHub repository contains processed manifests, evaluation code,
historical API ReAct baseline code, download/preparation tools, and the frozen
source lock. artifact_manifest.json records every uploaded object's size,
SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.khatme-nubuwwat-ocr-dataset
Khatme-Nubuwwat Urdu OCR Corpus
This is a structure-aware, fully OCR'd text dataset of around 215 Urdu Khatme Nubuwat books/volumes (approximately 86,557 pages of text) The majority of the books were in Urdu Nastaliq font, with Arabic Naskh and English text present minimally as well. The text dataset is paried with source-page scans.
The books are composed of Nastaliq prose with heavy references to Quran and Hadith. Effort was made to ensure that the OCR pipeline transcribed the… See the full description on the dataset page: https://huggingface.co/datasets/nubuwwat/khatme-nubuwwat-ocr-dataset.LLaVA-OneVision-Mid-Data
Dataset Card for LLaVA-OneVision
Due to unknow reasons, we are unable to process dataset with large amount into required HF format. So we directly upload the json files and image folders (compressed into tar.gz files).
You can use the following link to directly download and decompress them.
https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Mid-Data/tree/main/evol_instruct
We provide the whole details of LLaVA-OneVision Dataset. In this dataset, we include the data splits… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Mid-Data.engsvg-dataset-v2
EngSVG: engineering drawings with solver-verified ground truth
140,000 tasks over 10,000 procedurally generated 2D engineering designs, each drawing
verified by a solver before it was allowed into the dataset.
What makes it unusual
Every design is certified before use. Trusses are solved with a linear axial finite-element
method and admitted only when the equilibrium residual falls below 1e-5; plates pass manufacturing
geometry checks. 10,000 of 10,000 designs… See the full description on the dataset page: https://huggingface.co/datasets/Conditiondark/engsvg-dataset-v2.InfiGUIAgent-DataThis repository contains trajectory data related to reasoning that was used in the second stage of training in InfiGUIAgent.
For more information, please refer to our repo.
62k-images-khmer-printed-dataset
62k Khmer-English Printed Dataset
This repository contains a dataset of Khmer and English printed text images for training, validation, and testing. The dataset is stored in parquet format and managed using Git Large File Storage (LFS).
Installation
Prerequisites
Before cloning this repository, make sure you have Git LFS installed:
Install Git LFS
Linux/macOS:curl -s https://packagecloud.io/install/repositories/github/git-lfs/script.deb.sh | sudo… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/62k-images-khmer-printed-dataset.glsl-opengl-educational-dataset
GLSL/OpenGL & WebGPU Universal Educational Dataset for AI Training
Curated, statically validated, and educational dataset of GLSL, WGSL, and HLSL shaders, OpenGL/WebGL programs, and real-time graphics pipelines.
Key Features
Multi-Stage Shaders: Vertex, Fragment, Compute, Geometry, Tessellation.
Static Validation: Validated against Khronos glslangValidator.
Universal Shading Targets: Multi-target transpilation (WGSL, HLSL, MSL).
Rich Annotations: Includes… See the full description on the dataset page: https://huggingface.co/datasets/gllllll/glsl-opengl-educational-dataset.iq-terrain-vlm-dataset
IQ Terrain VLM Dataset
A high-fidelity, mathematically pristine Vision-Language Model (VLM) dataset designed specifically to teach models the procedural graphics and raymarching techniques of Inigo Quilez.
Dataset Summary
Most coding datasets rely on broadly scraped, often buggy code from GitHub or StackOverflow. This dataset takes a highly targeted approach:
Mathematical Ground Truth: All GLSL code and mathematical concepts are sourced directly from Inigo… See the full description on the dataset page: https://huggingface.co/datasets/True2456/iq-terrain-vlm-dataset.pcb-defect-multi-modal-dataset
印制电路板焊接及装联缺陷多模态数据集
面向 PCB 焊接与装联缺陷识别、资料辅助根因分析和维修建议生成的中文多模态研究数据集。数据集将光学、扫描声学显微(SAM)和 X 射线等缺陷图像与来源标签、中文任务文本,以及可用的检测测量和试验工况关联,支持视觉语言模型的数据准备、指令微调实验和检索增强问答。
版本 v3 包含 5,000 条图像/缺陷区域样本、15,000 条任务指令,覆盖焊点连接质量、焊料分布与形态、焊接层结构及元件装联位置等方面的 14 类目标缺陷:
目标缺陷
含义与标注范围
虚焊
焊接界面未形成可靠的冶金结合,可能表现为接触不稳或间歇导通。
冷焊
热过程不足导致焊料未充分熔融或未形成合格连接,结合来源的热过程与外观信息标注。
少锡
焊料量或有效焊脚不足,焊接部位的填充或覆盖不充分。
多锡
焊料过量,形成过大的焊脚或明显堆积。
连锡
焊料连接本应相互隔离的端子或焊盘,形成焊料桥接。
开焊
应有焊接连接的部位缺少完整连接,或连接部位发生分离。
润湿不良… See the full description on the dataset page: https://huggingface.co/datasets/dezoe/pcb-defect-multi-modal-dataset.
