datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
textvqa
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of TextVQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{singh2019towards,
title={Towards vqa models that can read},
author={Singh, Amanpreet and… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/textvqa.ProLong-TextFullhle_text_only
Humanity's Last Exam - (Text only)
🌐 Website | 📄 Paper | GitHub
Center for AI Safety & Scale AI
Humanity's Last Exam (HLE) is a multi-modal benchmark at the frontier of human knowledge, designed to be the final closed-ended academic benchmark of its kind with broad subject coverage. Humanity's Last Exam consists of 3,000 questions across dozens of subjects, including mathematics, humanities, and the natural sciences. HLE is developed globally by subject-matter experts and… See the full description on the dataset page: https://huggingface.co/datasets/macabdul9/hle_text_only.ptb_text_onlyThis is the Penn Treebank Project: Release 2 CDROM, featuring a million words of 1989 Wall Street Journal material. This corpus has been annotated for part-of-speech (POS) information. In addition, over half of it has been annotated for skeletal syntactic structure.text-to-image-promptsIf you have questions about this dataset , feel free to ask them on the fusion-discord : https://discord.gg/8TVHPf6Edn
This collection contains sets from the fusion-t2i-ai-generator on perchance.
This datset is used in this notebook: https://huggingface.co/datasets/codeShare/text-to-image-prompts/tree/main/Google%20Colab%20Notebooks
To see the full sets, please use the url "https://perchance.org/" + url
, where the urls are listed below:
_generator
gen_e621
fusion-t2i-e621-tags-1… See the full description on the dataset page: https://huggingface.co/datasets/codeShare/text-to-image-prompts.ipo-text
SEC IPO Filings Dataset
A large-scale, comprehensive dataset of 100,000+ filings (S-1 and F-1 filings) filed with the SEC EDGAR system, spanning 1994–2026 and over 20,000 unique registrants.
Every filing has been downloaded and then parsed using the IPO-Mine Python Package. We have extracted three common sections found in these documents (Prospectus Summary, Risk Factors, Legal Matters), and then used an LLM classifier to group them into three categories. For this dataset, we have… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/ipo-text.quail
Dataset Card for "quail"
Dataset Summary
QuAIL is a reading comprehension dataset. QuAIL contains 15K multi-choice questions in texts 300-350 tokens long 4 domains (news, user stories, fiction, blogs).QuAIL is balanced and annotated for question types.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
quail
Size of downloaded dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/textmachinelab/quail.muse_textbooksTextAtlas5M
TextAtlas5M
This dataset is a training set for TextAtlas.
Paper: https://huggingface.co/papers/2502.07870
(All the data in this repo is uploaded :>)
Dataset subsets
Subsets in this dataset are CleanTextSynth, PPT2Details, PPT2Structured,LongWordsSubset-A,LongWordsSubset-M,Cover Book,Paper2Text,TextVisionBlend,StyledTextSynth and TextScenesHQ. The dataset features are as follows:
Dataset Features
image (img): The GT image.
annotation (string): The input prompt… See the full description on the dataset page: https://huggingface.co/datasets/CSU-JPG/TextAtlas5M.dummy_image_text_data
Dataset Card for "dummy_image_text_data"
More Information needed
Sindhi-texts-big-dataset
Sindhi Texts (big dataset)
A large plain-text corpus of Sindhi (سنڌي), assembled for pretraining language models. It
combines material digitized by Sindhi literary institutions and forums, a Sindhi encyclopedia,
newspaper archives, a classical dictionary, and the Sindhi portions of two web-crawl corpora.
3.19 GB, ~1.81 billion characters, ~390,000 documents across 9 sources. With a
Sindhi-specific 12k SentencePiece tokenizer that is roughly 530M tokens (3.2–3.5
characters per… See the full description on the dataset page: https://huggingface.co/datasets/arnizamani/Sindhi-texts-big-dataset.Reverse-Text-RL
Reverse-Text-RL
A small, scrappy RL dataset used in prime-rl's CI to debug RL training asking a model to reverse small sentences character-by-character. Follows the general format of PrimeIntellect/Reverse-Text-SFT
The following script was used to generate the dataset.
from datasets import Dataset, load_dataset
dataset = load_dataset("willcb/R1-reverse-wikipedia-paragraphs-v1-1000", split="train")
prompt = "Reverse the text character-by-character. Put your answer in… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Reverse-Text-RL.Textground4MTextGround4M: A Prompt-Aligned Dataset for Layout-Aware Text Rendering
TextGround4M is a large-scale dataset for prompt-grounded, layout-aware text rendering in text-to-image (T2I) generation, introduced in our AAAI 2026 paper.
Dataset Summary
TextGround4M contains 4.1 million prompt-image pairs, each annotated with:
A natural language caption where all rendered text spans are explicitly quoted
Span-level bounding boxes linking each quoted… See the full description on the dataset page: https://huggingface.co/datasets/CSU-JPG/Textground4M.td02_urban-surface-texturesThe Dataset Teaser is now enabled instead! Isn't this better?
TD 02: Urban Surface Textures
This dataset contains multi-photo texture captures in outdoor urban scenes — many focusing on the ground and the others are walls. Each set has different photos that showcase texture variety, making them ideal for training a domain-specific image generator!
Overall information about this dataset:
Format — JPEG-XL, lossless RGB
Resolution — 4032 × 2268
Device — mobile camera
Technique —… See the full description on the dataset page: https://huggingface.co/datasets/texturedesign/td02_urban-surface-textures.textsformat-textTextEdit
TextEdit: A High-Quality, Multi-Scenario Text Editing Benchmark for Generation Models
Danni Yang,
Sitao Chen,
Changyao Tian
If you find our work helpful, please give us a ⭐ or cite our paper. See the InternVL-U technical report appendix for more details.
🎉 News
[2026/03/06] TextEdit benchmark released.
[2026/03/06] Evaluation code and initial baselines released.
[2026/03/06] Leaderboard updated with latest models.
📖 Introduction… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/TextEdit.WTO-Text
Dataset Card for WTO Documents Dataset
Dataset Overview
Title: WTO Documents Dataset
Source: World Trade Organization Documents Online
Description: The WTO Documents Dataset is a comprehensive collection of official documentation from the World Trade Organization (WTO). This dataset is sourced from the WTO's official Documents Online platform, which provides access to documents in the three official languages (English, French, and Spanish) from 1995 onwards. The… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/WTO-Text.persian-asr-audio-text-2.69M-chizzled
🗂️ persian-asr-audio-text-2.69M-chizzled
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Phase A-scale audio/text dataset.
پیکرهٔ بزرگ جفتهای صوت و متنِ پالایششده برای آموزش در مقیاس فاز A.
🧩 Role
Persian text and linguistic asset
مصنوع متنی و زبانی فارسی
📦 Snapshot
417 files; approximately 236.86 GB
417 فایل؛ حدود 236.86 GB
🧱 Packaging
414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.turkish-raw-text-cleaned
Turkish Raw Text Cleaned
turkish-raw-text-cleaned, Türkçe dil modeli çalışmaları için hazırlanmış temizlenmiş ham metin veri kümesidir. Veri kümesi, turkish-nlp-suite çatısı altında yayımlanan Türkçe metin kaynaklarının temizlenmesi, filtrelenmesi ve model eğitimine daha uygun hale getirilmesiyle oluşturulmuştur.
Bu çalışma özellikle Türkçe LLM ön-eğitimi, continual pre-training (CPT), tokenizer analizi, embedding modeli eğitimi, alan bağımsız Türkçe metin modelleme ve veri… See the full description on the dataset page: https://huggingface.co/datasets/serda-dev/turkish-raw-text-cleaned.frodobots-mini-text-500g
FrodoBots-Mini-4K
~4,000 hours of real-world teleoperation data from Earth Rover Mini / Mini+ sidewalk robots,
driven by a global operator network across 29 countries. Each ride bundles synchronized camera video
(front, and rear when available), two-way audio, and time-aligned GPS, IMU, and
drive/control (DRV) streams.
Third public FrodoBots dataset, after
BitRobot/FrodoBots-2K and
BitRobot/Berkeley-FrodoBots-7K.
Like the 2K release it ships raw, unannotated per-ride folders —… See the full description on the dataset page: https://huggingface.co/datasets/Wolfie-Jr/frodobots-mini-text-500g.hle_text_onlycode_x_glue_ct_code_to_text
Dataset Card for "code_x_glue_ct_code_to_text"
Dataset Summary
CodeXGLUE code-to-text dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Text/code-to-text
The dataset we use comes from CodeSearchNet and we filter the dataset as the following:
Remove examples that codes cannot be parsed into an abstract syntax tree.
Remove examples that #tokens of documents is < 3 or >256
Remove examples that documents contain special tokens (e.g. <img ...> or… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_ct_code_to_text.Total-Text-DatasetTotal Text Dataset.
It consists of 1555 images with more than 3 different text orientations: Horizontal, Multi-Oriented, and Curved, one of a kind.
Original github repo; https://github.com/cs-chan/Total-Text-Dataset
Forked repo; https://github.com/yunusserhat/Total-Text-Dataset
gaia_filtered_text_onlybeamit-full-texts-dataset
Dataset Card for "beamit-full-texts-dataset"
More Information needed
Java-Code-Large-text-onlyJava-Code-Large
Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis.
By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/XEUIPR/Java-Code-Large-text-only.TextPecker-1.5M
TextPecker-1.5M: A Dataset for Training and evaluating TextPecker
This repository contains the TextPecker-1.5M dataset, a new benchmark proposed in the paper "TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering".
Code and Project Page
The official implementation and project details for the TextPecker and TextPecker-1.5M dataset can be found on the GitHub repository:
https://github.com/CIawevy/TextPecker
Sample Usage
You… See the full description on the dataset page: https://huggingface.co/datasets/CIawevy/TextPecker-1.5M.text-sft-questions-answers-only
text-sft: Questions and Answers
This dataset consists of question-and-answer pairs generated from short excerpts drawn from Wikipedia, Cosmopedia, and FineWeb-Edu. It is an adapted version of agentlans/text-sft.
Overview
The dataset provides compact examples of English question-and-answer relationships that can help models learn linguistic patterns, syntactic structures, and semantic associations between questions and their corresponding answers.
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-sft-questions-answers-only.HLE_text_200
language:
- en
tags:
- chemistry
- biology
- math
HLEのうち、text形式のものを抽出した200問
