Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lmms-lab-encoder /textvqa Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of TextVQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{singh2019towards, title={Towards vqa models that can read}, author={Singh, Amanpreet and… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/textvqa.image10K<n<100K25 likes49k downloads3y agoHugging Face02orionweller /ProLong-TextFull0 likes32k downloads2y agoHugging Face03macabdul9 /hle_text_only Humanity's Last Exam - (Text only) 🌐 Website | 📄 Paper | GitHub Center for AI Safety & Scale AI Humanity's Last Exam (HLE) is a multi-modal benchmark at the frontier of human knowledge, designed to be the final closed-ended academic benchmark of its kind with broad subject coverage. Humanity's Last Exam consists of 3,000 questions across dozens of subjects, including mathematics, humanities, and the natural sciences. HLE is developed globally by subject-matter experts and… See the full description on the dataset page: https://huggingface.co/datasets/macabdul9/hle_text_only.image1K<n<10K4 likes32k downloads2y agoHugging Face04ptb-text-only /ptb_text_onlyThis is the Penn Treebank Project: Release 2 CDROM, featuring a million words of 1989 Wall Street Journal material. This corpus has been annotated for part-of-speech (POS) information. In addition, over half of it has been annotated for skeletal syntactic structure.text-generation10K<n<100K20 likes29k downloads3y agoHugging Face05codeShare /text-to-image-promptsIf you have questions about this dataset , feel free to ask them on the fusion-discord : https://discord.gg/8TVHPf6Edn This collection contains sets from the fusion-t2i-ai-generator on perchance. This datset is used in this notebook: https://huggingface.co/datasets/codeShare/text-to-image-prompts/tree/main/Google%20Colab%20Notebooks To see the full sets, please use the url "https://perchance.org/" + url , where the urls are listed below: _generator gen_e621 fusion-t2i-e621-tags-1… See the full description on the dataset page: https://huggingface.co/datasets/codeShare/text-to-image-prompts.text-to-image100K<n<1M11 likes19k downloads2y agoHugging Face06gtfintechlab /ipo-text SEC IPO Filings Dataset A large-scale, comprehensive dataset of 100,000+ filings (S-1 and F-1 filings) filed with the SEC EDGAR system, spanning 1994–2026 and over 20,000 unique registrants. Every filing has been downloaded and then parsed using the IPO-Mine Python Package. We have extracted three common sections found in these documents (Prospectus Summary, Risk Factors, Legal Matters), and then used an LLM classifier to group them into three categories. For this dataset, we have… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/ipo-text.tabulartext-classification100K<n<1M6 likes14k downloads8mo agoHugging Face07textmachinelab /quail Dataset Card for "quail" Dataset Summary QuAIL is a reading comprehension dataset. QuAIL contains 15K multi-choice questions in texts 300-350 tokens long 4 domains (news, user stories, fiction, blogs).QuAIL is balanced and annotated for question types. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances quail Size of downloaded dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/textmachinelab/quail.textmultiple-choice10K<n<100K8 likes13k downloads3y agoHugging Face08amongglue /muse_textbookstext1M<n<10M3 likes12k downloads3y agoHugging Face09CSU-JPG /TextAtlas5M TextAtlas5M This dataset is a training set for TextAtlas. Paper: https://huggingface.co/papers/2502.07870 (All the data in this repo is uploaded :>) Dataset subsets Subsets in this dataset are CleanTextSynth, PPT2Details, PPT2Structured,LongWordsSubset-A,LongWordsSubset-M,Cover Book,Paper2Text,TextVisionBlend,StyledTextSynth and TextScenesHQ. The dataset features are as follows: Dataset Features image (img): The GT image. annotation (string): The input prompt… See the full description on the dataset page: https://huggingface.co/datasets/CSU-JPG/TextAtlas5M.imagetext-to-image1M<n<10M39 likes10k downloads1y agoHugging Face10hf-internal-testing /dummy_image_text_data Dataset Card for "dummy_image_text_data" More Information needed imagen<1K1 likes9.1k downloads4y agoHugging Face11arnizamani /Sindhi-texts-big-dataset Sindhi Texts (big dataset) A large plain-text corpus of Sindhi (سنڌي), assembled for pretraining language models. It combines material digitized by Sindhi literary institutions and forums, a Sindhi encyclopedia, newspaper archives, a classical dictionary, and the Sindhi portions of two web-crawl corpora. 3.19 GB, ~1.81 billion characters, ~390,000 documents across 9 sources. With a Sindhi-specific 12k SentencePiece tokenizer that is roughly 530M tokens (3.2–3.5 characters per… See the full description on the dataset page: https://huggingface.co/datasets/arnizamani/Sindhi-texts-big-dataset.texttext-generation100K<n<1M4 likes8.8k downloads2mo agoHugging Face12PrimeIntellect /Reverse-Text-RL Reverse-Text-RL A small, scrappy RL dataset used in prime-rl's CI to debug RL training asking a model to reverse small sentences character-by-character. Follows the general format of PrimeIntellect/Reverse-Text-SFT The following script was used to generate the dataset. from datasets import Dataset, load_dataset dataset = load_dataset("willcb/R1-reverse-wikipedia-paragraphs-v1-1000", split="train") prompt = "Reverse the text character-by-character. Put your answer in… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Reverse-Text-RL.textquestion-answering1K<n<10K2 likes7.8k downloads4d agoHugging Face13CSU-JPG /Textground4MTextGround4M: A Prompt-Aligned Dataset for Layout-Aware Text Rendering TextGround4M is a large-scale dataset for prompt-grounded, layout-aware text rendering in text-to-image (T2I) generation, introduced in our AAAI 2026 paper. Dataset Summary TextGround4M contains 4.1 million prompt-image pairs, each annotated with: A natural language caption where all rendered text spans are explicitly quoted Span-level bounding boxes linking each quoted… See the full description on the dataset page: https://huggingface.co/datasets/CSU-JPG/Textground4M.image1M<n<10M2 likes7.3k downloads5mo agoHugging Face14texturedesign /td02_urban-surface-texturesThe Dataset Teaser is now enabled instead! Isn't this better? TD 02: Urban Surface Textures This dataset contains multi-photo texture captures in outdoor urban scenes — many focusing on the ground and the others are walls. Each set has different photos that showcase texture variety, making them ideal for training a domain-specific image generator! Overall information about this dataset: Format — JPEG-XL, lossless RGB Resolution — 4032 × 2268 Device — mobile camera Technique —… See the full description on the dataset page: https://huggingface.co/datasets/texturedesign/td02_urban-surface-textures.unconditional-image-generationn<1K4 likes7k downloads3y agoHugging Face15polinaeterna /textstextn<1K0 likes6.2k downloads4y agoHugging Face16chupei /format-texttextn<1K0 likes6.1k downloads2y agoHugging Face17opencompass /TextEdit TextEdit: A High-Quality, Multi-Scenario Text Editing Benchmark for Generation Models Danni Yang, Sitao Chen, Changyao Tian If you find our work helpful, please give us a ⭐ or cite our paper. See the InternVL-U technical report appendix for more details. 🎉 News [2026/03/06] TextEdit benchmark released. [2026/03/06] Evaluation code and initial baselines released. [2026/03/06] Leaderboard updated with latest models. 📖 Introduction… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/TextEdit.imageimage-to-image1K<n<10K9 likes6k downloads7mo agoHugging Face18PleIAs /WTO-Text Dataset Card for WTO Documents Dataset Dataset Overview Title: WTO Documents Dataset Source: World Trade Organization Documents Online Description: The WTO Documents Dataset is a comprehensive collection of official documentation from the World Trade Organization (WTO). This dataset is sourced from the WTO's official Documents Online platform, which provides access to documents in the three official languages (English, French, and Spanish) from 1995 onwards. The… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/WTO-Text.tabular100K<n<1M9 likes5k downloads2y agoHugging Face19Reza2kn /persian-asr-audio-text-2.69M-chizzled 🗂️ persian-asr-audio-text-2.69M-chizzled English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Phase A-scale audio/text dataset. پیکرهٔ بزرگ جفت‌های صوت و متنِ پالایش‌شده برای آموزش در مقیاس فاز A. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 417 files; approximately 236.86 GB 417 فایل؛ حدود 236.86 GB 🧱 Packaging 414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.tabular1M<n<10M2 likes4.7k downloads2mo agoHugging Face20serda-dev /turkish-raw-text-cleaned Turkish Raw Text Cleaned turkish-raw-text-cleaned, Türkçe dil modeli çalışmaları için hazırlanmış temizlenmiş ham metin veri kümesidir. Veri kümesi, turkish-nlp-suite çatısı altında yayımlanan Türkçe metin kaynaklarının temizlenmesi, filtrelenmesi ve model eğitimine daha uygun hale getirilmesiyle oluşturulmuştur. Bu çalışma özellikle Türkçe LLM ön-eğitimi, continual pre-training (CPT), tokenizer analizi, embedding modeli eğitimi, alan bağımsız Türkçe metin modelleme ve veri… See the full description on the dataset page: https://huggingface.co/datasets/serda-dev/turkish-raw-text-cleaned.text-generation1M<n<10M0 likes4.7k downloads3mo agoHugging Face21Wolfie-Jr /frodobots-mini-text-500g FrodoBots-Mini-4K ~4,000 hours of real-world teleoperation data from Earth Rover Mini / Mini+ sidewalk robots, driven by a global operator network across 29 countries. Each ride bundles synchronized camera video (front, and rear when available), two-way audio, and time-aligned GPS, IMU, and drive/control (DRV) streams. Third public FrodoBots dataset, after BitRobot/FrodoBots-2K and BitRobot/Berkeley-FrodoBots-7K. Like the 2K release it ships raw, unannotated per-ride folders —… See the full description on the dataset page: https://huggingface.co/datasets/Wolfie-Jr/frodobots-mini-text-500g.videorobotics0 likes4.6k downloads3mo agoHugging Face22DongfuJiang /hle_text_onlyimage1K<n<10K0 likes4.5k downloads1y agoHugging Face23google /code_x_glue_ct_code_to_text Dataset Card for "code_x_glue_ct_code_to_text" Dataset Summary CodeXGLUE code-to-text dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Text/code-to-text The dataset we use comes from CodeSearchNet and we filter the dataset as the following: Remove examples that codes cannot be parsed into an abstract syntax tree. Remove examples that #tokens of documents is < 3 or >256 Remove examples that documents contain special tokens (e.g. <img ...> or… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_ct_code_to_text.texttranslation1M<n<10M79 likes4.4k downloads3y agoHugging Face24yunusserhat /Total-Text-DatasetTotal Text Dataset. It consists of 1555 images with more than 3 different text orientations: Horizontal, Multi-Oriented, and Curved, one of a kind. Original github repo; https://github.com/cs-chan/Total-Text-Dataset Forked repo; https://github.com/yunusserhat/Total-Text-Dataset imagetext-retrieval1K<n<10K0 likes4.1k downloads2y agoHugging Face25sayan1101 /gaia_filtered_text_onlytextn<1K0 likes3.8k downloads3y agoHugging Face26acmc /beamit-full-texts-dataset Dataset Card for "beamit-full-texts-dataset" More Information needed text10K<n<100K0 likes3.6k downloads3y agoHugging Face27XEUIPR /Java-Code-Large-text-onlyJava-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/XEUIPR/Java-Code-Large-text-only.texttext-generation10M<n<100M0 likes3.6k downloads3d agoHugging Face28CIawevy /TextPecker-1.5M TextPecker-1.5M: A Dataset for Training and evaluating TextPecker This repository contains the TextPecker-1.5M dataset, a new benchmark proposed in the paper "TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering". Code and Project Page The official implementation and project details for the TextPecker and TextPecker-1.5M dataset can be found on the GitHub repository: https://github.com/CIawevy/TextPecker Sample Usage You… See the full description on the dataset page: https://huggingface.co/datasets/CIawevy/TextPecker-1.5M.imageimage-to-text1M<n<10M0 likes3.4k downloads7mo agoHugging Face29agentlans /text-sft-questions-answers-only text-sft: Questions and Answers This dataset consists of question-and-answer pairs generated from short excerpts drawn from Wikipedia, Cosmopedia, and FineWeb-Edu. It is an adapted version of agentlans/text-sft. Overview The dataset provides compact examples of English question-and-answer relationships that can help models learn linguistic patterns, syntactic structures, and semantic associations between questions and their corresponding answers. Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-sft-questions-answers-only.texttext-generation100K<n<1M2 likes3.3k downloads11mo agoHugging Face30Ta1k1 /HLE_text_200 language: - en tags: - chemistry - biology - math HLEのうち、text形式のものを抽出した200問 textn<1K0 likes3.3k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.