Team Ai
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ysn-rfd /text-dataset-tiny-code-script-py-format USED of tahamajs/medicine_ds_persian for .parquet file USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file USED of Abirate/english_quotes for .jsonl file NEW FILES (05/12/2025) NEW FILES (12/26/2025) NEW FILES (02/15/2026) texttext-generation10K<n<100K3 likes1.7k downloads4mo agoHugging Face02artefactory /Argimi-Ardian-Finance-10k-text-image The ArGiMI Ardian datasets : text and images The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.imagetext-retrieval1M<n<10M14 likes625 downloads8mo agoHugging Face03EvanOLeary /Text2WetLab Text2WetLab Natural language → Opentrons OT-2 robot protocol benchmark Text2WetLab evaluates whether large language model agents can translate plain-English wet-lab procedure descriptions into correct, safe, executable liquid-handling protocols for the Opentrons OT-2. Reference episode — MuJoCo 3D physics render Results (pass@1, 2026-10-04) 3 Claude models evaluated as Claude Code agents in Harbor sandboxes on Modal. All 21 trials passed the… See the full description on the dataset page: https://huggingface.co/datasets/EvanOLeary/Text2WetLab.imagetext-generationn<1K1 likes510 downloads3d agoHugging Face04jrzhang /TextVQA_GT_bbox TextVQA validation set with grounding truth bounding box The dataset used in the paper MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs for studying MLLMs' attention patterns. The dataset is sourced from TextVQA and annotated manually with ground-truth bounding boxes. We consider questions with a single area of interest in the image so that 4370 out of 5000 samples are kept. Citation If you find our paper and code useful… See the full description on the dataset page: https://huggingface.co/datasets/jrzhang/TextVQA_GT_bbox.imagequestion-answering1K<n<10K4 likes226 downloads2y agoHugging Face05yonilev /Text2Receipt Text2Receipt Messy free-text Hebrew income notes -> valid, complete Israeli fiscal documents (receipts & tax invoices). Live demo (Space): yonilev/Text2Receipt Dataset: yonilev/Text2Receipt Dataset Creation A synthetic corpus from a deterministic, rule-based generator plus a bounded LLM-paraphrase layer, so the ground truth is exact by construction. Pipeline Scenario sampling - category, issuer status, document type, client type, year, payment… See the full description on the dataset page: https://huggingface.co/datasets/yonilev/Text2Receipt.imagetext-generation10K<n<100K0 likes147 downloads4mo agoHugging Face06minn4 /Text2CGL Text2CGL: A Clean Zero-Leakage Dataset for Vietnamese Geometry to CGL This dataset contains formal mappings from natural Vietnamese geometry problem texts to Constructive Geometry Language (CGL) S-expressions for automated differentiable diagram synthesis. 📌 Dataset Releases & Versions v1.0 (Paper Benchmark): 3,797 instruction–CGL pairs (3,037 train, 379 validation, 381 test) across 1,045 root problem lineages, as reported in the DASA 2026 paper. v1.1 (Current… See the full description on the dataset page: https://huggingface.co/datasets/minn4/Text2CGL.imagetext-generation1K<n<10K0 likes69 downloads3d agoHugging Face07REILX /Chinese-Image-Text-Corpus-dataset REILX/Chinese-Image-Text-Corpus-dataset [ English | 中文 ] Introduction The REILX/Chinese-Image-Text-Corpus-dataset is a multimodal dataset that pairs Chinese textual data with corresponding images. This dataset is derived from the Chinese-Xinhua Dictionary Database, which includes idioms, single characters, words, and aphorisms. Dataset Structure The dataset is organized into the following categories: Idioms: Traditional Chinese idioms with explanations and… See the full description on the dataset page: https://huggingface.co/datasets/REILX/Chinese-Image-Text-Corpus-dataset.imagequestion-answering100K<n<1M0 likes52 downloads2y agoHugging Face08crawlfeeds /Curated-Fox-News-Headlines-and-Full-Text Curated Fox News Headlines and Full Text This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis. 📁 Dataset Format Format: CSV Encoding: UTF-8 Fields: headline: The article title or headline publish_date: Date the article was published (YYYY-MM-DD) content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.imagetext-classification1K<n<10K2 likes43 downloads1y agoHugging Face09Windsao /eis-text250 EIS-Text250: 1970s U.S. Environmental Impact Statements (text-only) Per-page OCR/extraction text for 250 scanned 1970s U.S. federal Environmental Impact Statements (EIS) from the Northwestern University Library collection — the text-only companion to Windsao/eis-subset50 (which carries full page images for a 50-doc subset). Built to test how current models handle long, dense, historical government text: mean ~300 pages/doc, 1970s typewriter prose, OCR noise from degraded… See the full description on the dataset page: https://huggingface.co/datasets/Windsao/eis-text250.imagetext-generation10K<n<100K0 likes19 downloads3mo agoHugging Face10hkust-gz-w2 /PDD3_text_rendered_v2gatedimagetext-generation0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.