Team Ai
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ysn-rfd /text-dataset-tiny-code-script-py-format USED of tahamajs/medicine_ds_persian for .parquet file USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file USED of Abirate/english_quotes for .jsonl file NEW FILES (05/12/2025) NEW FILES (12/26/2025) NEW FILES (02/15/2026) texttext-generation10K<n<100K3 likes1.7k downloads4mo agoHugging Face02ed001 /ds-coder-instruct-v1 Dataset Card for DS Coder Instruct Dataset DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in R and Python. The goal of this dataset is to enable creation of small-scale, specialized language model assistants for data science projects. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v1.imagetext-generation10K<n<100K5 likes345 downloads3y agoHugging Face03sagnik-mukherjee /codex-opencomic-dev OpenComic-Continue Corpus This versioned research dataset is intended for leakage-safe comic understanding and continuation. Each example retains work/series/page identity and a machine-readable provenance link. Sources and licences Only per-work records accepted by data/licenses/license_manifest.parquet may enter the core split. Public Domain, CC0, and CC BY are the default classes. Source-specific licence and attribution metadata govern every item; this card… See the full description on the dataset page: https://huggingface.co/datasets/sagnik-mukherjee/codex-opencomic-dev.imageimage-to-textn<1K0 likes95 downloads2mo agoHugging Face04Coder109 /C3BEnglish | 简体中文 C³B: Comics Cross-Cultural Benchmark Culture In a Frame: C³B as a Comic-Based Benchmark for Multimodal Cultural Awareness ICLR 2026 About C³B C³B (Comics Cross-Cultural Benchmark) is a multicultural, multitask, and multilingual benchmark for evaluating cultural awareness capabilities of Multimodal Large Language Models (MLLMs). Progressive task difficulty: From basic visual recognition, to higher-level cultural conflict understanding, to cultural content… See the full description on the dataset page: https://huggingface.co/datasets/Coder109/C3B.imagevisual-question-answering1K<n<10K1 likes73 downloads7mo agoHugging Face05Maksonchek /codeocr-dataset CodeOCR Dataset (Python Code Images + Ground Truth) This dataset is designed for Optical Character Recognition (OCR) of source code.Each example pairs Python code (ground-truth text) with image renderings of that code (light/dark themes) and a real photo. Dataset Summary Language: Python (text ground truth), images of code Splits: easy, medium, hard Total examples: 1,000 easy: 700 medium: 200 hard: 100 Modalities: image + text What is “ground truth”… See the full description on the dataset page: https://huggingface.co/datasets/Maksonchek/codeocr-dataset.imageimage-to-text1K<n<10K1 likes57 downloads10mo agoHugging Face06codecainecowboy /Nemotron-Personas-Korea Nemotron-Personas-Korea 우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템 A compound AI approach to personas grounded in real-world distributions 데이터셋 개요 (Overview) Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 통계청(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다. Nemotron-Personas-Korea는… See the full description on the dataset page: https://huggingface.co/datasets/codecainecowboy/Nemotron-Personas-Korea.imagetext-generation1M<n<10M0 likes54 downloads5mo agoHugging Face07Raiff1982 /codette_training Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/Raiff1982/codette_training.imagetext-generationn<1K0 likes49 downloads8mo agoHugging Face08codelucas /ceo-quotes-verified-sample 🎙️ CEO Transcripts — Verified Executive Interviews The World's Largest Database of Verified C-Suite Transcripts 20,000+ Executives · 100,000+ Transcripts · 400,000+ Quotes · S&P 500 + NASDAQ + Global Leaders 🔥 What's In This Sample? This is a free evaluation sample from CEOInterviews.ai featuring 9 of the most market-moving voices in finance, tech, and policy. Executive Role Why They Matter Jensen Huang CEO, NVIDIA Every AI… See the full description on the dataset page: https://huggingface.co/datasets/codelucas/ceo-quotes-verified-sample.imagetext-generation1K<n<10K2 likes27 downloads11mo agoHugging Face09Ruler138 /CodeAnything-1.835Mgated CodeAnything 1.835M SFT and evaluation release Gated public release containing the clean 1,835,476-sample training set, the 800-sample/16-domain evaluation set, paper-model predictions and rendered outputs, and raw per-sample rating records. Layout training/ manifest/ all_training_v5.jsonl all_training_v5.jsonl.idx all_training_v5.jsonl.true_lengths.u32 shards/<domain>/ exact media/code closure (tar shards) evaluation/ benchmark/… See the full description on the dataset page: https://huggingface.co/datasets/Ruler138/CodeAnything-1.835M.imageimage-to-text10K<n<100K0 likes23 downloads25d agoHugging Face10AnonymousSubmissionASE /pixels_vs_code Pattern Over Pixels Screenshot-to-Code This dataset contains controlled counterfactual screenshot-to-code examples built from 30 real-world webpages from Design2Code. Each example preserves a repeated UI pattern while introducing a single localized deviation, allowing researchers to test whether multimodal models follow the pixels or simply restore the dominant template. Contents 720 perturbed HTML instances 360 structural-card examples 360 text-style examples 2… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousSubmissionASE/pixels_vs_code.imageimage-to-textn<1K0 likes11 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.