datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-dataset-tiny-code-script-py-format
USED of tahamajs/medicine_ds_persian for .parquet file
USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file
USED of Abirate/english_quotes for .jsonl file
NEW FILES (05/12/2025)
NEW FILES (12/26/2025)
NEW FILES (02/15/2026)
ds-coder-instruct-v1
Dataset Card for DS Coder Instruct Dataset
DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on
data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in R and Python.
The goal of this dataset is to enable creation of small-scale, specialized language model assistants for data science projects.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v1.codex-opencomic-dev
OpenComic-Continue Corpus
This versioned research dataset is intended for leakage-safe comic understanding and continuation. Each example retains work/series/page identity and a machine-readable provenance link.
Sources and licences
Only per-work records accepted by data/licenses/license_manifest.parquet may enter the core split. Public Domain, CC0, and CC BY are the default classes. Source-specific licence and attribution metadata govern every item; this card… See the full description on the dataset page: https://huggingface.co/datasets/sagnik-mukherjee/codex-opencomic-dev.C3BEnglish | 简体中文
C³B: Comics Cross-Cultural Benchmark
Culture In a Frame: C³B as a Comic-Based Benchmark for Multimodal Cultural Awareness
ICLR 2026
About C³B
C³B (Comics Cross-Cultural Benchmark) is a multicultural, multitask, and multilingual benchmark for evaluating cultural awareness capabilities of Multimodal Large Language Models (MLLMs).
Progressive task difficulty: From basic visual recognition, to higher-level cultural conflict understanding, to cultural content… See the full description on the dataset page: https://huggingface.co/datasets/Coder109/C3B.codeocr-dataset
CodeOCR Dataset (Python Code Images + Ground Truth)
This dataset is designed for Optical Character Recognition (OCR) of source code.Each example pairs Python code (ground-truth text) with image renderings of that code (light/dark themes) and a real photo.
Dataset Summary
Language: Python (text ground truth), images of code
Splits: easy, medium, hard
Total examples: 1,000
easy: 700
medium: 200
hard: 100
Modalities: image + text
What is “ground truth”… See the full description on the dataset page: https://huggingface.co/datasets/Maksonchek/codeocr-dataset.Nemotron-Personas-Korea
Nemotron-Personas-Korea
우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템
A compound AI approach to personas grounded in real-world distributions
데이터셋 개요 (Overview)
Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 통계청(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다.
Nemotron-Personas-Korea는… See the full description on the dataset page: https://huggingface.co/datasets/codecainecowboy/Nemotron-Personas-Korea.codette_training
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/Raiff1982/codette_training.ceo-quotes-verified-sample
🎙️ CEO Transcripts — Verified Executive Interviews
The World's Largest Database of Verified C-Suite Transcripts
20,000+ Executives · 100,000+ Transcripts · 400,000+ Quotes · S&P 500 + NASDAQ + Global Leaders
🔥 What's In This Sample?
This is a free evaluation sample from CEOInterviews.ai featuring 9 of the most market-moving voices in finance, tech, and policy.
Executive
Role
Why They Matter
Jensen Huang
CEO, NVIDIA
Every AI… See the full description on the dataset page: https://huggingface.co/datasets/codelucas/ceo-quotes-verified-sample.CodeAnything-1.835M
CodeAnything 1.835M SFT and evaluation release
Gated public release containing the clean 1,835,476-sample training set, the
800-sample/16-domain evaluation set, paper-model predictions and rendered
outputs, and raw per-sample rating records.
Layout
training/
manifest/
all_training_v5.jsonl
all_training_v5.jsonl.idx
all_training_v5.jsonl.true_lengths.u32
shards/<domain>/ exact media/code closure (tar shards)
evaluation/
benchmark/… See the full description on the dataset page: https://huggingface.co/datasets/Ruler138/CodeAnything-1.835M.pixels_vs_code
Pattern Over Pixels Screenshot-to-Code
This dataset contains controlled counterfactual screenshot-to-code examples built
from 30 real-world webpages from Design2Code. Each example preserves a repeated
UI pattern while introducing a single localized deviation, allowing researchers
to test whether multimodal models follow the pixels or simply restore the
dominant template.
Contents
720 perturbed HTML instances
360 structural-card examples
360 text-style examples
2… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousSubmissionASE/pixels_vs_code.
