datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-coding-plan-prices
AI coding plan prices (tokens and requests per dollar)
What the flat rate AI coding subscriptions cost, and what you actually get for the money: published
quotas, per model spending caps, included models, and the resulting tokens and requests per dollar
paid. This is the dataset behind vibeplan.cc.
Snapshot of this revision: 2026-10-06 10:22 UTC, 80 plans from 26 providers,
53 of them directly comparable, built from 30 official sources.
Disclosure split: 51 disclosed, 21… See the full description on the dataset page: https://huggingface.co/datasets/harrytyp/ai-coding-plan-prices.mimo-coding-synthetic-5k
MiMo Coding Synthetic 5.4K
MiMo Coding Synthetic 5.4K is a purely synthetic coding instruction dataset generated with Xiaomi MiMo mimo-v2.5-pro.
It contains 5,411 validated examples across programming languages, coding task types, and difficulty levels. The dataset is provided in two formats:
A canonical rich JSONL format with metadata and labels.
An OpenAI chat messages JSONL format for supervised fine-tuning pipelines.
The generation run used 20 parallel workers for roughly… See the full description on the dataset page: https://huggingface.co/datasets/himanshunakrani9/mimo-coding-synthetic-5k.coding-master-dataset
Coding Master Dataset
Overview
A large-scale coding instruction-tuning dataset in ShareGPT conversational format, assembled from multiple open sources and deduplicated.
Records: 766,987
Format: JSONL / ShareGPT
License: Apache 2.0
Sources
CodeX-2M-Thinking (430,542 records)
python-code-dataset-500k (559,515 records)
StackPulse high-quality subset (20,205 records)
CodeFeedback-Filtered-Instruction (156,525 records)
secure_programming_dpo (4,656… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/coding-master-dataset.MrGrammaticalOntology_clinical_coding
Mr. Grammatical Ontology: Clinical Coding
This dataset was created from a motivation to train Medical Large Language Models for improved fluency in clinical
coding, as measurable by MedConceptsQA, an open-source medical coding
evaluation benchmark designed to evaluate the understanding and reasoning capabilities of LLMs on medical concepts. It was extracted from the Centers for Medicare & Medicaid Services'
International Classification of Diseases, Tenth Revision, Clinical… See the full description on the dataset page: https://huggingface.co/datasets/cogbuji/MrGrammaticalOntology_clinical_coding.sota-codingCodingQuestionDatabaseCodeLlamaThe questions, responses, and topics were generated with the codellama 7b model.
They may be empty data points due to the AI generation.
Main.json: CodeLlama 7b
moneywordmath.json: pplx-7b-chat (Math Word Problems)
SNU_Thunder-synthetic-coding
Dataset Card for SNU Thunder Synthetic Coding
Dataset Summary
This dataset was used as part of the post-training corpus for SnuLLM(to_fill).
This dataset consists of Korean and English question-answer pairs. Questions are sourced from publicly available datasets, and answers were generated using open large language models (Exaone 3.5, LLaMA 3.3, Qwen 2.5). It is intended for research and non-commercial use.
Supported Tasks
Tasks: Python coding
Languages… See the full description on the dataset page: https://huggingface.co/datasets/thunder-research-group/SNU_Thunder-synthetic-coding.
