datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
orbis-coder
Orbis Coder Dataset (10K)
A coding-first instruction dataset to train or fine-tune assistants that behave like Orbis Coder — friendly, concise, practical, and focused on helping people build, debug, and ship software.
This dataset is intended for:
instruction-tuning / SFT
LoRA / QLoRA fine-tunes
“persona + skill” alignment for coding assistants
quick experiments + dataset viewer testing
What this dataset contains
Most rows are coding help across many languages and… See the full description on the dataset page: https://huggingface.co/datasets/xlelords/orbis-coder.TILO.RA_CODER_Dataset
TILO.RA CODER Dataset
Объединённый русско-английский датасет для обучения и поиска по коду.
Формат — пары question / code: вопрос на естественном языке → готовый код-ответ.
Датасет собран для локального ассистента TILO.RA CODER — офлайн-помощника по программированию
Скачать по ссылке
https://github.com/thetemirbolatov/TILO.RA_CODER_Dataset/releases/download/v1.0.0/tilora_knowledge_merged.jsonl
Состав
Источник
Язык
Записей
English coding… See the full description on the dataset page: https://huggingface.co/datasets/thetemirbolatov/TILO.RA_CODER_Dataset.qwen3-coder-480b-distill-mini
qwen3-coder-480b-distill-mini
Short Description
This dataset is distilled using Qwen3-Coder-480B-A35B-Instruct.We extracted 10,000 code questions from microsoft/rStar-Coder as seed problems, distilled them with 32K context, and after cleaning and filtering, 9,543 samples remain.License: Apache-2.0.
Dataset Overview
Seed Source: 10,000 code reasoning problems sampled from microsoft/rStar-Coder.
Distillation Model: Qwen3-Coder-480B-A35B-Instruct (480B… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/qwen3-coder-480b-distill-mini.cleand_microsoft_rStar-Coder元データ: https://huggingface.co/datasets/microsoft/rStar-Coder
データ件数: 269,863
平均トークン数: 11674
最大トークン数: 31,184
合計トークン数: 3,150,447,484
ファイル形式: JSONL
ファイルサイズ: 不明
加工内容
synthetic_sftを使用
トークン処理が重たいので、文字数でフィルター
seed_question < 6000
generation < 80000
thinkタグ除去 が中途半端なものを除外
トークナイズ処理(速度向上アップデート
繰り返し除去
Python_GOD_Coder_Omniforge_AI_12k
Python GOD Coder Omniforge AI 12k
Creator: Within Us AI
A 12,000-row mixed-format Python coding dataset designed as a sharpening corpus for building a small but dangerous Python specialist.
This dataset is intentionally focused on the practical behaviors that matter for a modern Python coding model:
implementation with tests
strict code-only instruction following
debugging and repair
refactoring for readability and production readiness
next-token code completion… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Python_GOD_Coder_Omniforge_AI_12k.prometech_inc_basic_coder
Prometech Inc Basic Coder Dataset
Dataset Overview
Filename: prometech_inc_basic_coder.jsonlTotal Entries: 263,903File Size: ~854 MBProvider: Prometech Bilgisayar Bilimleri AŞ
This dataset is a unified collection of high-quality coding instruction-following records, designed for fine-tuning Large Language Models (LLMs) or for use in Retrieval-Augmented Generation (RAG) systems. It aggregates data from multiple open-source high-quality datasets, synthetic documentation… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/prometech_inc_basic_coder.aihub-korean-education-instruct-sample
Korean Education Instruction Dataset (Sample)
Note: 이 데이터셋은 전체 데이터셋의 샘플 버전입니다 (카테고리별 최대 1,000건).
개요
AI Hub의 한국어 교육 데이터셋 13종을 sLLM 지시학습(Instruction Tuning)용으로 변환한 데이터셋입니다.
초등학교부터 고등학교까지의 다양한 교육 콘텐츠를 포함합니다.
데이터셋 통계
카테고리
데이터 수
math (수학)
1000
korean (국어)
1000
writing (글쓰기)
1000
career (진로)
1000
curriculum (교과)
1000
tutor (튜터링)
1000
총계
6000
사용 방법
from datasets import load_dataset
# 데이터셋 로드
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/neuralfoundry-coder/aihub-korean-education-instruct-sample.
