datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DBPedia_ClassesAbout Dataset
DBpedia (from "DB" for "database") is a project aiming to extract structured content from the information created in Wikipedia.
This is an extract of the data (after cleaning, kernel included) that provides taxonomic, hierarchical categories ("classes") for 342,782 wikipedia articles. There are 3 levels, with 9, 70 and 219 classes respectively.
A version of this dataset is a popular baseline for NLP/text classification tasks. This version of the dataset is much tougher… See the full description on the dataset page: https://huggingface.co/datasets/DeveloperOats/DBPedia_Classes.kaidol-character-dataset
KAIdol Character Chat Dataset
한국어 캐릭터 롤플레이 대화 데이터셋
📋 목차
개요
데이터셋 통계
데이터 형식
캐릭터 목록
품질 지표
사용 방법
학습 가이드
제한사항
라이선스
🎯 개요
KAIdol Character Chat Dataset은 41개 고유 캐릭터의 롤플레이 대화 데이터셋입니다. 각 캐릭터는 독특한 **음성 프로필(Voice Profile)**을 가지고 있으며, 이를 기반으로 일관된 성격과 말투를 유지합니다.
주요 특징
특징
설명
🎭 41개 캐릭터
다양한 성격, 배경, 말투를 가진 캐릭터
🗣️ 음성 프로필
시그니처 표현, 종결어미, 금지 표현 정의
📊 3가지 형식
SFT, DPO, Multiturn 학습 지원
✅ 품질 검증
A등급 음성 프로필 일치율 (0.805)
🇰🇷 100% 한국어
자연스러운… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/kaidol-character-dataset.CMAPSS_Jet_Engine_Simulated_Dataclaude-fable-5-claude-code-etheroi
claude-fable-5 Agent Traces
It's worth noting that our team was working with Glint-Research to collect as much fable data as possible.
These are just the anonymized raw traces of both of our teams combined. This means that Glint-Research/Fable-5-traces was created from formatting and splitting up this same dataset. If you use one for your tune, don't use the other (it's the same exact data).
For training on this dataset I recommend using the teich package to convert to openai… See the full description on the dataset page: https://huggingface.co/datasets/developerjeremylive/claude-fable-5-claude-code-etheroi.kimi-idol-dataset
Kimi K2 Idol Character Dataset
한국어 아이돌 캐릭터 대화 데이터셋 - Kimi K2 Teacher-Student Distillation 용
Overview
이 데이터셋은 moonshotai/Kimi-K2-Instruct-0905 모델을 Teacher로 사용하여 생성한 고품질 한국어 아이돌 캐릭터 대화 데이터셋입니다. 5명의 독특한 캐릭터가 6단계 thinking (CoT) 과정을 통해 썸남/썸녀 관계의 밀당 대화를 생성합니다.
Dataset Details
Total Samples: 942 (Train: 847, Eval: 95)
Characters: 5명 (강율, 서이안, 이지후, 차도하, 최민)
Teacher Model: moonshotai/Kimi-K2-Instruct-0905
Format: Messages format with </think> tags for thinking
Quality:… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/kimi-idol-dataset.stack-overflow-developer-surveycyber-cve2cwe-extension
cyber-cve2cwe-extension
Overview
The CVE-to-CWE classification task suffers from low macro-averaged F1 scores because many CWE categories appear only a handful of times in the training data. This dataset supplies additional examples for 36 low-frequency (tail) CWE classes with the aim of improving model performance on those categories and providing a reproducible record of how the training data for the companion model was extended. It is intended as a transparency… See the full description on the dataset page: https://huggingface.co/datasets/luca-software-developer/cyber-cve2cwe-extension.HanIFEval
HanIFEval: a checker-consistent Korean translation of IFEval
HanIFEval is a Korean translation of 429 items of
google/IFEval (revision
966cd89). Every Korean prompt is translated together with the arguments
of the rule-based checker that scores it, and the result is verified by
code.
Current version: v1.1 (2026-10-03). v1 is kept as the v1 config and
as the v1 revision tag.
What "checker-consistent" means here
Wherever the checker scores a constraint, the Korean… See the full description on the dataset page: https://huggingface.co/datasets/developer0hye/HanIFEval.ai-developer-dataset
AI Developer Dataset
A large-scale instruction-tuning dataset for fine-tuning an open-weight LLM into a universal AI developer assistant.
The model trained on this dataset should be especially good at:
PROGRAMMING + WEB DEVELOPMENT + UI/UX + ANIMATIONS + BOTS + SCRIPTS + AUTOMATION + BACKEND + API + DATABASES + DEBUGGING + LINUX + DEPLOYMENT + AI DEVELOPMENT + SECURITY.
📊 Statistics
Total examples: 1,124,699
File size: 2.13 GB
Format: JSONL (conversational)… See the full description on the dataset page: https://huggingface.co/datasets/VIRUS374/ai-developer-dataset.kaidol-phase2-rp-base-v0.3
KAIDOL Phase 2 RP Base Dataset v0.3
Dataset Description
KAIDOL Phase 2 RP Base v0.3 is a Korean-English bilingual conversational dataset designed for fine-tuning large language models (LLMs) for roleplay and character-based dialogue systems. This version includes GPT-Slop filtering to remove AI-sounding patterns and improve response quality.
What's New in v0.3
GPT-Slop Filtering: Removed 1,529 samples containing AI-sounding patterns
Cleaner Responses: Filtered… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/kaidol-phase2-rp-base-v0.3.github-top-developers
GitHub Top Developers by Year (2015-2025)
A derived dataset showing the top-ranked GitHub trending developers for each year, based on weighted scoring of their trending appearances across 41,841 raw data points from the Wayback Machine.
📊 Dataset Overview
Total Entries: 8,125 ranked developers
Years Covered: 2015 - 2025 (11 years)
Unique Developers: 4,763
Source: Derived from Wayback Machine snapshots of GitHub trending developers
Data Order: Sorted by year (descending:… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-developers.developer-reference-datasets
Developer Reference Datasets
Open, reproducible lookup tables that web and app developers reach for constantly — computed from first principles, not scraped, so every value is exact and re-runnable. CC BY 4.0.
Quick answers (straight from the data)
What is 16:9 in pixels? 1920×1080, 1280×720, 3840×2160. 9:16 (Stories, Reels, TikTok) is those flipped. → aspect-ratios, resolutions
What contrast ratio does WCAG require? 4.5:1 for normal text (AA), 3:1 for large… See the full description on the dataset page: https://huggingface.co/datasets/cleanorlabs/developer-reference-datasets.Million_News_HeadlinesAbout Dataset
Context
This contains data of news headlines published over a period of nineteen years.
Sourced from the reputable Australian news source ABC (Australian Broadcasting Corporation)
Agency Site: (http://www.abc.net.au)
Content
Format: CSV ; Single File
publish_date: Date of publishing for the article in yyyyMMdd format
headline_text: Text of the headline in Ascii , English , lowercase
Start Date: 2003-02-19 ; End Date: 2021-12-31
Inspiration
I look at this news dataset as a… See the full description on the dataset page: https://huggingface.co/datasets/DeveloperOats/Million_News_Headlines.korean-character-roleplay-sft
Korean Character Roleplay SFT Dataset
Character-based Korean roleplay conversation dataset for fine-tuning language models.
Dataset Description
This dataset contains high-quality Korean roleplay conversations between users and AI characters. Each conversation follows a specific character's personality, speech patterns, and voice profile.
Dataset Statistics
Split
Samples
Train
965
Test
108
Total
1,073
Quality Metrics
Overall… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/korean-character-roleplay-sft.aws-lambda-developer-guide-docsafrica-synth-realestate-developer-profiles-nigeria
Africa Synth Realestate Developer Profiles Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: infrastructure_transport - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-realestate-developer-profiles-nigeria.github-top-developers
GitHub Top Developers by Year (2015-2025)
A derived dataset showing the top-ranked GitHub trending developers for each year, based on weighted scoring of their trending appearances across 41,841 raw data points from the Wayback Machine.
📊 Dataset Overview
Total Entries: 8,125 ranked developers
Years Covered: 2015 - 2025 (11 years)
Unique Developers: 4,763
Source: Derived from Wayback Machine snapshots of GitHub trending developers
Data Order: Sorted by year (descending:… See the full description on the dataset page: https://huggingface.co/datasets/diamond-in/github-top-developers.github-top-developers
GitHub Top Developers by Year (2015-2025)
A derived dataset showing the top-ranked GitHub trending developers for each year, based on weighted scoring of their trending appearances across 41,841 raw data points from the Wayback Machine.
📊 Dataset Overview
Total Entries: 8,125 ranked developers
Years Covered: 2015 - 2025 (11 years)
Unique Developers: 4,763
Source: Derived from Wayback Machine snapshots of GitHub trending developers
Data Order: Sorted by year (descending:… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/github-top-developers.software-developer-hourly-rates-2026
Software Developer Hourly Rate Benchmarks 2026 (by Platform, Region & Tier)
Open benchmark of software/web developer hourly rates in 2026, segmented by technology platform, geographic region, and delivery tier (independent freelancer vs. team-based agency). Two files:
developer_hourly_rates_2026.csv — 120 rows: platform, region, tier, hourly_low_usd, hourly_median_usd, hourly_high_usd. Platforms include WordPress, Shopify, Magento, WooCommerce, custom (React/Next/Node), and… See the full description on the dataset page: https://huggingface.co/datasets/zimzum1984/software-developer-hourly-rates-2026.tablesChronos-Thinking-v1-mini
English:
🌌 Chronos-Thinking-v1-mini: The Genesis of Structured Reasoning
Chronos-Thinking-v1-mini is a fundamental, high—density dataset designed to initialize deep reasoning processes in large language models (LLM). This dataset is the first step in the Chronos Super-AI project. Unlike mass datasets generated automatically, v1-mini relies on absolute quality and density of knowledge. He trains the model not just to answer questions, but to think like a system… See the full description on the dataset page: https://huggingface.co/datasets/KZ-Media-Developers/Chronos-Thinking-v1-mini.encompass-developer-connect-index
Encompass Developer Connect — Hybrid Retrieval Index
Unofficial. Not affiliated with ICE Mortgage Technology. A community-built retrieval index over the publicly available Encompass Developer Connect documentation, intended as a developer reference and for educational/research use. For canonical, up-to-date documentation, always defer to the official source: https://developer.icemortgagetechnology.com/.
What this is
A hybrid (semantic + lexical) retrieval index over… See the full description on the dataset page: https://huggingface.co/datasets/Richie-rk/encompass-developer-connect-index.cowdbdeveloper-portfolio-ragChronos-Reasoning-v1
🌌 Chronos Omega Reasoning v1 (Alpha)
📝 Описание
Chronos Omega Reasoning v1 — это высококачественный синтетический датасет, разработанный командой KZ Media Developers. Он предназначен для обучения языковых моделей глубокому логическому рассуждению (Chain-of-Thought) и формированию осознанного внутреннего монолога перед выдачей ответа.
Датасет сфокусирован на сложных задачах в области математики, программирования, физики и лингвистического анализа… See the full description on the dataset page: https://huggingface.co/datasets/KZ-Media-Developers/Chronos-Reasoning-v1.developers-high-quality-mozgach
developers-high-quality-mozgach
Описание
Высококачественные примеры для разработчиков, сгенерированные mozgach108.
Датасет содержит отборные примеры для различных задач программирования:
Написание кода
Отладка
Рефакторинг
Архитектурные решения
Code review
Тестирование
Особенность: высокое качество ответов, сгенерированных специализированной моделью mozgach108.
Сгенерировано через Ollama (mozgach108:latest).
Статистика
Всего примеров: 1200… See the full description on the dataset page: https://huggingface.co/datasets/nativemind/developers-high-quality-mozgach.aicoolies-developer-tools-knowledge-graph
aicoolies-developer-tools-knowledge-graph
Public catalog dump from aicoolies.com: tools, comparisons, and scored reviews as JSON.
This dataset is not a coding agent. It does not edit repositories, run tools, or execute code. It is a machine-readable snapshot of the public aicoolies Developer Tools Knowledge Graph so humans and research agents can reuse the catalog without scraping HTML.
Canonical open-data page: https://aicoolies.com/data
Homepage: https://aicoolies.com… See the full description on the dataset page: https://huggingface.co/datasets/rasitakyol/aicoolies-developer-tools-knowledge-graph.test_el_talar
Dataset Card for "test_el_talar"
More Information needed
qiniu-developer-faq_ShareGPTgithub-top-developers
GitHub Top Developers by Year (2015-2025)
A derived dataset showing the top-ranked GitHub trending developers for each year, based on weighted scoring of their trending appearances across 41,841 raw data points from the Wayback Machine.
📊 Dataset Overview
Total Entries: 8,125 ranked developers
Years Covered: 2015 - 2025 (11 years)
Unique Developers: 4,763
Source: Derived from Wayback Machine snapshots of GitHub trending developers
Data Order: Sorted by year (descending:… See the full description on the dataset page: https://huggingface.co/datasets/Rendra8631/github-top-developers.
