Team Ai
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01developer-lunark /kaidol-character-dataset KAIdol Character Chat Dataset 한국어 캐릭터 롤플레이 대화 데이터셋 📋 목차 개요 데이터셋 통계 데이터 형식 캐릭터 목록 품질 지표 사용 방법 학습 가이드 제한사항 라이선스 🎯 개요 KAIdol Character Chat Dataset은 41개 고유 캐릭터의 롤플레이 대화 데이터셋입니다. 각 캐릭터는 독특한 **음성 프로필(Voice Profile)**을 가지고 있으며, 이를 기반으로 일관된 성격과 말투를 유지합니다. 주요 특징 특징 설명 🎭 41개 캐릭터 다양한 성격, 배경, 말투를 가진 캐릭터 🗣️ 음성 프로필 시그니처 표현, 종결어미, 금지 표현 정의 📊 3가지 형식 SFT, DPO, Multiturn 학습 지원 ✅ 품질 검증 A등급 음성 프로필 일치율 (0.805) 🇰🇷 100% 한국어 자연스러운… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/kaidol-character-dataset.texttext-generation1K<n<10K0 likes442 downloads9mo agoHugging Face02developerjeremylive /claude-fable-5-claude-code-etheroi claude-fable-5 Agent Traces It's worth noting that our team was working with Glint-Research to collect as much fable data as possible. These are just the anonymized raw traces of both of our teams combined. This means that Glint-Research/Fable-5-traces was created from formatting and splitting up this same dataset. If you use one for your tune, don't use the other (it's the same exact data). For training on this dataset I recommend using the teich package to convert to openai… See the full description on the dataset page: https://huggingface.co/datasets/developerjeremylive/claude-fable-5-claude-code-etheroi.tabulartext-generationn<1K0 likes272 downloads4mo agoHugging Face03developer0hye /HanIFEval HanIFEval: a checker-consistent Korean translation of IFEval HanIFEval is a Korean translation of 429 items of google/IFEval (revision 966cd89). Every Korean prompt is translated together with the arguments of the rule-based checker that scores it, and the result is verified by code. Current version: v1.1 (2026-10-03). v1 is kept as the v1 config and as the v1 revision tag. What "checker-consistent" means here Wherever the checker scores a constraint, the Korean… See the full description on the dataset page: https://huggingface.co/datasets/developer0hye/HanIFEval.texttext-generationn<1K0 likes111 downloads7d agoHugging Face04VIRUS374 /ai-developer-dataset AI Developer Dataset A large-scale instruction-tuning dataset for fine-tuning an open-weight LLM into a universal AI developer assistant. The model trained on this dataset should be especially good at: PROGRAMMING + WEB DEVELOPMENT + UI/UX + ANIMATIONS + BOTS + SCRIPTS + AUTOMATION + BACKEND + API + DATABASES + DEBUGGING + LINUX + DEPLOYMENT + AI DEVELOPMENT + SECURITY. 📊 Statistics Total examples: 1,124,699 File size: 2.13 GB Format: JSONL (conversational)… See the full description on the dataset page: https://huggingface.co/datasets/VIRUS374/ai-developer-dataset.texttext-generation1M<n<10M0 likes70 downloads7d agoHugging Face05developer-lunark /kaidol-phase2-rp-base-v0.3 KAIDOL Phase 2 RP Base Dataset v0.3 Dataset Description KAIDOL Phase 2 RP Base v0.3 is a Korean-English bilingual conversational dataset designed for fine-tuning large language models (LLMs) for roleplay and character-based dialogue systems. This version includes GPT-Slop filtering to remove AI-sounding patterns and improve response quality. What's New in v0.3 GPT-Slop Filtering: Removed 1,529 samples containing AI-sounding patterns Cleaner Responses: Filtered… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/kaidol-phase2-rp-base-v0.3.texttext-generation10K<n<100K0 likes39 downloads9mo agoHugging Face06developer-lunark /korean-character-roleplay-sft Korean Character Roleplay SFT Dataset Character-based Korean roleplay conversation dataset for fine-tuning language models. Dataset Description This dataset contains high-quality Korean roleplay conversations between users and AI characters. Each conversation follows a specific character's personality, speech patterns, and voice profile. Dataset Statistics Split Samples Train 965 Test 108 Total 1,073 Quality Metrics Overall… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/korean-character-roleplay-sft.texttext-generation1K<n<10K0 likes26 downloads9mo agoHugging Face07KZ-Media-Developers /Chronos-Reasoning-v1 🌌 Chronos Omega Reasoning v1 (Alpha) 📝 Описание Chronos Omega Reasoning v1 — это высококачественный синтетический датасет, разработанный командой KZ Media Developers. Он предназначен для обучения языковых моделей глубокому логическому рассуждению (Chain-of-Thought) и формированию осознанного внутреннего монолога перед выдачей ответа. Датасет сфокусирован на сложных задачах в области математики, программирования, физики и лингвистического анализа… See the full description on the dataset page: https://huggingface.co/datasets/KZ-Media-Developers/Chronos-Reasoning-v1.texttext-generation1K<n<10K1 likes16 downloads6mo agoHugging Face08nativemind /developers-high-quality-mozgach developers-high-quality-mozgach Описание Высококачественные примеры для разработчиков, сгенерированные mozgach108. Датасет содержит отборные примеры для различных задач программирования: Написание кода Отладка Рефакторинг Архитектурные решения Code review Тестирование Особенность: высокое качество ответов, сгенерированных специализированной моделью mozgach108. Сгенерировано через Ollama (mozgach108:latest). Статистика Всего примеров: 1200… See the full description on the dataset page: https://huggingface.co/datasets/nativemind/developers-high-quality-mozgach.texttext-generation1K<n<10K0 likes14 downloads1y agoHugging Face09KZ-Media-Developers /Chronos-Thinking-v1-mini English: 🌌 Chronos-Thinking-v1-mini: The Genesis of Structured Reasoning Chronos-Thinking-v1-mini is a fundamental, high—density dataset designed to initialize deep reasoning processes in large language models (LLM). This dataset is the first step in the Chronos Super-AI project. Unlike mass datasets generated automatically, v1-mini relies on absolute quality and density of knowledge. He trains the model not just to answer questions, but to think like a system… See the full description on the dataset page: https://huggingface.co/datasets/KZ-Media-Developers/Chronos-Thinking-v1-mini.texttext-generationn<1K1 likes13 downloads5mo agoHugging Face10nativemind /developers-108-perfect developers-108-perfect Описание Датасет для обучения AI помощников разработчиков (continue.ai). Включает 7 специализированных сфер × 108 примеров: 073: Developer - написание кода 074: Code Reviewer - проверка и ревью кода 075: Architect - проектирование систем 076: DevOps Engineer - CI/CD и инфраструктура 077: QA Tester - тестирование и качество 078: Technical Writer - документация Плюс духовная сфера 001 + 1080 примеров Alpaca. Сгенерировано через Ollama… See the full description on the dataset page: https://huggingface.co/datasets/nativemind/developers-108-perfect.texttext-generationn<1K0 likes7 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.