Team Ai
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01xlelords /orbis-coder Orbis Coder Dataset (10K) A coding-first instruction dataset to train or fine-tune assistants that behave like Orbis Coder — friendly, concise, practical, and focused on helping people build, debug, and ship software. This dataset is intended for: instruction-tuning / SFT LoRA / QLoRA fine-tunes “persona + skill” alignment for coding assistants quick experiments + dataset viewer testing What this dataset contains Most rows are coding help across many languages and… See the full description on the dataset page: https://huggingface.co/datasets/xlelords/orbis-coder.texttext-generation10K<n<100K0 likes112 downloads9mo agoHugging Face02thetemirbolatov /TILO.RA_CODER_Dataset TILO.RA CODER Dataset Объединённый русско-английский датасет для обучения и поиска по коду. Формат — пары question / code: вопрос на естественном языке → готовый код-ответ. Датасет собран для локального ассистента TILO.RA CODER — офлайн-помощника по программированию Скачать по ссылке https://github.com/thetemirbolatov/TILO.RA_CODER_Dataset/releases/download/v1.0.0/tilora_knowledge_merged.jsonl Состав Источник Язык Записей English coding… See the full description on the dataset page: https://huggingface.co/datasets/thetemirbolatov/TILO.RA_CODER_Dataset.texttext-generation100K<n<1M1 likes112 downloads26d agoHugging Face03Jackrong /qwen3-coder-480b-distill-mini qwen3-coder-480b-distill-mini Short Description This dataset is distilled using Qwen3-Coder-480B-A35B-Instruct.We extracted 10,000 code questions from microsoft/rStar-Coder as seed problems, distilled them with 32K context, and after cleaning and filtering, 9,543 samples remain.License: Apache-2.0. Dataset Overview Seed Source: 10,000 code reasoning problems sampled from microsoft/rStar-Coder. Distillation Model: Qwen3-Coder-480B-A35B-Instruct (480B… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/qwen3-coder-480b-distill-mini.texttext-classification1K<n<10K14 likes72 downloads1y agoHugging Face04LLMTeamAkiyama /cleand_microsoft_rStar-Coder元データ: https://huggingface.co/datasets/microsoft/rStar-Coder データ件数: 269,863 平均トークン数: 11674 最大トークン数: 31,184 合計トークン数: 3,150,447,484 ファイル形式: JSONL ファイルサイズ: 不明 加工内容 synthetic_sftを使用 トークン処理が重たいので、文字数でフィルター seed_question < 6000 generation < 80000 thinkタグ除去 が中途半端なものを除外 トークナイズ処理(速度向上アップデート 繰り返し除去 tabularquestion-answering100K<n<1M0 likes66 downloads1y agoHugging Face05WithinUsAI /Python_GOD_Coder_Omniforge_AI_12k Python GOD Coder Omniforge AI 12k Creator: Within Us AI A 12,000-row mixed-format Python coding dataset designed as a sharpening corpus for building a small but dangerous Python specialist. This dataset is intentionally focused on the practical behaviors that matter for a modern Python coding model: implementation with tests strict code-only instruction following debugging and repair refactoring for readability and production readiness next-token code completion… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Python_GOD_Coder_Omniforge_AI_12k.texttext-generation10K<n<100K1 likes59 downloads7mo agoHugging Face06pthinc /prometech_inc_basic_coder Prometech Inc Basic Coder Dataset Dataset Overview Filename: prometech_inc_basic_coder.jsonlTotal Entries: 263,903File Size: ~854 MBProvider: Prometech Bilgisayar Bilimleri AŞ This dataset is a unified collection of high-quality coding instruction-following records, designed for fine-tuning Large Language Models (LLMs) or for use in Retrieval-Augmented Generation (RAG) systems. It aggregates data from multiple open-source high-quality datasets, synthetic documentation… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/prometech_inc_basic_coder.tabulartext-generation100K<n<1M0 likes21 downloads9mo agoHugging Face07neuralfoundry-coder /aihub-korean-education-instruct-sample Korean Education Instruction Dataset (Sample) Note: 이 데이터셋은 전체 데이터셋의 샘플 버전입니다 (카테고리별 최대 1,000건). 개요 AI Hub의 한국어 교육 데이터셋 13종을 sLLM 지시학습(Instruction Tuning)용으로 변환한 데이터셋입니다. 초등학교부터 고등학교까지의 다양한 교육 콘텐츠를 포함합니다. 데이터셋 통계 카테고리 데이터 수 math (수학) 1000 korean (국어) 1000 writing (글쓰기) 1000 career (진로) 1000 curriculum (교과) 1000 tutor (튜터링) 1000 총계 6000 사용 방법 from datasets import load_dataset # 데이터셋 로드 dataset =… See the full description on the dataset page: https://huggingface.co/datasets/neuralfoundry-coder/aihub-korean-education-instruct-sample.texttext-generation1K<n<10K0 likes19 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.