datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swe-bench-multi-file-refactoring-sft-dpo-2026
💻 Enterprise Autonomous SWE-bench AI & Multi-File Code Refactoring SFT/DPO Dataset (2026)
High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step call-stack Chain-of-Thought (<thought>) reasoning trees for fine-tuning LLMs (Qwen-2.5-Coder, Llama-3.3, DeepSeek-R1-Distill, Mistral) into Autonomous Software Engineers and SWE-bench Benchmark Agents.
📊 Dataset Architecture & Highlights
Multi-Turn Code Reviews:… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/swe-bench-multi-file-refactoring-sft-dpo-2026.Refactor-Dialogue-1.4k-Multi-turn-Refactoring-Conversations
Refactor-Dialogue-1.4k — Multi-turn Refactoring Conversations
Synthetic dataset for fine-tuning coding-focused LLMs on multi-turn
refactoring dialogues. Generated with
Dataset Generator —
an open-source pipeline for building high-quality training data.
Overview
1,414 multi-turn conversations across 3 refactoring categories. Each example
is a 4-message dialogue: user pastes real code → assistant refactors with a
short explanation → user follows up with a constraint or… See the full description on the dataset page: https://huggingface.co/datasets/AronDaron/Refactor-Dialogue-1.4k-Multi-turn-Refactoring-Conversations.Refactor-Dialogue-1.4k-Multi-turn-Refactoring-Conversations-Reasoning
Refactor-Dialogue-1.4k-Reasoning — Multi-turn Refactoring Conversations with <think> Reasoning
Synthetic dataset for fine-tuning reasoning-style coding LLMs on
multi-turn refactoring dialogues. Every assistant turn carries a
first-person <think>...</think> internal monologue before the actual
response — DeepSeek-R1 / Qwen3-thinking convention, broadest trainer
compatibility out of the box.
Generated with
Dataset Generator —
an open-source pipeline for building high-quality… See the full description on the dataset page: https://huggingface.co/datasets/AronDaron/Refactor-Dialogue-1.4k-Multi-turn-Refactoring-Conversations-Reasoning.
