Team Ai
21 results

refactoring

AISE-TUDelft /MOSAIC-Refactoring Agentic Pull Request Dataset Dataset Overview The dataset contains 4,910,698 Pull Requests in total, consisting of 4,392,818 agent-authored PRs from 10 agents and 517,880 human-authored PRs. The agent-authored PRs come from Claude, Codegen, Codex, Copilot, Cosine, Cursor, Devin, Jules, Junie, and OpenHands. A summary of the dataset is presented below. Cohort Pull Requests Merged Pull Requests Repositories Sum of Additions Sum of Deletions Humans 517880… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/MOSAIC-Refactoring.tabular10M<n<100M4 likes38k downloads3mo agoHugging Facebeatsprom /swe-bench-multi-file-refactoring-sft-dpo-2026 💻 Enterprise Autonomous SWE-bench AI & Multi-File Code Refactoring SFT/DPO Dataset (2026) High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step call-stack Chain-of-Thought (<thought>) reasoning trees for fine-tuning LLMs (Qwen-2.5-Coder, Llama-3.3, DeepSeek-R1-Distill, Mistral) into Autonomous Software Engineers and SWE-bench Benchmark Agents. 📊 Dataset Architecture & Highlights Multi-Turn Code Reviews:… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/swe-bench-multi-file-refactoring-sft-dpo-2026.texttext-generationn<1K0 likes203 downloads1mo agoHugging Faceinaesh-joshi /MOSAIC-Refactoring-copy Post-Processed Pull Request Dataset Dataset Overview The dataset contains a total of 2393 Pull Requests from OpenHands. It also includes additional activity metadata such as repository state snapshots and modified-file records. A summary of the dataset is presented below. Cohort Number of PRs Number of Merged PRs Unique Repositories Sum of Additions Sum of Deletions OpenHands 2393 1737 667 14186956 1732609 Total 2393 1737 667 14186956 1732609… See the full description on the dataset page: https://huggingface.co/datasets/inaesh-joshi/MOSAIC-Refactoring-copy.1 likes81 downloads4mo agoHugging FaceAronDaron /Refactor-Dialogue-1.4k-Multi-turn-Refactoring-Conversations Refactor-Dialogue-1.4k — Multi-turn Refactoring Conversations Synthetic dataset for fine-tuning coding-focused LLMs on multi-turn refactoring dialogues. Generated with Dataset Generator — an open-source pipeline for building high-quality training data. Overview 1,414 multi-turn conversations across 3 refactoring categories. Each example is a 4-message dialogue: user pastes real code → assistant refactors with a short explanation → user follows up with a constraint or… See the full description on the dataset page: https://huggingface.co/datasets/AronDaron/Refactor-Dialogue-1.4k-Multi-turn-Refactoring-Conversations.texttext-generation1K<n<10K0 likes79 downloads5mo agoHugging Facesukosmos /repro-codetaste-can-llms-generate-human-level-code-refactorings-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes50 downloads2mo agoHugging FaceAronDaron /Refactor-Dialogue-1.4k-Multi-turn-Refactoring-Conversations-Reasoning Refactor-Dialogue-1.4k-Reasoning — Multi-turn Refactoring Conversations with <think> Reasoning Synthetic dataset for fine-tuning reasoning-style coding LLMs on multi-turn refactoring dialogues. Every assistant turn carries a first-person <think>...</think> internal monologue before the actual response — DeepSeek-R1 / Qwen3-thinking convention, broadest trainer compatibility out of the box. Generated with Dataset Generator — an open-source pipeline for building high-quality… See the full description on the dataset page: https://huggingface.co/datasets/AronDaron/Refactor-Dialogue-1.4k-Multi-turn-Refactoring-Conversations-Reasoning.texttext-generation1K<n<10K0 likes47 downloads5mo agoHugging Face