datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gordon-ramsay-code-review-v2
Gordon Ramsay Code Review & Auditor Corpus v2 (dcmutlu/gordon-ramsay-code-review-v2)
A high-density synthetic dataset of 10,000 multi-turn code review pairs designed to fine-tune open-weight reasoners (specifically Qwen2.5-Coder-7B-Instruct) into Chef Gordon Ramsay: Sovereign Executive Code Auditor and Supreme Software Gastronomer.
🍳 Dataset Overview
This dataset merges rigorous computer science diagnostics (Abstract Syntax Tree inspection, concurrency lifecycle… See the full description on the dataset page: https://huggingface.co/datasets/dcmutlu/gordon-ramsay-code-review-v2.code-reviewA Scrape of the codereview stack exchange, good for high quality code
code-review-qwen35
Code Review Corpus
Description
Teaches domain-specific instruction following and code generation for this expert.
Source
ronantakizawa/github-codereview - Other
code-review-bench/code-review-bench - CC-BY-4.0
Formatted for the MoE-orchestrator project
(https://github.com/michaelowusuntim6/MoE-orchestrator). Expert target:
debug_review.
Format
Each record is a JSON object with a messages field formatted for Qwen3.5's
native chat… See the full description on the dataset page: https://huggingface.co/datasets/michaelowusuntim6/code-review-qwen35.code-review-dpo-3k
Code Review DPO Pairs (3K)
DPO preference pairs for training LLMs to produce specific, actionable, educational code reviews.
Dataset Description
3,000 preference pairs across 4 programming languages:
Language
Examples
Python
~64%
JavaScript
~12%
TypeScript
~12%
Go
~12%
8 review scenarios covering real-world code quality issues:
SQL injection & security vulnerabilities
XSS via innerHTML
Hardcoded credentials
Resource leaks (unclosed… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/code-review-dpo-3k.gordon-ramsay-code-review
gordon-ramsay-code-review
Autonomous synthetic pretraining dataset synthesized by JESUS Sovereign Forge.
Synthesized via JESUS Sovereign Cloud Model Forge (hf-colab-forge) for native byte-level micro-transformers (Atom GPT) and LLM fine-tuning.
Dataset Summary
Metric
Value
Total Scenarios
500
Train Samples
450
Validation Samples
50
Total Byte Tokens
819,927
Train Tokens
737,852
Val Tokens
82,075
Vocab Size
258 (UTF-8 Bytes + BOS/PAD)… See the full description on the dataset page: https://huggingface.co/datasets/dcmutlu/gordon-ramsay-code-review.code-review-findings-samples
Code Review Findings Samples
Curated synthetic examples for evaluating automated code review pipelines — especially the
AI Code Reviewer MCP stack built with
Qwen3.6-27B.
Each row contains a short code snippet, the analysis type, and a structured JSON output that
matches the review contract used by ImTamsi/qwen3.6-27b-code-reviewer.
Dataset structure
Column
Description
id
Stable sample identifier
analysis_type
review, bugs, security, performance… See the full description on the dataset page: https://huggingface.co/datasets/ImTamsi/code-review-findings-samples.lunamax-multilingual-code-review-50
LunaMax Multilingual Code Review 50
A 50-record synthetic multilingual code-review dataset generated with ChatGPT LunaMax.
Every record is a code-review task in user / assistant format. The set spans multiple languages and review scenarios, including correctness, debugging, API usage, security, and implementation behavior.
Dataset Size
Metric
Count
Final records
50
Unique records
50
Fresh GPT-5.6 Sol audit coverage
50
Accepted unchanged
48… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/lunamax-multilingual-code-review-50.
