Team Ai
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fasterinnerlooper /codereviewertabular100K<n<1M1 likes210 downloads3y agoHugging Face02auphong2707 /dl4se-code-review-datasettext100K<n<1M0 likes86 downloads2y agoHugging Face03dcmutlu /gordon-ramsay-code-review-v2 Gordon Ramsay Code Review & Auditor Corpus v2 (dcmutlu/gordon-ramsay-code-review-v2) A high-density synthetic dataset of 10,000 multi-turn code review pairs designed to fine-tune open-weight reasoners (specifically Qwen2.5-Coder-7B-Instruct) into Chef Gordon Ramsay: Sovereign Executive Code Auditor and Supreme Software Gastronomer. 🍳 Dataset Overview This dataset merges rigorous computer science diagnostics (Abstract Syntax Tree inspection, concurrency lifecycle… See the full description on the dataset page: https://huggingface.co/datasets/dcmutlu/gordon-ramsay-code-review-v2.texttext-generation10K<n<100K0 likes69 downloads1mo agoHugging Face04VatsaDev /code-reviewA Scrape of the codereview stack exchange, good for high quality code texttext-generation10K<n<100K3 likes67 downloads3y agoHugging Face05ewhk9887 /korean_code_reviews_from_githubtext10K<n<100K1 likes47 downloads2y agoHugging Face06provie17 /react_native_code_review-reasoning-SFTtext1K<n<10K3 likes46 downloads1y agoHugging Face07Tomo-Melb /CodeReviewQAgated CodeReviewQA: The Code Review Comprehension Assessment for Large Language Models The task of automated code refinement aims to automate the developer's perspective in resolving an actionable code review comment provided by a reviewer. This is a generative task, where the LLM is required to revise a pre-review code submission with respect to the natural language code review comment to produce an intended post-review code revision. CodeReviewQA further breaks down this generative task… See the full description on the dataset page: https://huggingface.co/datasets/Tomo-Melb/CodeReviewQA.textmultiple-choicen<1K5 likes46 downloads7mo agoHugging Face08michaelowusuntim6 /code-review-qwen35 Code Review Corpus Description Teaches domain-specific instruction following and code generation for this expert. Source ronantakizawa/github-codereview - Other code-review-bench/code-review-bench - CC-BY-4.0 Formatted for the MoE-orchestrator project (https://github.com/michaelowusuntim6/MoE-orchestrator). Expert target: debug_review. Format Each record is a JSON object with a messages field formatted for Qwen3.5's native chat… See the full description on the dataset page: https://huggingface.co/datasets/michaelowusuntim6/code-review-qwen35.texttext-generation100K<n<1M0 likes35 downloads6d agoHugging Face09TaskPuppyAI /qwen3.8-targeted-code-review-50 Qwen3.8 Max Targeted Code Review 50 A 50-record synthetic code-review dataset generated with Qwen3.8 Max and reviewed with ChatGPT 5.6 Sol High. The dataset is designed around precise defect identification, especially resisting unsupported or speculative code-review findings. Dataset Summary The publication artifact contains 50 unique records using the schema: { "user": "...", "assistant": "..." } Recovered final-artifact characteristics: records: 50 unique… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/qwen3.8-targeted-code-review-50.textn<1K0 likes30 downloads1mo agoHugging Face10stindardlogic /code-review-dpo-3k Code Review DPO Pairs (3K) DPO preference pairs for training LLMs to produce specific, actionable, educational code reviews. Dataset Description 3,000 preference pairs across 4 programming languages: Language Examples Python ~64% JavaScript ~12% TypeScript ~12% Go ~12% 8 review scenarios covering real-world code quality issues: SQL injection & security vulnerabilities XSS via innerHTML Hardcoded credentials Resource leaks (unclosed… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/code-review-dpo-3k.texttext-generation1K<n<10K0 likes27 downloads3mo agoHugging Face11abdo1819 /arabic-english-code-switching-review-annotations Review Annotations for Arabic-English Code-Switching Speech This metadata-only dataset publishes review decisions and transcript-correction deltas for MohamedRashad/arabic-english-code-switching. It contains no human audio, no local file paths, no raw review notes, and no copies of unchanged upstream transcripts. The annotations are pinned to upstream revision 4a3bffc45219c35949470de32b8d4cb328b0ce11 and join by upstream_row_index. Coverage and outcomes The… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-review-annotations.tabularautomatic-speech-recognition10K<n<100K0 likes26 downloads2mo agoHugging Face12TaskPuppyAI /qwen3.8-contrastive-code-review-12 Qwen3.8 Max Contrastive Code Review 12 A 12-record synthetic contrastive code-review dataset generated with Qwen3.8 Max and reviewed with ChatGPT 5.6 Sol High. The dataset combines direct code review, proposed-finding adjudication, and review rewriting in a compact set intended to train precise, evidence-grounded review behavior. Dataset Summary The publication artifact contains 12 unique records using the schema: { "instruction": "...", "input": "..."… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/qwen3.8-contrastive-code-review-12.textn<1K0 likes26 downloads1mo agoHugging Face13ewhk9887 /korean_code_reviewtextn<1K1 likes22 downloads2y agoHugging Face14gram-chan-jp /code-review-dataset-ja Japanese Code Review Dataset (500 Samples) A dataset of 500 code review pairs (buggy code + fixed code) with Japanese review comments. Designed for training and evaluating code review assistance models. Total samples: 500 Languages: Python (220), JavaScript (136), Go (54), Rust (52), TypeScript (38) Difficulties: Easy (135), Medium (259), Hard (106) Bug types: Logic Error (120), Null Pointer (81), Off-by-One (80), Edge Case (77), Type Error (60), Security (42), Performance (40)… See the full description on the dataset page: https://huggingface.co/datasets/gram-chan-jp/code-review-dataset-ja.textn<1K0 likes22 downloads5mo agoHugging Face15dcmutlu /gordon-ramsay-code-review gordon-ramsay-code-review Autonomous synthetic pretraining dataset synthesized by JESUS Sovereign Forge. Synthesized via JESUS Sovereign Cloud Model Forge (hf-colab-forge) for native byte-level micro-transformers (Atom GPT) and LLM fine-tuning. Dataset Summary Metric Value Total Scenarios 500 Train Samples 450 Validation Samples 50 Total Byte Tokens 819,927 Train Tokens 737,852 Val Tokens 82,075 Vocab Size 258 (UTF-8 Bytes + BOS/PAD)… See the full description on the dataset page: https://huggingface.co/datasets/dcmutlu/gordon-ramsay-code-review.texttext-generationn<1K0 likes21 downloads1mo agoHugging Face16ImTamsi /code-review-findings-samples Code Review Findings Samples Curated synthetic examples for evaluating automated code review pipelines — especially the AI Code Reviewer MCP stack built with Qwen3.6-27B. Each row contains a short code snippet, the analysis type, and a structured JSON output that matches the review contract used by ImTamsi/qwen3.6-27b-code-reviewer. Dataset structure Column Description id Stable sample identifier analysis_type review, bugs, security, performance… See the full description on the dataset page: https://huggingface.co/datasets/ImTamsi/code-review-findings-samples.texttext-generationn<1K0 likes15 downloads3mo agoHugging Face17TaskPuppyAI /lunamax-multilingual-code-review-50 LunaMax Multilingual Code Review 50 A 50-record synthetic multilingual code-review dataset generated with ChatGPT LunaMax. Every record is a code-review task in user / assistant format. The set spans multiple languages and review scenarios, including correctness, debugging, API usage, security, and implementation behavior. Dataset Size Metric Count Final records 50 Unique records 50 Fresh GPT-5.6 Sol audit coverage 50 Accepted unchanged 48… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/lunamax-multilingual-code-review-50.texttext-generationn<1K0 likes15 downloads1mo agoHugging Face18amarshiv86 /p06-code-review-datasettextn<1K0 likes10 downloads4mo agoHugging Face19muttakin2301 /contextual_code_review_fulltext100K<n<1M0 likes8 downloads2y agoHugging Face20AshokChakravarthy /ai-code-review-datasettextn<1K0 likes6 downloads5mo agoHugging Face21AzerChakir /CodeReviewWithSummaryQAgatedtextn<1K0 likes5 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.