michaelowusuntim6/code-review-qwen35
Code Review Corpus Description Teaches domain-specific instruction following and code generation for this expert. Source ronantakizawa/github-codereview - Other code-review-bench/code-review-bench - CC-BY-4.0 Formatted for the MoE-orchestrator project (https://github.com/michaelowusuntim6/MoE-orchestrator). Expert target: debug_review. Format Each record is a JSON object with a messages field formatted for Qwen3.5's native chat… See the full description on the dataset page: https://huggingface.co/datasets/michaelowusuntim6/code-review-qwen35.
Code Review Corpus
Description
Teaches domain-specific instruction following and code generation for this expert.
Source
- ronantakizawa/github-codereview - Other
- code-review-bench/code-review-bench - CC-BY-4.0
Formatted for the MoE-orchestrator project (https://github.com/michaelowusuntim6/MoE-orchestrator). Expert target: debug_review.
Format
Each record is a JSON object with a messages field formatted for Qwen3.5's native chat template:
{"messages": [
{"role": "system", "content": "..."},
{"role": "user", "content": "..."},
{"role": "assistant", "content": "..."}
]}The records are consumed via tokenizer.apply_chat_template(). Special tokens (<|im_start|>, <|im_end|>) are added by the template, never embedded in content.
Splits
train: 228,452 recordsval: 4,783 records
Usage
from datasets import load_dataset
ds = load_dataset("michaelowusuntim6/code-review-qwen35", split="train")
print(ds[0]["messages"])License
mixed (other / cc-by-4.0). Upstream sources keep their own licences - see the source list above and docs/DATASET_SOURCES.md in the MoE-orchestrator repository for per-source detail.
