atlas-institute/code-trainer-v9-mixed
code-trainer-v9-mixed 40,401-row mixed training dataset for supervised fine-tuning (SFT) in the Code-Trainer / RTPI pipeline. Used for both Qwen 14B (V9 SFT) and Gemma 26B (aggressive-full1 SFT) training. Composition Slice Source Rows (train) Purpose A -- Code generation cmndcntrlcyber/code-trainer-offsec-dataset (8K subsample) 7,074 Preserve code-gen quality B -- Tool calling glaiveai/glaive-function-calling-v2 (19K cap) ~15,125 High-density tool… See the full description on the dataset page: https://huggingface.co/datasets/atlas-institute/code-trainer-v9-mixed.
code-trainer-v9-mixed
40,401-row mixed training dataset for supervised fine-tuning (SFT) in the Code-Trainer / RTPI pipeline. Used for both Qwen 14B (V9 SFT) and Gemma 26B (aggressive-full1 SFT) training.
Composition
- Tool coverage: 25,964 / 40,401 train rows (64.3%) contain tool definitions
- Format: Unified ChatML, tool calls formatted via
apply_chat_template(tools=...) - Split: 90/10 train/validation (seed 42)
Schema
Each row has a messages list in ChatML format, plus metadata columns:
slice: which data slice (A, B, B+, C, D)source: original dataset namecategory: content categoryhas_tools: whether the row contains tool definitionsn_turns: number of conversation turns
How to use
from datasets import load_dataset
ds = load_dataset("cmndcntrlcyber/code-trainer-v9-mixed")
print(f"Train: {len(ds['train'])}, Val: {len(ds['validation'])}")
print(ds["train"][0]["messages"][:2])Reproducibility
python -m src.phase2_preprocessing.scripts.build_v9_mixed_dataset --config src/config/config.yaml