Team Ai
Datasetpublic

atlas-institute/code-trainer-v9-mixed

code-trainer-v9-mixed 40,401-row mixed training dataset for supervised fine-tuning (SFT) in the Code-Trainer / RTPI pipeline. Used for both Qwen 14B (V9 SFT) and Gemma 26B (aggressive-full1 SFT) training. Composition Slice Source Rows (train) Purpose A -- Code generation cmndcntrlcyber/code-trainer-offsec-dataset (8K subsample) 7,074 Preserve code-gen quality B -- Tool calling glaiveai/glaive-function-calling-v2 (19K cap) ~15,125 High-density tool… See the full description on the dataset page: https://huggingface.co/datasets/atlas-institute/code-trainer-v9-mixed.

sourceHugging Faceapache-2.0updated 22d agoView on Hugging Face
0likes74downloads
Dataset Card

code-trainer-v9-mixed

40,401-row mixed training dataset for supervised fine-tuning (SFT) in the Code-Trainer / RTPI pipeline. Used for both Qwen 14B (V9 SFT) and Gemma 26B (aggressive-full1 SFT) training.

Composition

SliceSourceRows (train)Purpose
A -- Code generationcmndcntrlcyber/code-trainer-offsec-dataset (8K subsample)7,074Preserve code-gen quality
B -- Tool callingglaiveai/glaive-function-calling-v2 (19K cap)~15,125High-density tool calling
B+ -- Multi-tool syntheticSynthetic from Slice B pairs (~2K)~2,000Multi-call per turn
C -- Agentic multi-turngreghavens/fable-5-coding-and-debugging-traces (10K cap)8,994Multi-step agent behaviour
D -- English instructionteknium/OpenHermes-2.5 (8K cap)7,208Language anchor
  • —Tool coverage: 25,964 / 40,401 train rows (64.3%) contain tool definitions
  • —Format: Unified ChatML, tool calls formatted via apply_chat_template(tools=...)
  • —Split: 90/10 train/validation (seed 42)

Schema

Each row has a messages list in ChatML format, plus metadata columns:

  • —slice: which data slice (A, B, B+, C, D)
  • —source: original dataset name
  • —category: content category
  • —has_tools: whether the row contains tool definitions
  • —n_turns: number of conversation turns

How to use

python
from datasets import load_dataset
ds = load_dataset("cmndcntrlcyber/code-trainer-v9-mixed")
print(f"Train: {len(ds['train'])}, Val: {len(ds['validation'])}")
print(ds["train"][0]["messages"][:2])

Reproducibility

bash
python -m src.phase2_preprocessing.scripts.build_v9_mixed_dataset --config src/config/config.yaml