Team Ai
Datasetpublic

michaelowusuntim6/cpp-qwen35

C / C++ Code Corpus Description Teaches domain-specific instruction following and code generation for this expert. Source shareAI/CodeChat Mxode/StackOverflow-QA-C-Language-40k dumb-dev/cpp-10k AmareshHebbar/leetcode-codegen-cpp Formatted for the MoE-orchestrator project (https://github.com/michaelowusuntim6/MoE-orchestrator). Expert target: code_cpp. Format Each record is a JSON object with a messages field formatted for Qwen3.5's… See the full description on the dataset page: https://huggingface.co/datasets/michaelowusuntim6/cpp-qwen35.

sourceHugging Faceotherupdated 3d agoView on Hugging Face
0likes30downloads
Dataset Card

C / C++ Code Corpus

Description

Teaches domain-specific instruction following and code generation for this expert.

Source

  • —shareAI/CodeChat
  • —Mxode/StackOverflow-QA-C-Language-40k
  • —dumb-dev/cpp-10k
  • —AmareshHebbar/leetcode-codegen-cpp

Formatted for the MoE-orchestrator project (https://github.com/michaelowusuntim6/MoE-orchestrator). Expert target: code_cpp.

Format

Each record is a JSON object with a messages field formatted for Qwen3.5's native chat template:

json
{"messages": [
  {"role": "system", "content": "..."},
  {"role": "user", "content": "..."},
  {"role": "assistant", "content": "..."}
]}

The records are consumed via tokenizer.apply_chat_template(). Special tokens (<|im_start|>, <|im_end|>) are added by the template, never embedded in content.

Splits

  • —train: 33,419 records
  • —val: 712 records

Usage

python
from datasets import load_dataset
ds = load_dataset("michaelowusuntim6/cpp-qwen35", split="train")
print(ds[0]["messages"])

License

mixed. Upstream sources keep their own licences - see the source list above and docs/DATASET_SOURCES.md in the MoE-orchestrator repository for per-source detail.