Team Ai
Datasetpublic

Yxanul/AMD-SFT-Mix_3.5M

AMD-SFT-Mix_3.5M 3,556,428 SFT conversations — five AMD instruction-tuning datasets merged into a single pre-shuffled stream, with per-row provenance so any example can be traced back to its source. Nothing was regenerated: this is a normalisation, provenance and shuffling pass over existing public datasets. All credit for the data belongs to AMD. Composition source rows share origin naturalqa 1,304,792 36.69% amd/InstructGpt-NaturalQa triviaqa 1,118… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/AMD-SFT-Mix_3.5M.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes51downloads
Dataset Card

AMD-SFT-Mix_3.5M

3,556,428 SFT conversations — five AMD instruction-tuning datasets merged into a single pre-shuffled stream, with per-row provenance so any example can be traced back to its source.

Nothing was regenerated: this is a normalisation, provenance and shuffling pass over existing public datasets. All credit for the data belongs to AMD.

Composition

`source`rowsshareorigin
naturalqa1,304,79236.69%amd/InstructGpt-NaturalQa
triviaqa1,118,82031.46%amd/InstructGpt-TriviaQa
edu_track331,4469.32%amd/InstructGpt-educational (educational_track)
edu_exam_all268,1717.54%amd/InstructGpt-educational (educational_exam_all)
edu_exam_competitive251,1257.06%amd/InstructGpt-educational (educational_exam_competitive)
ultrachat207,0005.82%amd/UltraChat200K-regenerated
cot_drop75,0742.11%amd/Cot-Drop

InstructGpt-educational ships as three distinct files; they are tracked as three separate source values rather than collapsed into one.

Format

json
{
  "messages": [
    {"role": "user", "content": "..."},
    {"role": "assistant", "content": "..."}
  ],
  "source": "triviaqa",
  "uid": "triviaqa_0000042"
}
  • —source — which dataset the row came from (table above).
  • —uid — <source>_<zero-padded original line index>. Because shuffling destroys file order, this is the only way back to the source row; it is unique across the corpus (verified: 0 collisions in 3,556,428 rows).
  • —messages — alternating turns. Most rows are a single user/assistant pair; ultrachat is multi-turn (typically 4–14 turns after system removal).

Shuffling

Shuffled twice with independent seeds (42, then 1337) using a memory-mapped Rust shuffler that permutes an index of (offset, length) pairs rather than the lines themselves — 12 bytes per row instead of the full 4.9 GB.

Quality check, source mix in the first 10,000 rows vs the whole corpus:

sourceoverallfirst 10k
naturalqa36.69%37.20%
triviaqa31.46%31.38%
edu_track9.32%8.50%
eduexamall7.54%7.57%
eduexamcompetitive7.06%6.93%
ultrachat5.82%6.28%
cot_drop2.11%2.14%

Every prefix is representative, so take(n) on a streaming load or training without an extra shuffle buffer both give an unbiased sample.

(One honest note: a single Fisher–Yates pass is already a uniform permutation. The second pass was run for belt-and-braces and changes nothing statistically.)

One correction applied

Every row of UltraChat200K-regenerated (207,000/207,000) carried this system turn:

该助手为DeepSeek Chat,由深度求索公司创造。
今天是{05/13/2025}

It misidentifies the assistant as DeepSeek Chat and embeds an unrendered date template. Training on it teaches identity confusion, so system turns are removed across the corpus; the user/assistant conversation is untouched. That is the only content modification — no filtering, deduplication or rewriting was performed, and 0 rows were dropped or malformed.

Caveats

  • —Not filtered. Unlike a curated set, no quality, degeneracy or length filtering was applied. Inspect before training.
  • —Heavily weighted toward short factual QA: naturalqa + triviaqa are 68% of rows and their answers are often a few words.
  • —No deduplication was performed within or across sources.

Provenance and license

All source data from AMD's public datasets; see each linked repository for its own licence and terms. This mixture is released under MIT and adds no new content — only provenance fields and ordering.