Yxanul/AMD-SFT-Mix_3.5M
AMD-SFT-Mix_3.5M 3,556,428 SFT conversations — five AMD instruction-tuning datasets merged into a single pre-shuffled stream, with per-row provenance so any example can be traced back to its source. Nothing was regenerated: this is a normalisation, provenance and shuffling pass over existing public datasets. All credit for the data belongs to AMD. Composition source rows share origin naturalqa 1,304,792 36.69% amd/InstructGpt-NaturalQa triviaqa 1,118… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/AMD-SFT-Mix_3.5M.
AMD-SFT-Mix_3.5M
3,556,428 SFT conversations — five AMD instruction-tuning datasets merged into a single pre-shuffled stream, with per-row provenance so any example can be traced back to its source.
Nothing was regenerated: this is a normalisation, provenance and shuffling pass over existing public datasets. All credit for the data belongs to AMD.
Composition
InstructGpt-educational ships as three distinct files; they are tracked as three separate source values rather than collapsed into one.
Format
{
"messages": [
{"role": "user", "content": "..."},
{"role": "assistant", "content": "..."}
],
"source": "triviaqa",
"uid": "triviaqa_0000042"
}source— which dataset the row came from (table above).uid—<source>_<zero-padded original line index>. Because shuffling destroys file order, this is the only way back to the source row; it is unique across the corpus (verified: 0 collisions in 3,556,428 rows).messages— alternating turns. Most rows are a single user/assistant pair;ultrachatis multi-turn (typically 4–14 turns after system removal).
Shuffling
Shuffled twice with independent seeds (42, then 1337) using a memory-mapped Rust shuffler that permutes an index of (offset, length) pairs rather than the lines themselves — 12 bytes per row instead of the full 4.9 GB.
Quality check, source mix in the first 10,000 rows vs the whole corpus:
Every prefix is representative, so take(n) on a streaming load or training without an extra shuffle buffer both give an unbiased sample.
(One honest note: a single Fisher–Yates pass is already a uniform permutation. The second pass was run for belt-and-braces and changes nothing statistically.)
One correction applied
Every row of UltraChat200K-regenerated (207,000/207,000) carried this system turn:
该助手为DeepSeek Chat,由深度求索公司创造。
今天是{05/13/2025}It misidentifies the assistant as DeepSeek Chat and embeds an unrendered date template. Training on it teaches identity confusion, so system turns are removed across the corpus; the user/assistant conversation is untouched. That is the only content modification — no filtering, deduplication or rewriting was performed, and 0 rows were dropped or malformed.
Caveats
- Not filtered. Unlike a curated set, no quality, degeneracy or length filtering was applied. Inspect before training.
- Heavily weighted toward short factual QA:
naturalqa+triviaqaare 68% of rows and their answers are often a few words. - No deduplication was performed within or across sources.
Provenance and license
All source data from AMD's public datasets; see each linked repository for its own licence and terms. This mixture is released under MIT and adds no new content — only provenance fields and ordering.
