FlexiSLM/FlexiSLM-Data-4M-s2s
FlexiSLM-Data — Speech-to-Speech Part (4M) Paper: https://arxiv.org/abs/2606.31247 Demo page: https://flexislm.github.io/ Code: https://github.com/AmphionTeam/FlexiSLM FlexiSLM-Data is a large-scale, single-turn English speech-to-speech dialogue dataset for training FlexiSLM, a spoken language model. This repository contains the paired prompt-and-response audio portion of the release in WebDataset format. Related data releases FlexiSLM/FlexiSLM-Data-4M-s2s (this repo)… See the full description on the dataset page: https://huggingface.co/datasets/FlexiSLM/FlexiSLM-Data-4M-s2s.
FlexiSLM-Data — Speech-to-Speech Part (4M)
  
- Paper: https://arxiv.org/abs/2606.31247
- Demo page: https://flexislm.github.io/
- Code: https://github.com/AmphionTeam/FlexiSLM
FlexiSLM-Data is a large-scale, single-turn English speech-to-speech dialogue dataset for training FlexiSLM, a spoken language model. This repository contains the paired prompt-and-response audio portion of the release in WebDataset format.
Related data releases
- **FlexiSLM/FlexiSLM-Data-4M-s2s** (this repo) contains 4M samples of speech-to-speech dialog data, Its question audios are stored in 44k wav, while response audios are stored in mp3. Size is around 2.8T.
- **FlexiSLM/FlexiSLM-Data-2M-s2s-compact** is the filtered, mp3-compressed subset of FlexiSLM/FlexiSLM-Data-4M-s2s that takes only 385G storage while maintaining high data quality.
- **FlexiSLM/FlexiSLM-Data-5M-t2t** provides the raw text input-output pairs without audio.
Data construction pipeline
- Prompt collection and response generation (released in **FlexiSLM/FlexiSLM-Data-5M-t2t**) Text prompts are collected from public QA, instruction-following, and dialogue datasets (see the table below). For multi-turn datasets, the user's first-turn utterance is used as the prompt; samples whose first turn is a generic greeting (e.g. "Hello") are skipped. Then, all text responses are generated with **Qwen3-Omni-30B-A3B**; the responses shipped with the source datasets are discarded. We use a spoken language model rather than a text-only LLM because Qwen3-Omni produces speech-friendly responses: short, conversational, and free of formatting that does not transfer to speech (bullet points, code blocks, long enumerations). The text version of the prompt is fed to the model rather than its synthesized speech, since text-input responses are typically more accurate.
- Speech synthesis. (released in **FlexiSLM/FlexiSLM-Data-4M-s2s**) Responses are synthesized with Qwen3-TTS using the fixed speaker "Ryan". Prompts are synthesized with **Fish-Audio TTS**, with speaker prompts randomly sampled from English Emilia utterances longer than 5 seconds. This yields the 4,222,459 samples and 26,735 hours of audio shipped here.
- Quality filtering and mp3-format compression. (released in **FlexiSLM/FlexiSLM-Data-2M-s2s-compact**) Before any shards are written, this release applies prompt deduplication and response-text filtering. The kept subset uses normalized prompt deduplication first, then drops responses that contain
!, contain?, look non-English, or look like code. From 4,222,459 raw rows, 1,790,681 are filtered out and 2,431,778 remain. The punctuation-based filtering is empirical: we find that questions with more knowledge density are responded more formally, while casual questions tend to be responded with!or?.
Prompt sources
Only the user prompts come from these datasets; their original answers are not used.
Statistics
Layout
Each sample is three tar members sharing an eight-digit key, unique across every shard and both splits:
00000001.question.wav # user prompt, 44.1 kHz 16-bit mono, Fish-Audio TTS
00000001.response.mp3 # assistant reply, Qwen3-TTS speaker "Ryan"
00000001.json # transcripts and per-sample metadatatrain: 4216459 samples across 1406 shards (shards/train-{00000..01405}-of-01406.tar)validation: 6000 samples across 2 shards (shards/validation-{00000..00001}-of-00002.tar)
Exact per-shard sample counts and byte offsets live in shards.json; per-sample records live in manifest.jsonl.
Each metadata record carries:
Loading
from datasets import load_dataset
dataset = load_dataset("FlexiSLM/FlexiSLM-Data-4M-s2s", split="train", streaming=True)Or straight from WebDataset, using the brace pattern recorded in shards.json:
import webdataset as wds
url = "https://huggingface.co/datasets/FlexiSLM/FlexiSLM-Data-4M-s2s/resolve/main/shards/train-{00000..00099}-of-01406.tar"
dataset = wds.WebDataset(f"pipe:curl -sL {url}")Limitations
- Single-turn only. Multi-turn sources are truncated to their first user turn.
- Unfiltered. These are the raw synthesized pairs before the format, correctness, and ASR-based filtering described above; some samples contain artifacts, mismatched audio, or non-English text.
- Synthetic speech. All audio is TTS-generated, with a single fixed response voice; it does not reflect real conversational acoustics, noise, or speaker diversity on the response side.
- Model-generated text. Every response comes from Qwen3-Omni-30B-A3B and inherits its biases and factual errors; none are human-verified.
- Derived data. Prompts inherit the licenses and terms of their source datasets; please check each source before redistribution.
References
- Qwen3-Omni — Xu et al., 2025
- Qwen3-TTS — Hu et al., 2026
- Fish-Audio TTS — Liao et al., 2024
- Emilia — He et al., 2024
- MLS — Pratap et al., 2020
- LibriSpeech — Panayotov et al., 2015
- LLaSO-Instruct — Sun et al., 2025
- TriviaQA — Joshi et al., 2017
- WebQuestions — Berant et al., 2013
- TyDiQA — Clark et al., 2020
- Alpaca — Taori et al., 2023
- SODA — Kim et al., 2023
- Magpie — Xu et al., 2025
- UltraChat — Ding et al., 2023
- HH-RLHF — Bai et al., 2022
- WildChat — Zhao et al., 2024
