Team Ai
Datasetpublic

FlexiSLM/FlexiSLM-Data-4M-s2s

FlexiSLM-Data — Speech-to-Speech Part (4M) Paper: https://arxiv.org/abs/2606.31247 Demo page: https://flexislm.github.io/ Code: https://github.com/AmphionTeam/FlexiSLM FlexiSLM-Data is a large-scale, single-turn English speech-to-speech dialogue dataset for training FlexiSLM, a spoken language model. This repository contains the paired prompt-and-response audio portion of the release in WebDataset format. Related data releases FlexiSLM/FlexiSLM-Data-4M-s2s (this repo)… See the full description on the dataset page: https://huggingface.co/datasets/FlexiSLM/FlexiSLM-Data-4M-s2s.

sourceHugging Faceupdated 2mo agoView on Hugging Face
1likes330downloads
Dataset Card

FlexiSLM-Data — Speech-to-Speech Part (4M)

![Paper](https://arxiv.org/abs/2606.31247) ![Demo](https://flexislm.github.io/) ![Code](https://github.com/AmphionTeam/FlexiSLM)

  • —Paper: https://arxiv.org/abs/2606.31247
  • —Demo page: https://flexislm.github.io/
  • —Code: https://github.com/AmphionTeam/FlexiSLM

FlexiSLM-Data is a large-scale, single-turn English speech-to-speech dialogue dataset for training FlexiSLM, a spoken language model. This repository contains the paired prompt-and-response audio portion of the release in WebDataset format.

Related data releases

Data construction pipeline

  1. 1.Prompt collection and response generation (released in **FlexiSLM/FlexiSLM-Data-5M-t2t**) Text prompts are collected from public QA, instruction-following, and dialogue datasets (see the table below). For multi-turn datasets, the user's first-turn utterance is used as the prompt; samples whose first turn is a generic greeting (e.g. "Hello") are skipped. Then, all text responses are generated with **Qwen3-Omni-30B-A3B**; the responses shipped with the source datasets are discarded. We use a spoken language model rather than a text-only LLM because Qwen3-Omni produces speech-friendly responses: short, conversational, and free of formatting that does not transfer to speech (bullet points, code blocks, long enumerations). The text version of the prompt is fed to the model rather than its synthesized speech, since text-input responses are typically more accurate.
  2. 2.Speech synthesis. (released in **FlexiSLM/FlexiSLM-Data-4M-s2s**) Responses are synthesized with Qwen3-TTS using the fixed speaker "Ryan". Prompts are synthesized with **Fish-Audio TTS**, with speaker prompts randomly sampled from English Emilia utterances longer than 5 seconds. This yields the 4,222,459 samples and 26,735 hours of audio shipped here.
  3. 3.Quality filtering and mp3-format compression. (released in **FlexiSLM/FlexiSLM-Data-2M-s2s-compact**) Before any shards are written, this release applies prompt deduplication and response-text filtering. The kept subset uses normalized prompt deduplication first, then drops responses that contain !, contain ?, look non-English, or look like code. From 4,222,459 raw rows, 1,790,681 are filtered out and 2,431,778 remain. The punctuation-based filtering is empirical: we find that questions with more knowledge density are responded more formally, while casual questions tend to be responded with ! or ?.

Prompt sources

Only the user prompts come from these datasets; their original answers are not used.

DatasetType# Prompts
TriviaQAQA138K
WebQuestionsQA3.8K
TyDiQAMultilingual QA167K
AlpacaInstruction52K
SmolTalk2Instruction / Dialogue385K
SODADialogue1.48M
Magpie-ProInstruction / Dialogue1M
UltraChatMulti-turn Dialogue949K
HH-RLHFDialogue / Preference167K
WildChatReal-user Dialogue159K

Statistics

value
Samples (train)4,216,459 across 1,406 shards
Samples (validation)6,000 across 2 shards
Samples (total)4,222,459
Prompt audio7,317.3 hours (26,342,351 s)
Response audio19,418.2 hours (69,905,434 s)
Total audio26,735.5 hours (96,247,785 s)
Mean prompt duration6.24 s
Mean response duration16.56 s

Layout

Each sample is three tar members sharing an eight-digit key, unique across every shard and both splits:

00000001.question.wav   # user prompt, 44.1 kHz 16-bit mono, Fish-Audio TTS
00000001.response.mp3   # assistant reply, Qwen3-TTS speaker "Ryan"
00000001.json           # transcripts and per-sample metadata
  • —train: 4216459 samples across 1406 shards (shards/train-{00000..01405}-of-01406.tar)
  • —validation: 6000 samples across 2 shards (shards/validation-{00000..00001}-of-00002.tar)

Exact per-shard sample counts and byte offsets live in shards.json; per-sample records live in manifest.jsonl.

Each metadata record carries:

fielddescription
keyEight-digit WebDataset key
uuidStable identifier, shared with the t2t release
splittrain or validation
shardShard the sample lives in
question_text / response_textTranscripts of the two audio members
question_duration / response_durationSeconds
total_audio_durationSeconds
num_tokens_estEstimated audio tokens at 12 tokens/s

Loading

python
from datasets import load_dataset

dataset = load_dataset("FlexiSLM/FlexiSLM-Data-4M-s2s", split="train", streaming=True)

Or straight from WebDataset, using the brace pattern recorded in shards.json:

python
import webdataset as wds

url = "https://huggingface.co/datasets/FlexiSLM/FlexiSLM-Data-4M-s2s/resolve/main/shards/train-{00000..00099}-of-01406.tar"
dataset = wds.WebDataset(f"pipe:curl -sL {url}")

Limitations

  • —Single-turn only. Multi-turn sources are truncated to their first user turn.
  • —Unfiltered. These are the raw synthesized pairs before the format, correctness, and ASR-based filtering described above; some samples contain artifacts, mismatched audio, or non-English text.
  • —Synthetic speech. All audio is TTS-generated, with a single fixed response voice; it does not reflect real conversational acoustics, noise, or speaker diversity on the response side.
  • —Model-generated text. Every response comes from Qwen3-Omni-30B-A3B and inherits its biases and factual errors; none are human-verified.
  • —Derived data. Prompts inherit the licenses and terms of their source datasets; please check each source before redistribution.

References

  • —Qwen3-Omni — Xu et al., 2025
  • —Qwen3-TTS — Hu et al., 2026
  • —Fish-Audio TTS — Liao et al., 2024
  • —Emilia — He et al., 2024
  • —MLS — Pratap et al., 2020
  • —LibriSpeech — Panayotov et al., 2015
  • —LLaSO-Instruct — Sun et al., 2025
  • —TriviaQA — Joshi et al., 2017
  • —WebQuestions — Berant et al., 2013
  • —TyDiQA — Clark et al., 2020
  • —Alpaca — Taori et al., 2023
  • —SODA — Kim et al., 2023
  • —Magpie — Xu et al., 2025
  • —UltraChat — Ding et al., 2023
  • —HH-RLHF — Bai et al., 2022
  • —WildChat — Zhao et al., 2024