flexislm
Datasets
All datasets matching “flexislm”FlexiSLM-Data-2M-s2s-compact
FlexiSLM-Data — Speech-to-Speech Part (2.43M filtered samples, 385G in size)
Paper: https://arxiv.org/abs/2606.31247
Demo page: https://flexislm.github.io/
Code: https://github.com/AmphionTeam/FlexiSLM
FlexiSLM-Data is a large-scale, single-turn English speech-to-speech dialogue dataset
for training FlexiSLM, a spoken language model.
This repository contains the paired prompt-and-response audio portion of the release in
WebDataset format.
Related data releases… See the full description on the dataset page: https://huggingface.co/datasets/FlexiSLM/FlexiSLM-Data-2M-s2s-compact.FlexiSLM-Data-4M-s2s
FlexiSLM-Data — Speech-to-Speech Part (4M)
Paper: https://arxiv.org/abs/2606.31247
Demo page: https://flexislm.github.io/
Code: https://github.com/AmphionTeam/FlexiSLM
FlexiSLM-Data is a large-scale, single-turn English speech-to-speech dialogue dataset
for training FlexiSLM, a spoken language model.
This repository contains the paired prompt-and-response audio portion of the release in
WebDataset format.
Related data releases
FlexiSLM/FlexiSLM-Data-4M-s2s (this repo)… See the full description on the dataset page: https://huggingface.co/datasets/FlexiSLM/FlexiSLM-Data-4M-s2s.asrtts_packed_webdataset
ASR+TTS Repacked Data (3.56M samples, mp3)
This dataset is a WebDataset repack prepared for FlexiSLM training (ASR+TTS tasks).
Paper: https://arxiv.org/abs/2606.31247
Demo page: https://flexislm.github.io/
Code: https://github.com/AmphionTeam/FlexiSLM
FlexiSLM-Data is a large-scale, single-turn English speech-to-speech dialogue dataset
for training FlexiSLM, a spoken language model.
This repository contains the paired prompt-and-response audio portion of the release in… See the full description on the dataset page: https://huggingface.co/datasets/FlexiSLM/asrtts_packed_webdataset.FlexiSLM-Data-5M-t2t
FlexiSLM-Data — Text-to-Text Part (5M)
Paper: https://arxiv.org/abs/2606.31247
Demo page: https://flexislm.github.io/
Code: https://github.com/AmphionTeam/FlexiSLM
FlexiSLM-Data is a large-scale, single-turn English speech-to-speech dialogue dataset
for training FlexiSLM, a spoken language model.
This repository contains the paired prompt-and-response audio portion of the release in
WebDataset format.
Related data releases
FlexiSLM/FlexiSLM-Data-5M-t2t (this repo) provides… See the full description on the dataset page: https://huggingface.co/datasets/FlexiSLM/FlexiSLM-Data-5M-t2t.flexislm_dialoguedataset_en_t2t
