datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
serena-synthetic-it-28h
Qwen3-TTS Italian Synthetic Speech (27h)
Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with
Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV,
Piper-ready metadata.
Dataset summary
Property
Value
Clips (train / val)
26,523 / 2,947
Total duration
~27.3 h (98,099 s)
Sample rate
22,050 Hz mono, 16-bit WAV
Loudness
Normalized to -23 LUFS, silence-trimmed
Language
Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-28h.knesset-committees-panel-hq
Knesset Committees Panel (high quality)
The audio of the adaptation stage's training plan v3 (docs/training_plan_v3.md in hadasy-tau/deep-learning-project),
as 16 kHz 16-bit mono WAV under panel_audio/<speaker_id>/<chunk_id>.wav: 10,972 chunks (40.8 h)
of 12 Knesset members, cut from knesset-asr/knesset-committees-chunks.
Two tables describe it, one row per chunk with the reference text, session, date and split:
panel_plan_v2.parquet -- train, dev and test (6,887 chunks). Every… See the full description on the dataset page: https://huggingface.co/datasets/knesset-asr/knesset-committees-panel-hq.serena-synthetic-it-27h
Qwen3-TTS Italian Synthetic Speech (27h)
Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with
Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV,
Piper-ready metadata.
Dataset summary
Property
Value
Clips (train / val)
26,523 / 2,947
Total duration
~27.3 h (98,099 s)
Sample rate
22,050 Hz mono, 16-bit WAV
Loudness
Normalized to -23 LUFS, silence-trimmed
Language
Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-27h.knesset-committees-panel
Knesset Committees Panel
The audio the adaptation stage trains and tests on: 8,113 chunks (28.8 h) of
11 Knesset members, cut from Hadasy/knesset-committees-chunks at alignment quality >= 0.7,
as 16 kHz 16-bit mono WAV under panel_audio/<speaker_id>/<chunk_id>.wav. panel_plan.parquet is the
table: one row per chunk with the reference text, session and date, and the split (part: test = the
speaker's newest sessions, then dev, the rest train; session-disjoint by date; train… See the full description on the dataset page: https://huggingface.co/datasets/knesset-asr/knesset-committees-panel.
