Team Ai
Datasetpublicgated

ivrit-ai/audio-v2-transcripts

Overview This dataset provides full, machine-generated transcriptions for the entire audio-v2 dataset, containing >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license. It was released on May 18th, 2025. You can find the full list of sources in this dataset under the audio-v2 dataset's sources.txt. All files were transcribed using the process.py pipeline, performing: Frame-level VAD Machine transcription using ivrit.ai's whisper-large-v3-turbo engine with the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2-transcripts.

sourceHugging Faceotherupdated 11mo agoView on Hugging Face
1likes1.4kdownloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.

ivrit-ai/audio-v2-transcripts · Team Ai