datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
linustechtips-transcript-audio
Dataset Card for "linustechtips"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Linus Tech Tips.
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel: Name of the… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/linustechtips-transcript-audio.lex-fridman-podcast-transcript-audio
Dataset Card for "lexFridmanPodcast-transcript-audio"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast.
Data Fields
The dataset is composed by:
id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast-transcript-audio.yannick-kilcher-transcript-audio
Dataset Card for "yannic-kilcher-transcript-audio"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Yannic Kilcher. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Yannic Kilcher.
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel:… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/yannick-kilcher-transcript-audio.whisper-transcripts-the-ai-epiphany
Dataset Card for "the-ai-epiphany"
More Information needed
lex-fridman-podcast
Dataset Card for "lexFridmanPodcast-transcript-audio"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast.
Data Fields
The dataset is composed by:
id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast.vi-fdb-v1-gpt-realtime
GPT-Realtime on Vi-FDB v1: full reference outputs
Public reference outputs from GPT-Realtime on all 400 Vi-FDB event cases and
their 200 available clean controls. This repository contains model outputs
and evaluation evidence, not benchmark inputs.
Canonical benchmark: https://huggingface.co/datasets/tuanamz/vi-fdb-v1
Benchmark revision: c9730eddf78b094610be09796afe8e60c178caef
Interactive explorer: https://huggingface.co/spaces/tuanamz/vi-fdb-v1-explorer
Contents… See the full description on the dataset page: https://huggingface.co/datasets/tuanamz/vi-fdb-v1-gpt-realtime.whisper-transcripts-linustechtips
Dataset Card for "whisper-transcripts-linustechtips"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel: Name of the channel.
channel_id: Id of the youtube channel.
title: Title given to… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/whisper-transcripts-linustechtips.whisper-transcripts-ml-street-talk
Dataset Card for "whisper-transcripts-mlst"
More Information needed
two-minute-papers
Dataset Card for "two-minute-papers"
More Information needed
yannic-kilcher-transcript
Dataset Card for "yannic-kilcher-transcript"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Yannic Kilcher. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Yannic Kilcher.
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel: Name… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/yannic-kilcher-transcript.gptsovits_dataset
bhyuan/gptsovits_dataset
GPT-SoVITS speech dataset, packed as WebDataset tar shards.
Layout
data/
train/
metadata.csv
audio/
train-000.tar
train-001.tar
...
validation/
metadata.csv
audio/
validation-000.tar
...
test/
metadata.csv
audio/
test-000.tar
...
Shard counts:
youshengshu_v5_test: 6536 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/bhyuan/gptsovits_dataset.
