datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aihub-464-preprocessed-680GB-set-52aihub-464-preprocessed-680GB-set-56aihub-464-preprocessed-680GB-set-57gramvaani_preprocessed_hi_train
Dataset Card for "gramvaani_preprocessed_hi_train"
More Information needed
aihub-464-preprocessed-680GB-set-53mozilla_commonvoice_hackathon_preprocessed_train_batch_3
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_3"
More Information needed
sada-train-wav2vec2-xls-r-300m-ar-preprocessedsada-train-preprocessedmozilla_commonvoice_hackathon_preprocessed_train_batch_2
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_2"
More Information needed
MUSIC-AVQA_cls-preprocessedaihub-464-preprocessed-680GB-set-45preprocessed_speech_datasetsvin100h-preprocessed-v2vin100h-preprocessed-16k-whisper-large-v3ltx25_preprocessed
LTX-Video 2.5 Preprocessed Dataset
Preprocessed training data for LTX-Video 2.5 (joint audio + video), stored as PyTorch tensors. Every file is a torch.save'd dict and can be loaded with:
import torch
d = torch.load("latents/10s/clip000001.pt", map_location="cpu", weights_only=True)
Structure
Clips are split by duration bucket (5s, 10s) and modality. File names are shared across folders: clipNNNNNN.pt refers to the same source clip in every folder that contains… See the full description on the dataset page: https://huggingface.co/datasets/colinb83/ltx25_preprocessed.mozilla_commonvoice_naijaHausa1_preprocessed_train_batch_1mozilla_commonvoice_Swahili_preprocessed_train_batch_1ds007808-sub01-speechopen-pangolin-preprocessed
ds007808 sub-01 / speechopen / pangolin — preprocessed EEG↔speech windows
Ready-to-train EEG↔speech windows for replicating the scaling experiment of
Sato et al. 2024, "Scaling Law in Neural Data: Non-Invasive Speech Decoding with 175 Hours
of EEG Data" (arXiv:2407.07595), built from the public
ds007808 dataset (arXiv:2606.01264).
Slice = subject sub-01, task speechopen (overt speech), device pangolin (128-ch
g.Pangolin) — the rig matching the 175 h paper. Each example is one… See the full description on the dataset page: https://huggingface.co/datasets/tankalapavankalyan/ds007808-sub01-speechopen-pangolin-preprocessed.mozilla_commonvoice_Swahili_preprocessed_train_batch_4mozilla_commonvoice_Swahili_preprocessed_train_batch_3mozilla_commonvoice_hackathon_preprocessed_train_batch_6
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_6"
More Information needed
mozilla_commonvoice_Swahili_preprocessed_train_batch_2mozilla_commonvoice_hackathon_preprocessed_train_batch_4
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_4"
More Information needed
mozilla_commonvoice_hackathon_preprocessed_train_batch_1
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_1"
More Information needed
preprocessed-whisper-btb-cv-cvad-wlga-ca-2607
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2605
5bfe2d098c8486d97fac8be76d86ec9146435245
train
56:46:32
50,557
589,095
11.7
31.9
techiaith/corpws-clllc-wlga
5d00294c31c78b1d7937bb2c2bc6cc70bc18d410
clips
48:20:49
27,579… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607.mozilla_commonvoice_hackathon_preprocessed_train_batch_5
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_5"
More Information needed
aihub-464-preprocessed-680GB-set-0
필요 없는 컬럼 제거가 필요해요.
다음의 코드를 통해 만들었어요.
import os
import json
from pydub import AudioSegment
from tqdm import tqdm
import re
from datasets import Audio, Dataset, DatasetDict, load_from_disk, concatenate_datasets
from transformers import WhisperFeatureExtractor, WhisperTokenizer
import pandas as pd
# 사용자 지정 변수를 설정해요.
# DATA_DIR = '/mnt/a/maxseats/(주의-원본-680GB)주요 영역별 회의 음성인식 데이터' # 데이터셋이 저장된 폴더
DATA_DIR = '/mnt/a/maxseats/(주의-원본)split_files/set_0' # 첫 10GB 테스트
# 원천, 라벨링 데이터 폴더 지정… See the full description on the dataset page: https://huggingface.co/datasets/maxseats/aihub-464-preprocessed-680GB-set-0.MELD-Preprocessed
MELD Preprocessed for SER
This dataset is the manually preprocessed audio only version of MELD, only audio IDs, utterance transcriptions, dialogue IDs and Utterance IDs were extracted.
S. Poria, D. Hazarika, N. Majumder, G. Naik, R. Mihalcea,
E. Cambria. MELD: A Multimodal Multi-Party Dataset
for Emotion Recognition in Conversation. (2018)
Chen, S.Y., Hsu, C.C., Kuo, C.C. and Ku, L.W.
EmotionLines: An Emotion Corpus of Multi-Party
Conversations. arXiv preprint arXiv:1802.08379… See the full description on the dataset page: https://huggingface.co/datasets/Vano04/MELD-Preprocessed.mozilla_commonvoice_Swahili_preprocessed_train_batch_6sada-validation-preprocessed
Details
This is the SADA 2022 dataset with the input_features whish are log mels and the cleaned_labels which is the tokenized version of the cleaned_text. You can directly use this as the validation dataset when training Whisper Tiny, Small, Base & Medium models, as they all use the same tokenizer. Please double check this as well from the original model repo.
In addtition, the following filters were applied to this data:
All audios are less than 30 seconds and greater than 0… See the full description on the dataset page: https://huggingface.co/datasets/mosama/sada-validation-preprocessed.
