mkrausio/audiosnippets-cleaned
Dataset Summary This dataset is a processed version of mitermix/audiosnippets. The dataset contains audio snippets that have been cleaned and resampled, making it suitable for tasks like audio captioning, audio classification, or other audio-based machine learning applications. Processing Details Transcriptions and broken characters were removed. All MP3 audio files were resampled to 16kHz for consistency. The accompanying JSON metadata was made consistent.… See the full description on the dataset page: https://huggingface.co/datasets/mkrausio/audiosnippets-cleaned.
Dataset Summary
This dataset is a processed version of mitermix/audiosnippets. The dataset contains audio snippets that have been cleaned and resampled, making it suitable for tasks like audio captioning, audio classification, or other audio-based machine learning applications.
Processing Details
- Transcriptions and broken characters were removed.
- All MP3 audio files were resampled to 16kHz for consistency.
- The accompanying JSON metadata was made consistent.
- Entries with empty captions were removed.
Dataset Structure
The dataset is divided into three configurations:
train/: Training setval/: Validation settest/: Test set
Each configuration contains audio files organized as WebDataset tar files.
Data Format
The dataset is in WebDataset format, where each tar file contains:
- Audio files (
.mp3format) - Caption (
.jsonformat)
Acknowledgments
Original dataset: mitermix/audiosnippets.
