Team Ai
Datasetpublic

Speech-data/Cebuano-Speech-Dataset

🎧 Cebuano Speech Dataset The Cebuano Speech Dataset is a high-quality speech audio dataset designed to deliver structured and diverse audio data for AI-powered voice applications. It includes 108 hours of audio data distributed across 807 files, provided in MP3 and WAV formats, with a total size of 135 MB. This well-organized audio dataset ensures balanced voice data, with 49% female and 51% male speakers, and a broad age range from 18 to 50+ years. The dataset language is… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Cebuano-Speech-Dataset.

sourceHugging Facecc-by-nc-nd-4.0updated 6mo agoView on Hugging Face
1likes22downloads
Dataset Card

🎧 Cebuano Speech Dataset

The Cebuano Speech Dataset is a high-quality speech audio dataset designed to deliver structured and diverse audio data for AI-powered voice applications. It includes 108 hours of audio data distributed across 807 files, provided in MP3 and WAV formats, with a total size of 135 MB. This well-organized audio dataset ensures balanced voice data, with 49% female and 51% male speakers, and a broad age range from 18 to 50+ years. The dataset language is Cebuano, with speakers from the Philippines, including Cebu, Mindanao, Bohol, and Leyte, making it a representative language speech dataset for regional linguistic diversity.


πŸ”— Learn more: https://speech-data.ai/datasets/cebuano/


πŸš€ Use Cases

This Cebuano speech dataset supports a wide range of AI applications, including speech recognition, voice assistant development, and natural language processing systems. The structured speech data enables efficient acoustic modeling, speaker identification, and scalable AI training workflows. It functions as a reliable speech recognition dataset for both research and production environments, particularly in multilingual and low-resource language scenarios.


πŸ“Š Dataset Metadata

FieldValue
πŸ“œ LicenseCC BY-NC-ND 4.0
🎯 Task CategoriesAutomatic Speech Recognition
🏷️ TagsAudio, Machine, Machine Learning, Speech, Dataset, Speech Recognition
πŸ“¦ Size Categoryn < 1K

⭐ Key Value

The key value of this speech dataset lies in its regional authenticity, balanced speaker distribution, and production-ready structure. It provides high-quality audio data that enhances model performance in real-world applications. This voice dataset is particularly valuable for building scalable, accurate, and inclusive voice-enabled AI systems.