datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Agri_STT_Benchmarking_Dataset
Agri STT Benchmarking Dataset
10,808 farmer voice queries in Hindi, Telugu and Odia, with reference transcripts, for benchmarking automatic speech recognition in agricultural contexts. The audio is included in this repository.
Every recording is a smallholder farmer speaking a question to Farmer.Chat, an AI advisory service run by Digital Green. Reference transcripts were produced by human annotators. Nothing here is read from a script or recorded in a studio, so the audio… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/Agri_STT_Benchmarking_Dataset.ASR-Benchmarking-Dataset
Hindi STT Benchmarking Eval
Overview
This dataset packages the Hindi eval split used for STT benchmarking across six Vistaar-derived parts: IndicTTS, FLEURS, CommonVoice, Kathbath, Kathbath noisy, and MUCS. Each row contains the audio, original reference transcript, and raw plus normalized transcripts from Ringg, ElevenLabs, Deepgram, and Sarvam.
The dataset contains 10,000 utterances and about 15.5 hours of 16 kHz mono WAV audio.
The dataset is published as part-specific… See the full description on the dataset page: https://huggingface.co/datasets/RinggAI/ASR-Benchmarking-Dataset.Agri_STT_Benchmarking_DatasetThis is a domain-specific, multilingual agricultural speech dataset with a primary focus on Hindi, Telugu, and Odia, designed for speech-to-text and automatic speech recognition (ASR) tasks. It features human-annotated transcriptions and is intended for benchmarking ASR model performance in real-world agricultural scenarios.
This paper presents a comprehensive benchmark of 10 ASR models for agricultural advisory use across Hindi, Telugu, and Odia, using 10,934 real-world Farmer.Chat audio… See the full description on the dataset page: https://huggingface.co/datasets/bullseye-4/Agri_STT_Benchmarking_Dataset.
