Team Ai
Datasetpublic

WhissleAI/indicvoices_hi_tagged_transcripts

Dataset Card for indicvoices_hi_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_hi_tagged_transcripts.

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
0likes192downloads
Dataset Card

Dataset Card for indicvoiceshitagged_transcripts

Dataset Description

This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks.

Languages

The dataset is primarily in Hindi.

Data Collection

The dataset was collected through automated processes and manual transcription.

Dataset Structure

The dataset contains:

  • —Audio files (.wav format)
  • —Transcriptions
  • —Duration information
  • —Additional task annotations

Data Fields

  • —audio: Path to the audio file
  • —text: Transcription of the audio
  • —duration: Length of the audio in seconds
  • —tasks: List of associated tasks

Additional Information

  • —License: CC-BY 4.0
  • —Version: 1.0.0
  • —Publisher: WhissleAI