Team Ai
Datasetpublicgated

kapturecx/bolAIndia

bolAIndia Human-side speech from production call recordings, cut into utterance-level chunks by a two-engine VAD (Silero + TEN) and transcribed by third-party ASR providers. Each row keeps the transcript, the provider's confidence, and full provenance back to the source recording. Alongside it, open Indian-language speech corpora converted to the same schema (16 kHz mono FLAC, one utterance per row), each in a config of its own and tagged with where it came from.… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/bolAIndia.

sourceHugging Faceupdated 19h agoView on Hugging Face
2likes17kdownloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.

kapturecx/bolAIndia · Team Ai