Team Ai
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PrakharGoonj /60M-Tokens-Multilingual-Indian-Languages-LLM-Training-Corpora 60M+ Words Gold-Standard Multilingual Text Corpora Package for LLM Training Developed by Prakhar Goonj Publications (D&B D-U-N-S Certified Tier-1 Media House, No. 64-125-5366 | FIP Member). 🚀 Commercial Licensing & Access Notice The full-volume, production-ready dataset (~60,000,000+ human-authored tokens) and all specialized domain sub-buckets are available for immediate global, non-exclusive bulk commercial licensing on our Opendatabay Corporate Storefront: 👉… See the full description on the dataset page: https://huggingface.co/datasets/PrakharGoonj/60M-Tokens-Multilingual-Indian-Languages-LLM-Training-Corpora.text-generation0 likes56 downloads20h agoHugging Face02praneeetha /tts-indian-languages TTS Indian Languages Dataset A curated Text-to-Speech training dataset of 168 segments (~66 minutes) of clean, single-speaker audio in Indian English, Hindi, and Telugu — sourced from YouTube and processed using Sarvam AI's ASR and LLM APIs. Dataset Summary Language Code Duration Indian English en-IN 32.4 min Hindi hi-IN 24.0 min Telugu te-IN 10.0 min Total 66.4 min Dataset Structure Each row contains: Field Type… See the full description on the dataset page: https://huggingface.co/datasets/praneeetha/tts-indian-languages.audiotext-to-speechn<1K0 likes30 downloads4mo agoHugging Face03monsterapi /Indian-languagestext100K<n<1M1 likes28 downloads3y agoHugging Face04sohomghosh /IndicFinNLP_FinancialNatural_Language_Processing_for_Indian_Languages IndicFinNLP This repository contains dataset mentioned in the paper, "IndicFinNLP: Financial Natural Language Processing for Indian Languages" @ LREC-COLING 2024 Tasks: Exaggerated Numeral Detection Sustainability Assessment, ESG Theme Determination Languages: Hindi, Bengali, Telugu Source: Budget speeches of Hindi, Bengali, and Telugu speaking states (Punjab, Uttarakhand, Haryana, West Bengal, Telangana, and Andhra Pradesh) starting from the year 2011 till 2023. Financial texts… See the full description on the dataset page: https://huggingface.co/datasets/sohomghosh/IndicFinNLP_FinancialNatural_Language_Processing_for_Indian_Languages.text-classification10K<n<100K0 likes13 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.