Team Ai
Datasetpublicgated

BashkirNLPWorld/bashkir-web-corpus

Dataset Card for Bashkir Web Corpus Dataset Details Dataset Description The Bashkir Web Corpus is a collection of 71,567 documents and approximately 46.9 million tokens in the Bashkir language (a Turkic language spoken in Bashkortostan, Russia). The corpus was compiled from 16 Bashkir‑language online sources, including news websites, literary magazines, social media, books, and Wikipedia. It is designed for language modeling, text classification… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-web-corpus.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes36downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.