bashkir
Datasets
All datasets matching “bashkir”bashkir-frequency-index
Bashkir Frequency Index v13.0-af1
Word-frequency index for the Bashkir language, computed over 152.16 million tokens of clean monolingual text across diverse public domains (periodicals, news media, literary publications, encyclopedic texts, books, and general web archives). Non-Bashkir language admixture and scanning artifacts were filtered using automated language-filtering pipelines.
Configurations
Config
Rows
Cutoff
Use Case
public (recommended)
589… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-frequency-index.bashkir-ngram-index
Bashkir Word N-gram Index v13.0-af1
Exact within-sentence word n-gram counts for Bashkir: unigrams, bigrams and
trigrams for spellchecking, OCR post-processing and lightweight language modelling.
Overview
Exact word n-gram counts derived from a monolingual Bashkir-language dataset. The
release provides unigram, bigram and trigram indexes for corpus processing,
spellchecking, OCR post-processing, autocomplete and lightweight language-model
experiments. The… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-ngram-index.bashkir-wikipedia-parallel
Bashkir-Russian Wikipedia Parallel Corpus
Sentence-level Bashkir-Russian parallel text from Wikipedia, scored and filtered
for machine translation.
Overview
Sentence-level Bashkir-Russian parallel dataset extracted from the corresponding
Bashkir and Russian Wikipedia dumps dated 2026-08-01. Candidate pairs are scored
for semantic alignment with multilingual sentence encoders (Meta LASER3, Google
LaBSE) and the in-domain Bashkir-Russian Pair Scorer. The filtered… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-parallel.bashkir-multilingual-phrasebooks
Bashkir-Russian Phrasebook Corpus
Edited Bashkir-Russian words, expressions and conversational phrases from
university phrasebooks, annotated by entry type.
Overview
Edited Bashkir-Russian pairs derived from the original
bashkorttele/trilingual-parallel-phrasebooks-bgpu
dataset, published by Bashkorttele from phrasebooks of M. Akmulla Bashkir State
Pedagogical University. The cleaned configuration is the deduplicated default;
reviewed is the edited edition… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-multilingual-phrasebooks.bashkir-wikipedia-monolingual
Bashkir Wikipedia Monolingual Corpus
Cleaned sentence-level Bashkir text from Wikipedia for pretraining, tokenizer
training and linguistic research.
Overview
Sentence-level text extracted from the Bashkir Wikipedia dump (bawiki-20260801),
cleaned and filtered with automated language identification. The cleaned
configuration is the recommended default for language modelling, tokenization and
linguistic research; precleaned is an earlier, lighter extraction kept… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-monolingual.bashkir-lora-qlora-benchmark
Bashkir LoRA/QLoRA Benchmark
📊 Description
This benchmark contains the complete results of fine-tuning various language models (from 82M to 7B parameters) on the Bashkir language. The study compares the effectiveness of LoRA/QLoRA against full fine-tuning, evaluating model quality (perplexity), GPU memory usage, and training time.
Key Findings
Mistral-7B with QLoRA (r=16) achieved the best performance among 7B models (perplexity 3.79)
LoRA drastically reduces… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-lora-qlora-benchmark.
