datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cyp3a4_marimomari-monolingual-corpusA monolingual corpus of the Mari language in various genres, containing over 20 million word occurrences.
The presented genres:
Genre
Russian
English
мутер
словарь
dictionary
газетысе увер
газетные новости
periodical news
прозо
проза
prose
фольклор
фольклор
folklore
публицистике
публицистика
publicistic literature
поэзий
поэзия
poetry
трагикомедийтрагикомедия
tragicomedy
пьесе
пьеса
play
драме
драма
drama
комедий-водевиль
водевиль
vaudeville
комедий
комедия… See the full description on the dataset page: https://huggingface.co/datasets/mari-lab/mari-monolingual-corpus.marimo-manim-sft-data
