datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ShIO-bash-26.1
ShIO-bash-26.1
Shell input-output (ShIO) Bash dataset produced by ShIOEnv, a Gymnasium-compatible Bash environment designed to collect execution-annotated command interactions in a Linux system.
Dataset summary
The dataset consists of command-line inputs paired with their execution artifacts, including observable outputs and a structured representation of environment state changes. Samples are produced by executing synthesized Bash inputs inside a… See the full description on the dataset page: https://huggingface.co/datasets/jragsdale1/ShIO-bash-26.1.bashkir-periodicals-izddom
Bashkir Periodicals Cleaned (Izddom)
High-quality cleaned and structured Bashkir periodicals (Respublika Bashkortostan Publishing House) for LLM pretraining and fine-tuning.
Overview
A curated, deduplicated and filtered edition of modern Bashkir periodicals derived from the original bashkorttele/periodicals-izddom dataset published by the Foundation for the Preservation and Development of the Bashkir Language.
The dataset includes material from 13 leading… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-periodicals-izddom.bashkir-books-kitap
Bashkir Books & Calendars Cleaned (Kitap)
Curated, cleaned and structured Bashkir literature and annual calendars (Kitap Publishing House) for LLM pretraining and fine-tuning.
Overview
A curated, deduplicated and filtered edition of modern Bashkir book publications and annual cultural calendars derived from the original bashkorttele/books-kitap dataset published by the Foundation for the Preservation and Development of the Bashkir Language.
The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-books-kitap.rlvr-bash-terminal-bench
rlvr-bash-terminal-bench
RLVR (Reinforcement Learning with Verifiable Rewards) dataset for bash scripting, generated from Terminal-Bench tasks.
Stats
Metric
Value
Total samples
1,120
Unique tasks
88
Avg samples/task
12.7
Average reward
0.249
Perfect solutions (reward=1.0)
10.4%
Partial solutions (0<reward<1)
28.8%
Zero reward
60.8%
Tasks fully solved
13.6%
Format
{
"task_id": "string",
"prompt": "string",
"completion":… See the full description on the dataset page: https://huggingface.co/datasets/lisayan/rlvr-bash-terminal-bench.books-kitap
Bashkir Books — "Kitap" Publishing House
9,019,878 characters of Bashkir-language text (6,081 records) from publications of "Kitap" Publishing House. A record is an article or a work where the publication marks its boundaries (a table of contents, a day of a tear-off calendar); elsewhere it is cut at a heading or at a page. Every record carries metadata sufficient to locate it in the original publication.
🌐 Languages of this card: English · Башҡортса · Русский
Part of the… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/books-kitap.periodicals-izddom
Bashkir Periodicals — Respublika Bashkortostan Publishing House
163,240,215 characters of Bashkir-language text (39,711 records) from publications of Respublika Bashkortostan Publishing House. A record is an article or a work where the publication marks its boundaries (a table of contents, a day of a tear-off calendar); elsewhere it is cut at a heading or at a page. Every record carries metadata sufficient to locate it in the original publication.
🌐 Languages of this card:… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/periodicals-izddom.bashkir-russian-parallel
Dataset Card for Bashkir-Russian Parallel Corpus
Dataset Details
Dataset Description
Bashkir-Russian Parallel Corpus is a large-scale sentence-aligned parallel corpus for the Bashkir–Russian language pair, assembled from authentic human-created translations. It contains 3,040,085 unique parallel sentence pairs, where each Bashkir sentence is aligned with its Russian counterpart.
The corpus combines data from three open parallel corpora: TIL-MT… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-russian-parallel.
