datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ShIO-bash-26.1
ShIO-bash-26.1
Shell input-output (ShIO) Bash dataset produced by ShIOEnv, a Gymnasium-compatible Bash environment designed to collect execution-annotated command interactions in a Linux system.
Dataset summary
The dataset consists of command-line inputs paired with their execution artifacts, including observable outputs and a structured representation of environment state changes. Samples are produced by executing synthesized Bash inputs inside a… See the full description on the dataset page: https://huggingface.co/datasets/jragsdale1/ShIO-bash-26.1.bash-instruct-III-55k
Bash Instruct III — 54,360 verified natural-language → Bash pairs
Bash Instruct III is a synthetic instruction-tuning dataset that maps natural-language
requests to correct Bash: single commands, short pipelines, and multi-line scripts. It
is built for supervised fine-tuning of small and mid-size LLMs that must turn a plain
request into shell code that actually runs.
Every row is a three-turn chat conversation (system / user / assistant) with metadata
for slicing (category… See the full description on the dataset page: https://huggingface.co/datasets/Frost2o24/bash-instruct-III-55k.bash_command_data_6K
📦 Bash Command Dataset v1
A high-quality dataset of natural language instructions paired with their equivalent Bash commands, designed for training and fine-tuning large language models (LLMs) that translate English tasks into shell commands.
This dataset is ideal for researchers, developers, and machine learning engineers interested in natural language to Bash command translation, command-line automation, and building intelligent terminal assistants.
📁 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/emirkaanozdemr/bash_command_data_6K.bashkir-wikipedia-monolingual
Bashkir Wikipedia Monolingual Corpus
Cleaned sentence-level Bashkir text from Wikipedia for pretraining, tokenizer
training and linguistic research.
Overview
Sentence-level text extracted from the Bashkir Wikipedia dump (bawiki-20260801),
cleaned and filtered with automated language identification. The cleaned
configuration is the recommended default for language modelling, tokenization and
linguistic research; precleaned is an earlier, lighter extraction kept… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-monolingual.bash-instruct-II-55k
Bash Instruct II — 54,803 verified natural-language → Bash pairs
⚠️ Superseded by Bash Instruct III
III is a corrected rebuild of this dataset with shell antipatterns removed at the
generator level (99.8% ShellCheck-clean vs 98.5% here). New work should use III.
The two share 92.5% of their (request, command) pairs, so they must never be
concatenated. This card is kept for reproducibility and citation of published results.
Bash Instruct II is a synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Frost2o24/bash-instruct-II-55k.bash_codeThis dataset is a collection of bash programs from various GitHub repositories and open source projects.
The dataset might contain harmful code.
bash-instruct-55k
Bash Instruct I — 55,000 verified natural-language → Bash pairs
⚠️ Superseded by Bash Instruct III
III doubles the utility vocabulary (89 → 182), adds grouped equivalent answers, real
human phrasing from tldr-pages, validation by execution on real Linux, and a style pass
that removes shell antipatterns. New work should use III.
This version remains useful for one specific purpose: it overlaps III by only ~29%, so
it is the one generation that can be deduplicated… See the full description on the dataset page: https://huggingface.co/datasets/Frost2o24/bash-instruct-55k.bashkir-lora-qlora-benchmark
Bashkir LoRA/QLoRA Benchmark
📊 Description
This benchmark contains the complete results of fine-tuning various language models (from 82M to 7B parameters) on the Bashkir language. The study compares the effectiveness of LoRA/QLoRA against full fine-tuning, evaluating model quality (perplexity), GPU memory usage, and training time.
Key Findings
Mistral-7B with QLoRA (r=16) achieved the best performance among 7B models (perplexity 3.79)
LoRA drastically reduces… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-lora-qlora-benchmark.bashkir-periodicals-izddom
Bashkir Periodicals Cleaned (Izddom)
High-quality cleaned and structured Bashkir periodicals (Respublika Bashkortostan Publishing House) for LLM pretraining and fine-tuning.
Overview
A curated, deduplicated and filtered edition of modern Bashkir periodicals derived from the original bashkorttele/periodicals-izddom dataset published by the Foundation for the Preservation and Development of the Bashkir Language.
The dataset includes material from 13 leading… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-periodicals-izddom.bashkir-books-kitap
Bashkir Books & Calendars Cleaned (Kitap)
Curated, cleaned and structured Bashkir literature and annual calendars (Kitap Publishing House) for LLM pretraining and fine-tuning.
Overview
A curated, deduplicated and filtered edition of modern Bashkir book publications and annual cultural calendars derived from the original bashkorttele/books-kitap dataset published by the Foundation for the Preservation and Development of the Bashkir Language.
The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-books-kitap.rlvr-bash-terminal-bench
rlvr-bash-terminal-bench
RLVR (Reinforcement Learning with Verifiable Rewards) dataset for bash scripting, generated from Terminal-Bench tasks.
Stats
Metric
Value
Total samples
1,120
Unique tasks
88
Avg samples/task
12.7
Average reward
0.249
Perfect solutions (reward=1.0)
10.4%
Partial solutions (0<reward<1)
28.8%
Zero reward
60.8%
Tasks fully solved
13.6%
Format
{
"task_id": "string",
"prompt": "string",
"completion":… See the full description on the dataset page: https://huggingface.co/datasets/lisayan/rlvr-bash-terminal-bench.books-kitap
Bashkir Books — "Kitap" Publishing House
9,019,878 characters of Bashkir-language text (6,081 records) from publications of "Kitap" Publishing House. A record is an article or a work where the publication marks its boundaries (a table of contents, a day of a tear-off calendar); elsewhere it is cut at a heading or at a page. Every record carries metadata sufficient to locate it in the original publication.
🌐 Languages of this card: English · Башҡортса · Русский
Part of the… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/books-kitap.bash-agent-grpo-pairs
Bash Agent GRPO Pairs
Single-turn (intent → shell command) pairs for training a small, local, Claude-Code-style
bash agent with SFT or GRPO. Each record pairs a natural-language objective with exactly
one verifiable bash command, framed as a single-tool bash(command, description) call.
The dataset is designed to be rewardable: the ground-truth command is a deterministic target,
so a shell-equivalence reward (canonical program + flag set + argument comparison) can score… See the full description on the dataset page: https://huggingface.co/datasets/adeelahmad/bash-agent-grpo-pairs.bashcraft-data
Bashcraft v1 training data
Bashcraft is a small, original, agent-assisted dataset for mapping an English
request and explicit environment context to a Bash command or a clarification
question. It supports learning about supervised fine-tuning and outcome-based
shell evaluation. It does not establish general Bash competence or certify
commands as safe to run on a real machine.
This release contains the full 2,000-record training split. The project's
110 validation records and 220… See the full description on the dataset page: https://huggingface.co/datasets/nima1/bashcraft-data.periodicals-izddom
Bashkir Periodicals — Respublika Bashkortostan Publishing House
163,240,215 characters of Bashkir-language text (39,711 records) from publications of Respublika Bashkortostan Publishing House. A record is an article or a work where the publication marks its boundaries (a table of contents, a day of a tear-off calendar); elsewhere it is cut at a heading or at a page. Every record carries metadata sufficient to locate it in the original publication.
🌐 Languages of this card:… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/periodicals-izddom.bashkir-web-corpus
Dataset Card for Bashkir Web Corpus
Dataset Details
Dataset Description
The Bashkir Web Corpus is a collection of 71,567 documents and approximately 46.9 million tokens in the Bashkir language (a Turkic language spoken in Bashkortostan, Russia). The corpus was compiled from 16 Bashkir‑language online sources, including news websites, literary magazines, social media, books, and Wikipedia. It is designed for language modeling, text classification… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-web-corpus.bashkir-russian-parallel
Dataset Card for Bashkir-Russian Parallel Corpus
Dataset Details
Dataset Description
Bashkir-Russian Parallel Corpus is a large-scale sentence-aligned parallel corpus for the Bashkir–Russian language pair, assembled from authentic human-created translations. It contains 3,040,085 unique parallel sentence pairs, where each Bashkir sentence is aligned with its Russian counterpart.
The corpus combines data from three open parallel corpora: TIL-MT… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-russian-parallel.bash-commands-dataset-rev-ja
trtd56/bash-commands-dataset-rev-ja
bash_commands_ja.csv から作成した、Bash コマンドを入力して日本語説明を出力するためのデータセットです。
Columns
prompt_ja: 日本語説明
prompt_en: 英語説明
response: Bash コマンド
prompt: 学習用入力。response と同じ
completion: 学習用出力。prompt_ja と同じ
task: タスク識別子
language: 出力言語
Splits
train: 756
test: 84
Usage
from datasets import load_dataset
ds = load_dataset("trtd56/bash-commands-dataset-rev-ja")
print(ds["train"][0]["prompt"])
print(ds["train"][0]["completion"])
bashkir-wiki-corpus
Dataset Card for Bashkir Wikipedia Corpus
Dataset Details
Dataset Description
The Bashkir Wikipedia Corpus is a collection of 43,926 articles from Bashkir Wikipedia and Wikibooks, totaling approximately 10,617,785 XLM-R tokens and 8,862,807 words. The data has been extracted from official Wikimedia dumps and processed to provide clean, well‑structured text suitable for NLP tasks. The corpus includes article titles, full content, categories, source… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-wiki-corpus.BashCoder
🌀 Claude-opus-4.7-TraceInversion-5000x
v1.0 Release
A High-Fidelity Reconstructed CoT Dataset Saturated with the 'Opus Deep Logic Style' via Trace Inversion
📊 5,000 Samples
🧬 Trace Inversion & Negentropy
🛠 SFT & DPO Ready
🔥 Claude 4.7-Max Distillation
🌐 English & Multilingual
💡 What is Trace Inversion?
In Large Language Model (LLM) reasoning distillation, proprietary API models (such as GPT-4/5 and Claude)… See the full description on the dataset page: https://huggingface.co/datasets/Coding-With-Bashir/BashCoder.bashkir-lexicon
Dataset Card for Bashkir Lexicon (Machine Fund)
Dataset Details
Dataset Description
Bashkir Lexicon (Machine Fund) is a machine-readable lexical database of the Bashkir language, compiled from the open digital resource Machine Fund of the Bashkir Language (mfbl2.ru). It contains 34,008 unique lexical entries covering dialectal word forms with their part of speech, dialect, subdialect, literary norm, and Russian translation.
The dataset preserves… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-lexicon.bash-reference-manual-general-QAs
Dataset generated from bash reference manual.
book information like date and bash version are available within the very first rows of the dataset
this dataset is pretty small in general, but covering almost all of the definition and technical terms, commands and flags in the book
columns : "Question", "Answer"
bash-org-archive.comDetails
This is an unofficial dataset of an archive mirror of Bash.org: https://bash-org-archive.com/
Bash.org was a website launched in 1999 dedicated to archiving funny quotes from IRC other chat platforms over the years.
It offers a look into jokes, memes, and often inappropriate content that was quite commonplace at the time.
This dataset has been cleaned with a custom parser, aiming to preserve the original format of the content.
The parquet file contains the following columns:
qid… See the full description on the dataset page: https://huggingface.co/datasets/Taranosaurus/bash-org-archive.com.diverse-bash-dataset
Diverse bash terminal dataset
A dataset containing Question&Code pairs for most of the standard bash commands.
columns : "Question", "Code answer".
Commands covered : echo, cat, cd, rm, mkdir, top, free, du, df, ps, head, tail, grep, cp, cut, sort, touch, ls, groupadd, ifconfig, ip, ln, ping, scp, ssh, sudo, systemctl, tar, useradd, userdel, usermod, wc
specs : high quality dataset with precise coding examples for each command, covering (almost) all the flags of every command… See the full description on the dataset page: https://huggingface.co/datasets/datasetter458/diverse-bash-dataset.bashkir-news-cluster
Dataset Card for Bashkir News Cluster Dataset
Dataset Details
Dataset Description
This dataset contains 24,428 Bashkir-language news and analytical articles collected from various online sources. It is intended for clustering, representation learning, and unsupervised NLP tasks. Each text is accompanied by metadata such as title, source, date, and original category. The corpus is part of the BashkirNLP project, aiming to support low-resource… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-cluster.diverse-bash-dataset-extended
Bash diverse dataset extended
This dataset is a continious work from datasetter458/bash-diverse-dataset
Commands covered: bash, apt, crontab, date, diff, docker, env, export, fdisk, file, fsck, gcc, git, history, join, kubectl, lsof, man, makefs, mount, nice, node, nohup, npm,
passwd, paste, patch, pkill, python, renice, split, ss, time, umount, uname, uniq, vi, whereis, which
Note : For each command, a big amount of Q&As, and each set of them covers pretty much all the flags for… See the full description on the dataset page: https://huggingface.co/datasets/datasetter458/diverse-bash-dataset-extended.
