datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PlantMetricDepth
PlantMetricDepth
PlantMetricDepth is a multimodal plant dataset designed for metric monocular depth estimation (MDE) and related plant analysis tasks.
The dataset provides paired stereo RGB images, disparity maps, generated metric depth maps, and plant segmentation masks collected across 15 acquisition days.
The metric depth maps provide dense per-pixel depth supervision in centimetres, enabling models to learn metric depth from a single RGB image at inference time without… See the full description on the dataset page: https://huggingface.co/datasets/BashayerAA/PlantMetricDepth.mcpc-modMichaelYitzchak
UAV Fault Symptom Reports
How to read this project
Problem. At the moment of a UAV incident, the operator describes what is happening or enters the live readings;
the system finds the most similar past faults and shows the class guidance recorded for that fault type, withheld
when the evidence is uncertain; each component has a pass mark set before it was scored. App: MichaelYitzchak/uav-similar-incident-workbench.
Step
Notebook
Course part
1… See the full description on the dataset page: https://huggingface.co/datasets/Bashifu/MichaelYitzchak.ShIO-bash-26.1
ShIO-bash-26.1
Shell input-output (ShIO) Bash dataset produced by ShIOEnv, a Gymnasium-compatible Bash environment designed to collect execution-annotated command interactions in a Linux system.
Dataset summary
The dataset consists of command-line inputs paired with their execution artifacts, including observable outputs and a structured representation of environment state changes. Samples are produced by executing synthesized Bash inputs inside a… See the full description on the dataset page: https://huggingface.co/datasets/jragsdale1/ShIO-bash-26.1.bashbench2
BashBench2
A successor in spirit to the original BashBench, this dataset is intended for high-stakes agentic control research using current and near-future frontier models.
The code required to set up and run these tasks is located in ControlArena.
broadcast-speech
Bashkir Broadcast Speech — Radio and Television
53.3 hours of speech in 293 recordings in the Bashkir language, from television and radio programmes produced by two public broadcasters of the Republic of Bashkortostan. Audio only — no transcripts in this release — which makes the set suitable for self-supervised speech pretraining for a low-resource Turkic language.
🌐 Languages of this card: English · Башҡортса · Русский
Part of the Bashkorttele dataset series — preservation… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/broadcast-speech.bash-instruct-III-55k
Bash Instruct III — 54,360 verified natural-language → Bash pairs
Bash Instruct III is a synthetic instruction-tuning dataset that maps natural-language
requests to correct Bash: single commands, short pipelines, and multi-line scripts. It
is built for supervised fine-tuning of small and mid-size LLMs that must turn a plain
request into shell code that actually runs.
Every row is a three-turn chat conversation (system / user / assistant) with metadata
for slicing (category… See the full description on the dataset page: https://huggingface.co/datasets/Frost2o24/bash-instruct-III-55k.bashkort_commands_omnivoice
Bashkort Commands OmniVoice
Partial eleven-label command snapshot generated with k2-fsa/OmniVoice
using the same cross-lingual voice-cloning recipe as
AigizK/homai_wake_word_omnivoice. Generation was stopped at the user's
request after 41,525 complete reference groups had been committed.
For every included reference row from the train split of:
bond005/sova_rudevices
the dataset contains one recording of every command:
Айвика — Russian
Айвикә — Bashkir
Айһылыу — Bashkir… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_commands_omnivoice.english-handwriting-diffusionbashkir-frequency-index
Bashkir Frequency Index v11.6
Word-frequency index for the Bashkir language, computed over 166.46 million tokens of clean monolingual text across diverse public domains (periodicals, news media, literary publications, encyclopedic texts, books, and general web archives). Non-Bashkir language admixture and scanning artifacts were filtered using automated language-filtering pipelines.
Configurations
Config
Rows
Cutoff
Use Case
public (recommended)
663,196… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-frequency-index.sample-text-file-combinedtrilingual-parallel-phrasebooks-bgpu
Bashkir Trilingual Parallel Phrasebooks
9,857 phrases aligned across three languages — Bashkir, Russian and one of Altai, Arabic, Kazakh, Yakut (Sakha), Chinese — from five phrasebooks published by M. Akmulla Bashkir State Pedagogical University. One row is one phrase in all three languages: a parallel corpus for machine translation and cross-lingual work with a low-resource Turkic language. Each phrasebook is a separate file and a separate config, because the third language… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/trilingual-parallel-phrasebooks-bgpu.bash_command_data_6K
📦 Bash Command Dataset v1
A high-quality dataset of natural language instructions paired with their equivalent Bash commands, designed for training and fine-tuning large language models (LLMs) that translate English tasks into shell commands.
This dataset is ideal for researchers, developers, and machine learning engineers interested in natural language to Bash command translation, command-line automation, and building intelligent terminal assistants.
📁 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/emirkaanozdemr/bash_command_data_6K.bashkir-multilingual-phrasebooks
Bashkir-Russian Phrasebook Corpus
Edited Bashkir-Russian words, expressions and conversational phrases from
university phrasebooks, annotated by entry type.
Overview
Edited Bashkir-Russian pairs derived from the original
bashkorttele/trilingual-parallel-phrasebooks-bgpu
dataset, published by Bashkorttele from phrasebooks of M. Akmulla Bashkir State
Pedagogical University. The cleaned configuration is the deduplicated default;
reviewed is the edited edition… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-multilingual-phrasebooks.bashkir-ngram-index
Bashkir Word N-gram Index v11.6
Exact within-sentence word n-gram counts for Bashkir: unigrams, bigrams and
trigrams for spellchecking, OCR post-processing and lightweight language modelling.
Overview
Exact word n-gram counts derived from a monolingual Bashkir-language dataset. The
release provides unigram, bigram and trigram indexes for corpus processing,
spellchecking, OCR post-processing, autocomplete and lightweight language-model
experiments. The unigrams… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-ngram-index.bashkort_voice
Bashkort Voice
🇬🇧 English Version
Dataset Description
This is a synthetic Bashkir audio dataset generated using the OmniVoice model. It is designed to expand the availability of spoken data for the Bashkir language.
Data Preparation Process
The dataset was constructed through a cross-lingual voice cloning and generation process, using the following methodology:
Target Text: Bashkir sentences were extracted from the AigizK/bashkir-russian-parallel-corpora dataset.… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_voice.bashkir-wikipedia-parallel
Bashkir-Russian Wikipedia Parallel Corpus
Sentence-level Bashkir-Russian parallel text from Wikipedia, scored and filtered
for machine translation.
Overview
Sentence-level Bashkir-Russian parallel dataset extracted from the corresponding
Bashkir and Russian Wikipedia dumps dated 2026-08-01. Candidate pairs are scored
for semantic alignment with multilingual sentence encoders (Meta LASER3, Google
LaBSE) and the in-domain Bashkir-Russian Pair Scorer. The filtered… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-parallel.bash-commands-dataset
🐧 Linux Command Automation Dataset
A dataset of natural language prompts paired with their corresponding Bash command-line equivalents, designed to train or fine-tune models for automating Linux tasks via natural language.
📁 Dataset Structure
The dataset is in JSON format, structured as a flat array of objects, where each object contains:
{
"prompt": "Natural language description of a task",
"response": "Equivalent Bash command"
}
✅ Example
{… See the full description on the dataset page: https://huggingface.co/datasets/aelhalili/bash-commands-dataset.shellminator-bash-sft100kbashkort_tts_dataset
Bashkort TTS Dataset
The largest open dataset for speech synthesis in the Bashkir language — featuring multi-speaker recordings and speaking styles.
📊 Dataset Overview
Total audio files: 62,852
Speakers: 7 female, 1 male
Speaking styles: friendly, question, neutral
Languages: Bashkir
Format: MP3 audio + transcription text
🎙 How It Was Collected
Initial recording: A female voice actor recorded ~15 hours of speech in Bashkir.
Voice cloning: Using ElevenLabs… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_tts_dataset.narrated-audiobooks-brsbs
Narrated Bashkir Audiobooks — BRSBS (Bashkorttele)
≈ 78.7 hours of human-narrated audiobooks, primarily in the Bashkir language, drawn from public-domain literary works and folk epics. Recorded as accessible "talking books" by the Bashkir Republican Special Library for the Blind (BRSBS) and released for language preservation and AI/ML research.
🌐 Languages of this card: English · Башҡортса · Русский
This dataset is part of a larger series published under the Bashkorttele… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/narrated-audiobooks-brsbs.bash_textbook_taskslabeled-bashBench
LLM Misbehavior Activation Dataset
Dataset of labeled agent trajectory steps for use with steering vector / activation extraction.
Source
This dataset labels the trajectories found in mandliya/basharena-synthetic-trajectories.
Structure
Each row is ONE specific step or flagged action from the full original agent trajectory.
Field
Description
id
Unique entry UUID
task_id
Original BashArena task_id
source_file
Path to the original trajectory file… See the full description on the dataset page: https://huggingface.co/datasets/AISafety-Student/labeled-bashBench.bash-instruct-II-55k
Bash Instruct II — 54,803 verified natural-language → Bash pairs
⚠️ Superseded by Bash Instruct III
III is a corrected rebuild of this dataset with shell antipatterns removed at the
generator level (99.8% ShellCheck-clean vs 98.5% here). New work should use III.
The two share 92.5% of their (request, command) pairs, so they must never be
concatenated. This card is kept for reproducibility and citation of published results.
Bash Instruct II is a synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Frost2o24/bash-instruct-II-55k.DCAgent2_terminal_bench_2_DCAgent_bash_textbook_tasks_traces_20251123_000222rl__40GPU_base_32b__exp_rpt_nemotron-bash__Qwen3-32Bbash_codeThis dataset is a collection of bash programs from various GitHub repositories and open source projects.
The dataset might contain harmful code.
bash-instruct-55k
Bash Instruct I — 55,000 verified natural-language → Bash pairs
⚠️ Superseded by Bash Instruct III
III doubles the utility vocabulary (89 → 182), adds grouped equivalent answers, real
human phrasing from tldr-pages, validation by execution on real Linux, and a style pass
that removes shell antipatterns. New work should use III.
This version remains useful for one specific purpose: it overlaps III by only ~29%, so
it is the one generation that can be deduplicated… See the full description on the dataset page: https://huggingface.co/datasets/Frost2o24/bash-instruct-55k.bashqort-raw
Bashqort Raw Corpus
Description
This dataset contains raw Bashkir text collected for continual training of large language models (LLMs). It is part of the project "Adapting Open-Source LLMs for the Bashkir Language", which aims to evaluate adaptation methods proposed by LlamaTurk (Toraman, 2024) and Persian adaptation (Mahdizadeh Sani et al., 2024).
The corpus is assembled from multiple sources to provide a diverse linguistic foundation for language modeling.… See the full description on the dataset page: https://huggingface.co/datasets/metuKKhud/bashqort-raw.Stack2Graph_KG_bash
Bash StackOverflow Knowledge Graph
Summary
This Hugging Face dataset repository contains the Bash shard of the Stack2Graph StackOverflow Knowledge Graph.
Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper.
The artifact is optimized for graph-based retrieval, SPARQL analytics, and retrieval-augmented question answering over Stack Overflow content.
Stack2Graph source:… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_KG_bash.
