datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MichaelYitzchak
UAV Fault Symptom Reports
How to read this project
Problem. At the moment of a UAV incident, the operator describes what is happening or enters the live readings;
the system finds the most similar past faults and shows the class guidance recorded for that fault type, withheld
when the evidence is uncertain; each component has a pass mark set before it was scored. App: MichaelYitzchak/uav-similar-incident-workbench.
Step
Notebook
Course part
1… See the full description on the dataset page: https://huggingface.co/datasets/Bashifu/MichaelYitzchak.ShIO-bash-26.1
ShIO-bash-26.1
Shell input-output (ShIO) Bash dataset produced by ShIOEnv, a Gymnasium-compatible Bash environment designed to collect execution-annotated command interactions in a Linux system.
Dataset summary
The dataset consists of command-line inputs paired with their execution artifacts, including observable outputs and a structured representation of environment state changes. Samples are produced by executing synthesized Bash inputs inside a… See the full description on the dataset page: https://huggingface.co/datasets/jragsdale1/ShIO-bash-26.1.bashkir-frequency-index
Bashkir Frequency Index v11.6
Word-frequency index for the Bashkir language, computed over 166.46 million tokens of clean monolingual text across diverse public domains (periodicals, news media, literary publications, encyclopedic texts, books, and general web archives). Non-Bashkir language admixture and scanning artifacts were filtered using automated language-filtering pipelines.
Configurations
Config
Rows
Cutoff
Use Case
public (recommended)
663,196… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-frequency-index.bashkir-ngram-index
Bashkir Word N-gram Index v11.6
Exact within-sentence word n-gram counts for Bashkir: unigrams, bigrams and
trigrams for spellchecking, OCR post-processing and lightweight language modelling.
Overview
Exact word n-gram counts derived from a monolingual Bashkir-language dataset. The
release provides unigram, bigram and trigram indexes for corpus processing,
spellchecking, OCR post-processing, autocomplete and lightweight language-model
experiments. The unigrams… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-ngram-index.bashkir-wikipedia-parallel
Bashkir-Russian Wikipedia Parallel Corpus
Sentence-level Bashkir-Russian parallel text from Wikipedia, scored and filtered
for machine translation.
Overview
Sentence-level Bashkir-Russian parallel dataset extracted from the corresponding
Bashkir and Russian Wikipedia dumps dated 2026-08-01. Candidate pairs are scored
for semantic alignment with multilingual sentence encoders (Meta LASER3, Google
LaBSE) and the in-domain Bashkir-Russian Pair Scorer. The filtered… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-parallel.labeled-bashBench
LLM Misbehavior Activation Dataset
Dataset of labeled agent trajectory steps for use with steering vector / activation extraction.
Source
This dataset labels the trajectories found in mandliya/basharena-synthetic-trajectories.
Structure
Each row is ONE specific step or flagged action from the full original agent trajectory.
Field
Description
id
Unique entry UUID
task_id
Original BashArena task_id
source_file
Path to the original trajectory file… See the full description on the dataset page: https://huggingface.co/datasets/AISafety-Student/labeled-bashBench.bimanual_so100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_aloha",
"total_episodes": 30,
"total_frames": 22836,
"total_tasks": 1,
"total_videos": 90,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bashar-Alhaffar/bimanual_so100.rlvr-bash-terminal-bench
rlvr-bash-terminal-bench
RLVR (Reinforcement Learning with Verifiable Rewards) dataset for bash scripting, generated from Terminal-Bench tasks.
Stats
Metric
Value
Total samples
1,120
Unique tasks
88
Avg samples/task
12.7
Average reward
0.249
Perfect solutions (reward=1.0)
10.4%
Partial solutions (0<reward<1)
28.8%
Zero reward
60.8%
Tasks fully solved
13.6%
Format
{
"task_id": "string",
"prompt": "string",
"completion":… See the full description on the dataset page: https://huggingface.co/datasets/lisayan/rlvr-bash-terminal-bench.bashkir-periodicals-izddom
Bashkir Periodicals Cleaned (Izddom)
High-quality cleaned and structured Bashkir periodicals (Respublika Bashkortostan Publishing House) for LLM pretraining and fine-tuning.
Overview
A curated, deduplicated and filtered edition of modern Bashkir periodicals derived from the original bashkorttele/periodicals-izddom dataset published by the Foundation for the Preservation and Development of the Bashkir Language.
The dataset includes material from 13 leading… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-periodicals-izddom.bashkir-books-kitap
Bashkir Books & Calendars Cleaned (Kitap)
Curated, cleaned and structured Bashkir literature and annual calendars (Kitap Publishing House) for LLM pretraining and fine-tuning.
Overview
A curated, deduplicated and filtered edition of modern Bashkir book publications and annual cultural calendars derived from the original bashkorttele/books-kitap dataset published by the Foundation for the Preservation and Development of the Bashkir Language.
The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-books-kitap.books-kitap
Bashkir Books — "Kitap" Publishing House
9,019,878 characters of Bashkir-language text (6,081 records) from publications of "Kitap" Publishing House. A record is an article or a work where the publication marks its boundaries (a table of contents, a day of a tear-off calendar); elsewhere it is cut at a heading or at a page. Every record carries metadata sufficient to locate it in the original publication.
🌐 Languages of this card: English · Башҡортса · Русский
Part of the… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/books-kitap.bash-reference
Bash Reference Manual — Cleaned Dataset
This dataset contains cleaned and structured text extracted from the GNU Bash Reference Manual, Edition 5.3.
The original manual was converted to plain text and processed to reduce PDF extraction artifacts, remove navigation references and page numbers, preserve paragraph structure, and split the content into topic-centered sections.
The dataset is intended primarily for language-model pretraining and continued pretraining, especially for… See the full description on the dataset page: https://huggingface.co/datasets/superAVTR/bash-reference.basharena-monitor-evalperiodicals-izddom
Bashkir Periodicals — Respublika Bashkortostan Publishing House
163,240,215 characters of Bashkir-language text (39,711 records) from publications of Respublika Bashkortostan Publishing House. A record is an article or a work where the publication marks its boundaries (a table of contents, a day of a tear-off calendar); elsewhere it is cut at a heading or at a page. Every record carries metadata sufficient to locate it in the original publication.
🌐 Languages of this card:… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/periodicals-izddom.bashkir-russian-parallel
Dataset Card for Bashkir-Russian Parallel Corpus
Dataset Details
Dataset Description
Bashkir-Russian Parallel Corpus is a large-scale sentence-aligned parallel corpus for the Bashkir–Russian language pair, assembled from authentic human-created translations. It contains 3,040,085 unique parallel sentence pairs, where each Bashkir sentence is aligned with its Russian counterpart.
The corpus combines data from three open parallel corpora: TIL-MT… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-russian-parallel.basharena_action_only_opus46_largeEDA_Assignment
🎓 Student Dropout Prediction Dataset — EDA Assignment
By Tomer Bash | Data Science Course — Assignment #1
📹 Presentation Video
Presentation Video link - https://youtu.be/KyafBx9W7Qg
📌 Dataset Overview
Property
Details
Source
Kaggle
Rows
4,424 students
Features
35 columns
Target Variable
Target — Graduate, Enrolled, Dropout
Task Type
Multi-class Classification
The dataset contains demographic, financial, academic, and… See the full description on the dataset page: https://huggingface.co/datasets/Bashifu/EDA_Assignment.basharena_action_only_sonnet45_largebashkir-news-multilabel
Dataset Card for Bashkir News Multilabel Classification Dataset
Dataset Details
Dataset Description
This dataset contains 22,318 Bashkir-language news and analytical articles annotated with 14 thematic labels for multi-label text classification tasks. Each article can belong to several categories simultaneously. The average number of labels per article is 3.6. The dataset is designed to support NLP research and applications for the Bashkir language… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-multilabel.Bwenge
Dataset Description
This dataset was created to develop a machine translation model for bidirectional translation between Kinyarwanda and English for education-based sentences, in particular for the Atingi learning platform.
Repository:link to the GitHub repository containing the code for training the model on this data, and the code for the collection of the monolingual data.
Data Format: TSV
Model: huggingface model link.
Dataset Summary
Data… See the full description on the dataset page: https://huggingface.co/datasets/Coding-With-Bashir/Bwenge.so100_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_aloha",
"total_episodes": 3,
"total_frames": 2024,
"total_tasks": 1,
"total_videos": 9,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bashar-Alhaffar/so100_test.bashkir-news-binary
Dataset Card for Bashkir News Binary Classification Dataset
Dataset Details
Dataset Description
This dataset contains 16,994 Bashkir-language news and analytical articles labeled for binary classification: news (label=1) vs analytics (label=0). The dataset is perfectly balanced with 8,497 examples in each class. It was created to support NLP research and applications for the Bashkir language, a low-resource Turkic language.
Curated by: Arabov… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-binary.css-colorsintercode-bash-qwen7b-activations50k_nepali_chatbot_datasettitanic_datasetbasharena_xml_with_assistant_textbasharena_awarenepali_chatbot_datasetbasharena_action_only
