Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Bashifu /MichaelYitzchak UAV Fault Symptom Reports How to read this project Problem. At the moment of a UAV incident, the operator describes what is happening or enters the live readings; the system finds the most similar past faults and shows the class guidance recorded for that fault type, withheld when the evidence is uncertain; each component has a pass mark set before it was scored. App: MichaelYitzchak/uav-similar-incident-workbench. Step Notebook Course part 1… See the full description on the dataset page: https://huggingface.co/datasets/Bashifu/MichaelYitzchak.imagetext-classification10K<n<100K0 likes1.1k downloads8d agoHugging Face02jragsdale1 /ShIO-bash-26.1 ShIO-bash-26.1 Shell input-output (ShIO) Bash dataset produced by ShIOEnv, a Gymnasium-compatible Bash environment designed to collect execution-annotated command interactions in a Linux system. Dataset summary The dataset consists of command-line inputs paired with their execution artifacts, including observable outputs and a structured representation of environment state changes. Samples are produced by executing synthesized Bash inputs inside a… See the full description on the dataset page: https://huggingface.co/datasets/jragsdale1/ShIO-bash-26.1.imagetext-generation1M<n<10M1 likes606 downloads4mo agoHugging Face03Eccentricity /bashbench2 BashBench2 A successor in spirit to the original BashBench, this dataset is intended for high-stakes agentic control research using current and near-future frontier models. The code required to set up and run these tasks is located in ControlArena. textn<1K1 likes474 downloads1y agoHugging Face04Frost2o24 /bash-instruct-III-55k Bash Instruct III — 54,360 verified natural-language → Bash pairs Bash Instruct III is a synthetic instruction-tuning dataset that maps natural-language requests to correct Bash: single commands, short pipelines, and multi-line scripts. It is built for supervised fine-tuning of small and mid-size LLMs that must turn a plain request into shell code that actually runs. Every row is a three-turn chat conversation (system / user / assistant) with metadata for slicing (category… See the full description on the dataset page: https://huggingface.co/datasets/Frost2o24/bash-instruct-III-55k.texttext-generation10K<n<100K0 likes350 downloads1mo agoHugging Face05failed09 /bashkir-frequency-index Bashkir Frequency Index v12.0-af1 Word-frequency index for the Bashkir language, computed over 110.94 million tokens of clean monolingual text across diverse public domains (periodicals, news media, literary publications, encyclopedic texts, books, and general web archives). Non-Bashkir language admixture and scanning artifacts were filtered using automated language-filtering pipelines. Configurations Config Rows Cutoff Use Case public (recommended) 516… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-frequency-index.tabulartext-classification1M<n<10M0 likes332 downloads7h agoHugging Face06bashkorttele /broadcast-speech Bashkir Broadcast Speech — Radio and Television 53.3 hours of speech in 293 recordings in the Bashkir language, from television and radio programmes produced by two public broadcasters of the Republic of Bashkortostan. Audio only — no transcripts in this release — which makes the set suitable for self-supervised speech pretraining for a low-resource Turkic language. 🌐 Languages of this card: English · Башҡортса · Русский Part of the Bashkorttele dataset series — preservation… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/broadcast-speech.audioautomatic-speech-recognitionn<1K2 likes278 downloads1mo agoHugging Face07AigizK /bashkort_commands_omnivoice Bashkort Commands OmniVoice Partial eleven-label command snapshot generated with k2-fsa/OmniVoice using the same cross-lingual voice-cloning recipe as AigizK/homai_wake_word_omnivoice. Generation was stopped at the user's request after 41,525 complete reference groups had been committed. For every included reference row from the train split of: bond005/sova_rudevices the dataset contains one recording of every command: Айвика — Russian Айвикә — Bashkir Айһылыу — Bashkir… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_commands_omnivoice.audioaudio-classification100K<n<1M0 likes241 downloads2mo agoHugging Face08emirkaanozdemr /bash_command_data_6K 📦 Bash Command Dataset v1 A high-quality dataset of natural language instructions paired with their equivalent Bash commands, designed for training and fine-tuning large language models (LLMs) that translate English tasks into shell commands. This dataset is ideal for researchers, developers, and machine learning engineers interested in natural language to Bash command translation, command-line automation, and building intelligent terminal assistants. 📁 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/emirkaanozdemr/bash_command_data_6K.texttext-generation1K<n<10K8 likes215 downloads6mo agoHugging Face09failed09 /bashkir-ngram-index Bashkir Word N-gram Index v11.6 Exact within-sentence word n-gram counts for Bashkir: unigrams, bigrams and trigrams for spellchecking, OCR post-processing and lightweight language modelling. Overview Exact word n-gram counts derived from a monolingual Bashkir-language dataset. The release provides unigram, bigram and trigram indexes for corpus processing, spellchecking, OCR post-processing, autocomplete and lightweight language-model experiments. The unigrams… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-ngram-index.tabulartext-classification10M<n<100M0 likes199 downloads8h agoHugging Face10failed09 /bashkir-wikipedia-parallel Bashkir-Russian Wikipedia Parallel Corpus Sentence-level Bashkir-Russian parallel text from Wikipedia, scored and filtered for machine translation. Overview Sentence-level Bashkir-Russian parallel dataset extracted from the corresponding Bashkir and Russian Wikipedia dumps dated 2026-08-01. Candidate pairs are scored for semantic alignment with multilingual sentence encoders (Meta LASER3, Google LaBSE) and the in-domain Bashkir-Russian Pair Scorer. The filtered… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-parallel.tabulartranslation100K<n<1M0 likes170 downloads9d agoHugging Face11ISB369 /shellminator-bash-sft100ktext100K<n<1M0 likes169 downloads1mo agoHugging Face12ai-and /termgrade-bandcentral-gemma4-31b-bash TermGrade Band-Central: the 151-task selection and its splits Part of TermGrade: graded environments and trajectories for terminal agents. Read the blog post. 271 executable environments · every task carries a measured base pass rate · the exact RL split we trained on The 151 training and 120 validation environments behind ai-and/termgrade-gemma4-31b-lora, with the verl config that produced it. Companion releases: ai-and/termgrade-environments (1,004 environments graded by six… See the full description on the dataset page: https://huggingface.co/datasets/ai-and/termgrade-bandcentral-gemma4-31b-bash.textreinforcement-learningn<1K3 likes164 downloads2d agoHugging Face13metuKKhud /bashqort-raw Bashqort Raw Corpus Description This dataset contains raw Bashkir text collected for continual training of large language models (LLMs). It is part of the project "Adapting Open-Source LLMs for the Bashkir Language", which aims to evaluate adaptation methods proposed by LlamaTurk (Toraman, 2024) and Persian adaptation (Mahdizadeh Sani et al., 2024). The corpus is assembled from multiple sources to provide a diverse linguistic foundation for language modeling.… See the full description on the dataset page: https://huggingface.co/datasets/metuKKhud/bashqort-raw.text100K<n<1M0 likes157 downloads5mo agoHugging Face14aelhalili /bash-commands-dataset 🐧 Linux Command Automation Dataset A dataset of natural language prompts paired with their corresponding Bash command-line equivalents, designed to train or fine-tune models for automating Linux tasks via natural language. 📁 Dataset Structure The dataset is in JSON format, structured as a flat array of objects, where each object contains: { "prompt": "Natural language description of a task", "response": "Equivalent Bash command" } ✅ Example {… See the full description on the dataset page: https://huggingface.co/datasets/aelhalili/bash-commands-dataset.textn<1K8 likes154 downloads1y agoHugging Face15AigizK /bashkort_voice Bashkort Voice 🇬🇧 English Version Dataset Description This is a synthetic Bashkir audio dataset generated using the OmniVoice model. It is designed to expand the availability of spoken data for the Bashkir language. Data Preparation Process The dataset was constructed through a cross-lingual voice cloning and generation process, using the following methodology: Target Text: Bashkir sentences were extracted from the AigizK/bashkir-russian-parallel-corpora dataset.… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_voice.audioautomatic-speech-recognition100K<n<1M2 likes149 downloads6mo agoHugging Face16failed09 /bashkir-multilingual-phrasebooks Bashkir-Russian Phrasebook Corpus Edited Bashkir-Russian words, expressions and conversational phrases from university phrasebooks, annotated by entry type. Overview Edited Bashkir-Russian pairs derived from the original bashkorttele/trilingual-parallel-phrasebooks-bgpu dataset, published by Bashkorttele from phrasebooks of M. Akmulla Bashkir State Pedagogical University. The cleaned configuration is the deduplicated default; reviewed is the edited edition… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-multilingual-phrasebooks.texttranslation10K<n<100K0 likes146 downloads9d agoHugging Face17bashkorttele /trilingual-parallel-phrasebooks-bgpu Bashkir Trilingual Parallel Phrasebooks 9,857 phrases aligned across three languages — Bashkir, Russian and one of Altai, Arabic, Kazakh, Yakut (Sakha), Chinese — from five phrasebooks published by M. Akmulla Bashkir State Pedagogical University. One row is one phrase in all three languages: a parallel corpus for machine translation and cross-lingual work with a low-resource Turkic language. Each phrasebook is a separate file and a separate config, because the third language… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/trilingual-parallel-phrasebooks-bgpu.texttranslation1K<n<10K2 likes129 downloads10d agoHugging Face18AISafety-Student /labeled-bashBench LLM Misbehavior Activation Dataset Dataset of labeled agent trajectory steps for use with steering vector / activation extraction. Source This dataset labels the trajectories found in mandliya/basharena-synthetic-trajectories. Structure Each row is ONE specific step or flagged action from the full original agent trajectory. Field Description id Unique entry UUID task_id Original BashArena task_id source_file Path to the original trajectory file… See the full description on the dataset page: https://huggingface.co/datasets/AISafety-Student/labeled-bashBench.tabulartext-classification1K<n<10K1 likes123 downloads7mo agoHugging Face19bashkorttele /narrated-audiobooks-brsbs Narrated Bashkir Audiobooks — BRSBS (Bashkorttele) ≈ 78.7 hours of human-narrated audiobooks, primarily in the Bashkir language, drawn from public-domain literary works and folk epics. Recorded as accessible "talking books" by the Bashkir Republican Special Library for the Blind (BRSBS) and released for language preservation and AI/ML research. 🌐 Languages of this card: English · Башҡортса · Русский This dataset is part of a larger series published under the Bashkorttele… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/narrated-audiobooks-brsbs.audioautomatic-speech-recognitionn<1K3 likes119 downloads3mo agoHugging Face20DCAgent /bash_textbook_taskstext10K<n<100K0 likes108 downloads10mo agoHugging Face21AigizK /bashkort_tts_dataset Bashkort TTS Dataset The largest open dataset for speech synthesis in the Bashkir language — featuring multi-speaker recordings and speaking styles. 📊 Dataset Overview Total audio files: 62,852 Speakers: 7 female, 1 male Speaking styles: friendly, question, neutral Languages: Bashkir Format: MP3 audio + transcription text 🎙 How It Was Collected Initial recording: A female voice actor recorded ~15 hours of speech in Bashkir. Voice cloning: Using ElevenLabs… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_tts_dataset.audiotext-to-speech10K<n<100K3 likes103 downloads1y agoHugging Face22failed09 /bashkir-wikipedia-monolingual Bashkir Wikipedia Monolingual Corpus Cleaned sentence-level Bashkir text from Wikipedia for pretraining, tokenizer training and linguistic research. Overview Sentence-level text extracted from the Bashkir Wikipedia dump (bawiki-20260801), cleaned and filtered with automated language identification. The cleaned configuration is the recommended default for language modelling, tokenization and linguistic research; precleaned is an earlier, lighter extraction kept… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-monolingual.texttext-generation1M<n<10M0 likes86 downloads9d agoHugging Face23Frost2o24 /bash-instruct-II-55k Bash Instruct II — 54,803 verified natural-language → Bash pairs ⚠️ Superseded by Bash Instruct III III is a corrected rebuild of this dataset with shell antipatterns removed at the generator level (99.8% ShellCheck-clean vs 98.5% here). New work should use III. The two share 92.5% of their (request, command) pairs, so they must never be concatenated. This card is kept for reproducibility and citation of published results. Bash Instruct II is a synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Frost2o24/bash-instruct-II-55k.texttext-generation10K<n<100K0 likes81 downloads1mo agoHugging Face24ISB369 /shellminator-bash-combinedtext10K<n<100K0 likes73 downloads2mo agoHugging Face25Euroswarms /bash Bash & Linux CLI Dataset What This Dataset Is A supervised fine-tuning (SFT) dataset that teaches a language model to produce correct Bash commands, shell scripts, and Linux CLI operations — including Arch Linux and Ubuntu/Debian specifics. The assistant responses are executable commands or scripts that directly accomplish the user's stated task. Explanatory rows describe what a given command does, its key flags, and caveats. Sources (provenance-tracked)… See the full description on the dataset page: https://huggingface.co/datasets/Euroswarms/bash.text1M<n<10M1 likes72 downloads6mo agoHugging Face26open-athena /stack-bash-v3-qwen3.5-122b-32k-tracestext10K<n<100K0 likes72 downloads4mo agoHugging Face27DCAgent2 /DCAgent2_terminal_bench_2_DCAgent_bash_textbook_tasks_traces_20251123_000222textn<1K0 likes70 downloads11mo agoHugging Face28open-athena /rl__24GPU_base__exp_rpt_stack-bash__Qwen3-8Btext10K<n<100K0 likes70 downloads7mo agoHugging Face29LangChainHub-Prompts /LLM_Bash Description of LLM Bash Prompt designed to convert natural language to bash command. Inputs This is a description of the inputs that the prompt expects. question: User question to be answered by writing a bash command. Usage Below is a code snippet for how to use the prompt. from langchain.prompts import load_prompt from langchain.chains import LLMBashChain llm = ... prompt = load_prompt('lc://prompts/llm_bash/<file-name>') chain = LLMBashChain(llm=llm… See the full description on the dataset page: https://huggingface.co/datasets/LangChainHub-Prompts/LLM_Bash.textn<1K9 likes68 downloads4y agoHugging Face30GunA-SD /bash_codeThis dataset is a collection of bash programs from various GitHub repositories and open source projects. The dataset might contain harmful code. texttext-generation100K<n<1M9 likes68 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.