Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Frost2o24 /bash-instruct-III-55k Bash Instruct III — 54,360 verified natural-language → Bash pairs Bash Instruct III is a synthetic instruction-tuning dataset that maps natural-language requests to correct Bash: single commands, short pipelines, and multi-line scripts. It is built for supervised fine-tuning of small and mid-size LLMs that must turn a plain request into shell code that actually runs. Every row is a three-turn chat conversation (system / user / assistant) with metadata for slicing (category… See the full description on the dataset page: https://huggingface.co/datasets/Frost2o24/bash-instruct-III-55k.texttext-generation10K<n<100K0 likes350 downloads1mo agoHugging Face02ISB369 /shellminator-bash-sft100ktext100K<n<1M0 likes169 downloads1mo agoHugging Face03aelhalili /bash-commands-dataset 🐧 Linux Command Automation Dataset A dataset of natural language prompts paired with their corresponding Bash command-line equivalents, designed to train or fine-tune models for automating Linux tasks via natural language. 📁 Dataset Structure The dataset is in JSON format, structured as a flat array of objects, where each object contains: { "prompt": "Natural language description of a task", "response": "Equivalent Bash command" } ✅ Example {… See the full description on the dataset page: https://huggingface.co/datasets/aelhalili/bash-commands-dataset.textn<1K8 likes154 downloads1y agoHugging Face04bashkorttele /trilingual-parallel-phrasebooks-bgpu Bashkir Trilingual Parallel Phrasebooks 9,857 phrases aligned across three languages — Bashkir, Russian and one of Altai, Arabic, Kazakh, Yakut (Sakha), Chinese — from five phrasebooks published by M. Akmulla Bashkir State Pedagogical University. One row is one phrase in all three languages: a parallel corpus for machine translation and cross-lingual work with a low-resource Turkic language. Each phrasebook is a separate file and a separate config, because the third language… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/trilingual-parallel-phrasebooks-bgpu.texttranslation1K<n<10K2 likes129 downloads10d agoHugging Face05Frost2o24 /bash-instruct-II-55k Bash Instruct II — 54,803 verified natural-language → Bash pairs ⚠️ Superseded by Bash Instruct III III is a corrected rebuild of this dataset with shell antipatterns removed at the generator level (99.8% ShellCheck-clean vs 98.5% here). New work should use III. The two share 92.5% of their (request, command) pairs, so they must never be concatenated. This card is kept for reproducibility and citation of published results. Bash Instruct II is a synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Frost2o24/bash-instruct-II-55k.texttext-generation10K<n<100K0 likes81 downloads1mo agoHugging Face06ISB369 /shellminator-bash-combinedtext10K<n<100K0 likes73 downloads2mo agoHugging Face07Euroswarms /bash Bash & Linux CLI Dataset What This Dataset Is A supervised fine-tuning (SFT) dataset that teaches a language model to produce correct Bash commands, shell scripts, and Linux CLI operations — including Arch Linux and Ubuntu/Debian specifics. The assistant responses are executable commands or scripts that directly accomplish the user's stated task. Explanatory rows describe what a given command does, its key flags, and caveats. Sources (provenance-tracked)… See the full description on the dataset page: https://huggingface.co/datasets/Euroswarms/bash.text1M<n<10M1 likes72 downloads6mo agoHugging Face08LangChainHub-Prompts /LLM_Bash Description of LLM Bash Prompt designed to convert natural language to bash command. Inputs This is a description of the inputs that the prompt expects. question: User question to be answered by writing a bash command. Usage Below is a code snippet for how to use the prompt. from langchain.prompts import load_prompt from langchain.chains import LLMBashChain llm = ... prompt = load_prompt('lc://prompts/llm_bash/<file-name>') chain = LLMBashChain(llm=llm… See the full description on the dataset page: https://huggingface.co/datasets/LangChainHub-Prompts/LLM_Bash.textn<1K9 likes68 downloads4y agoHugging Face09kth8 /bash-toolcallstext10K<n<100K0 likes66 downloads7d agoHugging Face10Frost2o24 /bash-instruct-55k Bash Instruct I — 55,000 verified natural-language → Bash pairs ⚠️ Superseded by Bash Instruct III III doubles the utility vocabulary (89 → 182), adds grouped equivalent answers, real human phrasing from tldr-pages, validation by execution on real Linux, and a style pass that removes shell antipatterns. New work should use III. This version remains useful for one specific purpose: it overlaps III by only ~29%, so it is the one generation that can be deduplicated… See the full description on the dataset page: https://huggingface.co/datasets/Frost2o24/bash-instruct-55k.texttext-generation10K<n<100K1 likes61 downloads1mo agoHugging Face11lisayan /rlvr-bash-terminal-bench rlvr-bash-terminal-bench RLVR (Reinforcement Learning with Verifiable Rewards) dataset for bash scripting, generated from Terminal-Bench tasks. Stats Metric Value Total samples 1,120 Unique tasks 88 Avg samples/task 12.7 Average reward 0.249 Perfect solutions (reward=1.0) 10.4% Partial solutions (0<reward<1) 28.8% Zero reward 60.8% Tasks fully solved 13.6% Format { "task_id": "string", "prompt": "string", "completion":… See the full description on the dataset page: https://huggingface.co/datasets/lisayan/rlvr-bash-terminal-bench.tabulartext-generation1K<n<10K0 likes57 downloads9mo agoHugging Face12Jawajawa /command-linux-bash-balanced-sfttext1K<n<10K1 likes54 downloads6mo agoHugging Face13ISB369 /shellminator-bash-datasettext10K<n<100K1 likes52 downloads2mo agoHugging Face14bashkorttele /books-kitap Bashkir Books — "Kitap" Publishing House 9,019,878 characters of Bashkir-language text (6,081 records) from publications of "Kitap" Publishing House. A record is an article or a work where the publication marks its boundaries (a table of contents, a day of a tear-off calendar); elsewhere it is cut at a heading or at a page. Every record carries metadata sufficient to locate it in the original publication. 🌐 Languages of this card: English · Башҡортса · Русский Part of the… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/books-kitap.tabulartext-generation1K<n<10K0 likes51 downloads10d agoHugging Face15BASH-Lab /SensorCaps SensorCaps SensorCaps is an LLM-assisted softly-labelled IMU sensor data captioning dataset with feature summarizations and narrations of human activities. Abstract Wearable systems can recognize activities from IMU data but often fail to explain their underlying causes or contextual significance. To address this limitation, we introduce two large-scale resources: SensorCap, comprising 35,960 IMU--caption pairs, and OpenSQA, with 199,701 question--answer pairs designed… See the full description on the dataset page: https://huggingface.co/datasets/BASH-Lab/SensorCaps.textsummarization10K<n<100K0 likes44 downloads1y agoHugging Face16superAVTR /bash-reference Bash Reference Manual — Cleaned Dataset This dataset contains cleaned and structured text extracted from the GNU Bash Reference Manual, Edition 5.3. The original manual was converted to plain text and processed to reduce PDF extraction artifacts, remove navigation references and page numbers, preserve paragraph structure, and split the content into topic-centered sections. The dataset is intended primarily for language-model pretraining and continued pretraining, especially for… See the full description on the dataset page: https://huggingface.co/datasets/superAVTR/bash-reference.tabularn<1K0 likes42 downloads15d agoHugging Face17adeelahmad /bash-agent-grpo-pairs Bash Agent GRPO Pairs Single-turn (intent → shell command) pairs for training a small, local, Claude-Code-style bash agent with SFT or GRPO. Each record pairs a natural-language objective with exactly one verifiable bash command, framed as a single-tool bash(command, description) call. The dataset is designed to be rewardable: the ground-truth command is a deterministic target, so a shell-equivalence reward (canonical program + flag set + argument comparison) can score… See the full description on the dataset page: https://huggingface.co/datasets/adeelahmad/bash-agent-grpo-pairs.texttext-generation10K<n<100K0 likes41 downloads3mo agoHugging Face18nima1 /bashcraft-data Bashcraft v1 training data Bashcraft is a small, original, agent-assisted dataset for mapping an English request and explicit environment context to a Bash command or a clarification question. It supports learning about supervised fine-tuning and outcome-based shell evaluation. It does not establish general Bash competence or certify commands as safe to run on a real machine. This release contains the full 2,000-record training split. The project's 110 validation records and 220… See the full description on the dataset page: https://huggingface.co/datasets/nima1/bashcraft-data.texttext-generation1K<n<10K0 likes41 downloads6d agoHugging Face19zonay /bash-data-nl2sh bash-data: 140k English → Bash pairs Synthetic, deterministic (scripts/generate.py, seed varies). Each row: id, instruction, input, output, category, risk. train 140k / eval 2k, bash -n clean, ~99% unique inputs. categories: shell core, net (ssh/dns/NAT), tailscale, macdev (xcode), AI tooling (ollama/lms/opencode/llamacpp/mlx/hf/vllm/agents/shellai), fullstack (js/py/db/devops/native), lab-authorized destructive+pentest (risk: destructive). destructive/pentest rows are… See the full description on the dataset page: https://huggingface.co/datasets/zonay/bash-data-nl2sh.texttranslation100K<n<1M0 likes39 downloads7d agoHugging Face20bashkorttele /periodicals-izddom Bashkir Periodicals — Respublika Bashkortostan Publishing House 163,240,215 characters of Bashkir-language text (39,711 records) from publications of Respublika Bashkortostan Publishing House. A record is an article or a work where the publication marks its boundaries (a table of contents, a day of a tear-off calendar); elsewhere it is cut at a heading or at a page. Every record carries metadata sufficient to locate it in the original publication. 🌐 Languages of this card:… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/periodicals-izddom.tabulartext-generation10K<n<100K0 likes38 downloads10d agoHugging Face21BashkirNLPWorld /bashkir-russian-parallelgated Dataset Card for Bashkir-Russian Parallel Corpus Dataset Details Dataset Description Bashkir-Russian Parallel Corpus is a large-scale sentence-aligned parallel corpus for the Bashkir–Russian language pair, assembled from authentic human-created translations. It contains 3,040,085 unique parallel sentence pairs, where each Bashkir sentence is aligned with its Russian counterpart. The corpus combines data from three open parallel corpora: TIL-MT… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-russian-parallel.tabulartext-generation1M<n<10M0 likes34 downloads23d agoHugging Face22ISB369 /shellminator-bash-sft105ktext100K<n<1M0 likes32 downloads1mo agoHugging Face23Coding-With-Bashir /BashCoder 🌀 Claude-opus-4.7-TraceInversion-5000x v1.0 Release A High-Fidelity Reconstructed CoT Dataset Saturated with the 'Opus Deep Logic Style' via Trace Inversion 📊 5,000 Samples 🧬 Trace Inversion & Negentropy 🛠 SFT & DPO Ready 🔥 Claude 4.7-Max Distillation 🌐 English & Multilingual 💡 What is Trace Inversion? In Large Language Model (LLM) reasoning distillation, proprietary API models (such as GPT-4/5 and Claude)… See the full description on the dataset page: https://huggingface.co/datasets/Coding-With-Bashir/BashCoder.texttext-generation1K<n<10K1 likes25 downloads4mo agoHugging Face24skarnati20 /intercode-bash-qwen7b-activationstabular1K<n<10K0 likes25 downloads4d agoHugging Face25BashkirNLPWorld /bashkir-lexicongated Dataset Card for Bashkir Lexicon (Machine Fund) Dataset Details Dataset Description Bashkir Lexicon (Machine Fund) is a machine-readable lexical database of the Bashkir language, compiled from the open digital resource Machine Fund of the Bashkir Language (mfbl2.ru). It contains 34,008 unique lexical entries covering dialectal word forms with their part of speech, dialect, subdialect, literary norm, and Russian translation. The dataset preserves… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-lexicon.texttext-generation10K<n<100K0 likes24 downloads26d agoHugging Face26Basharat78 /Medical_v1atext100K<n<1M0 likes17 downloads2y agoHugging Face27dravidmathavan07 /bash-commands-dataset 🐧 Linux Command Automation Dataset A dataset of natural language prompts paired with their corresponding Bash command-line equivalents, designed to train or fine-tune models for automating Linux tasks via natural language. 📁 Dataset Structure The dataset is in JSON format, structured as a flat array of objects, where each object contains: { "prompt": "Natural language description of a task", "response": "Equivalent Bash command" } ✅ Example… See the full description on the dataset page: https://huggingface.co/datasets/dravidmathavan07/bash-commands-dataset.textn<1K0 likes17 downloads2mo agoHugging Face28ISB369 /shellminator-bash-cleantext10K<n<100K0 likes13 downloads2mo agoHugging Face29georgiyozhegov /bashorgtext100K<n<1M0 likes7 downloads1y agoHugging Face30basher-rardon /lego_grey_test1imagen<1K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.