Team Ai
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jragsdale1 /ShIO-bash-26.1 ShIO-bash-26.1 Shell input-output (ShIO) Bash dataset produced by ShIOEnv, a Gymnasium-compatible Bash environment designed to collect execution-annotated command interactions in a Linux system. Dataset summary The dataset consists of command-line inputs paired with their execution artifacts, including observable outputs and a structured representation of environment state changes. Samples are produced by executing synthesized Bash inputs inside a… See the full description on the dataset page: https://huggingface.co/datasets/jragsdale1/ShIO-bash-26.1.imagetext-generation1M<n<10M1 likes606 downloads4mo agoHugging Face02Frost2o24 /bash-instruct-III-55k Bash Instruct III — 54,360 verified natural-language → Bash pairs Bash Instruct III is a synthetic instruction-tuning dataset that maps natural-language requests to correct Bash: single commands, short pipelines, and multi-line scripts. It is built for supervised fine-tuning of small and mid-size LLMs that must turn a plain request into shell code that actually runs. Every row is a three-turn chat conversation (system / user / assistant) with metadata for slicing (category… See the full description on the dataset page: https://huggingface.co/datasets/Frost2o24/bash-instruct-III-55k.texttext-generation10K<n<100K0 likes350 downloads1mo agoHugging Face03emirkaanozdemr /bash_command_data_6K 📦 Bash Command Dataset v1 A high-quality dataset of natural language instructions paired with their equivalent Bash commands, designed for training and fine-tuning large language models (LLMs) that translate English tasks into shell commands. This dataset is ideal for researchers, developers, and machine learning engineers interested in natural language to Bash command translation, command-line automation, and building intelligent terminal assistants. 📁 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/emirkaanozdemr/bash_command_data_6K.texttext-generation1K<n<10K8 likes215 downloads6mo agoHugging Face04failed09 /bashkir-wikipedia-monolingual Bashkir Wikipedia Monolingual Corpus Cleaned sentence-level Bashkir text from Wikipedia for pretraining, tokenizer training and linguistic research. Overview Sentence-level text extracted from the Bashkir Wikipedia dump (bawiki-20260801), cleaned and filtered with automated language identification. The cleaned configuration is the recommended default for language modelling, tokenization and linguistic research; precleaned is an earlier, lighter extraction kept… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-monolingual.texttext-generation1M<n<10M0 likes86 downloads9d agoHugging Face05Frost2o24 /bash-instruct-II-55k Bash Instruct II — 54,803 verified natural-language → Bash pairs ⚠️ Superseded by Bash Instruct III III is a corrected rebuild of this dataset with shell antipatterns removed at the generator level (99.8% ShellCheck-clean vs 98.5% here). New work should use III. The two share 92.5% of their (request, command) pairs, so they must never be concatenated. This card is kept for reproducibility and citation of published results. Bash Instruct II is a synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Frost2o24/bash-instruct-II-55k.texttext-generation10K<n<100K0 likes81 downloads1mo agoHugging Face06GunA-SD /bash_codeThis dataset is a collection of bash programs from various GitHub repositories and open source projects. The dataset might contain harmful code. texttext-generation100K<n<1M9 likes68 downloads2y agoHugging Face07Frost2o24 /bash-instruct-55k Bash Instruct I — 55,000 verified natural-language → Bash pairs ⚠️ Superseded by Bash Instruct III III doubles the utility vocabulary (89 → 182), adds grouped equivalent answers, real human phrasing from tldr-pages, validation by execution on real Linux, and a style pass that removes shell antipatterns. New work should use III. This version remains useful for one specific purpose: it overlaps III by only ~29%, so it is the one generation that can be deduplicated… See the full description on the dataset page: https://huggingface.co/datasets/Frost2o24/bash-instruct-55k.texttext-generation10K<n<100K1 likes61 downloads1mo agoHugging Face08failed09 /bashkir-periodicals-izddom Bashkir Periodicals Cleaned (Izddom) High-quality cleaned and structured Bashkir periodicals (Respublika Bashkortostan Publishing House) for LLM pretraining and fine-tuning. Overview A curated, deduplicated and filtered edition of modern Bashkir periodicals derived from the original bashkorttele/periodicals-izddom dataset published by the Foundation for the Preservation and Development of the Bashkir Language. The dataset includes material from 13 leading… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-periodicals-izddom.tabulartext-generation1M<n<10M0 likes60 downloads10d agoHugging Face09failed09 /bashkir-books-kitap Bashkir Books & Calendars Cleaned (Kitap) Curated, cleaned and structured Bashkir literature and annual calendars (Kitap Publishing House) for LLM pretraining and fine-tuning. Overview A curated, deduplicated and filtered edition of modern Bashkir book publications and annual cultural calendars derived from the original bashkorttele/books-kitap dataset published by the Foundation for the Preservation and Development of the Bashkir Language. The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-books-kitap.tabulartext-generation10K<n<100K0 likes60 downloads10d agoHugging Face10lisayan /rlvr-bash-terminal-bench rlvr-bash-terminal-bench RLVR (Reinforcement Learning with Verifiable Rewards) dataset for bash scripting, generated from Terminal-Bench tasks. Stats Metric Value Total samples 1,120 Unique tasks 88 Avg samples/task 12.7 Average reward 0.249 Perfect solutions (reward=1.0) 10.4% Partial solutions (0<reward<1) 28.8% Zero reward 60.8% Tasks fully solved 13.6% Format { "task_id": "string", "prompt": "string", "completion":… See the full description on the dataset page: https://huggingface.co/datasets/lisayan/rlvr-bash-terminal-bench.tabulartext-generation1K<n<10K0 likes57 downloads9mo agoHugging Face11bashkorttele /books-kitap Bashkir Books — "Kitap" Publishing House 9,019,878 characters of Bashkir-language text (6,081 records) from publications of "Kitap" Publishing House. A record is an article or a work where the publication marks its boundaries (a table of contents, a day of a tear-off calendar); elsewhere it is cut at a heading or at a page. Every record carries metadata sufficient to locate it in the original publication. 🌐 Languages of this card: English · Башҡортса · Русский Part of the… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/books-kitap.tabulartext-generation1K<n<10K0 likes51 downloads10d agoHugging Face12adeelahmad /bash-agent-grpo-pairs Bash Agent GRPO Pairs Single-turn (intent → shell command) pairs for training a small, local, Claude-Code-style bash agent with SFT or GRPO. Each record pairs a natural-language objective with exactly one verifiable bash command, framed as a single-tool bash(command, description) call. The dataset is designed to be rewardable: the ground-truth command is a deterministic target, so a shell-equivalence reward (canonical program + flag set + argument comparison) can score… See the full description on the dataset page: https://huggingface.co/datasets/adeelahmad/bash-agent-grpo-pairs.texttext-generation10K<n<100K0 likes41 downloads3mo agoHugging Face13nima1 /bashcraft-data Bashcraft v1 training data Bashcraft is a small, original, agent-assisted dataset for mapping an English request and explicit environment context to a Bash command or a clarification question. It supports learning about supervised fine-tuning and outcome-based shell evaluation. It does not establish general Bash competence or certify commands as safe to run on a real machine. This release contains the full 2,000-record training split. The project's 110 validation records and 220… See the full description on the dataset page: https://huggingface.co/datasets/nima1/bashcraft-data.texttext-generation1K<n<10K0 likes41 downloads6d agoHugging Face14bashkorttele /periodicals-izddom Bashkir Periodicals — Respublika Bashkortostan Publishing House 163,240,215 characters of Bashkir-language text (39,711 records) from publications of Respublika Bashkortostan Publishing House. A record is an article or a work where the publication marks its boundaries (a table of contents, a day of a tear-off calendar); elsewhere it is cut at a heading or at a page. Every record carries metadata sufficient to locate it in the original publication. 🌐 Languages of this card:… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/periodicals-izddom.tabulartext-generation10K<n<100K0 likes38 downloads10d agoHugging Face15BashkirNLPWorld /bashkir-web-corpusgated Dataset Card for Bashkir Web Corpus Dataset Details Dataset Description The Bashkir Web Corpus is a collection of 71,567 documents and approximately 46.9 million tokens in the Bashkir language (a Turkic language spoken in Bashkortostan, Russia). The corpus was compiled from 16 Bashkir‑language online sources, including news websites, literary magazines, social media, books, and Wikipedia. It is designed for language modeling, text classification… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-web-corpus.texttext-generation10K<n<100K0 likes36 downloads1mo agoHugging Face16BashkirNLPWorld /bashkir-russian-parallelgated Dataset Card for Bashkir-Russian Parallel Corpus Dataset Details Dataset Description Bashkir-Russian Parallel Corpus is a large-scale sentence-aligned parallel corpus for the Bashkir–Russian language pair, assembled from authentic human-created translations. It contains 3,040,085 unique parallel sentence pairs, where each Bashkir sentence is aligned with its Russian counterpart. The corpus combines data from three open parallel corpora: TIL-MT… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-russian-parallel.tabulartext-generation1M<n<10M0 likes34 downloads23d agoHugging Face17trtd56 /bash-commands-dataset-rev-ja trtd56/bash-commands-dataset-rev-ja bash_commands_ja.csv から作成した、Bash コマンドを入力して日本語説明を出力するためのデータセットです。 Columns prompt_ja: 日本語説明 prompt_en: 英語説明 response: Bash コマンド prompt: 学習用入力。response と同じ completion: 学習用出力。prompt_ja と同じ task: タスク識別子 language: 出力言語 Splits train: 756 test: 84 Usage from datasets import load_dataset ds = load_dataset("trtd56/bash-commands-dataset-rev-ja") print(ds["train"][0]["prompt"]) print(ds["train"][0]["completion"]) texttext-generationn<1K0 likes28 downloads7mo agoHugging Face18BashkirNLPWorld /bashkir-wiki-corpusgated Dataset Card for Bashkir Wikipedia Corpus Dataset Details Dataset Description The Bashkir Wikipedia Corpus is a collection of 43,926 articles from Bashkir Wikipedia and Wikibooks, totaling approximately 10,617,785 XLM-R tokens and 8,862,807 words. The data has been extracted from official Wikimedia dumps and processed to provide clean, well‑structured text suitable for NLP tasks. The corpus includes article titles, full content, categories, source… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-wiki-corpus.texttext-generation10K<n<100K0 likes28 downloads5d agoHugging Face19Coding-With-Bashir /BashCoder 🌀 Claude-opus-4.7-TraceInversion-5000x v1.0 Release A High-Fidelity Reconstructed CoT Dataset Saturated with the 'Opus Deep Logic Style' via Trace Inversion 📊 5,000 Samples 🧬 Trace Inversion & Negentropy 🛠 SFT & DPO Ready 🔥 Claude 4.7-Max Distillation 🌐 English & Multilingual 💡 What is Trace Inversion? In Large Language Model (LLM) reasoning distillation, proprietary API models (such as GPT-4/5 and Claude)… See the full description on the dataset page: https://huggingface.co/datasets/Coding-With-Bashir/BashCoder.texttext-generation1K<n<10K1 likes25 downloads4mo agoHugging Face20BashkirNLPWorld /bashkir-lexicongated Dataset Card for Bashkir Lexicon (Machine Fund) Dataset Details Dataset Description Bashkir Lexicon (Machine Fund) is a machine-readable lexical database of the Bashkir language, compiled from the open digital resource Machine Fund of the Bashkir Language (mfbl2.ru). It contains 34,008 unique lexical entries covering dialectal word forms with their part of speech, dialect, subdialect, literary norm, and Russian translation. The dataset preserves… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-lexicon.texttext-generation10K<n<100K0 likes24 downloads26d agoHugging Face21datasetter458 /bash-reference-manual-general-QAs Dataset generated from bash reference manual. book information like date and bash version are available within the very first rows of the dataset this dataset is pretty small in general, but covering almost all of the definition and technical terms, commands and flags in the book columns : "Question", "Answer" texttext-generation1K<n<10K0 likes18 downloads5mo agoHugging Face22Taranosaurus /bash-org-archive.comDetails This is an unofficial dataset of an archive mirror of Bash.org: https://bash-org-archive.com/ Bash.org was a website launched in 1999 dedicated to archiving funny quotes from IRC other chat platforms over the years. It offers a look into jokes, memes, and often inappropriate content that was quite commonplace at the time. This dataset has been cleaned with a custom parser, aiming to preserve the original format of the content. The parquet file contains the following columns: qid… See the full description on the dataset page: https://huggingface.co/datasets/Taranosaurus/bash-org-archive.com.texttext-generation10K<n<100K0 likes16 downloads2y agoHugging Face23datasetter458 /diverse-bash-dataset Diverse bash terminal dataset A dataset containing Question&Code pairs for most of the standard bash commands. columns : "Question", "Code answer". Commands covered : echo, cat, cd, rm, mkdir, top, free, du, df, ps, head, tail, grep, cp, cut, sort, touch, ls, groupadd, ifconfig, ip, ln, ping, scp, ssh, sudo, systemctl, tar, useradd, userdel, usermod, wc specs : high quality dataset with precise coding examples for each command, covering (almost) all the flags of every command… See the full description on the dataset page: https://huggingface.co/datasets/datasetter458/diverse-bash-dataset.texttext-generation1K<n<10K0 likes16 downloads5mo agoHugging Face24BashkirNLPWorld /bashkir-news-clustergated Dataset Card for Bashkir News Cluster Dataset Dataset Details Dataset Description This dataset contains 24,428 Bashkir-language news and analytical articles collected from various online sources. It is intended for clustering, representation learning, and unsupervised NLP tasks. Each text is accompanied by metadata such as title, source, date, and original category. The corpus is part of the BashkirNLP project, aiming to support low-resource… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-cluster.textfeature-extraction10K<n<100K0 likes10 downloads1mo agoHugging Face25datasetter458 /diverse-bash-dataset-extended Bash diverse dataset extended This dataset is a continious work from datasetter458/bash-diverse-dataset Commands covered: bash, apt, crontab, date, diff, docker, env, export, fdisk, file, fsck, gcc, git, history, join, kubectl, lsof, man, makefs, mount, nice, node, nohup, npm, passwd, paste, patch, pkill, python, renice, split, ss, time, umount, uname, uniq, vi, whereis, which Note : For each command, a big amount of Q&As, and each set of them covers pretty much all the flags for… See the full description on the dataset page: https://huggingface.co/datasets/datasetter458/diverse-bash-dataset-extended.texttext-generation1K<n<10K0 likes6 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.