Team Ai
Datasetpublic

Frost2o24/bash-instruct-II-55k

Bash Instruct II — 54,803 verified natural-language → Bash pairs ⚠️ Superseded by Bash Instruct III III is a corrected rebuild of this dataset with shell antipatterns removed at the generator level (99.8% ShellCheck-clean vs 98.5% here). New work should use III. The two share 92.5% of their (request, command) pairs, so they must never be concatenated. This card is kept for reproducibility and citation of published results. Bash Instruct II is a synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Frost2o24/bash-instruct-II-55k.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes82downloads
Dataset Card

Bash Instruct II — 54,803 verified natural-language → Bash pairs

### ⚠️ Superseded by **Bash Instruct III** III is a corrected rebuild of this dataset with shell antipatterns removed at the generator level (99.8% ShellCheck-clean vs 98.5% here). New work should use III. The two share 92.5% of their (request, command) pairs, so they must never be concatenated. This card is kept for reproducibility and citation of published results.

Bash Instruct II is a synthetic instruction-tuning dataset that maps natural-language requests to correct Bash: single commands, short pipelines, and multi-line scripts. It is built for supervised fine-tuning of small and mid-size LLMs that must turn a plain request into shell code that actually runs.

Every row is a three-turn chat conversation (system / user / assistant) with metadata for slicing (category, utility) and for grouping equivalent answers to the same request (variant_group).

This is the second-generation build. Over version I it introduced the chat format, grouped alternative answers, real human phrasing seeded from tldr-pages, a doubled utility vocabulary (89 → 182), and validation by real execution on Linux.

🤗 Dataset viewerhttps://huggingface.co/datasets/Frost2o24/bash-instruct-II-55k
🐙 Source, generator & validatorhttps://github.com/ya5h-P/bash-instruct-55k
📦 Other generations**III (recommended)** · I
⚖️ LicenseMIT
python
from datasets import load_dataset

ds = load_dataset("Frost2o24/bash-instruct-II-55k", split="train")

Version comparison

[I](https://huggingface.co/datasets/Frost2o24/bash-instruct-55k)**II** (this)[III](https://huggingface.co/datasets/Frost2o24/bash-instruct-III-55k)
Rows55,00054,80354,360
Distinct primary utilities89182182
variant_group (grouped equivalent answers)✗✓✓
Real phrasing seeded from tldr-pages✗✓✓
Validated by execution on real Linux✗✓✓
ShellCheck-clean rate (warning level)97.9%98.5%99.8%
Antipattern style pass (SC2002, SC2010)✗✗✓

Choose III unless you have a specific reason to reproduce II — for example, matching a previously published training run.


At a glance

All figures below are recomputed directly from the shipped `bash_dataset.jsonl`, not carried over from an earlier run.

MetricValue
Rows54,803
Distinct requests (variant_groups)45,142
Rows per request1.21
single / pipeline / script40.1% / 34.8% / 25.1%
Distinct primary utilities182
bash -n syntax-valid100% (0 / 54,803 failures)
ShellCheck-clean at warning level98.5% of a 6,000-row random sample
Exact-duplicate (request, command) pairs0
Unique request text (per distinct request)96.8%
Requests with ≥2 equivalent answers16.6%
Max share of any single utility6.0% of rows · 4.9% of requests
Most common 4-word request opening2.1% of rows
Mean command length87 chars · 13.8 tokens

Quickstart

Load

python
from datasets import load_dataset

ds = load_dataset("Frost2o24/bash-instruct-II-55k", split="train")
print(ds)
# Dataset({features: ['messages', 'category', 'utility', 'variant_group'], num_rows: 54803})

print(ds[0]["messages"])

Fine-tune (TRL SFTTrainer)

The messages column is already in the OpenAI-style chat format that TRL, Axolotl, Unsloth, and LLaMA-Factory consume directly — no preprocessing or column mapping needed.

python
from datasets import load_dataset
from trl import SFTTrainer, SFTConfig

ds = load_dataset("Frost2o24/bash-instruct-II-55k", split="train")

trainer = SFTTrainer(
    model="Qwen/Qwen2.5-Coder-1.5B-Instruct",
    train_dataset=ds,                       # `messages` is auto-detected
    args=SFTConfig(
        output_dir="bash-sft",
        max_length=1024,
        num_train_epochs=2,
        per_device_train_batch_size=8,
    ),
)
trainer.train()

Recommended: one answer per request

Rows sharing a variant_group answer the same request with equivalent commands. For a standard single-answer SFT run, collapse them first so no request is over-weighted:

python
seen = set()
canonical = ds.filter(
    lambda r: not (r["variant_group"] in seen or seen.add(r["variant_group"]))
)
print(len(canonical), "distinct requests")   # 45142

Held-out split without leakage

Split on variant_group, never on rows — otherwise variants of the same request land on both sides of the split:

python
groups = sorted({r["variant_group"] for r in ds})
import random; random.Random(0).shuffle(groups)
val = set(groups[:2000])

train_ds = ds.filter(lambda r: r["variant_group"] not in val)
val_ds   = ds.filter(lambda r: r["variant_group"] in val)

Inspect without Python

bash
head -n 1 bash_dataset.jsonl | jq .
jq -r '.utility'  bash_dataset.jsonl | sort | uniq -c | sort -rn | head -20
jq -r '.category' bash_dataset.jsonl | sort | uniq -c

Dataset structure

One JSON object per line in bash_dataset.jsonl:

json
{
  "messages": [
    {"role": "system",    "content": "You are a Bash expert."},
    {"role": "user",      "content": "Show the last 20 lines of error.log."},
    {"role": "assistant", "content": "tail -n 20 error.log"}
  ],
  "category": "single",
  "utility": "tail",
  "variant_group": "3f9a1c0b7e2d4a86"
}
FieldTypeDescription
messageslist[{role, content}]The training conversation, always exactly system → user → assistant.
categorystringsingle · pipeline · script. See below.
utilitystringPrimary command of the solution (grep, awk, find, for, …). Use it to slice, re-balance, or curriculum-order training.
variant_groupstring16-hex id shared by rows answering the same request with different-but-equivalent commands (e.g. wc -l f vs `cat f \wc -l`). Requests with a single answer get their own unique id.

Verified schema integrity: all 54,803 lines parse as JSON, all carry the same four keys, all have exactly the system/user/assistant role order, and no message content is empty.

Categories

`category`RowsShareWhat it contains
single21,96540.1%One command, no pipe or chaining. chmod 644 build.bak
pipeline19,08834.8%Pipes, &&/`\\, $(...), xargs. ps aux \awk '$3>50 {print $2}'`
script13,75025.1%Multi-line Bash (JSON-escaped \n), mean 9.8 lines.

What's in the scripts

FeatureShare of `script` rows
set -euo pipefail header100%
for loop49.0%
if [ … ] test13.7%
while loop11.6%
getopts argument parsing9.5%
trap cleanup0.1%
Here-docs (<<)0.1%

Scripts are short, safe-by-default operational snippets — not large programs. trap and here-doc usage is deliberately rare, and there are no function definitions.

System & info utility coverage

Small models routinely fumble the observability and process-management tools. These carry enforced minimum counts with correct, idiomatic flag usage:

UtilityRowsUtilityRowsUtilityRows
journalctl3,295du1,496lsof447
pstree405nice289vmstat268
ss263ps238renice213
netstat201iostat179w159
dmesg146free113systemctl111
df56uptime35hostname30

uptime and hostname are lower because their genuine idiomatic command space is small; they are covered with real flag variants rather than padded with synonymous phrasings.

System prompts

Five equivalent Bash-assistant system prompts are rotated near-uniformly (~10.7k–11.2k rows each) so the model does not overfit one string.


Quality assurance

Four layers, in increasing strength. Layers 1–2 are reproducible in minutes on any machine; layer 3 needs a disposable Linux box.

1. Syntax — bash -n on every command

0 failures out of 54,803 (100% pass), verified with GNU bash 5.2.37 on Debian.

2. Static analysis — ShellCheck at warning level

On a 6,000-row random sample (ShellCheck 0.10.0): 98.45% of commands are completely clean. 93 rows produced a finding:

CodeFindingsWhat it is
SC201080`ls \grep` instead of a glob — the dominant style nit
SC21647cd without `\\exit`
SC10833Literal {/} — a false positive on find -exec … {} \;
SC20622Unquoted grep pattern
SC20462Unquoted command substitution
SC21821printf with no format operand

This is the specific weakness that [Bash Instruct III](https://huggingface.co/datasets/Frost2o24/bash-instruct-III-55k) fixes. III rewrites the offending recipes to use file operands, redirects, and shell globs, cutting findings from 93 to 10 (99.8% clean). If you care about the model learning idiomatic style rather than merely valid syntax, prefer III.

3. Execution on real Linux

43,257 commands were executed in a disposable Debian 13 sandbox. After the cleaning passes described under Provenance, essentially all gradable commands run cleanly.

"Gradable" excludes failures that are environmental rather than command defects — a blank sandbox legitimately lacks the referenced files, users, services, or systemd (e.g. ps -C postgres when postgres is not running, chown ec2-user: when that user does not exist, rmdir on a populated fixture dir). Those commands are correct; only the sandbox lacks the target, so they are kept and reported separately.

4. Argument fidelity

Every pair is gated so the command actually references the concrete nouns the request names — the same filename, extension, user, group, service, process, and port. A wrong filename still parses and still lints clean, so neither bash -n nor ShellCheck can catch this class of error; argcheck.py exists specifically to close that gap. The shipped file has 0 mismatches.

Reproduce the audit yourself

bash
# static: bash -n on every command, plus ShellCheck if installed,
# broken down by category and utility
python validate.py --data bash_dataset.jsonl

# execute the SAFE self-contained subset in a disposable fixture dir
python validate.py --data bash_dataset.jsonl --execute --n 3000

# whole-set execution -- THROWAWAY Linux box only (WSL, container, VM)
python validate.py --data bash_dataset.jsonl --execute --permissive --n 60000 --timeout 5

# argument-fidelity audit over the shipped file
python argcheck.py --data bash_dataset.jsonl

Validation is decoupled from generation — nothing in generate.py calls validate.py, so the checks are an independent audit rather than a self-certification.

Two execution modes:

  • —strict (default) — only self-contained, read-only coreutils on relative paths.
  • —`--permissive` — for a disposable Linux box. Runs ~43k rows including scripts and system-info / filesystem-mutating tools, while hard-blocking genuinely dangerous commands (dd, mkfs, shutdown, reboot, mount, fork bombs, kill/pkill, writes to /etc·/var·$HOME, block devices) and skipping anything that would hang (tail -f, watch, while true, sleep, installers, network calls, yes). Every command runs in a temp sandbox with a timeout, stdin from /dev/null, sudo stripped, and stdout discarded so infinite-output commands cannot exhaust memory.

The report separates ran-clean, environmental, and genuine errors (bad flag/option/logic). --dump-errors FILE writes every genuine error for inspection; --yes / --no control the delete prompt non-interactively.


Provenance

How it is built

Generated by a recipe engine (generate.py, included and reproducible).

  • —Recipe engine — ~150 hand-written recipe families emit (request, command) pairs by combining varied phrasings (imperative / question / casual) with realistic parameter pools: plausible filenames, directories, extensions, ports, services, users, patterns. Every parameter is drawn once per example and reused in both the request and the command, so the two can never disagree.
  • —Real phrasing (seeds) — a share of requests is seeded from tldr-pages (CC-BY-4.0): the human-written description becomes the request and its command the solution, with {{placeholders}} filled fresh from the pools. Heavily quality-filtered — subcommand multiplexers and non-fillable placeholders dropped — down to ~600 clean coreutils/text/net templates, cached in seeds/_seedcache.json so generation runs offline.
  • —Answer diversity — a share of requests also emits 1–2 genuinely-equivalent alternates sharing a variant_group: flag reorderings, long/short options, wc -l f ↔ cat f | wc -l, sort -u ↔ sort | uniq. Each alternate is re-checked before it is kept.
  • —Guards — SHA-1 dedup of (request, command); a degeneracy guard rejecting no-op transforms (sed 's/x/x/', mv f f); category quotas (40/35/25); a 4% per-utility request cap and 6% row cap; minimum floors for the system/info utilities; the argument-fidelity gate; and a bash -n gate on every command before it is written.

Data cleaning history

Whole-set execution and a stricter validator found real defects, each fixed at the source rather than patched in the output. Reported here because the failure modes are instructive for anyone building a similar corpus:

Defect foundResolutionRows
43 malformed commands that fail on any system — bad operand counts from tldr placeholder fills (split 10 10 file, date --rfc-3339 20, uname java, head 20 -3 file, `echo "host" \git credential fill`)Removed55,000 → 54,957
Long-form flag substitution leaked across pipeline stages — grep's -n → --line-number landed on a downstream tail -n, producing invalid tail --line-numberGenerator restricted to the owning command's segment. (tail --lines / head --lines are valid on GNU and were kept.) Also removed 3 kill --CONT <name>−149
find -size N emitted for "larger than N" requests — bare -size N means exactly NRepaired to -size +N; template now sets +/- from the request direction113 repaired
journalctl <service> without -u — bare units, missed by the old exit-status checkValidator hardened with set -o pipefail + a per-stage stderr scan for invalid/unrecognized option; rows removed→ 54,803

The generation before this one (version I) had a subtler problem: ~12% of pairs referenced the wrong concrete noun — right command shape, wrong filename. Neither bash -n nor ShellCheck can detect that, because a wrong filename still parses and still lints. The argcheck.py gate was introduced for II and there have been 0 mismatches since.

The generation after this one (III) fixes the SC2002 / SC2010 antipatterns that this version's ShellCheck table still shows.

Regenerating from source

bash
python generate.py plan            # quota table + registered recipes
python generate.py run             # generate to ./bash_dataset.jsonl (resumable)
python generate.py status          # where a run stands
python generate.py validate-all    # re-run bash -n over the whole file
python generate.py seeds           # build/inspect the real-phrasing seed pool

bash_dataset.jsonl is ground truth and a derived bash_dataset.progress.json snapshot is rewritten after every batch, so an interrupted run continues where it stopped.

FlagDefaultEffect
--target55000Total rows (rescales quotas, caps, floors).
--seed-frac0.40Target share of seedable requests drawn from real tldr seeds.
--seed-reuse8Max reuse of one tldr template (higher = more real phrasing, less request variety).
--variant-frac0.30Share of requests that attempt alternate answers.
--row-cap-pct0.06Max total-row share per utility (trades off answer diversity).
--no-netoffNever download seeds; use the local cache only.

Seed and answer-diversity fractions are soft — they land under target when the clean seed pool or the caps bind. The hard invariants always hold: category split, per-utility caps, coverage floors, argument fidelity, bash -n, and dedup.


Files

FileDescription
bash_dataset.jsonlThe dataset — 54,803 rows, ~22 MB (no Git LFS needed).
generate.pyThe generator. Reproducible and resumable.
validate.pyStandalone validator: static checks plus strict / --permissive execution.
argcheck.pyArgument-fidelity checker — generation gate and standalone auditor.
seeds/_seedcache.jsonCleaned real-phrasing tldr templates, so generation runs offline.

generate.py, validate.py, and argcheck.py are pure Python 3 with no third-party dependencies. validate.py optionally calls shellcheck if it is on PATH, and generate.py seeds optionally fetches tldr-pages unless --no-net is passed. Loading the dataset requires only datasets (or nothing at all — it is plain JSONL).


Combining versions

Overlap of (request, command) pairs, measured across the shipped files:

PairShared pairsShare of the smaller set
II ∩ III50,71592.5% — never concatenate
I ∩ II18,20033.2%
I ∩ III15,89629.2%

II and III are two builds of the same corpus; mixing them near-doubles the weight of the shared majority. Version I is genuinely more distinct and can be deduplicated in for extra volume, though it is measurably lower quality (97.9% ShellCheck-clean, 89 utilities).


Intended uses & limitations

Intended uses

Teaching a small or mid-size model to map natural-language requests to correct single commands, short pipelines, and small operational scripts — including the observability and process-management utilities that such models routinely get wrong. Also usable as a held-out evaluation set for shell-command generation, and as a seed corpus for rejection-sampling or preference-data pipelines.

Limitations

Please read these before reporting results.

  • —Superseded. III fixes this version's style antipatterns. Use II only to reproduce existing results.
  • —Synthetic. Variety comes from recombining templates and token pools plus a quality-filtered slice of real tldr phrasings. Request shapes originate from a finite family set. Phrasing is diversified — 96.8% of distinct requests have unique text and no 4-word opening exceeds 2.1% of rows — but this is not human-authored data and should not be presented as such.
  • —Valid ≠ semantically correct. bash -n, ShellCheck, sandbox execution, and the argument gate together prove that commands parse, lint, run, and use the nouns the request named. None of them prove the command's logic satisfies the intent. Some pairs are plausible-but-approximate — an awk column index, for instance, assumes a particular log layout. There is no human correctness audit.
  • —Single-turn, single canonical answer. No explanations, no reasoning traces, no negative examples, no multi-turn dialogue, no error recovery. The only notion of alternatives is the equivalent commands grouped by variant_group.
  • —GNU/Linux-centric. Commands assume GNU coreutils and a systemd-based distribution. BSD/macOS flag differences (sed -i, stat, ps) are not covered, and neither are zsh/fish.
  • —Short scripts only. Mean 9.8 lines, no function definitions, minimal trap and here-doc coverage. Not a source for large structured shell programs.
  • —Not safety-tuned. The corpus contains destructive-capable commands (rm, chmod, chown, kill, dd) as legitimate answers to requests that ask for them. A model trained on it will emit such commands readily. Add refusal/confirmation behaviour separately if your deployment needs it, and never auto-execute model output.
  • —Potential eval contamination. Derived in part from tldr-pages, which is public and widely scraped. If you evaluate on a tldr-derived benchmark, check for leakage first.

Citation

If you use this dataset, please cite it:

bibtex
@misc{pimplapure2026bashinstruct2,
  title        = {Bash Instruct II: A Validated Synthetic Instruction-Tuning
                  Dataset for Natural Language to Bash Generation},
  author       = {Pimplapure, Yash},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/datasets/Frost2o24/bash-instruct-II-55k}},
  note         = {54,803 chat-formatted examples; 100\% \texttt{bash -n} valid,
                  98.5\% ShellCheck-clean. Superseded by Bash Instruct III}
}

Please also credit tldr-pages (CC-BY-4.0), the source of the real-phrasing seed templates.

License

MIT. The tldr-derived seed phrasings are used under CC-BY-4.0.


Supporting this work

This dataset — generation, whole-set Linux execution, and the fine-tuning runs used to check that it actually teaches the task — was built and validated on modest consumer hardware, which is the main thing limiting how far the next version can go.

If your team has an NVIDIA DGX Spark or an AMD Ryzen AI Max+ ("Strix Halo") AI dev kit to spare, it would go directly into larger validated corpora, real fine-tuned baselines, and published eval numbers for this dataset family. Reach out via the dataset discussions tab. No obligation either way — the data stays MIT and freely available regardless.