rajivmehtapy/shell-script-specialist-dataset
Full-Spectrum Shell Script Specialist Dataset This dataset contains 1,000 curated, unique ChatML conversation records engineered to fine-tune a specialist language model for Production-Grade Shell Scripting (Bash 5+, POSIX /bin/sh, jq, awk, sed). It was used to train the rajivmehtapy/gemma-4-e4b-shell-specialist model using Unsloth. Dataset Splits Split File Records Description train train.jsonl 900 Core training set across all 4 production modules… See the full description on the dataset page: https://huggingface.co/datasets/rajivmehtapy/shell-script-specialist-dataset.
Full-Spectrum Shell Script Specialist Dataset
This dataset contains 1,000 curated, unique ChatML conversation records engineered to fine-tune a specialist language model for Production-Grade Shell Scripting (Bash 5+, POSIX `/bin/sh`, `jq`, `awk`, `sed`).
It was used to train the rajivmehtapy/gemma-4-e4b-shell-specialist model using Unsloth.
Dataset Splits
Modules Covered
- Production Bash 5+ Automation (35%): Archival rotations, health probes with exponential backoff, systemd service watchdogs, disk/memory threshold monitors.
- Minimal POSIX `/bin/sh` Portability (20%): Alpine Linux / Docker entrypoints, socket polling (
nc -z), POSIX parameter expansions, zero Bashisms. - Advanced CLI Stream Wrangling (25%): Complex
jqqueries with atomic file replacements,awkcolumnar parsing,sedstream updates, null-delimitedfind/xargspipelines. - Defensive Hardening & Bug Refactoring (20%): Auditing fragile bash snippets, eliminating unquoted variable risks, adding
set -euo pipefailandtraphandlers.
Data Schema (ChatML JSONL)
{
"messages": [
{
"role": "user",
"content": "Write a bash script to archive and remove .log files older than 14 days in /var/log with dry-run support."
},
{
"role": "assistant",
"content": "#!/usr/bin/env bash\nset -euo pipefail\n\nDRY_RUN=false\n[[ \"${1:-}\" == \"--dry-run\" ]] && DRY_RUN=true\n..."
}
]
}Usage with Hugging Face Datasets
from datasets import load_dataset
dataset = load_dataset("rajivmehtapy/shell-script-specialist-dataset")
print(dataset)
print(dataset["train"][0])Next Steps: DPO & GRPO Alignment
When you spin up your next machine for DPO and GRPO, you can immediately resume using the following one-liners:
1. Pulling the Policy Model on the New Machine
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="rajivmehtapy/gemma-4-e4b-shell-specialist",
max_seq_length=1024,
load_in_4bit=True,
)2. Pulling the Dataset on the New Machine
from datasets import load_dataset
dataset = load_dataset("rajivmehtapy/shell-script-specialist-dataset")3. Ready for DPO
- Use the SFT model as the reference policy.
- Provide
(prompt, chosen, rejected)triplets (wherechosencontains defensive standards likeset -euo pipefailandrejectedcontains common bash antipatterns). - Train with
trl.DPOTrainer.
4. Ready for GRPO
- Use the SFT model as the actor model.
- Set up automated rule-based reward functions (
shellcheckreturncode, exit code in Docker sandbox, security parameter validation). - Train with
trl.GRPOTrainer.
Everything is backed up and ready for your next phase!
