datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shell-cmd-instruct
Used to train models that interact directly with shells
Note: This dataset is out-dated in the llm world, probably easier to just setup a tool with a decent model that supports tooling.
Follow-up details of my process
MacOS terminal commands for now. This dataset is still in alpha stages and will be modified.
Contains 500 somewhat unique training examples so far.
GPT4 seems like a good candidate for generating more data, licensing would need to be addressed.
I fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/byroneverson/shell-cmd-instruct.postcutoff-facts-qa
postcutoff-facts-qa
Closed-book QA over ~300 post-knowledge-cutoff facts (world events from
Wikipedia current events + ECB reference rates / index closes, June 2024 →
August 2026), built to verifiably measure whether fine-tuning teaches a model
new facts — and whether it destroys the model's calibration while doing so.
Companion to the adapter
evs-cmd/qwen2.5-1.5b-verifiable-facts-v8.
Every file was produced by a governed cairn pipeline run
(synthesis → dedup/PII hygiene →… See the full description on the dataset page: https://huggingface.co/datasets/evs-cmd/postcutoff-facts-qa.cmdguard-dataset
cmdguard-dataset
Hand-labeled CLI commands for training a binary classifier: exploring (read-only) vs mutating (changes state).
Format
Chat-style JSONL — each line is:
{"messages": [{"role": "user", "content": "Classify: git status"}, {"role": "assistant", "content": "exploring"}]}
Splits
Split
Examples
Balance
train.jsonl
354
168 exploring / 168 mutating + 18 targeted
eval.jsonl
20
10 exploring / 10 mutating
Coverage
git, docker… See the full description on the dataset page: https://huggingface.co/datasets/qmxme/cmdguard-dataset.customer_support_cmd_gen_basicsys-cmd-promptlinux_cmd_alpaca
linux_cmd_alpaca
This repository contains a dataset in Alpaca format, consisting of natural language instructions, shell commands, and corresponding responses of linux terminal. This dataset is made from some existing datasets with more data and in alpaca format
Helpsteer-Pref-Edit-CMDtest-datasetHelpsteer-3-Pref-CMDMisc-Med-CMDA
