datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shell-safety-v1.1
Shell Safety v1.1
A synthetic dataset of safe/unsafe shell commands with respective running session contexts.
v1 has only a safe column with boolean values.
This version has a label column with 3 values: allow, ask or deny; making it closer to permissions handler in coding harness.
alchemist-shell.ai-hackathon-2025This project is described in detail at this website:
https://alchemist-shellai-hackathon-2025.readthedocs.io/en/latest/
The codes and relevant materials are available here:
https://github.com/Sukantabasu/alchemist-shell.ai-hackathon-2025
The trained models (in pkl format) are stored in this HF repository.
shellminator-bash-sft100kshell-attack-evolution-dataset
Shell Honeypot Attack Request–Response Dataset
A standardized, MITRE ATT&CK–annotated dataset of post-login shell
attacks captured by Cowrie SSH/Telnet
honeypots across two collection periods — 2021–2022 and 2024. It pairs
attacker shell commands with real captured system responses, enabling both
longitudinal threat analysis and the training/evaluation of AI-driven honeypots.
This is the open-source release accompanying the paper “Unveiling Evolving
Threats: A Data Analysis… See the full description on the dataset page: https://huggingface.co/datasets/zyw-286/shell-attack-evolution-dataset.Shell-Code-Large
Shell-Code-Large
Shell-Code-Large is a large-scale corpus of Shell scripting source code comprising approximately 640,000 code samples stored in JSON Lines (.jsonl) format. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, DevOps automation, cloud infrastructure engineering, system administration, and software engineering automation.
By providing a high-volume, language-specific corpus focused exclusively on Shell scripting… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Shell-Code-Large.Shellcode_Exploit_Dataset
Shellcode Exploit Dataset for Red Team GPT Training
Dataset Overview
The Shellcode Exploit Dataset is a comprehensive collection of 700 unique shellcode exploits, spanning 2021–2025, designed for training machine learning models, particularly for red team and cybersecurity research. The dataset includes a diverse set of vulnerabilities, platforms, architectures, and payload goals, sourced from Exploit-DB, GitHub, CTF challenges, and CVE databases.
It is structured in JSON… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Shellcode_Exploit_Dataset.shell-safety-transcriptsConverted from tomngdev/shell-safety into conversations transcripts.
Structure is for my own training with static system prompt and changing <SessionContext> block
shellkeeper-data
shellkeeper-data
This dataset contains about 39k synthetic, context-dependent safety labels for shell commands proposed
by AI agents. Each example has:
the user's request(s)
the agent's session so far (previous commands and their outputs)
a proposed command
a safe or unsafe label
The data was used to train hizkifw/shellkeeper-0.6b.
Code is at https://github.com/hizkifw/shellkeeper.
Label definition. A command is "unsafe" if a careful human operator would want to be asked… See the full description on the dataset page: https://huggingface.co/datasets/hizkifw/shellkeeper-data.shell-cmd-instruct
Used to train models that interact directly with shells
Note: This dataset is out-dated in the llm world, probably easier to just setup a tool with a decent model that supports tooling.
Follow-up details of my process
MacOS terminal commands for now. This dataset is still in alpha stages and will be modified.
Contains 500 somewhat unique training examples so far.
GPT4 seems like a good candidate for generating more data, licensing would need to be addressed.
I fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/byroneverson/shell-cmd-instruct.Dans-Toolmaxx-ShellCommandsshellminator-bash-combinedshellminator-bash-datasetshell-safety
Shell Safety
A synthetic dataset of safe/unsafe shell commands with respective running session contexts.
shell-script-specialist-dataset
Full-Spectrum Shell Script Specialist Dataset
This dataset contains 1,000 curated, unique ChatML conversation records engineered to fine-tune a specialist language model for Production-Grade Shell Scripting (Bash 5+, POSIX /bin/sh, jq, awk, sed).
It was used to train the rajivmehtapy/gemma-4-e4b-shell-specialist model using Unsloth.
Dataset Splits
Split
File
Records
Description
train
train.jsonl
900
Core training set across all 4 production modules… See the full description on the dataset page: https://huggingface.co/datasets/rajivmehtapy/shell-script-specialist-dataset.shellminator-bash-sft105kcancelli-shell-actions
Shell/CLI agent actions: approve/deny labels + rubric teacher grades
2965 proposed agent tool calls that run a shell command or a command-line program, each with
an approve/deny label and a teacher model's answers to a 39-question safety rubric.
Labels: 1181 approve, 1784 deny.
Fields
field
type
meaning
id
string
opaque stable id
state
string
### PROPOSED ACTION (tool + args), ### USER REQUEST, and any agent history
label
string
approve or deny… See the full description on the dataset page: https://huggingface.co/datasets/mastinon/cancelli-shell-actions.shell-tasksshellminator-contrastiveshellsmith-commands
shellsmith-commands
Curated (natural-language instruction → shell command) pairs for macOS/Linux,
used to fine-tune Qwen2.5-Coder-1.5B-Shellsmith.
Format
JSONL in chat format (mlx-lm / OpenAI style):
{"messages": [
{"role": "system", "content": "You are a shell command generator ..."},
{"role": "user", "content": "list files sorted by size, largest first"},
{"role": "assistant", "content": "ls -lS"}
]}
Splits
File
Rows
train.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/ajayk007/shellsmith-commands.dataset_shellqwen3-235b-shell-tasksshell-safety-conversationsshellminator-bash-cleanV3-shell-formatShellLifedominican-banking-regulations-corpus
Dominican Banking Regulatory Corpus (ANIMUS)
142 regulatory documents from the Superintendencia de Bancos de la República Dominicana (SB), processed and cleaned as part of the ANIMUS project — an autonomous Rust-based knowledge system that builds and audits its own knowledge graph.
Contents
Each entry contains:
label: document identifier (circular/regulation number and title)
content: extracted text (UTF-8, cleaned of encoding artifacts)
source: original file… See the full description on the dataset page: https://huggingface.co/datasets/shellhack/dominican-banking-regulations-corpus.test20241217webDesignershellm-V3-simple-unix
