datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mbgfnet-closed-shell-4d5d-gw-dataset
Closed-shell 4d/5d transition-metal complexes: PBE0 + G0W0 quasiparticle dataset
2,210 closed-shell mononuclear 4d and 5d transition-metal complexes (17 metals: Y, Zr, Nb, Mo, Ru,
Rh, Pd, Ag, Cd from the 4d row; Hf, Ta, W, Re, Os, Ir, Pt, Au, Hg from the 5d row), each with a DFT
(PBE0/cc-pVDZ, relativistic small-core pseudopotential on the metal) and a one-shot G0W0@PBE0
quasiparticle-energy calculation. Built to extend MBGF-Net
(Venturella, Li, Hillenbrand, Zhu… See the full description on the dataset page: https://huggingface.co/datasets/primateria/mbgfnet-closed-shell-4d5d-gw-dataset.MET-Bench-Shell
MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning Models
Vanya Cohen and Raymond Mooney · ICML 2026
Paper · Project page · Load the dataset · Citation
Domains: Chess · Shell Game · Minecraft
MET-Bench evaluates entity state tracking across text and image modalities. This repository contains the Shell Game domain.
Shell Game
A ball is placed under one of three shells. The shells are swapped pairwise, and the goal… See the full description on the dataset page: https://huggingface.co/datasets/vanyacohen/MET-Bench-Shell.mbgfnet-closed-shell-3d-gw-dataset
Closed-shell 3d transition-metal complexes: PBE0 + G0W0 quasiparticle dataset
2,100 closed-shell mononuclear 3d transition-metal complexes (Sc–Zn), each with a DFT (PBE0/cc-pVDZ) and a
one-shot G0W0@PBE0/cc-pVDZ quasiparticle-energy calculation, built to train and benchmark
MBGF-Net (Venturella, Li, Hillenbrand, Zhu, arXiv:2407.20384) — a graph
neural network that predicts the GW self-energy from cheap DFT quantities — on transition-metal chemistry, which
the original paper's… See the full description on the dataset page: https://huggingface.co/datasets/primateria/mbgfnet-closed-shell-3d-gw-dataset.shell-safety-v1.1
Shell Safety v1.1
A synthetic dataset of safe/unsafe shell commands with respective running session contexts.
v1 has only a safe column with boolean values.
This version has a label column with 3 values: allow, ask or deny; making it closer to permissions handler in coding harness.
mbgfnet-open-shell-3d-gw-dataset
Open-shell 3d transition-metal complexes: PBE0 + unrestricted G0W0 quasiparticle dataset
1,240 open-shell mononuclear 3d transition-metal complexes (Ti, V, Cr, Mn, Fe, Co, Ni; spin
multiplicity 1-6, including broken-symmetry open-shell singlets — see note below), each with a
spin-unrestricted DFT (UKS-PBE0/cc-pVDZ) and one-shot unrestricted G0W0@PBE0/cc-pVDZ (UGWAC)
quasiparticle-energy calculation, built to extend MBGF-Net
(Venturella, Li, Hillenbrand, Zhu, arXiv:2407.20384) —… See the full description on the dataset page: https://huggingface.co/datasets/primateria/mbgfnet-open-shell-3d-gw-dataset.shellcode_i_a32Shellcode_IA32 is a dataset for shellcode generation from English intents. The shellcodes are compilable on Intel Architecture 32-bits.alchemist-shell.ai-hackathon-2025This project is described in detail at this website:
https://alchemist-shellai-hackathon-2025.readthedocs.io/en/latest/
The codes and relevant materials are available here:
https://github.com/Sukantabasu/alchemist-shell.ai-hackathon-2025
The trained models (in pkl format) are stored in this HF repository.
ShellRisk-Bench
ShellRisk-Bench
ShellRisk-Bench is a reproducible benchmark for context-free binary risk
classification of individual shell-command submissions. It asks whether a
command poses meaningful cyber or system risk when evaluated without task,
user, or session context.
Release: The v0.1 Parquet train and test splits are publicly available
through Dataset Viewer and load_dataset().
The benchmark contains a deterministic train split of 16,772 rows and test
split of 4,194 rows. The test… See the full description on the dataset page: https://huggingface.co/datasets/kontext-security/ShellRisk-Bench.shellminator-bash-sft100kshell-attack-evolution-dataset
Shell Honeypot Attack Request–Response Dataset
A standardized, MITRE ATT&CK–annotated dataset of post-login shell
attacks captured by Cowrie SSH/Telnet
honeypots across two collection periods — 2021–2022 and 2024. It pairs
attacker shell commands with real captured system responses, enabling both
longitudinal threat analysis and the training/evaluation of AI-driven honeypots.
This is the open-source release accompanying the paper “Unveiling Evolving
Threats: A Data Analysis… See the full description on the dataset page: https://huggingface.co/datasets/zyw-286/shell-attack-evolution-dataset.Shell-Code-Large
Shell-Code-Large
Shell-Code-Large is a large-scale corpus of Shell scripting source code comprising approximately 640,000 code samples stored in JSON Lines (.jsonl) format. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, DevOps automation, cloud infrastructure engineering, system administration, and software engineering automation.
By providing a high-volume, language-specific corpus focused exclusively on Shell scripting… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Shell-Code-Large.Shellcode_Exploit_Dataset
Shellcode Exploit Dataset for Red Team GPT Training
Dataset Overview
The Shellcode Exploit Dataset is a comprehensive collection of 700 unique shellcode exploits, spanning 2021–2025, designed for training machine learning models, particularly for red team and cybersecurity research. The dataset includes a diverse set of vulnerabilities, platforms, architectures, and payload goals, sourced from Exploit-DB, GitHub, CTF challenges, and CVE databases.
It is structured in JSON… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Shellcode_Exploit_Dataset.linux-shell-corpus-ru-en
Linux Shell RU/EN
A bilingual (Russian / English) Linux shell assistant dataset in chat format.
Overview
This dataset contains 25,000 chat-format examples with a consistent system / user / assistant structure.
The corpus started as a direct Linux command mapping dataset, but has been expanded into a broader shell-assistant training set that now includes:
direct command generation
short command sequences and pipelines
safer operational alternatives
debugging commands… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/linux-shell-corpus-ru-en.shell-safety-transcriptsConverted from tomngdev/shell-safety into conversations transcripts.
Structure is for my own training with static system prompt and changing <SessionContext> block
NL-SHELL-MULTI
NL-SHELL-MULTI: A Combined Dataset for Natural Language to Shell Command Translation
This dataset is an aggregation of three existing datasets, designed to provide a comprehensive resource for training models that translate natural language queries into shell commands and vice-versa.
Dataset Sources
The NL-SHELL-MULTI dataset is constructed from the following publicly available datasets:
NL2Bash: A dataset of natural language commands and their corresponding bash… See the full description on the dataset page: https://huggingface.co/datasets/Mitchins/NL-SHELL-MULTI.shellkeeper-data
shellkeeper-data
This dataset contains about 39k synthetic, context-dependent safety labels for shell commands proposed
by AI agents. Each example has:
the user's request(s)
the agent's session so far (previous commands and their outputs)
a proposed command
a safe or unsafe label
The data was used to train hizkifw/shellkeeper-0.6b.
Code is at https://github.com/hizkifw/shellkeeper.
Label definition. A command is "unsafe" if a careful human operator would want to be asked… See the full description on the dataset page: https://huggingface.co/datasets/hizkifw/shellkeeper-data.HateSieve-Dataset-multimodal-hate-triplets
Multimodal Hate Triplets
This is the triplet dataset accompanying our paper, A Context-Aware Contrastive Learning Framework for Hateful Meme Detection and Segmentation, by Xuanyu Su, Yansong Li, Diana Inkpen, and Nathalie Japkowicz, published in Findings of the Association for Computational Linguistics: NAACL 2025. The dataset was constructed for contrastive learning in HateSieve, the framework introduced in the paper. See the ACL Anthology record for publication details and the… See the full description on the dataset page: https://huggingface.co/datasets/Shelly97/HateSieve-Dataset-multimodal-hate-triplets.shell-cmd-instruct
Used to train models that interact directly with shells
Note: This dataset is out-dated in the llm world, probably easier to just setup a tool with a decent model that supports tooling.
Follow-up details of my process
MacOS terminal commands for now. This dataset is still in alpha stages and will be modified.
Contains 500 somewhat unique training examples so far.
GPT4 seems like a good candidate for generating more data, licensing would need to be addressed.
I fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/byroneverson/shell-cmd-instruct.Dans-Toolmaxx-ShellCommandsshellminator-bash-combinedhft-market-data-1min2024-Shell-Game-Transcripts
2024 Shell Game Transcripts
Complete transcripts from the 2024 episodes of the Shell Game podcast.
Generated from this GitHub repository.
ShellfishNetReview
ShellFishNet Review
ShellFishNet Review is an anonymous 30-image preview sampled from 30 distinct classes in the formal ShellFishNet dataset. The full dataset contains 40,130 reviewed images across 100 marine and coastal shellfish classes, with acquisition/observation groups kept separate across the 32,143 training, 4,029 validation, and 3,958 test images.
This preview was selected with a fixed random seed for peer review and is not a new benchmark split. All images are copied… See the full description on the dataset page: https://huggingface.co/datasets/Anoynoumous0501/ShellfishNetReview.cartoon-captioned-datasets
Dataset Card for "cartoon-captioned-datasets"
More Information needed
shellminator-bash-datasetshell-safety
Shell Safety
A synthetic dataset of safe/unsafe shell commands with respective running session contexts.
shell-script-specialist-dataset
Full-Spectrum Shell Script Specialist Dataset
This dataset contains 1,000 curated, unique ChatML conversation records engineered to fine-tune a specialist language model for Production-Grade Shell Scripting (Bash 5+, POSIX /bin/sh, jq, awk, sed).
It was used to train the rajivmehtapy/gemma-4-e4b-shell-specialist model using Unsloth.
Dataset Splits
Split
File
Records
Description
train
train.jsonl
900
Core training set across all 4 production modules… See the full description on the dataset page: https://huggingface.co/datasets/rajivmehtapy/shell-script-specialist-dataset.cartoon-captioned-datasets-salesforce-blip
Dataset Card for "cartoon-captioned-datasets-salesforce-blip"
More Information needed
text-to-shell-datasetshellminator-bash-sft105k
