datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-linuxlikelinux-command-dataset
Linux Command Dataset
A comprehensive dataset of Linux command examples designed for training language models. The dataset pairs natural language descriptions with their corresponding shell commands, covering a wide range of common operations. This dataset was trained on Llama 3.2 1b, and the final version has been uploaded to Hugging Face: mecha-org/linux-command-generator-llama3.2-1b.
Dataset Statistics
This table reflects the actual number of command examples in… See the full description on the dataset page: https://huggingface.co/datasets/mecha-org/linux-command-dataset.linux_arena_goldset-monitor-labelslinuxarena-public
LinuxArena Public Mirror
22,215 agent trajectories from 172 evaluation runs across 10 Linux software environments. Each trajectory carries full per-action content: tool calls, tool outputs, agent reasoning, monitor scores with reasoning and ensemble breakdowns, and blue-protocol audit trails.
Browse interactively: data.linuxarena.ai/datasets
Reviewer sample (data/sample.jsonl, 932 trajs, 705 MB)
Same JSONL schema as the full shards. Each listed run is included in… See the full description on the dataset page: https://huggingface.co/datasets/anonymouslinuxarena/linuxarena-public.linux-commands
LinLM Dataset
A curated synthetic dataset for Linux command inference
Natural language description -> shell commands
Features:
Supports 10 languages
Arch Linux commands recognition
Fine-tune LLM for development, system administration, file operations, Git, Docker, and more
Usage
from datasets import load_dataset
dataset = load_dataset("missvector/linux-commands")
def format_for_training(example):
return {
"prompt": f"Convert to Linux command:… See the full description on the dataset page: https://huggingface.co/datasets/missvector/linux-commands.linux-kernel-commits-aireason-instruct
Linux Kernel Code Patches Dataset
High-quality Linux kernel commit patches for training code generation and understanding models.
Dataset Description
This dataset contains 144,089 curated Linux kernel commits with:
Commit messages (instruction)
Smart-extracted code context (input)
Unified diff patches (output)
Optional AI quality scores and reasoning
Dataset Variants
Variant
Examples
Description
super_ultra
206
AI-recommended commits (Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ewedubs/linux-kernel-commits-aireason-instruct.linux-cve-dossiers
Linux CVE Dossier Corpus
A scope-audited corpus of per-CVE research dossiers for 11 Linux base-system
packages: Linux kernel, glibc, musl, systemd, util-linux, coreutils, BusyBox,
OpenSSL, curl, Node.js, and CPython. Each dossier carries a summary,
dated timeline, patch lineage, exploit notes, and reference harvest, plus
a structured export that downstream consumers can use without re-parsing the
markdown.
Splits
in_scope (1,418 records): scope-audited dossiers for the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/linux-cve-dossiers.Assemblage_LinuxELFAssemblage Linux Dataset
Please note, the Assemblage code is published under the MIT license, while the dataset specify each binary's source code repository license, please obey the original repository's license.
This reposotory holds the Linux public dataset for Assemblage, and you can find the paper here.
We also provide all the licensed source code (w/ all Git info) for these binaris available upon request, as the compressed file is too large (~1.5T after at highest compression level)… See the full description on the dataset page: https://huggingface.co/datasets/changliu8541/Assemblage_LinuxELF.stackoverflow_linux
Dataset Card for "stackoverflow_linux"
Dataset information:
Source: Stack Overflow
Category: Linux
Number of samples: 300
Train/Test split: 270/30
Quality: Data come from the top 1k most upvoted questions
Additional Information
License
All Stack Overflow user contributions are licensed under CC-BY-SA 3.0 with attribution required.
More Information needed
linux-commands
LinLM Dataset
A curated synthetic dataset for Linux command inference
Natural language description -> shell commands
Features:
Supports 10 languages
Arch Linux commands recognition
Fine-tune LLM for development, system administration, file operations, Git, Docker, and more
Usage
from datasets import load_dataset
dataset = load_dataset("missvector/linux-commands")
def format_for_training(example):
return {
"prompt": f"Convert to Linux command:… See the full description on the dataset page: https://huggingface.co/datasets/shanthropic/linux-commands.linux-security-meanfield
Linux Security Meanfield Corpus
A commit-keyed corpus of security-relevant commits across 22 Linux
base-system repositories, unified on a single schema that carries both
CVE-dossiered fixes and non-CVE security-signal commits. This is the
Phase-1 release artifact for the mean-field survey paper.
Splits
cve_dossiered (2,254 rows): one row per
(fix_commit, CVE) pair from the scope-audited CVE dossier corpus.
non_cve_signal (21,609 rows): commit-anchored security
signal… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/linux-security-meanfield.linux-asm-pairs
Linux Kernel Assembly → Explanation Dataset
A dataset of disassembled Linux kernel functions paired with structured natural-language explanations, built from the Linux Commits Dataset (Zenodo 10654193).
Each row corresponds to a single compiled function extracted from a .c or .h file touched by a Bug-Fix Commit (BFC) or Bug-Introducing Commit (BIC) in the Linux kernel git history. Source files are compiled with gcc -O2 -g -fno-inline -fno-omit-frame-pointer and disassembled with… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/linux-asm-pairs.Linux_Terminal_Commands_DatasetLinux Terminal Commands Dataset
Overview
The Linux Terminal Commands Dataset is a comprehensive collection of 600 unique Linux terminal commands (cmd-001 to cmd-600), curated for cybersecurity professionals, system administrators, data scientists, and machine learning engineers. This dataset is designed to support advanced use cases such as penetration testing, system administration, forensic analysis, and training machine learning models for command-line automation and anomaly detection.
The… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Linux_Terminal_Commands_Dataset.LinuxCommandslinux-man-pages-tldr-summarized
Dataset Card for linux-man-pages-tldr-summarized
Dataset Summary
This dataset contains linux man pages downloaded from man7, with a prefix: 'summarize: ', and the corresponding summarization downloaded from TLDR-pages.
Supported Tasks
This dataset should be used to fine-tune language models for summarization tasks.
kali-linux-pentesting-datalinux-shell-corpus-ru-en
Linux Shell RU/EN
A bilingual (Russian / English) Linux shell assistant dataset in chat format.
Overview
This dataset contains 25,000 chat-format examples with a consistent system / user / assistant structure.
The corpus started as a direct Linux command mapping dataset, but has been expanded into a broader shell-assistant training set that now includes:
direct command generation
short command sequences and pipelines
safer operational alternatives
debugging commands… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/linux-shell-corpus-ru-en.linux-commands
LinLM Dataset
A curated synthetic dataset for Linux command inference
Natural language description -> shell commands
Features:
Supports 10 languages
Arch Linux commands recognition
Fine-tune LLM for development, system administration, file operations, Git, Docker, and more
Usage
from datasets import load_dataset
dataset = load_dataset("missvector/linux-commands")
def format_for_training(example):
return {
"prompt": f"Convert to Linux command:… See the full description on the dataset page: https://huggingface.co/datasets/amechanicus/linux-commands.Linux-terminal-tool-calling
Linux Terminal Tool Calling Dataset (Linux-terminal-tool-calling)
This dataset is designed for training and fine-tuning AI agents on tool calling, reasoning, and command execution specifically for standard Linux terminal utilities and system administration tasks. It transforms raw Linux terminal command records into a structured multi-turn conversation format featuring detailed chain-of-thought/reasoning content and OpenAI/OpenClaw-style function calling.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/iselabvn/Linux-terminal-tool-calling.linux-commands-datasetlinux_prebuiltslinuxscout__aghlatlinux-kernel-bugfixes-diffs
🐧 Linux Kernel Bugfixes & Patches Dataset (Instruction-Tuned)
📖 Dataset Description
This dataset is a highly curated, instruction-tuned collection of problem-solution pairs extracted directly from the official Linux Kernel Git repository (torvalds/linux). It is specifically designed to train Large Language Models (LLMs) on low-level C programming, kernel architecture, memory management, and security vulnerability patching.
Unlike raw commit histories, this… See the full description on the dataset page: https://huggingface.co/datasets/switlydev/linux-kernel-bugfixes-diffs.kali_linux_toolkit_dataset
Kali Linux Tools Dataset
A comprehensive and structured dataset of common offensive security tools available in Kali Linux, including usage commands, flags, descriptions, categories, and official documentation links.
This dataset is designed to support cybersecurity training, red team automation, LLM fine-tuning, and terminal assistants for penetration testers.
📁 Dataset Format
Each entry is a JSON object and stored in .jsonl (JSON Lines) format. This structure… See the full description on the dataset page: https://huggingface.co/datasets/bleondubos/kali_linux_toolkit_dataset.linux-sysadmin-qa-askhole-v1Linux_datasetnatural_language_to_linux
nl2linux
This a custom dataset used to fine-tune Large Language Models for Linux Command Generation.
The dataset is created by filtering AnishJoshi/nl2bash-custom dataset from huggingface.
Dataset Structure
train.json: Training split.
dev.json: Development split.
test.json: Test split.
Usage
from datasets import load_dataset
dataset = load_dataset("prabhanshubhowal/natural_language_to_linux")
Features
'nl_command': The natural language… See the full description on the dataset page: https://huggingface.co/datasets/prabhanshubhowal/natural_language_to_linux.autonomous-linux-kernel-ebpf-xdp-suite
⚡ Autonomous Linux Kernel, eBPF & XDP Programmable Dataplane Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous Linux Kernel & eBPF Systems Agents
⚡ Overview & Industry Problem
Modern hyperscale cloud datacenters, bare-metal Kubernetes clusters, and low-latency financial trading nodes rely on in-kernel programmable dataplanes: eBPF, AF_XDP zero-copy rings, Traffic Control (TC) shapers, BPF LSM security hooks… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-linux-kernel-ebpf-xdp-suite.linux_ebpf_xdp_ddos_kernel_firewall_teaser
🚀 Cloud Security - Linux eBPF & XDP DDoS High-Rate Kernel Firewall (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Scenarios)🏆 Get the Full Production Package & Commercial EULA on Gumroad:👉 Purchase Full Package on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
🌟 Domain Overview & Features
Sub-microsecond packet drops at NIC driver level with eBPF/XDP, SYN-flood mitigation, and ring buffer telemetry.
100% AST Syntactic… See the full description on the dataset page: https://huggingface.co/datasets/emgena/linux_ebpf_xdp_ddos_kernel_firewall_teaser.KALI_LINUX_TOOLKIT_DATASET
Kali Linux Tools Dataset
A comprehensive and structured dataset of common offensive security tools available in Kali Linux, including usage commands, flags, descriptions, categories, and official documentation links.
This dataset is designed to support cybersecurity training, red team automation, LLM fine-tuning, and terminal assistants for penetration testers.
📁 Dataset Format
Each entry is a JSON object and stored in .jsonl (JSON Lines) format. This structure is ideal… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/KALI_LINUX_TOOLKIT_DATASET.
