datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
linux-command-dataset
Linux Command Dataset
A comprehensive dataset of Linux command examples designed for training language models. The dataset pairs natural language descriptions with their corresponding shell commands, covering a wide range of common operations. This dataset was trained on Llama 3.2 1b, and the final version has been uploaded to Hugging Face: mecha-org/linux-command-generator-llama3.2-1b.
Dataset Statistics
This table reflects the actual number of command examples in… See the full description on the dataset page: https://huggingface.co/datasets/mecha-org/linux-command-dataset.linux-commands
LinLM Dataset
A curated synthetic dataset for Linux command inference
Natural language description -> shell commands
Features:
Supports 10 languages
Arch Linux commands recognition
Fine-tune LLM for development, system administration, file operations, Git, Docker, and more
Usage
from datasets import load_dataset
dataset = load_dataset("missvector/linux-commands")
def format_for_training(example):
return {
"prompt": f"Convert to Linux command:… See the full description on the dataset page: https://huggingface.co/datasets/missvector/linux-commands.linux-commands
LinLM Dataset
A curated synthetic dataset for Linux command inference
Natural language description -> shell commands
Features:
Supports 10 languages
Arch Linux commands recognition
Fine-tune LLM for development, system administration, file operations, Git, Docker, and more
Usage
from datasets import load_dataset
dataset = load_dataset("missvector/linux-commands")
def format_for_training(example):
return {
"prompt": f"Convert to Linux command:… See the full description on the dataset page: https://huggingface.co/datasets/shanthropic/linux-commands.linux-asm-pairs
Linux Kernel Assembly → Explanation Dataset
A dataset of disassembled Linux kernel functions paired with structured natural-language explanations, built from the Linux Commits Dataset (Zenodo 10654193).
Each row corresponds to a single compiled function extracted from a .c or .h file touched by a Bug-Fix Commit (BFC) or Bug-Introducing Commit (BIC) in the Linux kernel git history. Source files are compiled with gcc -O2 -g -fno-inline -fno-omit-frame-pointer and disassembled with… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/linux-asm-pairs.linux-shell-corpus-ru-en
Linux Shell RU/EN
A bilingual (Russian / English) Linux shell assistant dataset in chat format.
Overview
This dataset contains 25,000 chat-format examples with a consistent system / user / assistant structure.
The corpus started as a direct Linux command mapping dataset, but has been expanded into a broader shell-assistant training set that now includes:
direct command generation
short command sequences and pipelines
safer operational alternatives
debugging commands… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/linux-shell-corpus-ru-en.autonomous-linux-kernel-ebpf-xdp-suite
⚡ Autonomous Linux Kernel, eBPF & XDP Programmable Dataplane Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous Linux Kernel & eBPF Systems Agents
⚡ Overview & Industry Problem
Modern hyperscale cloud datacenters, bare-metal Kubernetes clusters, and low-latency financial trading nodes rely on in-kernel programmable dataplanes: eBPF, AF_XDP zero-copy rings, Traffic Control (TC) shapers, BPF LSM security hooks… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-linux-kernel-ebpf-xdp-suite.linux-kernel-bugfixes-diffs
🐧 Linux Kernel Bugfixes & Patches Dataset (Instruction-Tuned)
📖 Dataset Description
This dataset is a highly curated, instruction-tuned collection of problem-solution pairs extracted directly from the official Linux Kernel Git repository (torvalds/linux). It is specifically designed to train Large Language Models (LLMs) on low-level C programming, kernel architecture, memory management, and security vulnerability patching.
Unlike raw commit histories, this… See the full description on the dataset page: https://huggingface.co/datasets/switlydev/linux-kernel-bugfixes-diffs.linuxarena-trajectories
LinuxArena Trajectories
Full agent trajectories from LinuxBench/LinuxArena evaluations
across 14 model/policy combinations and 10 environments.
Dataset Description
Each row is one complete evaluation trajectory — every tool call the agent made from start
to finish, with arguments, outputs, errors, and reasoning. Actions are represented as
parallel variable-length lists (one element per action).
Two granularity levels are provided per action:
Normalized… See the full description on the dataset page: https://huggingface.co/datasets/eval-aware/linuxarena-trajectories.linux-commands
LinLM Dataset
A curated synthetic dataset for Linux command inference
Natural language description -> shell commands
Features:
Supports 10 languages
Arch Linux commands recognition
Fine-tune LLM for development, system administration, file operations, Git, Docker, and more
Usage
from datasets import load_dataset
dataset = load_dataset("missvector/linux-commands")
def format_for_training(example):
return {
"prompt": f"Convert to Linux command:… See the full description on the dataset page: https://huggingface.co/datasets/amechanicus/linux-commands.linux-commandsstraj_linuxarena
straj_linuxarena (public-env subset)
Adversarial sabotage agentic SWE benchmark trajectories from the linuxarena project, run as part of the no-CoT time-horizons paper.
The dataset viewer above shows the per-cell outcome records (precomputed_results.csv, 257 rows). Full Inspect .eval trajectories for the 13 public-environment tasks (~2.4 GB, 150 files) are stored under evals/ — see "Inspect trajectories" below.
Per-cell schema
Each row in precomputed_results.csv (and… See the full description on the dataset page: https://huggingface.co/datasets/anonymouslinuxarena/straj_linuxarena.linux-commands
LinLM Dataset
A curated synthetic dataset for Linux command inference
Natural language description -> shell commands
Features:
Supports 10 languages
Arch Linux commands recognition
Fine-tune LLM for development, system administration, file operations, Git, Docker, and more
Usage
from datasets import load_dataset
dataset = load_dataset("missvector/linux-commands")
def format_for_training(example):
return {
"prompt": f"Convert to Linux command:… See the full description on the dataset page: https://huggingface.co/datasets/Rendra8631/linux-commands.Linux-terminal-tool-calling
Linux Terminal Tool Calling Dataset (Linux-terminal-tool-calling)
This dataset is designed for training and fine-tuning AI agents on tool calling, reasoning, and command execution specifically for standard Linux terminal utilities and system administration tasks. It transforms raw Linux terminal command records into a structured multi-turn conversation format featuring detailed chain-of-thought/reasoning content and OpenAI/OpenClaw-style function calling.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/iselabvn/Linux-terminal-tool-calling.linux-kernel-ioctl-qwen35
Linux ioctl Census
Description
Teaches domain-specific instruction following and code generation for this expert.
Source
mjbommar/linux-ioctl-census - CC-BY-4.0
Formatted for the MoE-orchestrator project
(https://github.com/michaelowusuntim6/MoE-orchestrator). Expert target:
linux_kernel.
Format
Each record is a JSON object with a messages field formatted for Qwen3.5's
native chat template:
{"messages": [
{"role": "system"… See the full description on the dataset page: https://huggingface.co/datasets/michaelowusuntim6/linux-kernel-ioctl-qwen35.linux-commands
LinLM Dataset
A curated synthetic dataset for Linux command inference
Natural language description -> shell commands
Features:
Supports 10 languages
Arch Linux commands recognition
Fine-tune LLM for development, system administration, file operations, Git, Docker, and more
Usage
from datasets import load_dataset
dataset = load_dataset("missvector/linux-commands")
def format_for_training(example):
return {
"prompt": f"Convert to Linux command:… See the full description on the dataset page: https://huggingface.co/datasets/jasmar2/linux-commands.linux-qwen35
Linux Shell and CLI Corpus
Description
Teaches domain-specific instruction following and code generation for this expert.
Source
b-mc2/cli-commands-explained
NickIBrody/linux-shell-corpus-ru-en
rajivmehtapy/shell-script-specialist-dataset
Formatted for the MoE-orchestrator project
(https://github.com/michaelowusuntim6/MoE-orchestrator). Expert target:
linux_kernel.
Format
Each record is a JSON object with a messages field formatted… See the full description on the dataset page: https://huggingface.co/datasets/michaelowusuntim6/linux-qwen35.linux-kernel-asm-qwen35
Linux Kernel Assembly Pairs
Description
Teaches domain-specific instruction following and code generation for this expert.
Source
theelderemo/linux-asm-pairs - GPL-2.0
Formatted for the MoE-orchestrator project
(https://github.com/michaelowusuntim6/MoE-orchestrator). Expert target:
linux_kernel.
Format
Each record is a JSON object with a messages field formatted for Qwen3.5's
native chat template:
{"messages": [
{"role": "system"… See the full description on the dataset page: https://huggingface.co/datasets/michaelowusuntim6/linux-kernel-asm-qwen35.linux-commands
LinLM Dataset
A curated synthetic dataset for Linux command inference
Natural language description -> shell commands
Features:
Supports 10 languages
Arch Linux commands recognition
Fine-tune LLM for development, system administration, file operations, Git, Docker, and more
Usage
from datasets import load_dataset
dataset = load_dataset("missvector/linux-commands")
def format_for_training(example):
return {
"prompt": f"Convert to Linux command:… See the full description on the dataset page: https://huggingface.co/datasets/chowmean/linux-commands.linux-kernel-commits-qwen35
Linux Kernel Commit Reasoning
Description
Teaches domain-specific instruction following and code generation for this expert.
Source
ewedubs/linux-kernel-commits-aireason-instruct - Apache-2.0
Formatted for the MoE-orchestrator project
(https://github.com/michaelowusuntim6/MoE-orchestrator). Expert target:
linux_kernel.
Format
Each record is a JSON object with a messages field formatted for Qwen3.5's
native chat template:… See the full description on the dataset page: https://huggingface.co/datasets/michaelowusuntim6/linux-kernel-commits-qwen35.linux-sysadmin-qa-askhole
Linux Sysadmin Q&A with Askhole Personality
This dataset contains 200 Linux system administration question-answer pairs with a snarky "Askhole" personality inspired by TARS (Interstellar), K-2SO (Star Wars), and Deadpool.
Dataset Description
The dataset provides helpful Linux sysadmin information delivered with sass, humor, and brutal honesty. Each answer combines technical accuracy with personality-driven commentary.
Features
200 Q&A pairs covering… See the full description on the dataset page: https://huggingface.co/datasets/crazycog/linux-sysadmin-qa-askhole.linuxarena-first5-trajectories
LinuxArena First-5 Trajectories
First-5 tool call trajectories from LinuxBench/LinuxArena evaluations across 14 model/policy combinations and 10 environments.
Dataset Description
Each row represents one evaluation trajectory with the first 5 tool calls extracted at two granularity levels:
Level 1 (normalized): Tool categories like text_editor:view, bash:find, bash:ls
Level 2 (exact): Full command strings like bash$ find /app/src -type f -name "*.ts" | sort
Designed for… See the full description on the dataset page: https://huggingface.co/datasets/eval-aware/linuxarena-first5-trajectories.mini_coco_linux
mini coco dataset files
Required dependencies
OpenCV (cv2)
matplotlib
ipywidgets
img_data.psv
Extract of the coco dataset containing the following labels: ["airplane", "backpack", "cell phone", "handbag", "suitcase", "knife", "laptop", "car"] (300 of each)
Structured as follows:
| Field | Description |
| --------------- |… See the full description on the dataset page: https://huggingface.co/datasets/iix/mini_coco_linux.smolified-linux-master
🤏 smolified-linux-master
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model draganite/smolified-linux-master.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 066632c9)
Records: 8610
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by draganite.
Generated via Smolify.ai.
linux-commands
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Ukoll/linux-commands.ru-linux-sysadmin-dialogues
Russian Linux Sysadmin Dialogues (Датасет для обучения ИИ)
Высококачественный структурированный набор данных (датасет) на русском языке, содержащий профессиональные инструкции, разборы технических проблем и сценарии общения в сфере системного администрирования операционных систем семейства Linux.
Этот датасет разработан специально для тонкой настройки (fine-tuning) больших языковых моделей (LLM), обучения диалоговых агентов, умных помощников технической поддержки и наполнения… See the full description on the dataset page: https://huggingface.co/datasets/CBERX/ru-linux-sysadmin-dialogues.Books-General-Linux
Linux Books Dataset
Dataset Description
The Linux Books Dataset is a curated text dataset derived from Linux-related books and learning materials. It focuses on Linux system administration, cybersecurity, networking, shell scripting, and operating system fundamentals.The dataset is designed to support training and evaluation of NLP models for technical domains, especially cybersecurity-aware language models and Linux-focused assistants.
This dataset is suitable for both… See the full description on the dataset page: https://huggingface.co/datasets/DexopT/Books-General-Linux.
