linux
Datasets
All datasets matching “linux”fineweb-linuxlikelinux-command-dataset
Linux Command Dataset
A comprehensive dataset of Linux command examples designed for training language models. The dataset pairs natural language descriptions with their corresponding shell commands, covering a wide range of common operations. This dataset was trained on Llama 3.2 1b, and the final version has been uploaded to Hugging Face: mecha-org/linux-command-generator-llama3.2-1b.
Dataset Statistics
This table reflects the actual number of command examples in… See the full description on the dataset page: https://huggingface.co/datasets/mecha-org/linux-command-dataset.linux_arena_goldset-monitor-labelslinuxarena-public
LinuxArena Public Mirror
22,215 agent trajectories from 172 evaluation runs across 10 Linux software environments. Each trajectory carries full per-action content: tool calls, tool outputs, agent reasoning, monitor scores with reasoning and ensemble breakdowns, and blue-protocol audit trails.
Browse interactively: data.linuxarena.ai/datasets
Reviewer sample (data/sample.jsonl, 932 trajs, 705 MB)
Same JSONL schema as the full shards. Each listed run is included in… See the full description on the dataset page: https://huggingface.co/datasets/anonymouslinuxarena/linuxarena-public.linux-commands
LinLM Dataset
A curated synthetic dataset for Linux command inference
Natural language description -> shell commands
Features:
Supports 10 languages
Arch Linux commands recognition
Fine-tune LLM for development, system administration, file operations, Git, Docker, and more
Usage
from datasets import load_dataset
dataset = load_dataset("missvector/linux-commands")
def format_for_training(example):
return {
"prompt": f"Convert to Linux command:… See the full description on the dataset page: https://huggingface.co/datasets/missvector/linux-commands.Assemblage_LinuxELFAssemblage Linux Dataset
Please note, the Assemblage code is published under the MIT license, while the dataset specify each binary's source code repository license, please obey the original repository's license.
This reposotory holds the Linux public dataset for Assemblage, and you can find the paper here.
We also provide all the licensed source code (w/ all Git info) for these binaris available upon request, as the compressed file is too large (~1.5T after at highest compression level)… See the full description on the dataset page: https://huggingface.co/datasets/changliu8541/Assemblage_LinuxELF.
