Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01coms4705-hewitt /fineweb-linuxliketext100K<n<1M0 likes2.9k downloads1y agoHugging Face02mecha-org /linux-command-dataset Linux Command Dataset A comprehensive dataset of Linux command examples designed for training language models. The dataset pairs natural language descriptions with their corresponding shell commands, covering a wide range of common operations. This dataset was trained on Llama 3.2 1b, and the final version has been uploaded to Hugging Face: mecha-org/linux-command-generator-llama3.2-1b. Dataset Statistics This table reflects the actual number of command examples in… See the full description on the dataset page: https://huggingface.co/datasets/mecha-org/linux-command-dataset.texttext-generation1K<n<10K14 likes444 downloads1y agoHugging Face03missvector /linux-commands LinLM Dataset A curated synthetic dataset for Linux command inference Natural language description -> shell commands Features: Supports 10 languages Arch Linux commands recognition Fine-tune LLM for development, system administration, file operations, Git, Docker, and more Usage from datasets import load_dataset dataset = load_dataset("missvector/linux-commands") def format_for_training(example): return { "prompt": f"Convert to Linux command:… See the full description on the dataset page: https://huggingface.co/datasets/missvector/linux-commands.texttext-generation10K<n<100K97 likes320 downloads9mo agoHugging Face04mjbommar /linux-cve-dossiers Linux CVE Dossier Corpus A scope-audited corpus of per-CVE research dossiers for 11 Linux base-system packages: Linux kernel, glibc, musl, systemd, util-linux, coreutils, BusyBox, OpenSSL, curl, Node.js, and CPython. Each dossier carries a summary, dated timeline, patch lineage, exploit notes, and reference harvest, plus a structured export that downstream consumers can use without re-parsing the markdown. Splits in_scope (1,418 records): scope-audited dossiers for the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/linux-cve-dossiers.tabulartext-classification1K<n<10K1 likes209 downloads6mo agoHugging Face05shanthropic /linux-commands LinLM Dataset A curated synthetic dataset for Linux command inference Natural language description -> shell commands Features: Supports 10 languages Arch Linux commands recognition Fine-tune LLM for development, system administration, file operations, Git, Docker, and more Usage from datasets import load_dataset dataset = load_dataset("missvector/linux-commands") def format_for_training(example): return { "prompt": f"Convert to Linux command:… See the full description on the dataset page: https://huggingface.co/datasets/shanthropic/linux-commands.texttext-generation10K<n<100K1 likes127 downloads9mo agoHugging Face06Romit2004 /LinuxCommandstext10K<n<100K29 likes113 downloads3y agoHugging Face07theelderemo /linux-asm-pairs Linux Kernel Assembly → Explanation Dataset A dataset of disassembled Linux kernel functions paired with structured natural-language explanations, built from the Linux Commits Dataset (Zenodo 10654193). Each row corresponds to a single compiled function extracted from a .c or .h file touched by a Bug-Fix Commit (BFC) or Bug-Introducing Commit (BIC) in the Linux kernel git history. Source files are compiled with gcc -O2 -g -fno-inline -fno-omit-frame-pointer and disassembled with… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/linux-asm-pairs.texttext-generation1K<n<10K0 likes111 downloads6mo agoHugging Face08NickIBrody /linux-shell-corpus-ru-en Linux Shell RU/EN A bilingual (Russian / English) Linux shell assistant dataset in chat format. Overview This dataset contains 25,000 chat-format examples with a consistent system / user / assistant structure. The corpus started as a direct Linux command mapping dataset, but has been expanded into a broader shell-assistant training set that now includes: direct command generation short command sequences and pipelines safer operational alternatives debugging commands… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/linux-shell-corpus-ru-en.texttext-generation10K<n<100K0 likes110 downloads6mo agoHugging Face09KonradSzafer /stackoverflow_linux Dataset Card for "stackoverflow_linux" Dataset information: Source: Stack Overflow Category: Linux Number of samples: 300 Train/Test split: 270/30 Quality: Data come from the top 1k most upvoted questions Additional Information License All Stack Overflow user contributions are licensed under CC-BY-SA 3.0 with attribution required. More Information needed textquestion-answeringn<1K9 likes97 downloads4y agoHugging Face10suryanshp1 /kali-linux-pentesting-datatextn<1K25 likes97 downloads2y agoHugging Face11beatsprom /autonomous-linux-kernel-ebpf-xdp-suite ⚡ Autonomous Linux Kernel, eBPF & XDP Programmable Dataplane Suite (2026) A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous Linux Kernel & eBPF Systems Agents ⚡ Overview & Industry Problem Modern hyperscale cloud datacenters, bare-metal Kubernetes clusters, and low-latency financial trading nodes rely on in-kernel programmable dataplanes: eBPF, AF_XDP zero-copy rings, Traffic Control (TC) shapers, BPF LSM security hooks… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-linux-kernel-ebpf-xdp-suite.tabulartext-generation1K<n<10K0 likes94 downloads22d agoHugging Face12tmskss /linux-man-pages-tldr-summarized Dataset Card for linux-man-pages-tldr-summarized Dataset Summary This dataset contains linux man pages downloaded from man7, with a prefix: 'summarize: ', and the corresponding summarization downloaded from TLDR-pages. Supported Tasks This dataset should be used to fine-tune language models for summarization tasks. textsummarizationn<1K10 likes93 downloads3y agoHugging Face13coms4705-hewitt /linuxlike-tokenizedtext10K<n<100K0 likes93 downloads1y agoHugging Face14hrsvrn /linux-commands-datasettext1M<n<10M5 likes80 downloads1y agoHugging Face15alielfilali01 /linuxscout__aghlattext10K<n<100K0 likes77 downloads3y agoHugging Face16emgena /linux_ebpf_xdp_ddos_kernel_firewall_teaser 🚀 Cloud Security - Linux eBPF & XDP DDoS High-Rate Kernel Firewall (Evaluation Teaser) ⚡ Official Free Evaluation Teaser (50 Verified Scenarios)🏆 Get the Full Production Package & Commercial EULA on Gumroad:👉 Purchase Full Package on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout! 🌟 Domain Overview & Features Sub-microsecond packet drops at NIC driver level with eBPF/XDP, SYN-flood mitigation, and ring buffer telemetry. 100% AST Syntactic… See the full description on the dataset page: https://huggingface.co/datasets/emgena/linux_ebpf_xdp_ddos_kernel_firewall_teaser.textn<1K0 likes77 downloads9d agoHugging Face17switlydev /linux-kernel-bugfixes-diffs 🐧 Linux Kernel Bugfixes & Patches Dataset (Instruction-Tuned) 📖 Dataset Description This dataset is a highly curated, instruction-tuned collection of problem-solution pairs extracted directly from the official Linux Kernel Git repository (torvalds/linux). It is specifically designed to train Large Language Models (LLMs) on low-level C programming, kernel architecture, memory management, and security vulnerability patching. Unlike raw commit histories, this… See the full description on the dataset page: https://huggingface.co/datasets/switlydev/linux-kernel-bugfixes-diffs.texttext-generation100K<n<1M0 likes74 downloads2mo agoHugging Face18bleondubos /kali_linux_toolkit_dataset Kali Linux Tools Dataset A comprehensive and structured dataset of common offensive security tools available in Kali Linux, including usage commands, flags, descriptions, categories, and official documentation links. This dataset is designed to support cybersecurity training, red team automation, LLM fine-tuning, and terminal assistants for penetration testers. 📁 Dataset Format Each entry is a JSON object and stored in .jsonl (JSON Lines) format. This structure… See the full description on the dataset page: https://huggingface.co/datasets/bleondubos/kali_linux_toolkit_dataset.textn<1K0 likes73 downloads24d agoHugging Face19nagrajn /synthetic_linux_commands Dataset Card for Dataset Name Linux Commands and Responses Data Set. Dataset Details Dataset Description Linux Commands and Responses Data Set. Curated by: Nagraj Naidu Language(s) (NLP): English License: MIT Dataset Sources [optional] Repository: https://huggingface.co/datasets/nagrajn/synthetic_linux_commands Uses Pre learning and Fine Tuning an LLM Direct Use Pre Learning Out-of-Scope Use None… See the full description on the dataset page: https://huggingface.co/datasets/nagrajn/synthetic_linux_commands.text1M<n<10M2 likes70 downloads2y agoHugging Face20prabhanshubhowal /natural_language_to_linux nl2linux This a custom dataset used to fine-tune Large Language Models for Linux Command Generation. The dataset is created by filtering AnishJoshi/nl2bash-custom dataset from huggingface. Dataset Structure train.json: Training split. dev.json: Development split. test.json: Test split. Usage from datasets import load_dataset dataset = load_dataset("prabhanshubhowal/natural_language_to_linux") Features 'nl_command': The natural language… See the full description on the dataset page: https://huggingface.co/datasets/prabhanshubhowal/natural_language_to_linux.text10K<n<100K12 likes67 downloads1y agoHugging Face21mjbommar /linux-ioctl-census Linux IOCTL Census -- public structural tier A source-derived census of the Linux kernel local ioctl/proc/sysfs handler surface: for each registered handler, its decoded _IOC command table, the permission gates on its path, and a capability-ungated reachability upper bound. The schema is identical to the Windows IOCTL Census (mjbommar/ioctl-census), so the two can be queried and compared together. This is the public structural tier: everything derivable from the already-public… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/linux-ioctl-census.tabular1K<n<10K1 likes65 downloads4mo agoHugging Face22eval-aware /linuxarena-trajectories LinuxArena Trajectories Full agent trajectories from LinuxBench/LinuxArena evaluations across 14 model/policy combinations and 10 environments. Dataset Description Each row is one complete evaluation trajectory — every tool call the agent made from start to finish, with arguments, outputs, errors, and reasoning. Actions are represented as parallel variable-length lists (one element per action). Two granularity levels are provided per action: Normalized… See the full description on the dataset page: https://huggingface.co/datasets/eval-aware/linuxarena-trajectories.tabulartext-generationn<1K1 likes64 downloads7mo agoHugging Face23amechanicus /linux-commands LinLM Dataset A curated synthetic dataset for Linux command inference Natural language description -> shell commands Features: Supports 10 languages Arch Linux commands recognition Fine-tune LLM for development, system administration, file operations, Git, Docker, and more Usage from datasets import load_dataset dataset = load_dataset("missvector/linux-commands") def format_for_training(example): return { "prompt": f"Convert to Linux command:… See the full description on the dataset page: https://huggingface.co/datasets/amechanicus/linux-commands.texttext-generation10K<n<100K0 likes63 downloads9mo agoHugging Face24mjbommar /linux-security-meanfield Linux Security Meanfield Corpus A commit-keyed corpus of security-relevant commits across 22 Linux base-system repositories, unified on a single schema that carries both CVE-dossiered fixes and non-CVE security-signal commits. This is the Phase-1 release artifact for the mean-field survey paper. Splits cve_dossiered (2,254 rows): one row per (fix_commit, CVE) pair from the scope-audited CVE dossier corpus. non_cve_signal (21,609 rows): commit-anchored security signal… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/linux-security-meanfield.text10K<n<100K1 likes62 downloads6mo agoHugging Face25dim /linux_man_pages_tldr_summarized Dataset Card for "linux_man_pages_tldr_summarized" More Information needed textn<1K3 likes58 downloads3y agoHugging Face26darkknight25 /KALI_LINUX_TOOLKIT_DATASET Kali Linux Tools Dataset A comprehensive and structured dataset of common offensive security tools available in Kali Linux, including usage commands, flags, descriptions, categories, and official documentation links. This dataset is designed to support cybersecurity training, red team automation, LLM fine-tuning, and terminal assistants for penetration testers. 📁 Dataset Format Each entry is a JSON object and stored in .jsonl (JSON Lines) format. This structure is ideal… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/KALI_LINUX_TOOLKIT_DATASET.texttext-classificationn<1K9 likes57 downloads1y agoHugging Face27software-si /linux-file-search Linux File Search Dataset Dataset Summary The Linux File Search NLI Dataset is a synthetic dataset designed to train and evaluate Natural Language Inference (NLI) models that map natural language file search queries into structured representations of file attributes. The dataset is intended to enable semantic file search on Linux systems by allowing models to extract structured constraints such as file type, extension, size, ownership, permissions, and other properties… See the full description on the dataset page: https://huggingface.co/datasets/software-si/linux-file-search.texttext-classification1K<n<10K1 likes57 downloads9mo agoHugging Face28mrheinen /linux-commandstexttext-generationn<1K0 likes56 downloads2y agoHugging Face29Jawajawa /command-linux-bash-balanced-sfttext1K<n<10K1 likes54 downloads6mo agoHugging Face30anonymouslinuxarena /straj_linuxarena straj_linuxarena (public-env subset) Adversarial sabotage agentic SWE benchmark trajectories from the linuxarena project, run as part of the no-CoT time-horizons paper. The dataset viewer above shows the per-cell outcome records (precomputed_results.csv, 257 rows). Full Inspect .eval trajectories for the 13 public-environment tasks (~2.4 GB, 150 files) are stored under evals/ — see "Inspect trajectories" below. Per-cell schema Each row in precomputed_results.csv (and… See the full description on the dataset page: https://huggingface.co/datasets/anonymouslinuxarena/straj_linuxarena.tabulartext-generationn<1K1 likes54 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.