Team Ai
Datasetpublic

iselabvn/Linux-terminal-tool-calling

Linux Terminal Tool Calling Dataset (Linux-terminal-tool-calling) This dataset is designed for training and fine-tuning AI agents on tool calling, reasoning, and command execution specifically for standard Linux terminal utilities and system administration tasks. It transforms raw Linux terminal command records into a structured multi-turn conversation format featuring detailed chain-of-thought/reasoning content and OpenAI/OpenClaw-style function calling. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/iselabvn/Linux-terminal-tool-calling.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
1likes83downloads
Dataset Card

Linux Terminal Tool Calling Dataset (Linux-terminal-tool-calling)

This dataset is designed for training and fine-tuning AI agents on tool calling, reasoning, and command execution specifically for standard Linux terminal utilities and system administration tasks. It transforms raw Linux terminal command records into a structured multi-turn conversation format featuring detailed chain-of-thought/reasoning content and OpenAI/OpenClaw-style function calling.

Dataset Details

Dataset Structure

Each record is formatted as a single-turn conversation with user and assistant roles, complemented by rich metadata for downstream filtering and analysis.

Field Descriptions

  • —`messages` (list): Conversation history.
  • —`role: "user"` (dict): The user request describing a Linux terminal or system administration task.
  • —`role: "assistant"` (dict): The assistant response containing:
  • —`reasoning_content` (str): A 2-3 sentence chain-of-thought explanation explaining the choice of command, flags, and parameter configurations.
  • —`tool_calls` (list): An array containing the function call. The tool utilizes the exec function to run the command on the target environment.
  • —`id` (str): A unique call identifier (e.g., call_exec_0).
  • —`type: "function"` (str): The type of tool call.
  • —`function` (dict): Details of the target function call.
  • —`name`: "exec"
  • —`arguments` (JSON string): Serialized dictionary containing the exact executable command ({"command": "..."}).
  • —`content` (str): Empty string (standard format for tool calling responses).
  • —`metadata` (dict): Metadata associated with the command execution.
  • —`id` (str): Unique command ID (e.g., cmd-001).
  • —`category` (str): Functional category of the command (e.g., File Management, Viewing, System Info).
  • —`example_output` (str): Sample execution output or return status.
  • —`man_reference` (str): Link to the official man page documentation.
  • —`original_description` (str): The original, generic user description before prompt synchronization.

Category Distribution

CategoryRecord CountDescription
File Management94File creation, copying, moving, deletion, and file system operations
Viewing90Displaying file contents, tailing log files, viewing headers & line counts
System Info79System resources, hardware details, uptime, kernel & OS information
Networking76Network interfaces, routing, connectivity tests, DNS lookups, ports
Process65Process listing, monitoring, signaling, background jobs, termination
Permissions50File and directory permissions, ownership, and ACL modifications
User Management46User/group management, password policies, and shell configurations
Package Management39APT/YUM/RPM package management, updates, and software installations
Navigation35Directory traversal, path printing, and directory listing
Editor26CLI text editor invocations and stream processing (nano, vim, sed, awk)

Sample Record

json
{
  "messages": [
    {
      "role": "user",
      "content": "Change the current working directory to /var/log."
    },
    {
      "role": "assistant",
      "reasoning_content": "The user wants to navigate to the /var/log directory. The 'cd' command is the standard way to change directories in Unix-like systems. The path /var/log is specified directly. Executing the tool.",
      "tool_calls": [
        {
          "id": "call_exec_0",
          "type": "function",
          "function": {
            "name": "exec",
            "arguments": "{\"command\": \"cd /var/log\"}"
          }
        }
      ],
      "content": ""
    }
  ],
  "metadata": {
    "id": "cmd-001",
    "category": "Navigation",
    "example_output": "[No output, changes directory]",
    "man_reference": "https://man7.org/linux/man-pages/man1/cd.1.html",
    "original_description": "Changes the current directory to /var/log."
  }
}

Intended Use

This dataset is ideal for:

  • —Fine-tuning language models to act as autonomous agent loops in Linux terminal environments.
  • —Training models on standard OpenAI tool calling conventions for shell commands.
  • —Supervised Fine Tuning (SFT) for system administration assistants, incorporating chain-of-thought (reasoning) before issuing commands.

Construction Method

The dataset was constructed by converting raw Linux terminal command records into a structured tool-use conversation trace.

To ensure consistency between user requests and executable commands, we utilized internal LLMs to perform Prompt Synchronization:

  1. 1.Target Injection: Generic references (e.g., "a directory", "a file") in user requests were automatically replaced or synchronized with specific target parameters found in the command (e.g., /var/log, logfile.txt).
  2. 2.Chain-of-Thought Synthesis: The LLM generated a 2-3 sentence reasoning_content to justify the selection of the command, flags, and arguments.
  3. 3.Metadata Preservation: Original command IDs, categories, man references, sample outputs, and descriptions are preserved in metadata for alignment and verification.