iselabvn/Linux-terminal-tool-calling
Linux Terminal Tool Calling Dataset (Linux-terminal-tool-calling) This dataset is designed for training and fine-tuning AI agents on tool calling, reasoning, and command execution specifically for standard Linux terminal utilities and system administration tasks. It transforms raw Linux terminal command records into a structured multi-turn conversation format featuring detailed chain-of-thought/reasoning content and OpenAI/OpenClaw-style function calling. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/iselabvn/Linux-terminal-tool-calling.
Linux Terminal Tool Calling Dataset (Linux-terminal-tool-calling)
This dataset is designed for training and fine-tuning AI agents on tool calling, reasoning, and command execution specifically for standard Linux terminal utilities and system administration tasks. It transforms raw Linux terminal command records into a structured multi-turn conversation format featuring detailed chain-of-thought/reasoning content and OpenAI/OpenClaw-style function calling.
Dataset Details
- Total Records: 600
- Language: English
- Format: JSONL (JSON Lines)
- License: Apache 2.0
- Repository: iselabvn/Linux-terminal-tool-calling
Dataset Structure
Each record is formatted as a single-turn conversation with user and assistant roles, complemented by rich metadata for downstream filtering and analysis.
Field Descriptions
- `messages` (list): Conversation history.
- `role: "user"` (dict): The user request describing a Linux terminal or system administration task.
- `role: "assistant"` (dict): The assistant response containing:
- `reasoning_content` (str): A 2-3 sentence chain-of-thought explanation explaining the choice of command, flags, and parameter configurations.
- `tool_calls` (list): An array containing the function call. The tool utilizes the
execfunction to run the command on the target environment. - `id` (str): A unique call identifier (e.g.,
call_exec_0). - `type: "function"` (str): The type of tool call.
- `function` (dict): Details of the target function call.
- `name`:
"exec" - `arguments` (JSON string): Serialized dictionary containing the exact executable command (
{"command": "..."}). - `content` (str): Empty string (standard format for tool calling responses).
- `metadata` (dict): Metadata associated with the command execution.
- `id` (str): Unique command ID (e.g.,
cmd-001). - `category` (str): Functional category of the command (e.g.,
File Management,Viewing,System Info). - `example_output` (str): Sample execution output or return status.
- `man_reference` (str): Link to the official man page documentation.
- `original_description` (str): The original, generic user description before prompt synchronization.
Category Distribution
Sample Record
{
"messages": [
{
"role": "user",
"content": "Change the current working directory to /var/log."
},
{
"role": "assistant",
"reasoning_content": "The user wants to navigate to the /var/log directory. The 'cd' command is the standard way to change directories in Unix-like systems. The path /var/log is specified directly. Executing the tool.",
"tool_calls": [
{
"id": "call_exec_0",
"type": "function",
"function": {
"name": "exec",
"arguments": "{\"command\": \"cd /var/log\"}"
}
}
],
"content": ""
}
],
"metadata": {
"id": "cmd-001",
"category": "Navigation",
"example_output": "[No output, changes directory]",
"man_reference": "https://man7.org/linux/man-pages/man1/cd.1.html",
"original_description": "Changes the current directory to /var/log."
}
}Intended Use
This dataset is ideal for:
- Fine-tuning language models to act as autonomous agent loops in Linux terminal environments.
- Training models on standard OpenAI tool calling conventions for shell commands.
- Supervised Fine Tuning (SFT) for system administration assistants, incorporating chain-of-thought (reasoning) before issuing commands.
Construction Method
The dataset was constructed by converting raw Linux terminal command records into a structured tool-use conversation trace.
To ensure consistency between user requests and executable commands, we utilized internal LLMs to perform Prompt Synchronization:
- Target Injection: Generic references (e.g., "a directory", "a file") in user requests were automatically replaced or synchronized with specific target parameters found in the command (e.g.,
/var/log,logfile.txt). - Chain-of-Thought Synthesis: The LLM generated a 2-3 sentence
reasoning_contentto justify the selection of the command, flags, and arguments. - Metadata Preservation: Original command IDs, categories, man references, sample outputs, and descriptions are preserved in
metadatafor alignment and verification.
