reiglabs/shellm-0.6B
SheLLM 0.6B
Turn a short English request into a Linux shell command.
Request: enter the root directory
Command: cd /SheLLM is an experimental English-to-shell model developed and trained by reiglabs, a two-person team. It adapts the language knowledge of Qwen/Qwen3-0.6B-Base with supervised LoRA fine-tuning. Its interface produces command text for a person or application to review.
The project explores training, tokenization, transformers, functional evaluation, and inference engineering on consumer hardware. Browser inference and custom WebGPU kernels are future project goals; this release uses PyTorch, Transformers, and PEFT.
Latest evaluated checkpoint: adapters/2026-10-10-pilot-v2-fresh-lora-expandable/epoch-3. It scores 104/112 (92.9%) on ShellBench v1 and 103/300 (34.3%) on ShellBench Extra v1. These are functional results from two Docker fixtures per case, with substantial remaining gaps in compositions and multi-step tasks.
Model details
Sources and contact
- Model and checkpoint archive: reiglabs/shellm-0.6B.
- Code, training data, documentation and benchmark definitions: matyas-toth/shellm on GitHub.
- Latest experiment notebook: pilot-v2 fresh LoRA experiment.
- Interactive local interface: `chat.py`.
- Paper: There is no dedicated SheLLM paper linked in the project. See the Qwen3 base model card for upstream model references.
- Hosted demo: No hosted SheLLM demo is documented for this release.
- Contact: Open a model discussion or a GitHub issue.
Uses
Direct use
Translate short requests for everyday Linux filesystem and text operations into candidate commands. Typical topics include navigation, listing files, creating or moving files and directories, searching content, reading text, counting, finding files, and changing permissions.
The intended interaction is request → generated command → human review. The inference example below prints text. Working directory, shell state, available utilities and filesystem contents are supplied by the eventual execution environment.
Downstream use
Use the adapters for educational experiments, local command suggestion tools, research into command composition, and further supervised fine-tuning with independently checked data. The source repository provides editable intent catalogs, rendered training examples, and functional evaluation tooling for contributors.
Out-of-scope use
General-purpose conversation, unrestricted system administration, Windows PowerShell, arbitrary shell dialects, production automation without review, and autonomous execution with elevated permissions have not been established by these experiments. Current benchmark results do not establish competence across all Linux tools or environments.
Bias, risks and limitations
- Command composition is still weak. On Extra v1, pipeline coverage passes 2/32 cases, multi-step coverage 0/18, and conditional coverage 1/13. Basic-command results should not be extrapolated to complex workflows.
- Syntax can be valid while the intent is wrong. All latest v1 commands pass syntax checks, but eight fail functional evaluation. The latest Extra checkpoint passes 292/300 syntax checks and 103/300 functional checks.
- Arguments require care. Filenames with spaces or shell metacharacters, path interpretation, quoting, exact flag requirements and combinations of operations can be mishandled.
- Generated commands can change or delete data. Review the command and arguments before executing it. Interactive confirmation and constrained execution environments are appropriate for tools that use the model.
- Training requests are synthetic and template-based. They reflect the authors' selected capabilities and phrasing. The model has not been evaluated across broad user demographics, languages or accessibility needs; robustness to arbitrary natural language is not established.
- Validation is limited. The new dataset portion holds out wording and argument roots together, while sharing authored intent templates with training. It does not independently hold out compositions. Inherited data uses weaker split separation.
- ShellBench v1 is a development diagnostic. Its aggregate gaps informed data expansion. It is not a blind generalization test. Extra v1 is a separate progress benchmark outside the training and validation splits.
- Functional checks cover finite fixtures. Matching two Docker environments does not establish correctness on every machine, directory tree, locale, utility version or shell state.
- Progress includes regressions. The latest adapter improves overall scores while losing some previously passing cases. Low completion loss and higher aggregate accuracy do not make every request reliable.
- Context and decoding differ from a chat assistant. Fine-tuning examples use a raw request prefix and short command completion. This checkpoint has not been evaluated as a multi-turn chat assistant or for long-context requests.
Recommendations
Use the exact prompt format below, begin with simple operations, and inspect the generated command. Build downstream checks for argument handling and filesystem effects. Keep validation and benchmark examples separate from training, and assess changes with functional tests as well as syntax and exact-string metrics.
Getting started
This repository stores adapters in run/epoch subfolders. Load the selected subfolder together with the pinned Qwen Base model. A generic root-level pipeline is not sufficient to select an archived checkpoint.
The project's recorded environment uses Python 3.12, Transformers 4.57.6, PEFT 0.18.1, and Hugging Face Hub 0.36.2, with a suitable PyTorch installation. Follow the project setup instructions for the CUDA environment.
from pathlib import Path
import torch
from huggingface_hub import snapshot_download
from peft import PeftConfig, PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
repo_id = "reiglabs/shellm-0.6B"
subfolder = "adapters/2026-10-10-pilot-v2-fresh-lora-expandable/epoch-3"
# Fetch only the selected adapter checkpoint.
snapshot = snapshot_download(repo_id, allow_patterns=[f"{subfolder}/*"])
adapter_path = Path(snapshot) / subfolder
config = PeftConfig.from_pretrained(str(adapter_path))
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.float16 if device == "cuda" else torch.float32
tokenizer = AutoTokenizer.from_pretrained(
config.base_model_name_or_path, revision=config.revision
)
base = AutoModelForCausalLM.from_pretrained(
config.base_model_name_or_path, revision=config.revision, dtype=dtype
).to(device)
model = PeftModel.from_pretrained(base, str(adapter_path)).eval()
request = "enter the root directory"
prompt = f"Request: {request}\nCommand:"
inputs = tokenizer(prompt, return_tensors="pt").to(device)
with torch.inference_mode():
output = model.generate(
**inputs,
do_sample=False,
max_new_tokens=64,
stop_strings=["\n"],
tokenizer=tokenizer,
pad_token_id=tokenizer.eos_token_id,
)
completion = tokenizer.decode(
output[0, inputs.input_ids.shape[-1]:], skip_special_tokens=True
)
command = completion.partition("\n")[0].strip()
print(command) # Command text for review; no shell execution here.Use a fixed Hub commit revision in snapshot_download when reproducing an experiment. The checkpoint metadata records the base revision and adapter hash. The existing tokenizer chat template is an upstream tokenizer artifact; inference for this task uses the raw prefix shown above.
Training details
Training data
The complete pilot-v2 release contains 22,784 examples: 18,392 training and 4,392 validation examples across 74 families, 4,001 scenario groups and 3,959 unique command strings. It preserves the previous release and adds 20,340 examples. Paraphrases from one scenario remain in the same split.
Labels are rendered from structured command intents. Natural-language requests are authored around those intents, and commands are checked against independent Python intent oracles in disposable Docker fixtures. All 4,001 groups passed two fixtures each: 8,002 label executions. These checks validate label behavior; they do not by themselves prove every English paraphrase expresses that behavior correctly.
New validation uses disjoint argument roots and separate sentence banks. Normalized benchmark request overlap is zero, and the Extra audit found zero exact existing command-label overlap. Literal separation does not establish semantic independence.
- Editable catalogs.
- Coverage, quality checks and review pack.
- Contributor guide.
- Dataset SHA-256:
d546e6075d539854706eecce78ce0e3667257c067a218722945ddff545265e93.
The broader language knowledge comes from the upstream pretrained Qwen model. Consult its model card for upstream training information.
Procedure and preprocessing
Each experiment starts with the pinned Base model and a fresh LoRA adapter. Latest training loads the complete pilot-v2 snapshot once. Requests use Request: <request>\nCommand:; command completions have a leading space and an EOS token. The supervised loss covers the completion and EOS, while prompt and padding tokens are masked. All examples passed tokenizer-boundary checks; the longest complete example is 86 tokens within the 128-token limit.
Latest training hyperparameters
This is a LoRA experiment with unquantized Base weights; QLoRA and browser-oriented quantization are possible future experiments.
Sizes and times
The latest adapter weights occupy 20,236,472 bytes (approximately 20.2 MB), excluding the Base model, tokenizer and other files. The successful training loop took 7,839.4 seconds / 130.7 minutes, including validation and checkpoint saving, excluding model loading. Peak live PyTorch allocation was 1.824 GiB, with 2.010 GiB peak reserved memory. These are training measurements, not browser memory or isolated inference benchmarks.
Validation completion loss was 1.137967 before training, 0.030199 after epoch one, 0.019838 after epoch two, and 0.019458 after epoch three. The final epoch was selected before benchmark scores. Earlier epochs are archived; only the final checkpoint was benchmarked for this latest milestone.
Evaluation
Testing data and protocol
ShellBench v1 contains 112 requests across 14 basic command families, with varied wording, arguments, quoting, options and some compositions. ShellBench Extra v1 contains 300 separate requests across 20 families, including practical, compositional and edge cases. Extra is a long-term progress measure outside training and validation.
Reference commands are validated independently. Each extracted model command runs on two disposable Docker fixtures. The evaluator compares the required stdout/stderr, exit status, working directory and observable filesystem effects under its case-specific rules. A functional case passes only when both fixtures pass.
Base and the latest adapter share the pinned model revision, raw prompt format, greedy decoding, FP16, generation batch size eight, a 64-token limit, and newline/EOS stopping. Suite, evaluator, Docker-image, prediction and adapter fingerprints are recorded with the results.
Metrics and latest results
- Functional accuracy: correct behavior on both fixtures.
- Exact string match: an extracted command matching a reference label; alternative valid commands can differ.
- Syntax: the command passes the shell syntax check.
- Format: the raw completion conforms to the command-only output contract.
- Usable: both functional and correctly formatted.
Historical final-checkpoint results
Against the previous final adapter, v1 has 11 newly passing cases and two regressions; Extra has 39 newly passing cases and 13 regressions. Data, batching and schedule length changed between runs, so these comparisons do not isolate the effect of dataset size.
Measured factors include command family, capability coverage and difficulty. The latest v1 run passes all basic, argument, informal and paraphrase groups, while quoting and composition remain imperfect. Extra exposes larger gaps in pipelines, conditionals and multi-step workflows.
Detailed machine-readable reports and charts, benchmark cases, and evaluator code are in the source repository. Raw completions, extracted commands, syntax checks and Docker functional outcomes are recorded separately.
Model examination
The project examines behavior through held-out examples, independent label oracles, functional fixtures, and per-case failure/regression reports. No dedicated mechanistic interpretability or attention-analysis study is reported for this checkpoint. The benchmark results measure observed behavior within the test contract.
Technical specifications and compute
The architecture is the upstream Qwen3 causal decoder transformer with LoRA updates on attention and MLP projections. Fine-tuning minimizes next-token cross-entropy on command completions and EOS. The pretrained tokenizer and Base weights are reused, with a raw task prefix instead of a chat interaction format.
Training ran locally on an NVIDIA GeForce RTX 2060 with 6 GiB VRAM, using Ubuntu in WSL2 on Windows. Recorded software: Python 3.12, PyTorch 2.11.0+cu130, Transformers 4.57.6, PEFT 0.18.1, Accelerate 1.15.0, and Hugging Face Hub 0.36.2. The successful run used PYTORCH_ALLOC_CONF=expandable_segments:True. Functional evaluation uses disposable Linux Docker containers.
The latest adapter SHA-256 is 11e702fb8d699cf88244d542bfab234a33588cfb0b4e9685239d78a5916cef3c. The pinned revision, dataset hash, per-epoch metadata and source fingerprints support reproduction of the recorded experiment.
Environmental impact
- Hardware: Local RTX 2060, 6 GiB VRAM.
- Successful latest training-loop duration: Approximately 2.18 hours; this excludes model loading, two interrupted preparation attempts, and benchmark execution.
- Cloud provider: Not applicable to the recorded local training run.
- Compute region: Physical location is not recorded in the training metadata; Europe/Budapest is the reporting timezone.
- Energy use and carbon emissions: Not measured. No quantitative emissions claim is made.
Licensing
reiglabs releases these LoRA adapter weights under the Apache License 2.0. The upstream Qwen3-0.6B-Base model is also distributed under Apache 2.0. Retain applicable license and attribution notices when redistributing the model or derivative artifacts.
Citation
There is no dedicated SheLLM publication. A suggested citation for this model release is:
@misc{reiglabs_shellm_2026,
author = {Tóth, Mátyás and Alattyányi, Csanád},
title = {SheLLM 0.6B: English-to-Linux-Shell LoRA Adapters},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/reiglabs/shellm-0.6B}
}Please also credit the upstream Qwen model when describing work based on its pretrained weights.
Glossary
- LoRA: Trainable low-rank updates attached to selected frozen model layers.
- SFT: Supervised fine-tuning using request/command examples.
- Scenario group: One structured intent with related request paraphrases kept in the same split.
- Functional evaluation: Executing a generated command in fixtures and comparing its required observable behavior.
- Checkpoint: The saved adapter state after an epoch, identified by run, epoch and weight hash.
Model card authors and maintenance
This card is published by reiglabs, the team of Mátyás Tóth and Csanád Alattyányi, and was prepared from the project's preserved training/data/evaluation records at the team's request. Snapshot date: 2026-10-10 (Europe/Budapest). Results refer to the explicitly identified checkpoint; future archive additions require fresh measurements before updating these claims.
