RedTeamLab/Gemma-4-E2B-Sol-Traces-v1
Gemma-4-E2B-Sol-Traces-v1
Repository coding-agent model fine-tuned from unsloth/gemma-4-E2B-it using LoRA on 25,000 verified deterministic reference trajectories.
The E2B variant is the smallest model in this four-run family. End-to-end tool-use benchmark comparisons are not yet published.
Sol Traces denotes tool-use traces compiled from Hermes Agent session logs; the traces do not originate from OpenCode.
Training Details
These are reported run metrics; the canonical training_stats.json artifact is not currently published for E2B.
Dataset
The training dataset consists of 25,000 executable trajectories built by a deterministic scenario generator and replayed against generated repositories. It uses 224 language/task/variant repository families with repository-family-balanced splits:
- 21,174 training records
- 1,324 validation records
- 2,502 test records (see
dataset_manifest.json)
Each trajectory is a full agent session containing:
- System instruction: Repository coding agent with tool-use guidelines
- User task: A well-scoped coding task from the deterministic fixture catalogue
- Assistant tool calls: Multi-step function-calling sequences using 5 tools:
list_files— glob-based file discoveryread_file— line-range file readingsearch_code— regex code search (defined in the schema; not emitted by the v1 reference policy)run_command— allowlisted shell executionapply_patch— unified diff application- Tool responses: Output, exit codes, truncation markers
- Verification: Post-task validation commands with pass/fail outcomes
Actual v1 task coverage
Repository fixtures cover TypeScript, JavaScript, Python, shell, configuration, Go, Rust, and JVM/Java.
Data generation and verification
Sol Traces are compiled from Hermes Agent session logs produced while running deterministic, seed-based coding scenarios through a reference executor. The scenarios define repository templates, task requirements, and verification commands; accepted records retain the corresponding tool-use events and verification outcomes. Records are included only when their configured post-task validation succeeds.
The v1 reference policy is intentionally narrow: it always lists files, reads the known implementation path, runs pre-patch verification, applies the reference patch, and reruns verification. search_code is included in the schema but has no v1 calls.
Key Statistics
Files
Note: The Q4KM file is the recommended deployment format. The F16 is provided for downstream quantization experiments.
Usage (llama.cpp)
# Q4_K_M — one file, ready to go
llama-cli \
-m gemma-4-e2b-sol-traces-v1-Q4_K_M.gguf \
-ngl 99 \
--prompt "List the files in the repository matching *.py"
# With conversation template
llama-cli \
-m gemma-4-e2b-sol-traces-v1-Q4_K_M.gguf \
-ngl 99 \
--temp 0.2 \
--chat-template gemma \
-p "Search the codebase for any TODO comments"Capabilities
The model excels at:
- Function calling: Selecting and populating the right tool from natural language
- Code navigation: Searching, reading, and listing files to understand codebases
- Shell execution: Running commands with proper flags and paths
- Patch application: Making small, correct code changes via unified diffs
- Deterministic verification flow: Reproducing the fixture failure, applying the reference patch, and rerunning configured checks
- Verification: Running tests and validating changes
Comparison with Other Sol-Traces Models
*E4B and 26B-A4B both activate 4 experts but have different base architectures (dedicated encoder vs unified).
Limitations
- Fine-tuned for repository coding agent scenarios — general chat or creative writing may not benefit
- Single-turn trajectories only — no conversational memory across separate turns
- Tool schemas are fixed to the 5 tools in the training set
- Trained on synthetic trajectories — real-world coding patterns may differ
Training Stats
{
"training_loss": 0.0229,
"eval_loss": 0.0248,
"steps": 377,
"train_tokens": 24,704,714,
"peak_vram_gib": 33.7,
"throughput_tok_s": 4587,
"runtime": "44m 46s"
}Disclaimer
Use at your own risk. This model is fine-tuned for coding-agent scenarios. The model owner accepts no liability for any damages or losses arising from its use. Users are responsible for compliance with applicable laws and regulations.
