Team Ai
Modelpublic

RedTeamLab/Gemma-4-E2B-Sol-Traces-v1

sourceHugging Facegemmaupdated 3mo agoView on Hugging Face
0likes243downloads
Model Card

Gemma-4-E2B-Sol-Traces-v1

Repository coding-agent model fine-tuned from unsloth/gemma-4-E2B-it using LoRA on 25,000 verified deterministic reference trajectories.

The E2B variant is the smallest model in this four-run family. End-to-end tool-use benchmark comparisons are not yet published.

Sol Traces denotes tool-use traces compiled from Hermes Agent session logs; the traces do not originate from OpenCode.

Training Details

ParameterValue
Base modelunsloth/gemma-4-E2B-it (MoE, 2 active experts)
Fine-tuningLoRA (r=16, alpha=16, dropout=0)
Target modulesLanguage + attention (k/q/v/o/gate/up/down projection)
Dataset21,174 train / 1,324 val (gemma-4-native-tools format)
Dataset provenanceoriginal-synthetic — 25,000 verified trajectories compiled from Hermes Agent session logs across 32,560 attempted scenarios
Epochs1
Learning rate1e-4, cosine scheduler with 3% warmup
Batch size8 (2 × 4 gradient accumulation)
Max sequence8,192 tokens
Loss typeAssistant-only (tool responses excluded from loss)
GPUModal H100 80GB
Training time~45 min (pilot 3min + full 42min)
Final train loss0.0229
Validation loss0.0248
Peak VRAM33.7 GiB / 80 GiB
Throughput4,587 tok/s

These are reported run metrics; the canonical training_stats.json artifact is not currently published for E2B.

Dataset

The training dataset consists of 25,000 executable trajectories built by a deterministic scenario generator and replayed against generated repositories. It uses 224 language/task/variant repository families with repository-family-balanced splits:

  • —21,174 training records
  • —1,324 validation records
  • —2,502 test records (see dataset_manifest.json)

Each trajectory is a full agent session containing:

  • —System instruction: Repository coding agent with tool-use guidelines
  • —User task: A well-scoped coding task from the deterministic fixture catalogue
  • —Assistant tool calls: Multi-step function-calling sequences using 5 tools:
  • —list_files — glob-based file discovery
  • —read_file — line-range file reading
  • —search_code — regex code search (defined in the schema; not emitted by the v1 reference policy)
  • —run_command — allowlisted shell execution
  • —apply_patch — unified diff application
  • —Tool responses: Output, exit codes, truncation markers
  • —Verification: Post-task validation commands with pass/fail outcomes

Actual v1 task coverage

TypeRecords
debugging5,424
feature4,709
refactoring3,582
testing3,607
build_config3,269
integration2,742
documentation_review1,667

Repository fixtures cover TypeScript, JavaScript, Python, shell, configuration, Go, Rust, and JVM/Java.

Data generation and verification

Sol Traces are compiled from Hermes Agent session logs produced while running deterministic, seed-based coding scenarios through a reference executor. The scenarios define repository templates, task requirements, and verification commands; accepted records retain the corresponding tool-use events and verification outcomes. Records are included only when their configured post-task validation succeeds.

The v1 reference policy is intentionally narrow: it always lists files, reads the known implementation path, runs pre-patch verification, applies the reference patch, and reruns verification. search_code is included in the schema but has no v1 calls.

Key Statistics

MetricValue
Trace sourceHermes Agent session logs (deterministic scenario generator + reference executor)
Attempted seeds32,560
Accepted trajectories25,000 (76.8% acceptance rate)
Rejections5,872 structural duplicates + 316 verification failures
Provenanceoriginal-synthetic
Repository families224 language/task/variant families across 8 fixture categories

Files

FileSizeDescription
gemma-4-e2b-sol-traces-v1-Q4_K_M.gguf3.18 GiBQuantized merged model (Q4KM) — recommended for deployment
gemma-4-e2b-sol-traces-v1-f16.gguf8.64 GiBFull F16 merged model — for custom quantization
dataset_manifest.json—Accepted-record counts, split ratios, and rejection summary
Note: The Q4KM file is the recommended deployment format. The F16 is provided for downstream quantization experiments.

Usage (llama.cpp)

bash
# Q4_K_M — one file, ready to go
llama-cli \
  -m gemma-4-e2b-sol-traces-v1-Q4_K_M.gguf \
  -ngl 99 \
  --prompt "List the files in the repository matching *.py"

# With conversation template
llama-cli \
  -m gemma-4-e2b-sol-traces-v1-Q4_K_M.gguf \
  -ngl 99 \
  --temp 0.2 \
  --chat-template gemma \
  -p "Search the codebase for any TODO comments"

Capabilities

The model excels at:

  • —Function calling: Selecting and populating the right tool from natural language
  • —Code navigation: Searching, reading, and listing files to understand codebases
  • —Shell execution: Running commands with proper flags and paths
  • —Patch application: Making small, correct code changes via unified diffs
  • —Deterministic verification flow: Reproducing the fixture failure, applying the reference patch, and rerunning configured checks
  • —Verification: Running tests and validating changes

Comparison with Other Sol-Traces Models

ModelActive ParamsQ4 SizeTraining LossSpeedBest For
E2B (this)~5B3.2 GB0.0229FastestEdge, CPU+GPU hybrid, low-resource
12B Unified12B6.8 GB0.0800FastBalanced performance
E4B~8B4.9 GB0.0096FastBest quality-size trade-off
26B-A4B~8B*15.6 GB0.0113ModerateMaximum capability

*E4B and 26B-A4B both activate 4 experts but have different base architectures (dedicated encoder vs unified).

Limitations

  • —Fine-tuned for repository coding agent scenarios — general chat or creative writing may not benefit
  • —Single-turn trajectories only — no conversational memory across separate turns
  • —Tool schemas are fixed to the 5 tools in the training set
  • —Trained on synthetic trajectories — real-world coding patterns may differ

Training Stats

json
{
  "training_loss": 0.0229,
  "eval_loss": 0.0248,
  "steps": 377,
  "train_tokens": 24,704,714,
  "peak_vram_gib": 33.7,
  "throughput_tok_s": 4587,
  "runtime": "44m 46s"
}

Disclaimer

Use at your own risk. This model is fine-tuned for coding-agent scenarios. The model owner accepts no liability for any damages or losses arising from its use. Users are responsible for compliance with applicable laws and regulations.