mailtotanvir/nano-agent-code-speed-1.5b
Nano-agent code-speed 1.5B: v5 and v6.1 experiment variants
This repository contains two equal experiment variants, v5/ and v6p1/. Neither is labeled the default or a general-purpose code optimizer. Their results differ and must be read together.
A plausible speedup can be a different program. These adapters propose patches; the verifier decides whether they still answer the same question.
The source repository, technical story, and paper PDF are the companion artifacts. The paper's separate Zenodo record (DOI 10.5281/zenodo.22922767) identifies this work. The existing Rust-repair DOI and model repository describe a different task and base.
Both variants are LoRA adapters for Qwen/Qwen2.5-Coder-1.5B-Instruct (Apache-2.0 per the base model's Hugging Face metadata). They propose SEARCH/REPLACE patches for restricted pure-Python optimization problems. Use them with the nano-agent controller and its isolated hidden-input correctness plus Cachegrind verifier; do not execute arbitrary proposals as trusted code.
One-shot behavioral results
Each problem allowed one greedy proposal through the same production Rust controller and verifier. A success passed an independent hidden-input battery and had positive Cachegrind instruction-reference reward. Model inputs did not include hidden tests or known-fast implementations.
On frozen family-heldout, v5 scored 7/10 string concatenation, 7/10 sort selection, and 0/10 indexed lookup; v6.1 scored 8/10, 8/10, and 0/10 respectively. The two 30-case frozen tracks test different questions and must not be merged into one headline. The 3-case diagnostic is too small to estimate broad transfer. v6.1's three proposals there dropped an in-function constant, producing candidate errors. Frozen results are descriptive and must not be used to tune another checkpoint.
Training and files
Both variants use LoRA rank 16, alpha 32, dropout 0.05 on attention and MLP projection modules. Both used one epoch, learning rate 5e-5, batch size 4, gradient accumulation 4, effective v1 replay ratio 0.4, and seed 7.
The v5 corpus contains 530 verifier-approved teacher turns, 502 extra focused copies in six training-only invariant families, and 115 v1 carryover rows. The v6.1 corpus contains 550 teacher turns, 504 focused copies, one repair trace, and 115 v1 carryover rows. Corpus SHA-256 values are v5 499338f9405e0892d08dd3faca72885c0c55822e70793f532dc9c33a3e19b610 and v6.1 fc9f92f2a3dcd5a2a0cf3dd2667ebdfa47b0cc842a73314a2d9364cc9f1ad4d2. Training loss and internal token accuracy are not behavioral success metrics.
The adapters are stored under v5/ and v6p1/; the shared tokenizer and task-specific chat template are stored once at repository root. A downloader can load either local subdirectory with PEFT after obtaining the base model:
from pathlib import Path
from huggingface_hub import snapshot_download
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base_id = "Qwen/Qwen2.5-Coder-1.5B-Instruct"
snapshot = Path(snapshot_download("mailtotanvir/nano-agent-code-speed-1.5b"))
variant = "v5" # or "v6p1"; neither is the default
base = AutoModelForCausalLM.from_pretrained(base_id, dtype="auto")
model = PeftModel.from_pretrained(base, str(snapshot / variant))
tokenizer = AutoTokenizer.from_pretrained(str(snapshot))Test both load paths before release. The repository's task-specific controller provides the prompt and verification contract; a generic chat pipeline example would be misleading. Exact file checksums and proposed layout are in the adjacent adapter release manifest.
Limits and evidence
The benchmark is a narrow pure-Python optimization task. Finite hidden tests are not proof of equivalence. Cachegrind Ir is simulated instruction references, not wall-clock latency. Both variants failed every indexed-lookup family-heldout case. v6.1 failed all three fresh residue-count-index diagnostics; v5 was not tested there. Neither adapter should be described as able to optimize arbitrary code.
The technical report and public source repository provide the protocol, training configuration, split-level results, and raw-report SHA-256 values. The raw reports and private test inputs are retained outside Git. The GCP training VMs were deleted after archive transfer and hash verification; price quotes are not actual billed amounts. The trainer-generated cards in both adapter extractions contain an unrelated sample prompt and incorrect license field; this joint card replaces them for release.
Provenance and citation
The adapter weight SHA-256 values are listed above; exact configuration and shared tokenizer hashes are in ADAPTER_RELEASE_MANIFEST.md. The frozen v6.1 full-report hashes are 00a41526db5fac9e963e6c41a4bd59a1888ba2fcd3ce2e2266a7d384f65b7dd2 (seen-family), adc0066bde045e9622654de6dca8bc16d2c7c42ea30f9d2257eb64e5a09ec383 (family-heldout), and 05a66dab3d26b7fbe2266e41eb1913b5faac74d59df5fb68fe4e08874d68b1c1 (three-case diagnostic). Cite the paper as: Tanvir Ahmed (2026), The Optimization That Forgot the Question: Verifier-Gated Code Speedups with a 1.5B Model, doi:10.5281/zenodo.22922767.
