Team Ai
Modelpublic

marzoukbaig14/committed-gguf

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
1likes45downloads
Model Card

Committed — Qwen3-1.7B (Q4KM GGUF) · higher-specificity option

A small model fine-tuned to write Conventional Commits messages from a git diff. It runs locally on CPU via llama.cpp, so your code never leaves your machine.

This is the larger of two sizes. Committed defaults to the smaller [0.6B GGUF](https://huggingface.co/marzoukbaig14/committed-gguf-0.6b) — it matches this 1.7B on commit-type accuracy and faithfulness at ~⅓ the size and a ~397 MB download. This 1.7B (~1 GB) is the bigger sibling: worth it when you want the extra specificity (0.67 vs the 0.6B's 0.55), i.e. more consistently concrete descriptions. Most users want the 0.6B; reach for this if concreteness matters more than footprint.

This repo holds the merged, quantized GGUF used for serving. The training dataset, LoRA adapter, and source code are linked below.

[Live Demo](https://my-portfolio-ten-rosy-36.vercel.app/committed) · [Gradio Space](https://huggingface.co/spaces/marzoukbaig14/committed-demo) · [0.6B GGUF](https://huggingface.co/marzoukbaig14/committed-gguf-0.6b) · 1.7B GGUF (this repo) · [0.6B adapter](https://huggingface.co/marzoukbaig14/committed-qwen3-0.6b-lora) · [1.7B adapter](https://huggingface.co/marzoukbaig14/committed-qwen3-1.7b-lora) · [Dataset](https://huggingface.co/datasets/marzoukbaig14/committed-train) · [GitHub](https://github.com/marzoukbaig14/Committed)

Details

  • —Base: Qwen/Qwen3-1.7B (Apache-2.0)
  • —Method: QLoRA fine-tune (PEFT LoRA + TRL SFTTrainer, vanilla transformers), merged into the base, converted to GGUF
  • —Quantization: Q4KM (~1 GB)
  • —Task: single-file git diff → one Conventional Commits subject line, type(scope): description
  • —Decoding: GBNF grammar-constrained, so output is always a well-formed CC line
  • —Trained on: marzoukbaig14/committed-train (~58k filtered CommitChronicle commits, 16 languages)

Usage

The trained behavior depends on the exact prompt rendering used in training plus the GBNF grammar applied at decode time, so a bare llama-cpp-python prompt will not reproduce the evaluated output. Run it through the project's inference path instead.

Via the CLI (no token needed; the GGUF is public). The CLI defaults to the 0.6B — pass --model 1.7b for this model:

bash
pip install --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cpu "committed @ git+https://github.com/marzoukbaig14/Committed.git"
git diff | committed --model 1.7b

Or call it through the repo's engine.py, the FastAPI endpoint, or the Gradio Space — all linked below. Full instructions and options are in the repo README.

Results

Evaluated against the un-tuned Qwen3-1.7B base on a 442-example test set, scored by an LLM judge (DeepSeek) on four axes; the judge was validated against 50 hand-rated examples. Headline numbers are reweighted to the true commit-type distribution.

The four-arm comparison — both base models, both fine-tunes, all judged by the same DeepSeek judge so the numbers are directly comparable:

Metric0.6B base0.6B fine-tune1.7B base**1.7B fine-tune**
Type accuracy0.1540.6010.1310.637
Type-correctness0.2960.7260.2960.778
Faithfulness0.2850.8100.4910.848
Completeness0.3530.7290.5430.776
Specificity0.4140.5450.8140.667
Conjunctive (all 4)0.1010.3590.1750.471
Graded mean (0–3)0.7772.0941.4472.139
feat-share of outputs86.7%9.7%95.5%8.4%

The finding: both base models "feat-collapse" — they label ~87–96% of all diffs feat regardless of content, scoring below a trivial always-guess-fix baseline (0.489) on type. Fine-tuning breaks the collapse on both. This 1.7B fine-tune is the stronger of the two overall, with its main edge in specificity (0.667 vs the 0.6B's 0.545) — but the margins on type, faithfulness, completeness, and the graded mean (2.139 vs 2.094) are small, which is why the 0.6B is the default and this is the higher-specificity option rather than the flagship.

Honest caveat on the judge: agreement with human raters is moderate-to-substantial on three axes (κ ≈ 0.56–0.61) but weakest on specificity (κ 0.34) — the very axis this model leads on — so that margin carries the most judge uncertainty. These DeepSeek-judged numbers are not comparable to any earlier Gemini-judged figures that may have appeared on this card previously; only the deltas within this all-DeepSeek table are valid.

Full before/after, the feat-collapse analysis, and the judge validation are in the eval writeup: [FINDINGS_v1.md](https://github.com/marzoukbaig14/Committed/blob/main/docs/eval/FINDINGS_v1.md).

Related

License

Apache-2.0, inherited from the Qwen3-1.7B base.

Citation

Trained with TRL. Dataset derived from CommitChronicle (Eliseeva et al., From Commit Message Generation to History-Aware Commit Message Generation, arXiv:2308.07655).