Team Ai
Modelpublic

ethanker/lfm2_350m_commit_diff_summarizer

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes19downloads
README.md136 linesDownload Raw Back to root
1---2base_model: unsloth/LFM2-350M-unsloth-bnb-4bit3library_name: peft4pipeline_tag: text-generation5tags:6  - "base_model:adapter:unsloth/LFM2-350M-unsloth-bnb-4bit"7  - lora8  - qlora9  - sft10  - transformers11  - trl12  - conventional-commits13  - code14---15 16 17# lfm2_350m_commit_diff_summarizer (LoRA)18 19A lightweight **helper model** that turns Git diffs into **Conventional Commit–style** messages.20It outputs **strict JSON** with a short `title` (≤ 65 chars) and up to 3 `bullets`, so your CLI/agents can parse it deterministically.21 22## Model Details23 24### Model Description25 26* **Purpose:** Summarize `git diff` patches into concise, Conventional Commit–compliant titles with optional bullets.27* **I/O format:**28 29  * **Input:** prompt containing the diff (plain text).30  * **Output:** JSON object: `{"title": "...", "bullets": ["...", "..."]}`.31* **Model type:** LoRA adapter for causal LM (text generation)32* **Language(s):** English (commit message conventions)33* **Finetuned from:** `unsloth/LFM2-350M-unsloth-bnb-4bit` (4-bit quantized base, trained with QLoRA)34 35### Model Sources36 37* **Repository:** This model card + adapter on the Hub under `ethanke/lfm2_350m_commit_diff_summarizer`38 39## Uses40 41### Direct Use42 43* Convert patch diffs into Conventional Commit messages for PR titles, commits, and changelogs.44* Provide human-readable summaries in agent UIs with guaranteed JSON structure.45 46### Recommendations47 48* Enforce JSON validation; if invalid, retry with a JSON-repair prompt.49* Keep a regex gate for Conventional Commit titles in your pipeline.50 51## How to Get Started52 53```python54from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig55from peft import PeftModel56import torch, json57 58BASE = "unsloth/LFM2-350M-unsloth-bnb-4bit"59ADAPTER = "ethanke/lfm2_350m_commit_diff_summarizer"  # replace with your repo id60 61bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",62                         bnb_4bit_use_double_quant=True, bnb_4bit_compute_dtype=torch.float16)63 64tok = AutoTokenizer.from_pretrained(BASE, use_fast=True)65mdl = AutoModelForCausalLM.from_pretrained(BASE, quantization_config=bnb, device_map="auto")66mdl = PeftModel.from_pretrained(mdl, ADAPTER)67 68diff = "...your git diff text..."69prompt = (70  "You are a commit message summarizer.\n"71  "Return a concise JSON object with fields 'title' (<=65 chars) and 'bullets' (0-3 items).\n"72  "Follow the Conventional Commit style for the title.\n\n"73  "### DIFF\n" + diff + "\n\n### OUTPUT JSON\n"74)75 76inputs = tok(prompt, return_tensors="pt").to(mdl.device)77with torch.no_grad():78    out = mdl.generate(**inputs, max_new_tokens=200, do_sample=False)79text = tok.decode(out[0], skip_special_tokens=True)80 81# naive JSON extraction82js = text[text.rfind("{"): text.rfind("}")+1]83obj = json.loads(js)84print(obj)85```86 87## Training Details88 89### Training Data90 91* **Dataset:** `Maxscha/commitbench` (diff → commit message).92* **Filtering:** kept only samples whose **first non-empty line** of the message matches Conventional Commits:93  `^(feat|fix|docs|style|refactor|perf|test|build|ci|chore|revert)(\([^)]+\))?(!)?:\s.+$`94* **Note:** The dataset card indicates non-commercial licensing. Confirm before commercial deployment.95 96### Training Procedure97 98* **Method:** Supervised fine-tuning (SFT) with TRL `SFTTrainer` + **QLoRA** (PEFT).99* **Prompting:** Instruction + `### DIFF` + `### OUTPUT JSON` target (title/bullets).100* **Precision:** fp16 compute on 4-bit base.101* **Hyperparameters (v0.1):**102 103  * `max_length=2048`, `per_device_train_batch_size=2`, `grad_accum=4`104  * `lr=2e-4`, `scheduler=cosine`, `warmup_ratio=0.03`105  * `epochs=1` over capped subset106  * LoRA: `r=16`, `alpha=32`, `dropout=0.05`, targets: q/k/v/o + MLP proj107 108### Evaluation109 110* **Validation:** filtered split from CommitBench.111* **Metrics (example run):**112 113  * `eval_loss ≈ 1.18`  → perplexity ≈ 3.26114  * `eval_mean_token_accuracy ≈ 0.77`115  * Suggested task metrics: JSON validity rate, CC-title compliance, title length ≤ 65 chars, bullets ≤ 3.116 117## Environmental Impact118 119* **Hardware:** 1× NVIDIA GTX 3060 12 GB (local)120* **Hours used:** ~2 h (prototype)121 122## Technical Specifications123 124* **Architecture:** LFM2-350M (decoder-only) + LoRA adapter125* **Libraries:** `transformers`, `trl`, `peft`, `bitsandbytes`, `datasets`, `unsloth`126 127## Contact128 129* Open an issue on the Hub repo or message `ethanke` on Hugging Face.130 131### Framework versions132 133* PEFT 0.17.1134* TRL (SFTTrainer)135* Transformers (recent version)136