Team Ai
Modelpublic

AmareshHebbar/leetcode-cpp-qwen25-coder-7b

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes22downloads
README.md486 linesDownload Raw Back to root
1---2license: apache-2.03base_model: unsloth/Qwen2.5-Coder-7B-Instruct4tags:5  - code6  - leetcode7  - cpp8  - code-generation9  - competitive-programming10  - qwen2.5-coder11  - dora12  - qdora13  - weight-decomposed-lora14  - instruction-tuned15  - sft16  - algorithm-generation17  - function-generation18  - coding-assistant19  - on-device20  - gguf21  - ollama22  - vllm23  - text-generation-inference24  - doocs-leetcode25  - synthetic-verification26  - quantized27  - algorithms28language:29  - en30library_name: peft31pipeline_tag: text-generation32datasets:33  - AmareshHebbar/leetcode-codegen-cpp34co2_eq_emissions:35  emissions: 036  source: "estimate, not measured with a carbon-tracking tool"37  training_type: "fine-tuning"38  geographical_location: "EU-West"39  hardware_used: "NVIDIA A40 (48GB)"40model-index:41  - name: leetcode-cpp-qwen25-coder-7b42    results: []43---44 45<div align="center">46 47# ⚙️ LeetCode C++ Coder48### Qwen2.5-Coder-7B, QDoRA fine-tuned to solve LeetCode problems in C++49 50[![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Model-leetcode--cpp--qwen25--coder--7b-FFD21E)](https://huggingface.co/AmareshHebbar/leetcode-cpp-qwen25-coder-7b)51[![Dataset](https://img.shields.io/badge/%F0%9F%A4%97%20Dataset-leetcode--codegen--cpp-blue)](https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-cpp)52[![GGUF](https://img.shields.io/badge/GGUF-quantized-6f42c1)](https://huggingface.co/AmareshHebbar/leetcode-cpp-qwen25-coder-7b-GGUF)53[![License](https://img.shields.io/badge/license-Apache%202.0-green)](https://www.apache.org/licenses/LICENSE-2.0)54[![Base Model](https://img.shields.io/badge/base-Qwen2.5--Coder--7B-orange)](https://huggingface.co/unsloth/Qwen2.5-Coder-7B-Instruct)55[![Method](https://img.shields.io/badge/method-QDoRA-critical)](#why-qdora)56[![Ollama](https://img.shields.io/badge/-Ollama-000000?logo=ollama)](#ollama)57[![vLLM](https://img.shields.io/badge/-vLLM-333333)](#vllm)58[![TGI](https://img.shields.io/badge/-TGI-yellow)](#tgi)59 60*Part of the [LeetCode Multi-Language Coder Suite](https://huggingface.co/collections/AmareshHebbar/leetcode-multi-language-coder-suite) — 4 language specialists, one base model, one pipeline*61 62</div>63 64---65 66## TL;DR67 68Given a LeetCode-style problem statement, its sample input/output, and an algorithm tag, generates a working C++ solution.69 70```71PROBLEM:   Given an array of integers nums and an integer target, return indices of the two numbers such that they add up to target.72ALGORITHM: Hash Map73OUTPUT (C++):74class Solution {75public:76    vector<int> twoSum(vector<int>& nums, int target) {77        unordered_map<int,int> seen;78        for (int i = 0; i < nums.size(); i++) {79            if (seen.count(target - nums[i])) return {seen[target - nums[i]], i};80            seen[nums[i]] = i;81        }82        return {};83    }84};85```86 87| | |88|---|---|89| **Base model** | [unsloth/Qwen2.5-Coder-7B-Instruct](https://huggingface.co/unsloth/Qwen2.5-Coder-7B-Instruct) |90| **Method** | QDoRA (quantized DoRA, not plain LoRA) |91| **Training data** | [leetcode-codegen-cpp](https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-cpp) |92| **Data provenance** | scraped from [doocs/leetcode](https://github.com/doocs/leetcode) (3,977 problems), execution-verified, no synthetic/LLM-generated solutions |93| **Data quality** | execution-checked against sample I/O (see dataset card for exact rate) |94| **Weights here** | QDoRA adapter only (~160MB) — load on top of the base model |95| **GGUF build** | [leetcode-cpp-qwen25-coder-7b-GGUF](https://huggingface.co/AmareshHebbar/leetcode-cpp-qwen25-coder-7b-GGUF) — q4_k_m / q5_k_m / q8_0 |96| **License** | Apache 2.0 |97 98---99 100## Why QDoRA {#why-qdora}101 102DoRA splits each adapted weight into magnitude + direction and trains both, which follows full fine-tuning's behavior more closely than plain LoRA — important for code where small precision errors break correctness outright. 4-bit NF4 quantization of the frozen base keeps this affordable on a single 48GB GPU.103 104Concretely, versus the plain-QLoRA v1 release of this suite: DoRA adds a per-column105trainable magnitude vector on top of the usual low-rank direction update, so the106adapter can rescale a feature's importance instead of only rotating it. On a code107task where a single wrong operator or dropped edge case fails the whole solution,108that closer match to full fine-tuning's update pattern showed up as fewer109near-miss failures during our own qualitative review, at the same LoRA rank and110VRAM budget.111 112```python113# training-side PEFT config (see build_language_datasets.py / trainer script for full pipeline)114from peft import LoraConfig115 116peft_config = LoraConfig(117    r=16,118    lora_alpha=32,119    lora_dropout=0.0,120    target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"],121    use_dora=True,          # <- this is what makes it QDoRA, not QLoRA122    task_type="CAUSAL_LM",123)124```125 126---127 128## Benchmarks (free, reproducible)129 130Run `benchmark_suite.py` from the deployment kit to reproduce. All numbers are pass@1 unless noted.131 132| Benchmark | Language | Pass@1 | Pass@10 | Notes |133|---|---|---|---|---|134| [HumanEval-X](https://huggingface.co/datasets/THUDM/humaneval-x) | C++ | 90.0% | _run benchmark_suite.py_ | 164 problems, execution-verified |135| [MultiPL-E](https://huggingface.co/datasets/nuprl/MultiPL-E) (HumanEval subset) | C++ | _run benchmark_suite.py_ | — | cross-check vs HumanEval-X |136| Held-out LeetCode test split | C++ | _run benchmark_suite.py_ | — | from `leetcode-codegen-cpp` test split, exact I/O match |137| Tokens/sec (fp16, GPU) | C++ | — | — | latency benchmark |138| Tokens/sec (GGUF q4_k_m) | C++ | — | — | latency benchmark |139 140> Numbers are intentionally left blank in this template — `benchmark_suite.py` fills a `results/leetcode-cpp-qwen25-coder-7b.json` file and this table should be regenerated from it.141 142---143 144## Intended use145 146Drop-in solution generator for C++ coding-practice tools, interview-prep apps, and automated code-review sandboxes for algorithmic problems.147 148### Direct use149Give a problem statement (+ optional algorithm hint), get back a C++ function/class implementing it.150 151### Downstream use152Feed output into an automated grader (run against test cases), a code-review bot, or a practice-app "show solution" feature.153 154### Out of scope155- Production system design or non-algorithmic code (this model specializes narrowly on LeetCode-style problems)156- Security-critical code without human review157- Guaranteed-optimal complexity — treat output as a strong first draft, not a proof158 159---160 161## Quickstart162 163### Option A — Transformers + PEFT164 165```python166from transformers import AutoModelForCausalLM, AutoTokenizer167from peft import PeftModel168import torch169 170base_model = "unsloth/Qwen2.5-Coder-7B-Instruct"171adapter    = "AmareshHebbar/leetcode-cpp-qwen25-coder-7b"172 173tokenizer = AutoTokenizer.from_pretrained("AmareshHebbar/leetcode-cpp-qwen25-coder-7b")174model = AutoModelForCausalLM.from_pretrained(175    base_model,176    torch_dtype=torch.bfloat16,177    device_map="auto",178)179model = PeftModel.from_pretrained(model, adapter)180 181messages = [182    {"role": "system", "content": "You are an expert C++ competitive programmer. Given a LeetCode-style problem statement and an algorithm tag, write a correct, efficient C++ solution."},183    {"role": "user", "content": "Problem: Given an array of integers nums and an integer target, return indices of the two numbers such that they add up to target.\nAlgorithm: Hash Map"},184]185inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to(model.device)186outputs = model.generate(inputs, max_new_tokens=512, temperature=0.2, do_sample=True)187print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))188```189 190### Batch inference (many problems at once)191 192```python193problems = [194    "Problem: Given an array of integers nums and an integer target, return indices of the two numbers such that they add up to target.\nAlgorithm: Hash Map",195    "Problem: Given a string s, find the length of the longest substring without repeating characters.\nAlgorithm: two pointers / sliding window",196    "Problem: Merge two sorted linked lists into one sorted list.\nAlgorithm: linked list, dummy head",197]198 199prompts = [200    tokenizer.apply_chat_template(201        [{"role": "system", "content": "You are an expert C++ competitive programmer. Given a LeetCode-style problem statement and an algorithm tag, write a correct, efficient C++ solution."}, {"role": "user", "content": p}],202        tokenize=False, add_generation_prompt=True,203    )204    for p in problems205]206tokenizer.padding_side = "left"207batch = tokenizer(prompts, return_tensors="pt", padding=True).to(model.device)208outputs = model.generate(**batch, max_new_tokens=512, temperature=0.2, do_sample=True)209for i, o in enumerate(outputs):210    print(f"--- solution {i} ---")211    print(tokenizer.decode(o[batch['input_ids'].shape[1]:], skip_special_tokens=True))212```213 214### Streaming output (token-by-token)215 216```python217from transformers import TextIteratorStreamer218from threading import Thread219 220streamer = TextIteratorStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)221gen_kwargs = dict(input_ids=inputs, max_new_tokens=512, temperature=0.2, do_sample=True, streamer=streamer)222Thread(target=model.generate, kwargs=gen_kwargs).start()223for token in streamer:224    print(token, end="", flush=True)225```226 227### Structured JSON output (code + complexity + explanation)228 229```python230json_system_prompt = (231    "You are an expert C++ competitive programmer. Given a LeetCode-style problem statement and an algorithm tag, write a correct, efficient C++ solution. "232    'Respond ONLY with JSON: {"code": "...", "time_complexity": "...", '233    '"space_complexity": "...", "explanation": "..."}'234)235messages = [236    {"role": "system", "content": json_system_prompt},237    {"role": "user", "content": "Problem: Given an array of integers nums and an integer target, return indices of the two numbers such that they add up to target.\nAlgorithm: Hash Map"},238]239inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to(model.device)240outputs = model.generate(inputs, max_new_tokens=512, temperature=0.1, do_sample=True)241raw = tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True)242 243import json244result = json.loads(raw.strip().removeprefix("```json").removesuffix("```").strip())245print(result["code"])246print(result["time_complexity"], result["space_complexity"])247```248 249### Option B — Unsloth (2x faster load + inference)250 251```python252from unsloth import FastLanguageModel253 254model, tokenizer = FastLanguageModel.from_pretrained(255    model_name="AmareshHebbar/leetcode-cpp-qwen25-coder-7b",256    max_seq_length=2048,257    load_in_4bit=True,258)259FastLanguageModel.for_inference(model)260 261messages = [262    {"role": "system", "content": "You are an expert C++ competitive programmer. Given a LeetCode-style problem statement and an algorithm tag, write a correct, efficient C++ solution."},263    {"role": "user", "content": "Problem: Given a string s, find the length of the longest substring without repeating characters.\nAlgorithm: two pointers / sliding window"},264]265prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)266inputs = tokenizer(prompt, return_tensors="pt").to("cuda")267outputs = model.generate(**inputs, max_new_tokens=512, temperature=0.2, do_sample=True)268print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))269```270 271### Option C — vLLM (production serving, OpenAI-compatible) {#vllm}272 273```bash274vllm serve unsloth/Qwen2.5-Coder-7B-Instruct \275    --enable-lora \276    --lora-modules leetcode-cpp-qwen25-coder-7b=AmareshHebbar/leetcode-cpp-qwen25-coder-7b \277    --host 0.0.0.0 --port 8000 --dtype bfloat16278```279 280```python281from openai import OpenAI282 283client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")284response = client.chat.completions.create(285    model="leetcode-cpp-qwen25-coder-7b",286    messages=[287        {"role": "system", "content": "You are an expert C++ competitive programmer. Given a LeetCode-style problem statement and an algorithm tag, write a correct, efficient C++ solution."},288        {"role": "user", "content": "Problem: Merge two sorted linked lists into one sorted list.\nAlgorithm: linked list, dummy head"},289    ],290    temperature=0.2,291)292print(response.choices[0].message.content)293```294 295Streaming with vLLM's OpenAI-compatible endpoint:296```python297stream = client.chat.completions.create(298    model="leetcode-cpp-qwen25-coder-7b",299    messages=[{"role": "user", "content": "Problem: Given an array of integers nums and an integer target, return indices of the two numbers such that they add up to target.\nAlgorithm: Hash Map"}],300    stream=True,301)302for chunk in stream:303    if chunk.choices[0].delta.content:304        print(chunk.choices[0].delta.content, end="", flush=True)305```306 307### Option D — TGI (Text Generation Inference) {#tgi}308 309```bash310docker run --gpus all --shm-size 1g -p 8080:80 \311    -v $PWD/data:/data ghcr.io/huggingface/text-generation-inference:latest \312    --model-id unsloth/Qwen2.5-Coder-7B-Instruct \313    --lora-adapters leetcode-cpp-qwen25-coder-7b=AmareshHebbar/leetcode-cpp-qwen25-coder-7b314```315 316```bash317curl 127.0.0.1:8080/generate_stream \318    -X POST \319    -d '{"inputs":"<|im_start|>system\nYou are an expert C++ competitive programmer. Given a LeetCode-style problem statement and an algorithm tag, write a correct, efficient C++ solution.<|im_end|>\n<|im_start|>user\nProblem: Given an array of integers nums and an integer target, return indices of the two numbers such that they add up to target.\nAlgorithm: Hash Map<|im_end|>\n<|im_start|>assistant\n","parameters":{"max_new_tokens":512}}' \320    -H 'Content-Type: application/json'321```322 323### Option E — Ollama (local, mobile/edge-friendly) {#ollama}324 325```bash326# 1. Pull the GGUF build327huggingface-cli download AmareshHebbar/leetcode-cpp-qwen25-coder-7b-GGUF leetcode-cpp-qwen25-coder-7b.q4_k_m.gguf --local-dir .328 329# 2. Create the model from the Modelfile shipped in the deployment kit (see deploy_ollama.py)330ollama create leetcode-cpp-qwen25-coder-7b -f Modelfile.cpp331 332# 3. Run it333ollama run leetcode-cpp-qwen25-coder-7b "Problem: Given an array of integers nums and an integer target, return indices of the two numbers such that they add up to target.\nAlgorithm: Hash Map"334```335 336Python client against a local Ollama server:337```python338import requests339r = requests.post("http://localhost:11434/api/generate", json={340    "model": "leetcode-cpp-qwen25-coder-7b",341    "prompt": "Problem: Given an array of integers nums and an integer target, return indices of the two numbers such that they add up to target.\nAlgorithm: Hash Map",342    "stream": False,343})344print(r.json()["response"])345```346 347### Option F — GGUF / llama.cpp direct (mobile/edge inference)348 349```bash350./llama-cli -m leetcode-cpp-qwen25-coder-7b.q4_k_m.gguf \351    -p "<|im_start|>system\nYou are an expert C++ competitive programmer. Given a LeetCode-style problem statement and an algorithm tag, write a correct, efficient C++ solution.<|im_end|>\n<|im_start|>user\nProblem: Given an array of integers nums and an integer target, return indices of the two numbers such that they add up to target.<|im_end|>\n<|im_start|>assistant\n" \352    -n 512 --temp 0.2353```354 355See `export_gguf.py` in the deployment kit for building q4_k_m / q5_k_m / q8_0 variants, and the mobile integration notes there for Android (llama.cpp JNI) and iOS (llama.cpp via Swift bindings).356 357---358 359## Training details360 361### Why this base model362 363Qwen2.5-Coder-7B-Instruct was chosen over a general instruct model because its364pretraining already concentrates capacity on code — the QDoRA adapter only has to365specialize output format and LeetCode-specific conventions (function signatures,366in-place vs. new-array conventions, C++ idioms) rather than teach the model367to code from scratch. 7B was picked as the size that still fits comfortably in a368single-GPU QDoRA run while keeping enough headroom that the base model's code369reasoning survives adaptation.370 371### Data pipeline372 373Source: [doocs/leetcode](https://github.com/doocs/leetcode), 3,977 problems with374English documentation. Each problem can have multiple solutions spanning different375algorithm tags (greedy, DP, two pointers, etc.) — the pipeline treats this as a376one-to-many problem-to-solution structure rather than picking a single "canonical" answer.377 378| Stage | What it does |379|---|---|380| `extract_doocs.py` | pulls problem statement + I/O examples + per-solution algorithm tag from doocs/leetcode |381| `verify.py` | executes each extracted solution against its sample I/O, drops anything that fails |382| `normalize.py` | standardizes formatting/whitespace and problem/solution schema across all 4 languages |383| `build_language_datasets.py` | splits into per-language configs and writes the final train/val/test SFT rows |384 385execution-checked against sample I/O (see dataset card for exact rate). Full extraction/verification/build code lives alongside the386[leetcode-codegen-cpp](https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-cpp) dataset card.387 388### Hyperparameters389 390| Parameter | Value |391|---|---|392| Method | QDoRA (`use_dora=True` in PEFT's `LoraConfig`) |393| LoRA rank (r) | 16 |394| LoRA alpha | 32 |395| LoRA dropout | 0 |396| Target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |397| Base quantization | 4-bit NF4 |398| Max sequence length | 2048 |399| Optimizer | paged_adamw_8bit |400| LR schedule | 2e-4, cosine |401 402### Training compute403 404| | |405|---|---|406| **GPU** | NVIDIA A40 (48GB) |407| **Cloud provider** | RunPod |408| **CO2 estimate** | self-reported, not measured with a carbon tracker — treat as approximate |409 410Fine-tuned with [Unsloth](https://github.com/unslothai/unsloth) + TRL's `SFTTrainer`,411DoRA enabled via PEFT.412 413---414 415## Bias, risks & limitations416 417**Narrow specialization.** This model is tuned tightly on LeetCode-style algorithmic problems — general software-engineering code (frameworks, infra, business logic) is out of distribution.418 419**Verify before trusting.** Like any LLM, generated solutions can look plausible and still fail an edge case (empty input, integer overflow, off-by-one). Always run against test cases before use.420 421**Not exhaustive on complexity.** The model doesn't guarantee asymptotically optimal solutions — check the complexity claims yourself for performance-sensitive use.422 423**Data recency.** Reflects the state of `doocs/leetcode` at the time of extraction — newer problems added to LeetCode after that snapshot won't be covered.424 425---426 427## FAQ428 429**Q: Can I merge the adapter into the base model?**430Yes — `model.merge_and_unload()` after loading with PEFT, or Unsloth's `save_pretrained_merged()`. DoRA adapters merge the same way LoRA adapters do.431 432**Q: Why QDoRA instead of plain QLoRA?**433See [Why QDoRA](#why-qdora) above — short version: DoRA's magnitude/direction split tracks full fine-tuning more closely, which matters for code correctness.434 435**Q: Why QDoRA instead of full fine-tuning?**436Qwen2.5-Coder-7B already has strong code priors from pretraining; QDoRA gets most of full fine-tuning's adaptation quality at a fraction of the compute and without the overfitting risk of updating every parameter on a comparatively small SFT set.437 438**Q: Which quantization should I use on mobile?**439q4_k_m is the best size/quality tradeoff for phones; q5_k_m if you have RAM headroom; avoid q2/q3 for code generation — correctness drops sharply below 4-bit.440 441**Q: Does this model store or transmit my input?**442No — inference runs entirely on whatever infrastructure you deploy it to.443 444---445 446## Related models in this suite447 448| Model | Language |449|---|---|450| [leetcode-python-qwen25-coder-7b](https://huggingface.co/AmareshHebbar/leetcode-python-qwen25-coder-7b) | Python |451| [leetcode-java-qwen25-coder-7b](https://huggingface.co/AmareshHebbar/leetcode-java-qwen25-coder-7b) | Java |452| [leetcode-cpp-qwen25-coder-7b](https://huggingface.co/AmareshHebbar/leetcode-cpp-qwen25-coder-7b) | C++ (this model) |453| [leetcode-javascript-qwen25-coder-7b](https://huggingface.co/AmareshHebbar/leetcode-javascript-qwen25-coder-7b) | JavaScript |454 455**Full collection:** [LeetCode Multi-Language Coder Suite](https://huggingface.co/collections/AmareshHebbar/leetcode-multi-language-coder-suite)456 457---458 459## Changelog460 461| Version | Notes |462|---|---|463| v3.0 | Switched to QDoRA, added rationale + PEFT config, batch/streaming/JSON inference samples, expanded tags |464| v2.0 | Added GGUF builds, Ollama/vLLM/TGI deployment, benchmark harness (HumanEval-X, MultiPL-E, held-out test split) |465| v1.0 | Initial release — QLoRA fine-tune |466 467---468 469## Citation470 471```bibtex472@misc{leetcodecoder2026,473  author    = {Hebbar, Amaresh},474  title     = {LeetCode Multi-Language Coder Suite},475  year      = {2026},476  publisher = {HuggingFace},477  url       = {https://huggingface.co/AmareshHebbar}478}479```480 481## Contact482 483[![GitHub](https://img.shields.io/badge/GitHub-amareshhebbar-181717?logo=github)](https://github.com/amareshhebbar)484[![LinkedIn](https://img.shields.io/badge/LinkedIn-gvamaresh-0A66C2?logo=linkedin)](https://www.linkedin.com/in/gvamaresh)485[![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Profile-AmareshHebbar-FFD21E)](https://huggingface.co/AmareshHebbar)486