Maaac/CodeLLaMA-Linux-BugFix
08
1---2license: mit3tags:4 - codellama5 - linux6 - bugfix7 - lora8 - qlora9 - git-diff10base_model: codellama/CodeLLaMA-7b-Instruct-hf11model_type: LlamaForCausalLM12library_name: peft13pipeline_tag: text-generation14 15model-index:16- name: CodeLLaMA-Linux-BugFix17 results:18 - task:19 type: text-generation20 name: Bug-fix Patch Generation21 dataset:22 type: custom23 name: Linux Kernel Bugfix Commits24 config: linux-bugfix-prompt-completion25 split: test26 metrics:27 - type: bleu28 value: 33.8729 name: BLEU30 - type: rouge131 value: 0.435532 name: ROUGE-1 F133 - type: rouge234 value: 0.345735 name: ROUGE-2 F136 - type: rougeL37 value: 0.361238 name: ROUGE-L F139---40 41 # CodeLLaMA-Linux-BugFix42 43 A fine-tuned version of `CodeLLaMA-7B-Instruct`, designed specifically for Linux kernel bug fixing using QLoRA (Quantized Low-Rank Adaptation). The model learns to generate Git diff patches based on buggy C code and commit messages.44 45 ---46 47 ## ๐ฏ Overview48 49 This project targets automated Linux kernel bug fixing by:50 51 - **Mining real commit data** from the kernel Git history52 - **Training a specialized QLoRA model** on diff-style fixes53 - **Generating Git patches** in response to bug-prone code54 - **Evaluating results** using BLEU, ROUGE, and human inspection55 56 The model achieves strong performance in generating accurate Linux kernel bug fixes, making it a valuable tool for automated code review and bug detection.57 58 ---59 60 ## ๐ Performance Results61 62 ### Evaluation Metrics63 64 โ
**BLEU Score**: 33.8765 66 โ
**ROUGE Scores**:67 - **ROUGE-1**: P=0.3775, R=0.7306, F1=0.435568 - **ROUGE-2**: P=0.2898, R=0.6096, F1=0.345769 - **ROUGE-L**: P=0.3023, R=0.6333, F1=0.361270 71 These results demonstrate the model's ability to:72 - Generate syntactically correct Git diff patches73 - Maintain semantic similarity to reference fixes74 - Produce meaningful code changes that address the underlying bugs75 76 ---77 78 ## ๐ง Model Configuration79 80 - **Base model**: `CodeLLaMA-7B-Instruct`81 - **Fine-tuning method**: QLoRA with 4-bit quantization82 - **Training setup**:83 - LoRA r=64, alpha=16, dropout=0.184 - Batch size: 64, LR: 2e-4, Epochs: 385 - Mixed precision (bfloat16), gradient checkpointing86 - **Hardware**: Optimized for NVIDIA H200 GPUs87 88 ---89 90 ## ๐ Training Progress91 The model was trained for 1000 steps with the following key metrics:92 ### Training Results93 - **Final Loss**: ~0.3335 (converged)94 - **Final Learning Rate**: 2.08304527802282E-0695 - **Training Steps**: 100096 - **Convergence**: Stable loss plateau achieved97 ### Training Curves98 99 *Training loss over 1000 steps showing convergence around 0.3335*100 101 *Learning rate decay schedule with final rate of 2.08304527802282E-06*102 103 ---104 105 ## ๐ Dataset106 107 Custom dataset extracted from Linux kernel Git history.108 109 ### Filtering Criteria110 Bug-fix commits containing:111 `fix`, `bug`, `crash`, `memory`, `null`, `panic`, `overflow`, `race`, `corruption`, etc.112 113 ### Structure114 - Language: C (`.c`, `.h`)115 - Context: 10 lines before/after the change116 - Format:117 118 ```json119 {120 "input": {121 "original code": "C code snippet with bug",122 "instruction": "Commit message or fix description"123 },124 "output": {125 "diff codes": "Git diff showing the fix"126 }127 }128 ```129 130 * **File**: `training_data_100k.jsonl` (100,000 samples)131 132 ---133 134 ## ๐ Quick Start135 136 ### Prerequisites137 138 - Python 3.8+139 - CUDA-compatible GPU (recommended)140 - 16GB+ RAM141 - 50GB+ disk space142 143 ### Install dependencies144 145 ```bash146 pip install -r requirements.txt147 ```148 149 ### 1. Build the Dataset150 151 ```bash152 cd dataset_builder153 python extract_linux_bugfixes_parallel.py154 python format_for_training.py155 ```156 157 ### 2. Fine-tune the Model158 159 ```bash160 cd train161 python train_codellama_qlora_linux_bugfix.py162 ```163 164 ### 3. Run Evaluation165 166 ```bash167 cd evaluate168 python evaluate_linux_bugfix_model.py169 ```170 171 ### 4. Use the Model172 173 ```python174 from transformers import AutoTokenizer, AutoModelForCausalLM175 from peft import PeftModel176 177 # Load the fine-tuned model178 model = AutoModelForCausalLM.from_pretrained("codellama/CodeLLaMA-7b-Instruct-hf")179 model = PeftModel.from_pretrained(model, "train/output/qlora-codellama-bugfix")180 tokenizer = AutoTokenizer.from_pretrained("codellama/CodeLLaMA-7b-Instruct-hf")181 182 # Generate a bug fix183 prompt = """184 Given the following original C code:185 if (!file->filter)186 return;187 188 Instruction: Fix the null pointer dereference189 190 Return the diff that fixes it:191 """192 193 inputs = tokenizer(prompt, return_tensors="pt")194 outputs = model.generate(**inputs, max_length=512, temperature=0.1)195 fix = tokenizer.decode(outputs[0], skip_special_tokens=True)196 print(fix)197 ```198 199 ---200 201 ## ๐ Project Structure202 203 ```204 CodeLLaMA-Linux-BugFix/205 โโโ dataset_builder/206 โ โโโ extract_linux_bugfixes_parallel.py # Parallel extraction of bug fixes207 โ โโโ format_for_training.py # Format data for training208 โ โโโ build_dataset.py # Main dataset builder209 โโโ dataset/210 โ โโโ training_data_100k.jsonl # 100K training samples211 โ โโโ training_data_prompt_completion.jsonl # Formatted training data212 โโโ train/213 โ โโโ train_codellama_qlora_linux_bugfix.py # Main training script214 โ โโโ train_codellama_qlora_simple.py # Simplified training215 โ โโโ download_codellama_model.py # Model download utility216 โ โโโ output/217 โ โโโ qlora-codellama-bugfix/ # Trained model checkpoints218 โโโ evaluate/219 โ โโโ evaluate_linux_bugfix_model.py # Evaluation script220 โ โโโ test_samples.jsonl # Test dataset221 โ โโโ output/ # Evaluation results222 โ โโโ eval_results.csv # Detailed results223 โ โโโ eval_results.json # JSON format results224 โโโ requirements.txt # Python dependencies225 โโโ README.md # This file226 โโโ PROJECT_STRUCTURE.md # Detailed project overview227 ```228 229 ---230 231 ## ๐งฉ Features232 233 * ๐ง **Efficient Fine-tuning**: QLoRA + 4-bit quant = massive memory savings234 * ๐ง **Real-world commits**: From actual Linux kernel development235 * ๐ก **Context-aware**: Code context extraction around bug lines236 * ๐ป **Output-ready**: Generates valid Git-style diffs237 * ๐ **Strong Performance**: BLEU score of 33.87 with good ROUGE metrics238 * ๐ **Production-ready**: Optimized for real-world deployment239 240 ---241 242 ## ๐ Evaluation Metrics243 244 * **BLEU**: Translation-style match to reference diffs245 * **ROUGE**: Overlap in fix content and semantic similarity246 * **Human Evaluation**: Subjective patch quality assessment247 248 ### Current Performance249 - **BLEU Score**: 33.87 (excellent for code generation tasks)250 - **ROUGE-1 F1**: 0.4355 (good semantic overlap)251 - **ROUGE-2 F1**: 0.3457 (reasonable bigram matching)252 - **ROUGE-L F1**: 0.3612 (good longest common subsequence)253 254 ---255 256 ## ๐งช Use Cases257 258 * **Automated kernel bug fixing**: Generate fixes for common kernel bugs259 * **Code review assistance**: Help reviewers identify potential issues260 * **Teaching/debugging kernel code**: Educational tool for kernel development261 * **Research in automated program repair (APR)**: Academic research applications262 * **CI/CD integration**: Automated testing and fixing in development pipelines263 264 ---265 266 ## ๐ฌ Technical Highlights267 268 ### Memory & Speed Optimizations269 270 * 4-bit quantization (NF4)271 * Gradient checkpointing272 * Mixed precision (bfloat16)273 * Gradient accumulation274 * LoRA parameter efficiency275 276 ### Training Efficiency277 278 * **QLoRA**: Reduces memory usage by ~75%279 * **4-bit quantization**: Further memory optimization280 * **Gradient checkpointing**: Trades compute for memory281 * **Mixed precision**: Faster training with maintained accuracy282 283 ---284 285 ## ๐ ๏ธ Advanced Usage286 287 ### Custom Training288 289 ```bash290 # Train with custom parameters291 python train_codellama_qlora_linux_bugfix.py \292 --learning_rate 1e-4 \293 --num_epochs 5 \294 --batch_size 32 \295 --lora_r 32 \296 --lora_alpha 16297 ```298 299 ### Evaluation on Custom Data300 301 ```bash302 # Evaluate on your own test set303 python evaluate_linux_bugfix_model.py \304 --test_file your_test_data.jsonl \305 --output_dir custom_eval_results306 ```307 308 ---309 310 ## ๐ค Contributing311 312 1. Fork this repo313 2. Create a feature branch (`git checkout -b feature/amazing-feature`)314 3. Commit your changes (`git commit -m 'Add amazing feature'`)315 4. Push to the branch (`git push origin feature/amazing-feature`)316 5. Open a Pull Request ๐317 318 ### Development Guidelines319 320 - Follow PEP 8 style guidelines321 - Add tests for new features322 - Update documentation for API changes323 - Ensure all tests pass before submitting PR324 325 ---326 327 ## ๐ License328 329 MIT License โ see `LICENSE` file for details.330 331 ---332 333 ## ๐ Acknowledgments334 335 * **Meta** for CodeLLaMA base model336 * **Hugging Face** for Transformers + PEFT libraries337 * **The Linux kernel community** for open access to commit data338 * **Microsoft** for introducing LoRA technique339 * **University of Washington** for QLoRA research340 341 ---342 343 ## ๐ References344 345 * [CodeLLaMA (Meta, 2023)](https://arxiv.org/abs/2308.12950)346 * [QLoRA (Dettmers et al., 2023)](https://arxiv.org/abs/2305.14314)347 * [LoRA (Hu et al., 2021)](https://arxiv.org/abs/2106.09685)348 * [Automated Program Repair: A Survey](https://ieeexplore.ieee.org/document/8449519)349 350 ---351 352 ## ๐ Support353 354 For questions, issues, or contributions:355 - Open an issue on GitHub356 - Check the project documentation357 - Review the evaluation results in `evaluate/output/`358 359 ---360 361 ## ๐ Version History362 363 - **v1.0.0**: Initial release with QLoRA training364 - **v1.1.0**: Added parallel dataset extraction365 - **v1.2.0**: Improved evaluation metrics and documentation366=======367---368license: mit369tags:370 - codellama371 - linux372 - bugfix373 - lora374 - qlora375 - git-diff376base_model: codellama/CodeLLaMA-7b-Instruct-hf377model_type: LlamaForCausalLM378library_name: peft379pipeline_tag: text-generation380---381 382# CodeLLaMA-Linux-BugFix383 384A fine-tuned version of `CodeLLaMA-7B-Instruct`, designed specifically for Linux kernel bug fixing using QLoRA (Quantized Low-Rank Adaptation). The model learns to generate Git diff patches based on buggy C code and commit messages.385 386---387 388## ๐ฏ Overview389 390This project targets automated Linux kernel bug fixing by:391 392- **Mining real commit data** from the kernel Git history393- **Training a specialized QLoRA model** on diff-style fixes394- **Generating Git patches** in response to bug-prone code395- **Evaluating results** using BLEU, ROUGE, and human inspection396 397The model achieves strong performance in generating accurate Linux kernel bug fixes, making it a valuable tool for automated code review and bug detection.398 399---400 401## ๐ Performance Results402 403### Evaluation Metrics404 405โ
**BLEU Score**: 33.87406 407โ
**ROUGE Scores**:408- **ROUGE-1**: P=0.3775, R=0.7306, F1=0.4355409- **ROUGE-2**: P=0.2898, R=0.6096, F1=0.3457410- **ROUGE-L**: P=0.3023, R=0.6333, F1=0.3612411 412These results demonstrate the model's ability to:413- Generate syntactically correct Git diff patches414- Maintain semantic similarity to reference fixes415- Produce meaningful code changes that address the underlying bugs416 417---418 419## ๐ง Model Configuration420 421- **Base model**: `CodeLLaMA-7B-Instruct`422- **Fine-tuning method**: QLoRA with 4-bit quantization423- **Training setup**:424 - LoRA r=64, alpha=16, dropout=0.1425 - Batch size: 64, LR: 2e-4, Epochs: 3426 - Mixed precision (bfloat16), gradient checkpointing427- **Hardware**: Optimized for NVIDIA H200 GPUs428 429---430 431## ๐ Dataset432 433Custom dataset extracted from Linux kernel Git history.434 435### Filtering Criteria436Bug-fix commits containing:437`fix`, `bug`, `crash`, `memory`, `null`, `panic`, `overflow`, `race`, `corruption`, etc.438 439### Structure440- Language: C (`.c`, `.h`)441- Context: 10 lines before/after the change442- Format:443 444```json445{446 "input": {447 "original code": "C code snippet with bug",448 "instruction": "Commit message or fix description"449 },450 "output": {451 "diff codes": "Git diff showing the fix"452 }453}454```455 456* **File**: `training_data_100k.jsonl` (100,000 samples)457 458---459 460## ๐ Quick Start461 462### Prerequisites463 464- Python 3.8+465- CUDA-compatible GPU (recommended)466- 16GB+ RAM467- 50GB+ disk space468 469### Install dependencies470 471```bash472pip install -r requirements.txt473```474 475### 1. Build the Dataset476 477```bash478cd dataset_builder479python extract_linux_bugfixes_parallel.py480python format_for_training.py481```482 483### 2. Fine-tune the Model484 485```bash486cd train487python train_codellama_qlora_linux_bugfix.py488```489 490### 3. Run Evaluation491 492```bash493cd evaluate494python evaluate_linux_bugfix_model.py495```496 497### 4. Use the Model498 499```python500from transformers import AutoTokenizer, AutoModelForCausalLM501from peft import PeftModel502 503# Load the fine-tuned model504model = AutoModelForCausalLM.from_pretrained("codellama/CodeLLaMA-7b-Instruct-hf")505model = PeftModel.from_pretrained(model, "train/output/qlora-codellama-bugfix")506tokenizer = AutoTokenizer.from_pretrained("codellama/CodeLLaMA-7b-Instruct-hf")507 508# Generate a bug fix509prompt = """510Given the following original C code:511if (!file->filter)512 return;513 514Instruction: Fix the null pointer dereference515 516Return the diff that fixes it:517"""518 519inputs = tokenizer(prompt, return_tensors="pt")520outputs = model.generate(**inputs, max_length=512, temperature=0.1)521fix = tokenizer.decode(outputs[0], skip_special_tokens=True)522print(fix)523```524 525---526 527## ๐ Project Structure528 529```530CodeLLaMA-Linux-BugFix/531โโโ dataset_builder/532โ โโโ extract_linux_bugfixes_parallel.py # Parallel extraction of bug fixes533โ โโโ format_for_training.py # Format data for training534โ โโโ build_dataset.py # Main dataset builder535โโโ dataset/536โ โโโ training_data_100k.jsonl # 100K training samples537โ โโโ training_data_prompt_completion.jsonl # Formatted training data538โโโ train/539โ โโโ train_codellama_qlora_linux_bugfix.py # Main training script540โ โโโ train_codellama_qlora_simple.py # Simplified training541โ โโโ download_codellama_model.py # Model download utility542โ โโโ output/543โ โโโ qlora-codellama-bugfix/ # Trained model checkpoints544โโโ evaluate/545โ โโโ evaluate_linux_bugfix_model.py # Evaluation script546โ โโโ test_samples.jsonl # Test dataset547โ โโโ output/ # Evaluation results548โ โโโ eval_results.csv # Detailed results549โ โโโ eval_results.json # JSON format results550โโโ requirements.txt # Python dependencies551โโโ README.md # This file552โโโ PROJECT_STRUCTURE.md # Detailed project overview553```554 555---556 557## ๐งฉ Features558 559* ๐ง **Efficient Fine-tuning**: QLoRA + 4-bit quant = massive memory savings560* ๐ง **Real-world commits**: From actual Linux kernel development561* ๐ก **Context-aware**: Code context extraction around bug lines562* ๐ป **Output-ready**: Generates valid Git-style diffs563* ๐ **Strong Performance**: BLEU score of 33.87 with good ROUGE metrics564* ๐ **Production-ready**: Optimized for real-world deployment565 566---567 568## ๐ Evaluation Metrics569 570* **BLEU**: Translation-style match to reference diffs571* **ROUGE**: Overlap in fix content and semantic similarity572* **Human Evaluation**: Subjective patch quality assessment573 574### Current Performance575- **BLEU Score**: 33.87 (excellent for code generation tasks)576- **ROUGE-1 F1**: 0.4355 (good semantic overlap)577- **ROUGE-2 F1**: 0.3457 (reasonable bigram matching)578- **ROUGE-L F1**: 0.3612 (good longest common subsequence)579 580---581 582## ๐งช Use Cases583 584* **Automated kernel bug fixing**: Generate fixes for common kernel bugs585* **Code review assistance**: Help reviewers identify potential issues586* **Teaching/debugging kernel code**: Educational tool for kernel development587* **Research in automated program repair (APR)**: Academic research applications588* **CI/CD integration**: Automated testing and fixing in development pipelines589 590---591 592## ๐ฌ Technical Highlights593 594### Memory & Speed Optimizations595 596* 4-bit quantization (NF4)597* Gradient checkpointing598* Mixed precision (bfloat16)599* Gradient accumulation600* LoRA parameter efficiency601 602### Training Efficiency603 604* **QLoRA**: Reduces memory usage by ~75%605* **4-bit quantization**: Further memory optimization606* **Gradient checkpointing**: Trades compute for memory607* **Mixed precision**: Faster training with maintained accuracy608 609---610 611## ๐ ๏ธ Advanced Usage612 613### Custom Training614 615```bash616# Train with custom parameters617python train_codellama_qlora_linux_bugfix.py \618 --learning_rate 1e-4 \619 --num_epochs 5 \620 --batch_size 32 \621 --lora_r 32 \622 --lora_alpha 16623```624 625### Evaluation on Custom Data626 627```bash628# Evaluate on your own test set629python evaluate_linux_bugfix_model.py \630 --test_file your_test_data.jsonl \631 --output_dir custom_eval_results632```633 634---635 636## ๐ค Contributing637 6381. Fork this repo6392. Create a feature branch (`git checkout -b feature/amazing-feature`)6403. Commit your changes (`git commit -m 'Add amazing feature'`)6414. Push to the branch (`git push origin feature/amazing-feature`)6425. Open a Pull Request ๐643 644### Development Guidelines645 646- Follow PEP 8 style guidelines647- Add tests for new features648- Update documentation for API changes649- Ensure all tests pass before submitting PR650 651---652 653## ๐ License654 655MIT License โ see `LICENSE` file for details.656 657---658 659## ๐ Acknowledgments660 661* **Meta** for CodeLLaMA base model662* **Hugging Face** for Transformers + PEFT libraries663* **The Linux kernel community** for open access to commit data664* **Microsoft** for introducing LoRA technique665* **University of Washington** for QLoRA research666 667---668 669## ๐ References670 671* [CodeLLaMA (Meta, 2023)](https://arxiv.org/abs/2308.12950)672* [QLoRA (Dettmers et al., 2023)](https://arxiv.org/abs/2305.14314)673* [LoRA (Hu et al., 2021)](https://arxiv.org/abs/2106.09685)674* [Automated Program Repair: A Survey](https://ieeexplore.ieee.org/document/8449519)675 676---677 678## ๐ Support679 680For questions, issues, or contributions:681- Open an issue on GitHub682- Check the project documentation683- Review the evaluation results in `evaluate/output/`684 685---686 687## ๐ Version History688 689- **v1.0.0**: Initial release with QLoRA training690- **v1.1.0**: Added parallel dataset extraction691- **v1.2.0**: Improved evaluation metrics and documentation692 