Team Ai
Modelpublic

Maaac/CodeLLaMA-Linux-BugFix

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes8downloads
README.md692 linesDownload Raw Back to root
1---2license: mit3tags:4  - codellama5  - linux6  - bugfix7  - lora8  - qlora9  - git-diff10base_model: codellama/CodeLLaMA-7b-Instruct-hf11model_type: LlamaForCausalLM12library_name: peft13pipeline_tag: text-generation14 15model-index:16- name: CodeLLaMA-Linux-BugFix17  results:18  - task:19      type: text-generation20      name: Bug-fix Patch Generation21    dataset:22      type: custom23      name: Linux Kernel Bugfix Commits24      config: linux-bugfix-prompt-completion25      split: test26    metrics:27      - type: bleu28        value: 33.8729        name: BLEU30      - type: rouge131        value: 0.435532        name: ROUGE-1 F133      - type: rouge234        value: 0.345735        name: ROUGE-2 F136      - type: rougeL37        value: 0.361238        name: ROUGE-L F139---40 41  # CodeLLaMA-Linux-BugFix42 43  A fine-tuned version of `CodeLLaMA-7B-Instruct`, designed specifically for Linux kernel bug fixing using QLoRA (Quantized Low-Rank Adaptation). The model learns to generate Git diff patches based on buggy C code and commit messages.44 45  ---46 47  ## ๐ŸŽฏ Overview48 49  This project targets automated Linux kernel bug fixing by:50 51  - **Mining real commit data** from the kernel Git history52  - **Training a specialized QLoRA model** on diff-style fixes53  - **Generating Git patches** in response to bug-prone code54  - **Evaluating results** using BLEU, ROUGE, and human inspection55 56  The model achieves strong performance in generating accurate Linux kernel bug fixes, making it a valuable tool for automated code review and bug detection.57 58  ---59 60  ## ๐Ÿ“Š Performance Results61 62  ### Evaluation Metrics63 64  โœ… **BLEU Score**: 33.8765 66  โœ… **ROUGE Scores**:67  - **ROUGE-1**: P=0.3775, R=0.7306, F1=0.435568  - **ROUGE-2**: P=0.2898, R=0.6096, F1=0.345769  - **ROUGE-L**: P=0.3023, R=0.6333, F1=0.361270 71  These results demonstrate the model's ability to:72  - Generate syntactically correct Git diff patches73  - Maintain semantic similarity to reference fixes74  - Produce meaningful code changes that address the underlying bugs75 76  ---77 78  ## ๐Ÿง  Model Configuration79 80  - **Base model**: `CodeLLaMA-7B-Instruct`81  - **Fine-tuning method**: QLoRA with 4-bit quantization82  - **Training setup**:83    - LoRA r=64, alpha=16, dropout=0.184    - Batch size: 64, LR: 2e-4, Epochs: 385    - Mixed precision (bfloat16), gradient checkpointing86  - **Hardware**: Optimized for NVIDIA H200 GPUs87 88  ---89 90  ## ๐Ÿ“ˆ Training Progress91  The model was trained for 1000 steps with the following key metrics:92  ### Training Results93  - **Final Loss**: ~0.3335 (converged)94  - **Final Learning Rate**: 2.08304527802282E-0695  - **Training Steps**: 100096  - **Convergence**: Stable loss plateau achieved97  ### Training Curves98  ![Training Loss](train/output/loss.png)99  *Training loss over 1000 steps showing convergence around 0.3335*100  ![Learning Rate Schedule](train/output/learning_rate.png)101  *Learning rate decay schedule with final rate of 2.08304527802282E-06*102 103  ---104 105  ## ๐Ÿ“Š Dataset106 107  Custom dataset extracted from Linux kernel Git history.108 109  ### Filtering Criteria110  Bug-fix commits containing:111  `fix`, `bug`, `crash`, `memory`, `null`, `panic`, `overflow`, `race`, `corruption`, etc.112 113  ### Structure114  - Language: C (`.c`, `.h`)115  - Context: 10 lines before/after the change116  - Format:117 118  ```json119  {120    "input": {121      "original code": "C code snippet with bug",122      "instruction": "Commit message or fix description"123    },124    "output": {125      "diff codes": "Git diff showing the fix"126    }127  }128  ```129 130  * **File**: `training_data_100k.jsonl` (100,000 samples)131 132  ---133 134  ## ๐Ÿš€ Quick Start135 136  ### Prerequisites137 138  - Python 3.8+139  - CUDA-compatible GPU (recommended)140  - 16GB+ RAM141  - 50GB+ disk space142 143  ### Install dependencies144 145  ```bash146  pip install -r requirements.txt147  ```148 149  ### 1. Build the Dataset150 151  ```bash152  cd dataset_builder153  python extract_linux_bugfixes_parallel.py154  python format_for_training.py155  ```156 157  ### 2. Fine-tune the Model158 159  ```bash160  cd train161  python train_codellama_qlora_linux_bugfix.py162  ```163 164  ### 3. Run Evaluation165 166  ```bash167  cd evaluate168  python evaluate_linux_bugfix_model.py169  ```170 171  ### 4. Use the Model172 173  ```python174  from transformers import AutoTokenizer, AutoModelForCausalLM175  from peft import PeftModel176 177  # Load the fine-tuned model178  model = AutoModelForCausalLM.from_pretrained("codellama/CodeLLaMA-7b-Instruct-hf")179  model = PeftModel.from_pretrained(model, "train/output/qlora-codellama-bugfix")180  tokenizer = AutoTokenizer.from_pretrained("codellama/CodeLLaMA-7b-Instruct-hf")181 182  # Generate a bug fix183  prompt = """184  Given the following original C code:185  if (!file->filter)186      return;187 188  Instruction: Fix the null pointer dereference189 190  Return the diff that fixes it:191  """192 193  inputs = tokenizer(prompt, return_tensors="pt")194  outputs = model.generate(**inputs, max_length=512, temperature=0.1)195  fix = tokenizer.decode(outputs[0], skip_special_tokens=True)196  print(fix)197  ```198 199  ---200 201  ## ๐Ÿ“ Project Structure202 203  ```204  CodeLLaMA-Linux-BugFix/205  โ”œโ”€โ”€ dataset_builder/206  โ”‚   โ”œโ”€โ”€ extract_linux_bugfixes_parallel.py    # Parallel extraction of bug fixes207  โ”‚   โ”œโ”€โ”€ format_for_training.py                # Format data for training208  โ”‚   โ””โ”€โ”€ build_dataset.py                      # Main dataset builder209  โ”œโ”€โ”€ dataset/210  โ”‚   โ”œโ”€โ”€ training_data_100k.jsonl              # 100K training samples211  โ”‚   โ””โ”€โ”€ training_data_prompt_completion.jsonl # Formatted training data212  โ”œโ”€โ”€ train/213  โ”‚   โ”œโ”€โ”€ train_codellama_qlora_linux_bugfix.py # Main training script214  โ”‚   โ”œโ”€โ”€ train_codellama_qlora_simple.py       # Simplified training215  โ”‚   โ”œโ”€โ”€ download_codellama_model.py           # Model download utility216  โ”‚   โ””โ”€โ”€ output/217  โ”‚       โ””โ”€โ”€ qlora-codellama-bugfix/           # Trained model checkpoints218  โ”œโ”€โ”€ evaluate/219  โ”‚   โ”œโ”€โ”€ evaluate_linux_bugfix_model.py        # Evaluation script220  โ”‚   โ”œโ”€โ”€ test_samples.jsonl                    # Test dataset221  โ”‚   โ””โ”€โ”€ output/                               # Evaluation results222  โ”‚       โ”œโ”€โ”€ eval_results.csv                  # Detailed results223  โ”‚       โ””โ”€โ”€ eval_results.json                 # JSON format results224  โ”œโ”€โ”€ requirements.txt                          # Python dependencies225  โ”œโ”€โ”€ README.md                                 # This file226  โ””โ”€โ”€ PROJECT_STRUCTURE.md                      # Detailed project overview227  ```228 229  ---230 231  ## ๐Ÿงฉ Features232 233  * ๐Ÿ”ง **Efficient Fine-tuning**: QLoRA + 4-bit quant = massive memory savings234  * ๐Ÿง  **Real-world commits**: From actual Linux kernel development235  * ๐Ÿ’ก **Context-aware**: Code context extraction around bug lines236  * ๐Ÿ’ป **Output-ready**: Generates valid Git-style diffs237  * ๐Ÿ“ˆ **Strong Performance**: BLEU score of 33.87 with good ROUGE metrics238  * ๐Ÿš€ **Production-ready**: Optimized for real-world deployment239 240  ---241 242  ## ๐Ÿ“ˆ Evaluation Metrics243 244  * **BLEU**: Translation-style match to reference diffs245  * **ROUGE**: Overlap in fix content and semantic similarity246  * **Human Evaluation**: Subjective patch quality assessment247 248  ### Current Performance249  - **BLEU Score**: 33.87 (excellent for code generation tasks)250  - **ROUGE-1 F1**: 0.4355 (good semantic overlap)251  - **ROUGE-2 F1**: 0.3457 (reasonable bigram matching)252  - **ROUGE-L F1**: 0.3612 (good longest common subsequence)253 254  ---255 256  ## ๐Ÿงช Use Cases257 258  * **Automated kernel bug fixing**: Generate fixes for common kernel bugs259  * **Code review assistance**: Help reviewers identify potential issues260  * **Teaching/debugging kernel code**: Educational tool for kernel development261  * **Research in automated program repair (APR)**: Academic research applications262  * **CI/CD integration**: Automated testing and fixing in development pipelines263 264  ---265 266  ## ๐Ÿ”ฌ Technical Highlights267 268  ### Memory & Speed Optimizations269 270  * 4-bit quantization (NF4)271  * Gradient checkpointing272  * Mixed precision (bfloat16)273  * Gradient accumulation274  * LoRA parameter efficiency275 276  ### Training Efficiency277 278  * **QLoRA**: Reduces memory usage by ~75%279  * **4-bit quantization**: Further memory optimization280  * **Gradient checkpointing**: Trades compute for memory281  * **Mixed precision**: Faster training with maintained accuracy282 283  ---284 285  ## ๐Ÿ› ๏ธ Advanced Usage286 287  ### Custom Training288 289  ```bash290  # Train with custom parameters291  python train_codellama_qlora_linux_bugfix.py \292      --learning_rate 1e-4 \293      --num_epochs 5 \294      --batch_size 32 \295      --lora_r 32 \296      --lora_alpha 16297  ```298 299  ### Evaluation on Custom Data300 301  ```bash302  # Evaluate on your own test set303  python evaluate_linux_bugfix_model.py \304      --test_file your_test_data.jsonl \305      --output_dir custom_eval_results306  ```307 308  ---309 310  ## ๐Ÿค Contributing311 312  1. Fork this repo313  2. Create a feature branch (`git checkout -b feature/amazing-feature`)314  3. Commit your changes (`git commit -m 'Add amazing feature'`)315  4. Push to the branch (`git push origin feature/amazing-feature`)316  5. Open a Pull Request ๐Ÿ™Œ317 318  ### Development Guidelines319 320  - Follow PEP 8 style guidelines321  - Add tests for new features322  - Update documentation for API changes323  - Ensure all tests pass before submitting PR324 325  ---326 327  ## ๐Ÿ“„ License328 329  MIT License โ€“ see `LICENSE` file for details.330 331  ---332 333  ## ๐Ÿ™ Acknowledgments334 335  * **Meta** for CodeLLaMA base model336  * **Hugging Face** for Transformers + PEFT libraries337  * **The Linux kernel community** for open access to commit data338  * **Microsoft** for introducing LoRA technique339  * **University of Washington** for QLoRA research340 341  ---342 343  ## ๐Ÿ“š References344 345  * [CodeLLaMA (Meta, 2023)](https://arxiv.org/abs/2308.12950)346  * [QLoRA (Dettmers et al., 2023)](https://arxiv.org/abs/2305.14314)347  * [LoRA (Hu et al., 2021)](https://arxiv.org/abs/2106.09685)348  * [Automated Program Repair: A Survey](https://ieeexplore.ieee.org/document/8449519)349 350  ---351 352  ## ๐Ÿ“ž Support353 354  For questions, issues, or contributions:355  - Open an issue on GitHub356  - Check the project documentation357  - Review the evaluation results in `evaluate/output/`358 359  ---360 361  ## ๐Ÿ”„ Version History362 363  - **v1.0.0**: Initial release with QLoRA training364  - **v1.1.0**: Added parallel dataset extraction365  - **v1.2.0**: Improved evaluation metrics and documentation366=======367---368license: mit369tags:370  - codellama371  - linux372  - bugfix373  - lora374  - qlora375  - git-diff376base_model: codellama/CodeLLaMA-7b-Instruct-hf377model_type: LlamaForCausalLM378library_name: peft379pipeline_tag: text-generation380---381 382# CodeLLaMA-Linux-BugFix383 384A fine-tuned version of `CodeLLaMA-7B-Instruct`, designed specifically for Linux kernel bug fixing using QLoRA (Quantized Low-Rank Adaptation). The model learns to generate Git diff patches based on buggy C code and commit messages.385 386---387 388## ๐ŸŽฏ Overview389 390This project targets automated Linux kernel bug fixing by:391 392- **Mining real commit data** from the kernel Git history393- **Training a specialized QLoRA model** on diff-style fixes394- **Generating Git patches** in response to bug-prone code395- **Evaluating results** using BLEU, ROUGE, and human inspection396 397The model achieves strong performance in generating accurate Linux kernel bug fixes, making it a valuable tool for automated code review and bug detection.398 399---400 401## ๐Ÿ“Š Performance Results402 403### Evaluation Metrics404 405โœ… **BLEU Score**: 33.87406 407โœ… **ROUGE Scores**:408- **ROUGE-1**: P=0.3775, R=0.7306, F1=0.4355409- **ROUGE-2**: P=0.2898, R=0.6096, F1=0.3457410- **ROUGE-L**: P=0.3023, R=0.6333, F1=0.3612411 412These results demonstrate the model's ability to:413- Generate syntactically correct Git diff patches414- Maintain semantic similarity to reference fixes415- Produce meaningful code changes that address the underlying bugs416 417---418 419## ๐Ÿง  Model Configuration420 421- **Base model**: `CodeLLaMA-7B-Instruct`422- **Fine-tuning method**: QLoRA with 4-bit quantization423- **Training setup**:424  - LoRA r=64, alpha=16, dropout=0.1425  - Batch size: 64, LR: 2e-4, Epochs: 3426  - Mixed precision (bfloat16), gradient checkpointing427- **Hardware**: Optimized for NVIDIA H200 GPUs428 429---430 431## ๐Ÿ“Š Dataset432 433Custom dataset extracted from Linux kernel Git history.434 435### Filtering Criteria436Bug-fix commits containing:437`fix`, `bug`, `crash`, `memory`, `null`, `panic`, `overflow`, `race`, `corruption`, etc.438 439### Structure440- Language: C (`.c`, `.h`)441- Context: 10 lines before/after the change442- Format:443 444```json445{446  "input": {447    "original code": "C code snippet with bug",448    "instruction": "Commit message or fix description"449  },450  "output": {451    "diff codes": "Git diff showing the fix"452  }453}454```455 456* **File**: `training_data_100k.jsonl` (100,000 samples)457 458---459 460## ๐Ÿš€ Quick Start461 462### Prerequisites463 464- Python 3.8+465- CUDA-compatible GPU (recommended)466- 16GB+ RAM467- 50GB+ disk space468 469### Install dependencies470 471```bash472pip install -r requirements.txt473```474 475### 1. Build the Dataset476 477```bash478cd dataset_builder479python extract_linux_bugfixes_parallel.py480python format_for_training.py481```482 483### 2. Fine-tune the Model484 485```bash486cd train487python train_codellama_qlora_linux_bugfix.py488```489 490### 3. Run Evaluation491 492```bash493cd evaluate494python evaluate_linux_bugfix_model.py495```496 497### 4. Use the Model498 499```python500from transformers import AutoTokenizer, AutoModelForCausalLM501from peft import PeftModel502 503# Load the fine-tuned model504model = AutoModelForCausalLM.from_pretrained("codellama/CodeLLaMA-7b-Instruct-hf")505model = PeftModel.from_pretrained(model, "train/output/qlora-codellama-bugfix")506tokenizer = AutoTokenizer.from_pretrained("codellama/CodeLLaMA-7b-Instruct-hf")507 508# Generate a bug fix509prompt = """510Given the following original C code:511if (!file->filter)512    return;513 514Instruction: Fix the null pointer dereference515 516Return the diff that fixes it:517"""518 519inputs = tokenizer(prompt, return_tensors="pt")520outputs = model.generate(**inputs, max_length=512, temperature=0.1)521fix = tokenizer.decode(outputs[0], skip_special_tokens=True)522print(fix)523```524 525---526 527## ๐Ÿ“ Project Structure528 529```530CodeLLaMA-Linux-BugFix/531โ”œโ”€โ”€ dataset_builder/532โ”‚   โ”œโ”€โ”€ extract_linux_bugfixes_parallel.py    # Parallel extraction of bug fixes533โ”‚   โ”œโ”€โ”€ format_for_training.py                # Format data for training534โ”‚   โ””โ”€โ”€ build_dataset.py                      # Main dataset builder535โ”œโ”€โ”€ dataset/536โ”‚   โ”œโ”€โ”€ training_data_100k.jsonl              # 100K training samples537โ”‚   โ””โ”€โ”€ training_data_prompt_completion.jsonl # Formatted training data538โ”œโ”€โ”€ train/539โ”‚   โ”œโ”€โ”€ train_codellama_qlora_linux_bugfix.py # Main training script540โ”‚   โ”œโ”€โ”€ train_codellama_qlora_simple.py       # Simplified training541โ”‚   โ”œโ”€โ”€ download_codellama_model.py           # Model download utility542โ”‚   โ””โ”€โ”€ output/543โ”‚       โ””โ”€โ”€ qlora-codellama-bugfix/           # Trained model checkpoints544โ”œโ”€โ”€ evaluate/545โ”‚   โ”œโ”€โ”€ evaluate_linux_bugfix_model.py        # Evaluation script546โ”‚   โ”œโ”€โ”€ test_samples.jsonl                    # Test dataset547โ”‚   โ””โ”€โ”€ output/                               # Evaluation results548โ”‚       โ”œโ”€โ”€ eval_results.csv                  # Detailed results549โ”‚       โ””โ”€โ”€ eval_results.json                 # JSON format results550โ”œโ”€โ”€ requirements.txt                          # Python dependencies551โ”œโ”€โ”€ README.md                                 # This file552โ””โ”€โ”€ PROJECT_STRUCTURE.md                      # Detailed project overview553```554 555---556 557## ๐Ÿงฉ Features558 559* ๐Ÿ”ง **Efficient Fine-tuning**: QLoRA + 4-bit quant = massive memory savings560* ๐Ÿง  **Real-world commits**: From actual Linux kernel development561* ๐Ÿ’ก **Context-aware**: Code context extraction around bug lines562* ๐Ÿ’ป **Output-ready**: Generates valid Git-style diffs563* ๐Ÿ“ˆ **Strong Performance**: BLEU score of 33.87 with good ROUGE metrics564* ๐Ÿš€ **Production-ready**: Optimized for real-world deployment565 566---567 568## ๐Ÿ“ˆ Evaluation Metrics569 570* **BLEU**: Translation-style match to reference diffs571* **ROUGE**: Overlap in fix content and semantic similarity572* **Human Evaluation**: Subjective patch quality assessment573 574### Current Performance575- **BLEU Score**: 33.87 (excellent for code generation tasks)576- **ROUGE-1 F1**: 0.4355 (good semantic overlap)577- **ROUGE-2 F1**: 0.3457 (reasonable bigram matching)578- **ROUGE-L F1**: 0.3612 (good longest common subsequence)579 580---581 582## ๐Ÿงช Use Cases583 584* **Automated kernel bug fixing**: Generate fixes for common kernel bugs585* **Code review assistance**: Help reviewers identify potential issues586* **Teaching/debugging kernel code**: Educational tool for kernel development587* **Research in automated program repair (APR)**: Academic research applications588* **CI/CD integration**: Automated testing and fixing in development pipelines589 590---591 592## ๐Ÿ”ฌ Technical Highlights593 594### Memory & Speed Optimizations595 596* 4-bit quantization (NF4)597* Gradient checkpointing598* Mixed precision (bfloat16)599* Gradient accumulation600* LoRA parameter efficiency601 602### Training Efficiency603 604* **QLoRA**: Reduces memory usage by ~75%605* **4-bit quantization**: Further memory optimization606* **Gradient checkpointing**: Trades compute for memory607* **Mixed precision**: Faster training with maintained accuracy608 609---610 611## ๐Ÿ› ๏ธ Advanced Usage612 613### Custom Training614 615```bash616# Train with custom parameters617python train_codellama_qlora_linux_bugfix.py \618    --learning_rate 1e-4 \619    --num_epochs 5 \620    --batch_size 32 \621    --lora_r 32 \622    --lora_alpha 16623```624 625### Evaluation on Custom Data626 627```bash628# Evaluate on your own test set629python evaluate_linux_bugfix_model.py \630    --test_file your_test_data.jsonl \631    --output_dir custom_eval_results632```633 634---635 636## ๐Ÿค Contributing637 6381. Fork this repo6392. Create a feature branch (`git checkout -b feature/amazing-feature`)6403. Commit your changes (`git commit -m 'Add amazing feature'`)6414. Push to the branch (`git push origin feature/amazing-feature`)6425. Open a Pull Request ๐Ÿ™Œ643 644### Development Guidelines645 646- Follow PEP 8 style guidelines647- Add tests for new features648- Update documentation for API changes649- Ensure all tests pass before submitting PR650 651---652 653## ๐Ÿ“„ License654 655MIT License โ€“ see `LICENSE` file for details.656 657---658 659## ๐Ÿ™ Acknowledgments660 661* **Meta** for CodeLLaMA base model662* **Hugging Face** for Transformers + PEFT libraries663* **The Linux kernel community** for open access to commit data664* **Microsoft** for introducing LoRA technique665* **University of Washington** for QLoRA research666 667---668 669## ๐Ÿ“š References670 671* [CodeLLaMA (Meta, 2023)](https://arxiv.org/abs/2308.12950)672* [QLoRA (Dettmers et al., 2023)](https://arxiv.org/abs/2305.14314)673* [LoRA (Hu et al., 2021)](https://arxiv.org/abs/2106.09685)674* [Automated Program Repair: A Survey](https://ieeexplore.ieee.org/document/8449519)675 676---677 678## ๐Ÿ“ž Support679 680For questions, issues, or contributions:681- Open an issue on GitHub682- Check the project documentation683- Review the evaluation results in `evaluate/output/`684 685---686 687## ๐Ÿ”„ Version History688 689- **v1.0.0**: Initial release with QLoRA training690- **v1.1.0**: Added parallel dataset extraction691- **v1.2.0**: Improved evaluation metrics and documentation692