Team Ai
Modelpublic

Maaac/CodeLLaMA-Linux-BugFix

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes8downloads
Model Card

# CodeLLaMA-Linux-BugFix

A fine-tuned version of CodeLLaMA-7B-Instruct, designed specifically for Linux kernel bug fixing using QLoRA (Quantized Low-Rank Adaptation). The model learns to generate Git diff patches based on buggy C code and commit messages.


## ๐ŸŽฏ Overview

This project targets automated Linux kernel bug fixing by:

  • โ€”Mining real commit data from the kernel Git history
  • โ€”Training a specialized QLoRA model on diff-style fixes
  • โ€”Generating Git patches in response to bug-prone code
  • โ€”Evaluating results using BLEU, ROUGE, and human inspection

The model achieves strong performance in generating accurate Linux kernel bug fixes, making it a valuable tool for automated code review and bug detection.


## ๐Ÿ“Š Performance Results

### Evaluation Metrics

โœ… BLEU Score: 33.87

โœ… ROUGE Scores:

  • โ€”ROUGE-1: P=0.3775, R=0.7306, F1=0.4355
  • โ€”ROUGE-2: P=0.2898, R=0.6096, F1=0.3457
  • โ€”ROUGE-L: P=0.3023, R=0.6333, F1=0.3612

These results demonstrate the model's ability to:

  • โ€”Generate syntactically correct Git diff patches
  • โ€”Maintain semantic similarity to reference fixes
  • โ€”Produce meaningful code changes that address the underlying bugs

## ๐Ÿง  Model Configuration

  • โ€”Base model: CodeLLaMA-7B-Instruct
  • โ€”Fine-tuning method: QLoRA with 4-bit quantization
  • โ€”Training setup:
  • โ€”LoRA r=64, alpha=16, dropout=0.1
  • โ€”Batch size: 64, LR: 2e-4, Epochs: 3
  • โ€”Mixed precision (bfloat16), gradient checkpointing
  • โ€”Hardware: Optimized for NVIDIA H200 GPUs

## ๐Ÿ“ˆ Training Progress The model was trained for 1000 steps with the following key metrics: ### Training Results

  • โ€”Final Loss: ~0.3335 (converged)
  • โ€”Final Learning Rate: 2.08304527802282E-06
  • โ€”Training Steps: 1000
  • โ€”Convergence: Stable loss plateau achieved ### Training Curves [image] Training loss over 1000 steps showing convergence around 0.3335 [image] Learning rate decay schedule with final rate of 2.08304527802282E-06

## ๐Ÿ“Š Dataset

Custom dataset extracted from Linux kernel Git history.

### Filtering Criteria Bug-fix commits containing: fix, bug, crash, memory, null, panic, overflow, race, corruption, etc.

### Structure

  • โ€”Language: C (.c, .h)
  • โ€”Context: 10 lines before/after the change
  • โ€”Format:
json
  {
    "input": {
      "original code": "C code snippet with bug",
      "instruction": "Commit message or fix description"
    },
    "output": {
      "diff codes": "Git diff showing the fix"
    }
  }
  • โ€”File: training_data_100k.jsonl (100,000 samples)

## ๐Ÿš€ Quick Start

### Prerequisites

  • โ€”Python 3.8+
  • โ€”CUDA-compatible GPU (recommended)
  • โ€”16GB+ RAM
  • โ€”50GB+ disk space

### Install dependencies

bash
  pip install -r requirements.txt

### 1. Build the Dataset

bash
  cd dataset_builder
  python extract_linux_bugfixes_parallel.py
  python format_for_training.py

### 2. Fine-tune the Model

bash
  cd train
  python train_codellama_qlora_linux_bugfix.py

### 3. Run Evaluation

bash
  cd evaluate
  python evaluate_linux_bugfix_model.py

### 4. Use the Model

python
  from transformers import AutoTokenizer, AutoModelForCausalLM
  from peft import PeftModel

  # Load the fine-tuned model
  model = AutoModelForCausalLM.from_pretrained("codellama/CodeLLaMA-7b-Instruct-hf")
  model = PeftModel.from_pretrained(model, "train/output/qlora-codellama-bugfix")
  tokenizer = AutoTokenizer.from_pretrained("codellama/CodeLLaMA-7b-Instruct-hf")

  # Generate a bug fix
  prompt = """
  Given the following original C code:
  if (!file->filter)
      return;

  Instruction: Fix the null pointer dereference

  Return the diff that fixes it:
  """

  inputs = tokenizer(prompt, return_tensors="pt")
  outputs = model.generate(**inputs, max_length=512, temperature=0.1)
  fix = tokenizer.decode(outputs[0], skip_special_tokens=True)
  print(fix)

## ๐Ÿ“ Project Structure

  CodeLLaMA-Linux-BugFix/
  โ”œโ”€โ”€ dataset_builder/
  โ”‚   โ”œโ”€โ”€ extract_linux_bugfixes_parallel.py    # Parallel extraction of bug fixes
  โ”‚   โ”œโ”€โ”€ format_for_training.py                # Format data for training
  โ”‚   โ””โ”€โ”€ build_dataset.py                      # Main dataset builder
  โ”œโ”€โ”€ dataset/
  โ”‚   โ”œโ”€โ”€ training_data_100k.jsonl              # 100K training samples
  โ”‚   โ””โ”€โ”€ training_data_prompt_completion.jsonl # Formatted training data
  โ”œโ”€โ”€ train/
  โ”‚   โ”œโ”€โ”€ train_codellama_qlora_linux_bugfix.py # Main training script
  โ”‚   โ”œโ”€โ”€ train_codellama_qlora_simple.py       # Simplified training
  โ”‚   โ”œโ”€โ”€ download_codellama_model.py           # Model download utility
  โ”‚   โ””โ”€โ”€ output/
  โ”‚       โ””โ”€โ”€ qlora-codellama-bugfix/           # Trained model checkpoints
  โ”œโ”€โ”€ evaluate/
  โ”‚   โ”œโ”€โ”€ evaluate_linux_bugfix_model.py        # Evaluation script
  โ”‚   โ”œโ”€โ”€ test_samples.jsonl                    # Test dataset
  โ”‚   โ””โ”€โ”€ output/                               # Evaluation results
  โ”‚       โ”œโ”€โ”€ eval_results.csv                  # Detailed results
  โ”‚       โ””โ”€โ”€ eval_results.json                 # JSON format results
  โ”œโ”€โ”€ requirements.txt                          # Python dependencies
  โ”œโ”€โ”€ README.md                                 # This file
  โ””โ”€โ”€ PROJECT_STRUCTURE.md                      # Detailed project overview

## ๐Ÿงฉ Features

  • โ€”๐Ÿ”ง Efficient Fine-tuning: QLoRA + 4-bit quant = massive memory savings
  • โ€”๐Ÿง  Real-world commits: From actual Linux kernel development
  • โ€”๐Ÿ’ก Context-aware: Code context extraction around bug lines
  • โ€”๐Ÿ’ป Output-ready: Generates valid Git-style diffs
  • โ€”๐Ÿ“ˆ Strong Performance: BLEU score of 33.87 with good ROUGE metrics
  • โ€”๐Ÿš€ Production-ready: Optimized for real-world deployment

## ๐Ÿ“ˆ Evaluation Metrics

  • โ€”BLEU: Translation-style match to reference diffs
  • โ€”ROUGE: Overlap in fix content and semantic similarity
  • โ€”Human Evaluation: Subjective patch quality assessment

### Current Performance

  • โ€”BLEU Score: 33.87 (excellent for code generation tasks)
  • โ€”ROUGE-1 F1: 0.4355 (good semantic overlap)
  • โ€”ROUGE-2 F1: 0.3457 (reasonable bigram matching)
  • โ€”ROUGE-L F1: 0.3612 (good longest common subsequence)

## ๐Ÿงช Use Cases

  • โ€”Automated kernel bug fixing: Generate fixes for common kernel bugs
  • โ€”Code review assistance: Help reviewers identify potential issues
  • โ€”Teaching/debugging kernel code: Educational tool for kernel development
  • โ€”Research in automated program repair (APR): Academic research applications
  • โ€”CI/CD integration: Automated testing and fixing in development pipelines

## ๐Ÿ”ฌ Technical Highlights

### Memory & Speed Optimizations

  • โ€”4-bit quantization (NF4)
  • โ€”Gradient checkpointing
  • โ€”Mixed precision (bfloat16)
  • โ€”Gradient accumulation
  • โ€”LoRA parameter efficiency

### Training Efficiency

  • โ€”QLoRA: Reduces memory usage by ~75%
  • โ€”4-bit quantization: Further memory optimization
  • โ€”Gradient checkpointing: Trades compute for memory
  • โ€”Mixed precision: Faster training with maintained accuracy

## ๐Ÿ› ๏ธ Advanced Usage

### Custom Training

bash
  # Train with custom parameters
  python train_codellama_qlora_linux_bugfix.py \
      --learning_rate 1e-4 \
      --num_epochs 5 \
      --batch_size 32 \
      --lora_r 32 \
      --lora_alpha 16

### Evaluation on Custom Data

bash
  # Evaluate on your own test set
  python evaluate_linux_bugfix_model.py \
      --test_file your_test_data.jsonl \
      --output_dir custom_eval_results

## ๐Ÿค Contributing

  1. 1.Fork this repo
  2. 2.Create a feature branch (git checkout -b feature/amazing-feature)
  3. 3.Commit your changes (git commit -m 'Add amazing feature')
  4. 4.Push to the branch (git push origin feature/amazing-feature)
  5. 5.Open a Pull Request ๐Ÿ™Œ

### Development Guidelines

  • โ€”Follow PEP 8 style guidelines
  • โ€”Add tests for new features
  • โ€”Update documentation for API changes
  • โ€”Ensure all tests pass before submitting PR

## ๐Ÿ“„ License

MIT License โ€“ see LICENSE file for details.


## ๐Ÿ™ Acknowledgments

  • โ€”Meta for CodeLLaMA base model
  • โ€”Hugging Face for Transformers + PEFT libraries
  • โ€”The Linux kernel community for open access to commit data
  • โ€”Microsoft for introducing LoRA technique
  • โ€”University of Washington for QLoRA research

## ๐Ÿ“š References


## ๐Ÿ“ž Support

For questions, issues, or contributions:

  • โ€”Open an issue on GitHub
  • โ€”Check the project documentation
  • โ€”Review the evaluation results in evaluate/output/

## ๐Ÿ”„ Version History

  • โ€”v1.0.0: Initial release with QLoRA training
  • โ€”v1.1.0: Added parallel dataset extraction
  • โ€”v1.2.0: Improved evaluation metrics and documentation ======= --- license: mit tags:
  • โ€”codellama
  • โ€”linux
  • โ€”bugfix
  • โ€”lora
  • โ€”qlora
  • โ€”git-diff basemodel: codellama/CodeLLaMA-7b-Instruct-hf modeltype: LlamaForCausalLM libraryname: peft pipelinetag: text-generation ---

CodeLLaMA-Linux-BugFix

A fine-tuned version of CodeLLaMA-7B-Instruct, designed specifically for Linux kernel bug fixing using QLoRA (Quantized Low-Rank Adaptation). The model learns to generate Git diff patches based on buggy C code and commit messages.


๐ŸŽฏ Overview

This project targets automated Linux kernel bug fixing by:

  • โ€”Mining real commit data from the kernel Git history
  • โ€”Training a specialized QLoRA model on diff-style fixes
  • โ€”Generating Git patches in response to bug-prone code
  • โ€”Evaluating results using BLEU, ROUGE, and human inspection

The model achieves strong performance in generating accurate Linux kernel bug fixes, making it a valuable tool for automated code review and bug detection.


๐Ÿ“Š Performance Results

Evaluation Metrics

โœ… BLEU Score: 33.87

โœ… ROUGE Scores:

  • โ€”ROUGE-1: P=0.3775, R=0.7306, F1=0.4355
  • โ€”ROUGE-2: P=0.2898, R=0.6096, F1=0.3457
  • โ€”ROUGE-L: P=0.3023, R=0.6333, F1=0.3612

These results demonstrate the model's ability to:

  • โ€”Generate syntactically correct Git diff patches
  • โ€”Maintain semantic similarity to reference fixes
  • โ€”Produce meaningful code changes that address the underlying bugs

๐Ÿง  Model Configuration

  • โ€”Base model: CodeLLaMA-7B-Instruct
  • โ€”Fine-tuning method: QLoRA with 4-bit quantization
  • โ€”Training setup:
  • โ€”LoRA r=64, alpha=16, dropout=0.1
  • โ€”Batch size: 64, LR: 2e-4, Epochs: 3
  • โ€”Mixed precision (bfloat16), gradient checkpointing
  • โ€”Hardware: Optimized for NVIDIA H200 GPUs

๐Ÿ“Š Dataset

Custom dataset extracted from Linux kernel Git history.

Filtering Criteria

Bug-fix commits containing: fix, bug, crash, memory, null, panic, overflow, race, corruption, etc.

Structure

  • โ€”Language: C (.c, .h)
  • โ€”Context: 10 lines before/after the change
  • โ€”Format:
json
{
  "input": {
    "original code": "C code snippet with bug",
    "instruction": "Commit message or fix description"
  },
  "output": {
    "diff codes": "Git diff showing the fix"
  }
}
  • โ€”File: training_data_100k.jsonl (100,000 samples)

๐Ÿš€ Quick Start

Prerequisites

  • โ€”Python 3.8+
  • โ€”CUDA-compatible GPU (recommended)
  • โ€”16GB+ RAM
  • โ€”50GB+ disk space

Install dependencies

bash
pip install -r requirements.txt

1. Build the Dataset

bash
cd dataset_builder
python extract_linux_bugfixes_parallel.py
python format_for_training.py

2. Fine-tune the Model

bash
cd train
python train_codellama_qlora_linux_bugfix.py

3. Run Evaluation

bash
cd evaluate
python evaluate_linux_bugfix_model.py

4. Use the Model

python
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel

# Load the fine-tuned model
model = AutoModelForCausalLM.from_pretrained("codellama/CodeLLaMA-7b-Instruct-hf")
model = PeftModel.from_pretrained(model, "train/output/qlora-codellama-bugfix")
tokenizer = AutoTokenizer.from_pretrained("codellama/CodeLLaMA-7b-Instruct-hf")

# Generate a bug fix
prompt = """
Given the following original C code:
if (!file->filter)
    return;

Instruction: Fix the null pointer dereference

Return the diff that fixes it:
"""

inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_length=512, temperature=0.1)
fix = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(fix)

๐Ÿ“ Project Structure

CodeLLaMA-Linux-BugFix/
โ”œโ”€โ”€ dataset_builder/
โ”‚   โ”œโ”€โ”€ extract_linux_bugfixes_parallel.py    # Parallel extraction of bug fixes
โ”‚   โ”œโ”€โ”€ format_for_training.py                # Format data for training
โ”‚   โ””โ”€โ”€ build_dataset.py                      # Main dataset builder
โ”œโ”€โ”€ dataset/
โ”‚   โ”œโ”€โ”€ training_data_100k.jsonl              # 100K training samples
โ”‚   โ””โ”€โ”€ training_data_prompt_completion.jsonl # Formatted training data
โ”œโ”€โ”€ train/
โ”‚   โ”œโ”€โ”€ train_codellama_qlora_linux_bugfix.py # Main training script
โ”‚   โ”œโ”€โ”€ train_codellama_qlora_simple.py       # Simplified training
โ”‚   โ”œโ”€โ”€ download_codellama_model.py           # Model download utility
โ”‚   โ””โ”€โ”€ output/
โ”‚       โ””โ”€โ”€ qlora-codellama-bugfix/           # Trained model checkpoints
โ”œโ”€โ”€ evaluate/
โ”‚   โ”œโ”€โ”€ evaluate_linux_bugfix_model.py        # Evaluation script
โ”‚   โ”œโ”€โ”€ test_samples.jsonl                    # Test dataset
โ”‚   โ””โ”€โ”€ output/                               # Evaluation results
โ”‚       โ”œโ”€โ”€ eval_results.csv                  # Detailed results
โ”‚       โ””โ”€โ”€ eval_results.json                 # JSON format results
โ”œโ”€โ”€ requirements.txt                          # Python dependencies
โ”œโ”€โ”€ README.md                                 # This file
โ””โ”€โ”€ PROJECT_STRUCTURE.md                      # Detailed project overview

๐Ÿงฉ Features

  • โ€”๐Ÿ”ง Efficient Fine-tuning: QLoRA + 4-bit quant = massive memory savings
  • โ€”๐Ÿง  Real-world commits: From actual Linux kernel development
  • โ€”๐Ÿ’ก Context-aware: Code context extraction around bug lines
  • โ€”๐Ÿ’ป Output-ready: Generates valid Git-style diffs
  • โ€”๐Ÿ“ˆ Strong Performance: BLEU score of 33.87 with good ROUGE metrics
  • โ€”๐Ÿš€ Production-ready: Optimized for real-world deployment

๐Ÿ“ˆ Evaluation Metrics

  • โ€”BLEU: Translation-style match to reference diffs
  • โ€”ROUGE: Overlap in fix content and semantic similarity
  • โ€”Human Evaluation: Subjective patch quality assessment

Current Performance

  • โ€”BLEU Score: 33.87 (excellent for code generation tasks)
  • โ€”ROUGE-1 F1: 0.4355 (good semantic overlap)
  • โ€”ROUGE-2 F1: 0.3457 (reasonable bigram matching)
  • โ€”ROUGE-L F1: 0.3612 (good longest common subsequence)

๐Ÿงช Use Cases

  • โ€”Automated kernel bug fixing: Generate fixes for common kernel bugs
  • โ€”Code review assistance: Help reviewers identify potential issues
  • โ€”Teaching/debugging kernel code: Educational tool for kernel development
  • โ€”Research in automated program repair (APR): Academic research applications
  • โ€”CI/CD integration: Automated testing and fixing in development pipelines

๐Ÿ”ฌ Technical Highlights

Memory & Speed Optimizations

  • โ€”4-bit quantization (NF4)
  • โ€”Gradient checkpointing
  • โ€”Mixed precision (bfloat16)
  • โ€”Gradient accumulation
  • โ€”LoRA parameter efficiency

Training Efficiency

  • โ€”QLoRA: Reduces memory usage by ~75%
  • โ€”4-bit quantization: Further memory optimization
  • โ€”Gradient checkpointing: Trades compute for memory
  • โ€”Mixed precision: Faster training with maintained accuracy

๐Ÿ› ๏ธ Advanced Usage

Custom Training

bash
# Train with custom parameters
python train_codellama_qlora_linux_bugfix.py \
    --learning_rate 1e-4 \
    --num_epochs 5 \
    --batch_size 32 \
    --lora_r 32 \
    --lora_alpha 16

Evaluation on Custom Data

bash
# Evaluate on your own test set
python evaluate_linux_bugfix_model.py \
    --test_file your_test_data.jsonl \
    --output_dir custom_eval_results

๐Ÿค Contributing

  1. 1.Fork this repo
  2. 2.Create a feature branch (git checkout -b feature/amazing-feature)
  3. 3.Commit your changes (git commit -m 'Add amazing feature')
  4. 4.Push to the branch (git push origin feature/amazing-feature)
  5. 5.Open a Pull Request ๐Ÿ™Œ

Development Guidelines

  • โ€”Follow PEP 8 style guidelines
  • โ€”Add tests for new features
  • โ€”Update documentation for API changes
  • โ€”Ensure all tests pass before submitting PR

๐Ÿ“„ License

MIT License โ€“ see LICENSE file for details.


๐Ÿ™ Acknowledgments

  • โ€”Meta for CodeLLaMA base model
  • โ€”Hugging Face for Transformers + PEFT libraries
  • โ€”The Linux kernel community for open access to commit data
  • โ€”Microsoft for introducing LoRA technique
  • โ€”University of Washington for QLoRA research

๐Ÿ“š References


๐Ÿ“ž Support

For questions, issues, or contributions:

  • โ€”Open an issue on GitHub
  • โ€”Check the project documentation
  • โ€”Review the evaluation results in evaluate/output/

๐Ÿ”„ Version History

  • โ€”v1.0.0: Initial release with QLoRA training
  • โ€”v1.1.0: Added parallel dataset extraction
  • โ€”v1.2.0: Improved evaluation metrics and documentation