Maaac/CodeLLaMA-Linux-BugFix
08
1# Project Structure
2
3This document provides a detailed overview of the CodeLLaMA-Linux-BugFix project structure, explaining the purpose and organization of each component.
4
5## ๐ Root Directory
6
7```
8CodeLLaMA-Linux-BugFix/
9โโโ dataset_builder/ # Dataset creation and processing
10โโโ dataset/ # Generated datasets and data files
11โโโ train/ # Model training scripts and outputs
12โโโ evaluate/ # Model evaluation and testing
13โโโ requirements.txt # Python dependencies
14โโโ README.md # Project documentation
15โโโ PROJECT_STRUCTURE.md # This file
16```
17
18## ๐ง Dataset Builder (`dataset_builder/`)
19
20The dataset builder extracts bug-fix data from the Linux kernel Git repository and converts it into training-ready format.
21
22### Files:
23- **`extract_linux_bugfixes.py`** - Main dataset extraction script
24 - Uses PyDriller to analyze Linux kernel Git history
25 - Filters commits using bug-fix keywords
26 - Extracts code context around bug locations
27 - Generates structured dataset entries
28
29- **`extract_linux_bugfixes_parallel.py`** - Parallelized version of dataset builder
30 - Multi-process implementation for faster processing
31 - Configurable worker count (default: 16 workers)
32 - Test mode with limited commit processing
33
34- **`format_for_training.py`** - Format conversion script
35 - Converts structured data to prompt-completion pairs
36 - Formats input for supervised fine-tuning
37 - Creates training-ready JSONL format
38
39### Key Features:
40- **Commit Filtering**: Identifies bug-fix commits using 17 keywords
41- **Code Context**: Extracts 10 lines before/after bug location
42- **File Filtering**: Focuses on C and header files (`.c`, `.h`)
43- **Diff Extraction**: Captures Git diff patches for fixes
44
45## ๐ Dataset (`dataset/`)
46
47Contains the generated datasets used for training and evaluation.
48
49### Files:
50- **`training_data_100k.jsonl`** - Main training dataset
51 - 100,000 bug-fix samples
52 - Structured format with input/output pairs
53 - Stored using Git LFS for large file handling
54
55- **`training_data_prompt_completion.jsonl`** - Converted training format
56 - Prompt-completion pairs for supervised learning
57 - Optimized for transformer model training
58 - Stored using Git LFS
59
60### Data Format:
61```json
62{
63 "input": {
64 "original code": "C code snippet with bug",
65 "instruction": "Bug fix instruction from commit message"
66 },
67 "output": {
68 "diff codes": "Git diff showing the fix"
69 }
70}
71```
72
73## ๐ Training (`train/`)
74
75Contains all training-related scripts, configurations, and model outputs.
76
77### Files:
78- **`train_codellama_qlora_linux_bugfix.py`** - Main training script
79 - QLoRA fine-tuning implementation
80 - Optimized for H200 GPU with bfloat16
81 - Includes Weights & Biases integration
82 - Comprehensive training configuration
83
84- **`train_codellama_qlora_simple.py`** - Alternative training script
85 - Simplified QLoRA implementation
86 - Basic training setup without advanced features
87 - Good for testing and development
88
89- **`download_codellama_model.py`** - Model download utility
90 - Downloads base CodeLLaMA-7B-Instruct model
91 - Ensures model availability before training
92
93### Output Directory (`train/output/`):
94- **`qlora-codellama-bugfix/`** - Main model output
95 - **`adapter_model.safetensors`** - LoRA adapter weights
96 - **`adapter_config.json`** - LoRA configuration
97 - **`tokenizer.json`** - Tokenizer files
98 - **`chat_template.jinja`** - Conversation template
99 - **`checkpoint-500/`** - Training checkpoint at step 500
100 - **`checkpoint-1000/`** - Training checkpoint at step 1000
101 - **`README.md`** - Model card and documentation
102
103### Training Configuration:
104- **Base Model**: `codellama/CodeLLaMA-7b-Instruct-hf`
105- **Method**: QLoRA with 4-bit quantization
106- **LoRA Config**: r=64, alpha=16, dropout=0.1
107- **Training**: 3 epochs, batch size 64, learning rate 2e-4
108- **Hardware**: Optimized for H200 GPU
109
110## ๐ Evaluation (`evaluate/`)
111
112Contains evaluation scripts and results for assessing model performance.
113
114### Files:
115- **`evaluate_linux_bugfix_model.py`** - Main evaluation script
116 - Loads fine-tuned model for inference
117 - Generates predictions on test data
118 - Computes BLEU and ROUGE metrics
119 - Saves results in multiple formats
120
121- **`test_samples.jsonl`** - Evaluation dataset
122 - Test samples for model evaluation
123 - Stored using Git LFS
124
125### Output Directory (`evaluate/output/`):
126- **`eval_results.json`** - Detailed evaluation results
127 - Complete predictions and references
128 - Stored using Git LFS
129
130- **`eval_results.csv`** - Tabular evaluation results
131 - CSV format for easy analysis
132 - Stored using Git LFS
133
134### Evaluation Metrics:
135- **BLEU Score**: Measures translation quality
136- **ROUGE Score**: Evaluates text generation accuracy
137- **Human Evaluation**: Qualitative assessment
138
139## ๐ง Dependencies (`requirements.txt`)
140
141Comprehensive list of Python packages required for the project:
142
143### Core ML Libraries:
144- `transformers==4.53.1` - Hugging Face transformers
145- `torch==2.7.1+cu128` - PyTorch with CUDA support
146- `peft==0.16.0` - Parameter-efficient fine-tuning
147- `accelerate==1.8.1` - Distributed training
148- `bitsandbytes==0.46.1` - Quantization support
149
150### Data Processing:
151- `datasets==3.6.0` - Dataset handling
152- `pandas==2.3.1` - Data manipulation
153- `numpy==2.3.1` - Numerical computing
154
155### Git Analysis:
156- `pydriller` - Git repository mining
157- `gitpython` - Git operations
158
159### Utilities:
160- `tqdm==4.67.1` - Progress bars
161- `wandb` - Experiment tracking
162- `evaluate==0.4.4` - Evaluation metrics
163
164## ๐ Workflow
165
166### 1. Dataset Creation
167```bash
168cd dataset_builder
169python extract_linux_bugfixes.py # Extract bug-fix data
170python format_for_training.py # Convert format
171```
172
173### 2. Model Training
174```bash
175cd train
176python train_codellama_qlora_linux_bugfix.py # Train with QLoRA
177```
178
179### 3. Model Evaluation
180```bash
181cd evaluate
182python evaluate_linux_bugfix_model.py # Evaluate performance
183```
184
185## ๐ฏ Key Design Principles
186
187### Modularity
188- Each component has a specific responsibility
189- Clear separation between data, training, and evaluation
190- Easy to modify or extend individual components
191
192### Efficiency
193- QLoRA for memory-efficient training
194- Parallel processing for dataset creation
195- Optimized for modern GPU hardware
196
197### Reproducibility
198- Version-controlled dependencies
199- Structured data formats
200- Comprehensive logging and evaluation
201
202### Scalability
203- Configurable parameters for different hardware
204- Support for distributed training
205- Efficient data handling with Git LFS
206
207## ๐ File Naming Conventions
208
209- **Scripts**: Descriptive names with clear purpose
210- **Datasets**: Include size/version information
211- **Models**: Include architecture and method
212- **Results**: Include timestamp or version
213- **Configs**: Use `.json` or `.yaml` format
214
215## ๐ Documentation
216
217- **README.md**: Project overview and quick start
218- **PROJECT_STRUCTURE.md**: This detailed structure guide
219- **Model README**: Generated model cards in output directories
220- **Code Comments**: Inline documentation in all scripts
221
222This structure ensures the project is organized, maintainable, and easy to understand for both users and contributors. 