Team Ai
Modelpublic

Maaac/CodeLLaMA-Linux-BugFix

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes8downloads
PROJECT_STRUCTURE.md222 linesDownload Raw Back to root
1# Project Structure
2
3This document provides a detailed overview of the CodeLLaMA-Linux-BugFix project structure, explaining the purpose and organization of each component.
4
5## ๐Ÿ“ Root Directory
6
7```
8CodeLLaMA-Linux-BugFix/
9โ”œโ”€โ”€ dataset_builder/          # Dataset creation and processing
10โ”œโ”€โ”€ dataset/                  # Generated datasets and data files
11โ”œโ”€โ”€ train/                    # Model training scripts and outputs
12โ”œโ”€โ”€ evaluate/                 # Model evaluation and testing
13โ”œโ”€โ”€ requirements.txt          # Python dependencies
14โ”œโ”€โ”€ README.md                 # Project documentation
15โ””โ”€โ”€ PROJECT_STRUCTURE.md      # This file
16```
17
18## ๐Ÿ”ง Dataset Builder (`dataset_builder/`)
19
20The dataset builder extracts bug-fix data from the Linux kernel Git repository and converts it into training-ready format.
21
22### Files:
23- **`extract_linux_bugfixes.py`** - Main dataset extraction script
24  - Uses PyDriller to analyze Linux kernel Git history
25  - Filters commits using bug-fix keywords
26  - Extracts code context around bug locations
27  - Generates structured dataset entries
28
29- **`extract_linux_bugfixes_parallel.py`** - Parallelized version of dataset builder
30  - Multi-process implementation for faster processing
31  - Configurable worker count (default: 16 workers)
32  - Test mode with limited commit processing
33
34- **`format_for_training.py`** - Format conversion script
35  - Converts structured data to prompt-completion pairs
36  - Formats input for supervised fine-tuning
37  - Creates training-ready JSONL format
38
39### Key Features:
40- **Commit Filtering**: Identifies bug-fix commits using 17 keywords
41- **Code Context**: Extracts 10 lines before/after bug location
42- **File Filtering**: Focuses on C and header files (`.c`, `.h`)
43- **Diff Extraction**: Captures Git diff patches for fixes
44
45## ๐Ÿ“Š Dataset (`dataset/`)
46
47Contains the generated datasets used for training and evaluation.
48
49### Files:
50- **`training_data_100k.jsonl`** - Main training dataset
51  - 100,000 bug-fix samples
52  - Structured format with input/output pairs
53  - Stored using Git LFS for large file handling
54
55- **`training_data_prompt_completion.jsonl`** - Converted training format
56  - Prompt-completion pairs for supervised learning
57  - Optimized for transformer model training
58  - Stored using Git LFS
59
60### Data Format:
61```json
62{
63  "input": {
64    "original code": "C code snippet with bug",
65    "instruction": "Bug fix instruction from commit message"
66  },
67  "output": {
68    "diff codes": "Git diff showing the fix"
69  }
70}
71```
72
73## ๐Ÿš€ Training (`train/`)
74
75Contains all training-related scripts, configurations, and model outputs.
76
77### Files:
78- **`train_codellama_qlora_linux_bugfix.py`** - Main training script
79  - QLoRA fine-tuning implementation
80  - Optimized for H200 GPU with bfloat16
81  - Includes Weights & Biases integration
82  - Comprehensive training configuration
83
84- **`train_codellama_qlora_simple.py`** - Alternative training script
85  - Simplified QLoRA implementation
86  - Basic training setup without advanced features
87  - Good for testing and development
88
89- **`download_codellama_model.py`** - Model download utility
90  - Downloads base CodeLLaMA-7B-Instruct model
91  - Ensures model availability before training
92
93### Output Directory (`train/output/`):
94- **`qlora-codellama-bugfix/`** - Main model output
95  - **`adapter_model.safetensors`** - LoRA adapter weights
96  - **`adapter_config.json`** - LoRA configuration
97  - **`tokenizer.json`** - Tokenizer files
98  - **`chat_template.jinja`** - Conversation template
99  - **`checkpoint-500/`** - Training checkpoint at step 500
100  - **`checkpoint-1000/`** - Training checkpoint at step 1000
101  - **`README.md`** - Model card and documentation
102
103### Training Configuration:
104- **Base Model**: `codellama/CodeLLaMA-7b-Instruct-hf`
105- **Method**: QLoRA with 4-bit quantization
106- **LoRA Config**: r=64, alpha=16, dropout=0.1
107- **Training**: 3 epochs, batch size 64, learning rate 2e-4
108- **Hardware**: Optimized for H200 GPU
109
110## ๐Ÿ“ˆ Evaluation (`evaluate/`)
111
112Contains evaluation scripts and results for assessing model performance.
113
114### Files:
115- **`evaluate_linux_bugfix_model.py`** - Main evaluation script
116  - Loads fine-tuned model for inference
117  - Generates predictions on test data
118  - Computes BLEU and ROUGE metrics
119  - Saves results in multiple formats
120
121- **`test_samples.jsonl`** - Evaluation dataset
122  - Test samples for model evaluation
123  - Stored using Git LFS
124
125### Output Directory (`evaluate/output/`):
126- **`eval_results.json`** - Detailed evaluation results
127  - Complete predictions and references
128  - Stored using Git LFS
129
130- **`eval_results.csv`** - Tabular evaluation results
131  - CSV format for easy analysis
132  - Stored using Git LFS
133
134### Evaluation Metrics:
135- **BLEU Score**: Measures translation quality
136- **ROUGE Score**: Evaluates text generation accuracy
137- **Human Evaluation**: Qualitative assessment
138
139## ๐Ÿ”ง Dependencies (`requirements.txt`)
140
141Comprehensive list of Python packages required for the project:
142
143### Core ML Libraries:
144- `transformers==4.53.1` - Hugging Face transformers
145- `torch==2.7.1+cu128` - PyTorch with CUDA support
146- `peft==0.16.0` - Parameter-efficient fine-tuning
147- `accelerate==1.8.1` - Distributed training
148- `bitsandbytes==0.46.1` - Quantization support
149
150### Data Processing:
151- `datasets==3.6.0` - Dataset handling
152- `pandas==2.3.1` - Data manipulation
153- `numpy==2.3.1` - Numerical computing
154
155### Git Analysis:
156- `pydriller` - Git repository mining
157- `gitpython` - Git operations
158
159### Utilities:
160- `tqdm==4.67.1` - Progress bars
161- `wandb` - Experiment tracking
162- `evaluate==0.4.4` - Evaluation metrics
163
164## ๐Ÿ”„ Workflow
165
166### 1. Dataset Creation
167```bash
168cd dataset_builder
169python extract_linux_bugfixes.py          # Extract bug-fix data
170python format_for_training.py  # Convert format
171```
172
173### 2. Model Training
174```bash
175cd train
176python train_codellama_qlora_linux_bugfix.py                  # Train with QLoRA
177```
178
179### 3. Model Evaluation
180```bash
181cd evaluate
182python evaluate_linux_bugfix_model.py               # Evaluate performance
183```
184
185## ๐ŸŽฏ Key Design Principles
186
187### Modularity
188- Each component has a specific responsibility
189- Clear separation between data, training, and evaluation
190- Easy to modify or extend individual components
191
192### Efficiency
193- QLoRA for memory-efficient training
194- Parallel processing for dataset creation
195- Optimized for modern GPU hardware
196
197### Reproducibility
198- Version-controlled dependencies
199- Structured data formats
200- Comprehensive logging and evaluation
201
202### Scalability
203- Configurable parameters for different hardware
204- Support for distributed training
205- Efficient data handling with Git LFS
206
207## ๐Ÿ” File Naming Conventions
208
209- **Scripts**: Descriptive names with clear purpose
210- **Datasets**: Include size/version information
211- **Models**: Include architecture and method
212- **Results**: Include timestamp or version
213- **Configs**: Use `.json` or `.yaml` format
214
215## ๐Ÿ“ Documentation
216
217- **README.md**: Project overview and quick start
218- **PROJECT_STRUCTURE.md**: This detailed structure guide
219- **Model README**: Generated model cards in output directories
220- **Code Comments**: Inline documentation in all scripts
221
222This structure ensures the project is organized, maintainable, and easy to understand for both users and contributors.