mahdin70/codebert-devign-code-vulnerability-detector
1104
1---2library_name: transformers3tags:4- Code5- Vulnerability6- Detection7datasets:8- DetectVul/devign9language:10- en11base_model:12- microsoft/codebert-base13license: mit14metrics:15- accuracy16- precision17- f118- recall19---20 21## CodeBERT for Code Vulnerability Detection22 23## Model Summary24This model is a fine-tuned version of **microsoft/codebert-base**, optimized for detecting vulnerabilities in code. It is trained on the **DetectVul/devign** dataset. The model takes in a code snippet and classifies it as either **safe (0)** or **vulnerable (1)**.25 26## Model Details27 28- **Developed by:** Mukit Mahdin29- **Finetuned from:** `microsoft/codebert-base`30- **Language(s):** English (for code comments & metadata), C/C++31- **License:** MIT32- **Task:** Code vulnerability detection33- **Dataset Used:** `DetectVul/devign`34- **Architecture:** Transformer-based sequence classification35 36## Uses37 38### Direct Use39This model can be used for **static code analysis**, security audits, and automatic vulnerability detection in software repositories. It is useful for:40- **Developers**: To analyze their code for potential security flaws.41- **Security Teams**: To scan repositories for known vulnerabilities.42- **Researchers**: To study vulnerability detection in AI-powered systems.43 44### Downstream Use45This model can be integrated into **IDE plugins**, **CI/CD pipelines**, or **security scanners** to provide real-time vulnerability detection.46 47### Out-of-Scope Use48- The model is **not meant to replace human security experts**.49- It may not generalize well to **languages other than C/C++**.50- False positives/negatives may occur due to dataset limitations.51 52## Bias, Risks, and Limitations53- **False Positives & False Negatives:** The model may flag safe code as vulnerable or miss actual vulnerabilities.54- **Limited to C/C++:** The model was trained on a dataset primarily composed of **C and C++ code**. It may not perform well on other languages.55- **Dataset Bias:** The training data may not cover all possible vulnerabilities.56 57### Recommendations58Users should **not rely solely on the model** for security assessments. Instead, it should be used alongside **manual code review and static analysis tools**.59 60## How to Get Started with the Model61Use the code below to load the model and run inference on a sample code snippet:62 63```python64from transformers import AutoTokenizer, AutoModelForSequenceClassification65import torch66 67# Load the fine-tuned model68tokenizer = AutoTokenizer.from_pretrained("microsoft/codebert-base")69model = AutoModelForSequenceClassification.from_pretrained("mahdin70/codebert-devign-code-vulnerability-detector")70 71# Sample code snippet72code_snippet = '''73void process(char *input) {74 char buffer[50];75 strcpy(buffer, input); // Potential buffer overflow76}77'''78 79# Tokenize the input80inputs = tokenizer(code_snippet, return_tensors="pt", truncation=True, padding="max_length", max_length=512)81 82# Run inference83with torch.no_grad():84 outputs = model(**inputs)85 predictions = torch.nn.functional.softmax(outputs.logits, dim=-1)86 predicted_label = torch.argmax(predictions, dim=1).item()87 88# Output the result89print("Vulnerable Code" if predicted_label == 1 else "Safe Code")90```91 92## Training Details93 94### Training Data95- **Dataset:** `DetectVul/devign`96- **Classes:** `0 (Safe)`, `1 (Vulnerable)`97- **Size:** `21800` Code Snippets98 99### Training Procedure100- **Optimizer:** AdamW101- **Loss Function:** CrossEntropyLoss102- **Batch Size:** 8103- **Learning Rate:** 2e-05104- **Epochs:** 3105- **Hardware Used:** 2x T4 GPU106 107### Metrics108| Metric | Score |109|------------|-------------|110| **Train Loss** | 0.5898 |111| **Evaluation Loss** | 0.6153 |112| **Accuracy** | 64.09% |113| **F1 Score** | 46.42% |114| **Precision** | 73.78% |115| **Recall** | 33.86% |116 117## Environmental Impact118 119| Factor | Value |120|-----------|----------|121| **GPU Used** | T4 GPU |122| **Training Time** | ~1 hour |