Team Ai
Modelpublic

azherali/CodeGenDetect-CodeBert

sourceHugging Faceupdated 10mo agoView on Hugging Face
0likes177downloads
README.md89 linesDownload Raw Back to root
1# CodeGenDetect-CodeBERT2 3**Model Name:** `azherali/CodeGenDetect-CodeBert`  4**Task:** Code Generation Detection (Human vs Machine Generated Code)  5**Languages Supported:** C++, Java, Python  6**Base Model:** CodeBERT  7**Author:** Azher Ali  8 9---10 11## ๐Ÿ“Œ Model Overview12 13`CodeGenDetect-CodeBert` is a transformer-based classification model designed to distinguish **human-written code** from **machine-generated code** produced by Large Language Models (LLMs). The model is fine-tuned on multilingual source code data spanning **C++**, **Java**, and **Python**, making it suitable for real-world, cross-language code analysis tasks.14 15Built on top of **CodeBERT**, the model leverages contextual and structural representations of source code to capture subtle stylistic, syntactic, and semantic patterns that differentiate human-authored code from AI-generated code.16 17---18 19## ๐ŸŽฏ Intended Use Cases20 21This model is well-suited for:22 23- **Academic integrity & plagiarism detection**24- **LLM-generated code identification**25- **Code authenticity verification**26- **Research on AI-generated programming artifacts**27- **Code forensics and auditing pipelines**28 29---30 31## ๐Ÿง  Model Details32 33- **Architecture:** Transformer-based (CodeBERT)34- **Task Type:** Binary Sequence Classification35- **Labels:**36  - `0` โ†’ Human-generated code37  - `1` โ†’ Machine-generated (LLM) code38- **Input:** Source code as plain text39- **Output:** Class probabilities and predicted label40 41---42 43## ๐ŸŒ Supported Programming Languages44 45The model has been trained and evaluated on code written in:46 47- **C++**48- **Java**49- **Python**50 51It generalizes across these languages by learning language-agnostic code patterns while still capturing language-specific constructs.52 53---54 55## ๐Ÿ‹๏ธ Training Summary56 57- **Training Objective:** Binary cross-entropy loss for classification58- **Tokenization:** CodeBERT tokenizer with fixed-length padding and truncation59- **Optimization:** Fine-tuned using modern deep learning best practices60- **Evaluation Metrics:** Accuracy, Precision, Recall, F1-score61 62The training data includes both human-written code and code generated by modern LLMs to ensure realistic detection performance.63 64---65 66## ๐Ÿš€ Example Usage67 68```python69from transformers import AutoTokenizer, AutoModelForSequenceClassification70import torch71 72model_name = "azherali/CodeGenDetect-CodeBert"73 74tokenizer = AutoTokenizer.from_pretrained(model_name)75model = AutoModelForSequenceClassification.from_pretrained(model_name)76 77code_snippet = """78def add(a, b):79    return a + b80"""81 82inputs = tokenizer(code_snippet, return_tensors="pt", truncation=True, padding=True)83outputs = model(**inputs)84 85prediction = torch.argmax(outputs.logits, dim=1).item()86label = "Machine-generated" if prediction == 1 else "Human-written"87 88print(label)89