Team Ai
Modelpublic

emekaphilians/trustchainai-codebert

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes19downloads
Model Card

trustchainai-codebert

Fine-tuned CodeBERT for Solidity smart contract vulnerability detection.

Part of the TrustChainAI project — an AI-powered smart contract auditor with explainability and ethics monitoring, built to make blockchain security accessible to African and emerging-market Web3 ecosystems.


Model Performance

MetricScore
F1 (weighted, test set)98.6%
Eval Loss0.0428
Test Samples1,032 contracts
Classes13 vulnerability categories

How to Use

python
from transformers import pipeline

classifier = pipeline(
    "text-classification",
    model="emekaphilians/trustchainai-codebert"
)

contract = """
pragma solidity ^0.8.0;
contract Vulnerable {
    mapping(address => uint) public balances;
    function withdraw() external {
        uint amt = balances[msg.sender];
        (bool ok,) = msg.sender.call{value: amt}("");
        balances[msg.sender] = 0;
    }
}
"""

result = classifier(contract[:512])
print(result)
# [{'label': 'reentrancy', 'score': 0.997}]

Label Schema

IDLabelDescription
0safeNo vulnerability detected
1reentrancyReentrancy attack (DAO-style)
2integer_overflowArithmetic overflow / underflow
3access_controlUnprotected ownership or selfdestruct
4txoriginphishingtx.origin used for authentication
5dos_gasUnbounded loop / gas exhaustion
6unchecked_callExternal call return value ignored
7frontrunningmevMempool-visible state / TOD
8timestamp_dependenceblock.timestamp manipulation
9proxystoragecollisionDelegatecall storage slot collision
10flashloanoracleOracle price manipulation via flash loan
11flashloansingle_blockSingle-block liquidity attack
12misnamed_constructorPre-Solidity-0.5 constructor naming bug
13otherMulti-class or miscellaneous vulnerability

Training Data

Assembled from four open-source sources using the prepare_datasets.py pipeline:

SourceContracts
SmartBugs Curated143
SolidiFI Benchmark1,700
DeFiHackLabs729
Not-So-Smart Contracts25
Synthetic augmentation3,600
Total (after dedup)6,879

Split: 70% train / 15% val / 15% test (stratified by label).


Training Details

ParameterValue
Base modelmicrosoft/codebert-base
Epochs5 (best checkpoint at epoch 2)
Batch size16
Learning rate2e-5
OptimizerAdamW (weight decay 0.01, warmup 100 steps)
Max token length512
Mixed precisionfp16
HardwareGoogle Colab T4 GPU

Intended Use

  • —Pre-deployment security screening of Solidity smart contracts
  • —Automated vulnerability triage for DeFi protocols
  • —Research baseline for smart contract security ML benchmarks
  • —Integration into the TrustChainAI multi-agent audit pipeline

Out-of-Scope Use

  • —This model is not a substitute for a full professional security audit on high-value contracts
  • —Performance on Vyper, Yul, or non-EVM contracts is untested
  • —The tx_origin_phishing class has limited real training samples (28); treat predictions for this class with extra caution

Limitations & Bias

  • —Synthetic augmentation was used for 9 of 13 classes to compensate for dataset scarcity. Synthetic contracts may not fully capture real-world obfuscation patterns.
  • —The tx_origin_phishing class had only 28 real-world training samples; model confidence for this class may be lower in practice.
  • —Training data skews toward older Solidity vulnerability patterns (pre-0.8). Newer attack vectors may be underrepresented.

Citation

bibtex
@misc{trustchainai2025,
  author = {Emeka Philian},
  title  = {TrustChainAI: AI-Powered Smart Contract Auditor},
  year   = {2025},
  url    = {https://github.com/emekaphilian/TrustChainAI}
}

Links