Team Ai
Modelpublic

sugiv/modernbert-us-stablecoin-encoder

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
0likes
Model Card

ModernBERT with LoRA - US Stablecoin Regulatory Encoder

Production-ready encoder model for semantic search over US stablecoin regulatory documents.

Fine-tuned from answerdotai/ModernBERT-base using LoRA on 10,260 synthetic query-document triplets.

๐ŸŽฏ Performance Highlights

MetricScorevs Base Model
NDCG@100.9236+517% (5.2x better) ๐Ÿ”ฅ
MRR@100.8991+714% (7.1x better) ๐Ÿ”ฅ
Recall@100.9961+250% (2.5x better) ๐Ÿ”ฅ
Recall@1001.0000Perfect - never misses docs
  • โ€”โœ… 92.3% of queries improved (9,468 out of 10,260)
  • โ€”โœ… Statistically significant: p < 0.001
  • โ€”โœ… 1ms/query inference on A100 GPU
  • โ€”โœ… 8.8MB adapter (not 149MB full model)

๐Ÿš€ Quick Start

bash
pip install "transformers>=4.48" "peft>=0.14" "huggingface_hub>=0.27" torch
Important: ModernBERT requires transformers >= 4.48. Older versions will fail with KeyError: 'modernbert'.

Requirements

LibraryMinimum VersionNotes
transformers>= 4.48ModernBERT architecture support
peft>= 0.14Compatible hf_hub_download API
huggingface_hub>= 0.27No deprecated use_auth_token
torch>= 2.0CUDA support
bash
pip install "transformers>=4.48" "peft>=0.14" "huggingface_hub>=0.27" torch

Loading Notes

Tokenizer: Load from the base model (answerdotai/ModernBERT-base), NOT from this adapter repo. The adapter repo stores LoRA weights only; the tokenizer lives with the base model.

UNEXPECTED keys on load: When loading AutoModel.from_pretrained("answerdotai/ModernBERT-base"), you may see warnings about UNEXPECTED keys (head.norm.weight, head.dense.weight, decoder.bias). These are the base model's MLM (masked language model) head weights that exist in the pretrained checkpoint but are not used by AutoModel (which loads only the encoder backbone). This is completely normal and safe to ignore.

python
from transformers import AutoModel, AutoTokenizer
from peft import PeftModel
import torch

# Load model
base_model = AutoModel.from_pretrained(
    "answerdotai/ModernBERT-base",
    trust_remote_code=True
)
model = PeftModel.from_pretrained(
    base_model, 
    "sugiv/modernbert-us-stablecoin-encoder"
)
model.eval()

# IMPORTANT: Load tokenizer from BASE MODEL, not from adapter repo
tokenizer = AutoTokenizer.from_pretrained(
    "answerdotai/ModernBERT-base",
    trust_remote_code=True
)

def encode(text, max_length=512):
    inputs = tokenizer(
        text, padding=True, truncation=True,
        max_length=max_length, return_tensors="pt"
    )
    with torch.no_grad():
        outputs = model(**inputs)
        embeddings = outputs.last_hidden_state.mean(dim=1)
        embeddings = embeddings / embeddings.norm(dim=1, keepdim=True)
    return embeddings

# Example
query = "What are reserve requirements for stablecoin issuers?"
query_emb = encode(query)

๐Ÿ“Š Training Details

Model Architecture

  • โ€”Base: answerdotai/ModernBERT-base (149M params)
  • โ€”Method: LoRA (rank=16, alpha=32, dropout=0.1)
  • โ€”Targets: Wqkv, Wo (attention projections)
  • โ€”Trainable: 2.3M params (1.52% of base)

Training Data

  • โ€”10,260 triplets (8,208 train / 2,052 val)
  • โ€”38 documents: US stablecoin regulations
  • โ€”Query types: Factual, policy, comparison, compliance, interpretive
  • โ€”Negatives: 3 hard negatives per query (BM25)

Training Config

  • โ€”GPU: NVIDIA A100-80GB
  • โ€”Time: 6 minutes (early stopped at Epoch 1)
  • โ€”Batch size: 16 ร— 2 grad accumulation = 32 effective
  • โ€”Learning rate: 2e-4 with cosine schedule
  • โ€”Loss: InfoNCE (temp=0.05)
  • โ€”Early stopping: NDCG@10 โ‰ฅ 0.75 (achieved 0.8472)

๐ŸŽ“ Domain Specialization

Trained to understand US stablecoin regulatory concepts:

  • โ€”Federal Reserve Act, Dodd-Frank, Bank Holding Company Act
  • โ€”CFTC, OCC, FSOC, NY DFS BitLicense
  • โ€”Payment Stablecoin Act, STABLE Act
  • โ€”Reserve requirements, redemption rights, custody standards
  • โ€”Qualified Stablecoin Issuer (QSI)

๐Ÿ“ˆ Use Cases

  1. 1.RAG for Regulatory Q&A - Retrieve context for LLMs
  2. 2.Compliance Search - Find relevant regulations
  3. 3.Legal Research - Cross-reference requirements
  4. 4.Policy Analysis - Compare regulatory frameworks

๐Ÿ” Evaluation

Metrics Explained:

  • โ€”NDCG@10 (0.9236): 92% as good as perfect ranking in top 10
  • โ€”MRR@10 (0.8991): First result at avg position 1.11
  • โ€”Recall@100 (1.0000): Never misses relevant docs

Validation: 10,260 queries ร— 38 docs = full corpus ranking

๐Ÿšจ Limitations

  • โ€”Domain-specific (US stablecoin regulations only)
  • โ€”Small corpus (38 documents)
  • โ€”English only
  • โ€”Snapshot from March 2026
  • โ€”1.9% queries slightly degraded vs base

๐Ÿ“œ License

Apache 2.0 (following ModernBERT-base)

๐Ÿ™ Acknowledgments


Status: ๐ŸŸข Production Ready Updated: March 10, 2026 Version: 1.0.0