TinyModels/Setfit-Banking-Spam
<div align="center">
<img src="https://readme-typing-svg.demolab.com?font=JetBrains+Mono&size=30&duration=2500&pause=700&color=FF0080¢er=true&vCenter=true&width=650&lines=%E2%9A%A1+TinyModels;SetFit+%C2%B7+Spam+Detection;16+Examples+%E2%86%92+Classifier;Small+Model.+Real+Task." alt="TinyModels animated header">
<br>
<img src="https://img.shields.io/badge/TinyModels-FF0080?style=for-the-badge&logoColor=white"> <img src="https://img.shields.io/badge/SetFit-7928CA?style=for-the-badge"> <img src="https://img.shields.io/badge/2--Class-Text%20Classification-FF0080?style=for-the-badge"> <img src="https://img.shields.io/badge/Few--Shot-Training-7928CA?style=for-the-badge"> <img src="https://img.shields.io/badge/Apache%202.0-000000?style=for-the-badge">
<br><br>
๐ง Tiny model. Tiny dataset. Real classification.
A compact few-shot spam classifier built with SetFit + BAAI/bge-small-en-v1.5.
</div>
<div align="center">
โก TINY โ FAST โ USEFUL
</div>
This model is a small experiment in few-shot text classification.
It learns to separate:
โ๏ธ HAM โ legitimate email ๐จ SPAM โ unwanted / suspicious email
The interesting part?
Only 16 labeled training examples.
โโโโโโโโโโโโโโโโโโโโโโโโ
โ 16 EXAMPLES โ
โ โ
โ 8 HAM + 8 SPAM โ
โโโโโโโโโโโโฌโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโ
โ BAAI/bge-small-en-v1.5โ
โโโโโโโโโโโโโฌโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโ
โ SetFit โ
โ Few-shot NLP โ
โโโโโโโโโฌโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโ
โ Logistic Regression โ
โโโโโโโโโโโโโฌโโโโโโโโโโโโ
โ
โโโโโโโโดโโโโโโโ
โผ โผ
HAM SPAM๐ก Model Status
<div align="center">
</div>
๐ Results
<div align="center">
91.4% Accuracy
91.4% Macro F1
<br>
Accuracy โโโโโโโโโโโโโโโโโโโโ 91.4%
Macro F1 โโโโโโโโโโโโโโโโโโโโ 91.4%</div>
These results come from a very small few-shot training setup. They should not be interpreted as a benchmark against production spam-filtering systems.
๐งฌ The TinyModels Recipe
โโโโโโโโโโโโโโโโโโโโโ
โ SetFit/enron โ
โ _spam โ
โโโโโโโโโโโฌโโโโโโโโโโ
โ
16 examples
โ
โผ
โโโโโโโโโโโโโโโโโโโโโ
โ BGE-small โ
โ text encoder โ
โโโโโโโโโโโฌโโโโโโโโโโ
โ
embeddings
โ
โผ
โโโโโโโโโโโโโโโโโโโโโ
โ SetFit โ
โ contrastive loss โ
โโโโโโโโโโโฌโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโ
โ LogisticRegressionโ
โโโโโโโโโโโฌโโโโโโโโโโ
โ
โโโโโโโโดโโโโโโโ
โผ โผ
โ๏ธ HAM ๐จ SPAMTraining configuration
backbone: BAAI/bge-small-en-v1.5
dataset: SetFit/enron_spam
examples:
total: 16
per_class: 8
setfit:
iterations: 20
epochs: 1
batch_size: 16
loss: CosineSimilarityLoss
distance_metric: cosine_distance
classifier:
type: LogisticRegression
seed: 42๐ Run It
pip install -q setfitfrom setfit import SetFitModel
model = SetFitModel.from_pretrained(
"TinyModels/setfit-banking-spam"
)
text = """
Congratulations! You have won $1,000,000.
Click here immediately to claim your prize.
"""
prediction = model(text)
print(prediction)Example output
1๐งช Quick Examples
๐จ Spam
CONGRATULATIONS!!!
You have been selected to receive
$1,000,000. Click the link below
to claim your prize immediately.โ SPAMโ๏ธ Legitimate
Please find attached the global markets
monitor for the week ending 12 January 2001.โ HAM๐ง Why SetFit?
Traditional supervised classification can require a large labeled dataset.
SetFit takes a different route:
Large Dataset
โ
โ
โ
โโโโโโโผโโโโโโ
โ SetFit โ
โโโโโโโฌโโโโโโ
โ
โผ
Few labeled examples
โ
โผ
Useful classifierThis makes the experiment useful for exploring:
- โก Few-shot learning
- ๐ง Sentence embeddings
- ๐ Text classification
- ๐จ Spam detection
- ๐ฌ Efficient training
- ๐ค SetFit
๐งฉ Architecture
<div align="center">
INPUT
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ BAAI/bge-small-en-v1.5 โ
โ โ
โ Sentence Transformer โ
โโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโ
โ
โผ
Dense Embedding
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Logistic Regression โ
โโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโ
โ
โโโโโโโโดโโโโโโโ
โผ โผ
HAM SPAM
โ๏ธ ๐จ</div>
โ ๏ธ Limitations
This is intentionally a tiny experimental model.
Because only 16 examples were used for training:
- Performance can vary on unseen data.
- Domain shift can significantly affect predictions.
- Unusual spam may be missed.
- Enron-style email does not represent every modern spam pattern.
- The reported score comes from a lightweight few-shot experiment.
Do not use this model as the sole component of a security-critical email filtering system.
๐ฆ Model Identity
<div align="center">
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ TINYMODEL โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ โ
โ MODEL SetFit Banking Spam โ
โ BACKBONE BGE-small โ
โ TASK Binary Classification โ
โ DATA Enron Spam โ
โ EXAMPLES 16 โ
โ RESULT 91.4% Accuracy โ
โ โ
โ STATUS โ EXPERIMENTAL โ
โ โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ</div>
<div align="center">
โก 16 Examples.
๐ง One Small Model.
๐จ One Real Task.
<br>
<img src="https://img.shields.io/badge/BUILT%20BY-TinyModels-FF0080?style=for-the-badge"> <img src="https://img.shields.io/badge/POWERED%20BY-SetFit-7928CA?style=for-the-badge">
<br><br>
TinyModels โ building small models that actually do things.
</div>
