Team Ai
Modelpublic

TinyModels/Setfit-Banking-Spam

sourceHugging Faceapache-2.0updated 11d agoView on Hugging Face
0likes9downloads
Model Card

<div align="center">

<img src="https://readme-typing-svg.demolab.com?font=JetBrains+Mono&size=30&duration=2500&pause=700&color=FF0080&center=true&vCenter=true&width=650&lines=%E2%9A%A1+TinyModels;SetFit+%C2%B7+Spam+Detection;16+Examples+%E2%86%92+Classifier;Small+Model.+Real+Task." alt="TinyModels animated header">

<br>

<img src="https://img.shields.io/badge/TinyModels-FF0080?style=for-the-badge&logoColor=white"> <img src="https://img.shields.io/badge/SetFit-7928CA?style=for-the-badge"> <img src="https://img.shields.io/badge/2--Class-Text%20Classification-FF0080?style=for-the-badge"> <img src="https://img.shields.io/badge/Few--Shot-Training-7928CA?style=for-the-badge"> <img src="https://img.shields.io/badge/Apache%202.0-000000?style=for-the-badge">

<br><br>

๐Ÿง  Tiny model. Tiny dataset. Real classification.

A compact few-shot spam classifier built with SetFit + BAAI/bge-small-en-v1.5.

</div>


<div align="center">

โšก TINY โ†’ FAST โ†’ USEFUL

</div>

This model is a small experiment in few-shot text classification.

It learns to separate:

โœ‰๏ธ HAM โ€” legitimate email ๐Ÿšจ SPAM โ€” unwanted / suspicious email

The interesting part?

Only 16 labeled training examples.

text
                 โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                 โ”‚      16 EXAMPLES     โ”‚
                 โ”‚                      โ”‚
                 โ”‚   8 HAM  +  8 SPAM   โ”‚
                 โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                            โ”‚
                            โ–ผ
                โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                โ”‚ BAAI/bge-small-en-v1.5โ”‚
                โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                            โ”‚
                            โ–ผ
                   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                   โ”‚     SetFit     โ”‚
                   โ”‚  Few-shot NLP  โ”‚
                   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                           โ”‚
                           โ–ผ
                โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                โ”‚ Logistic Regression   โ”‚
                โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                            โ”‚
                     โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”
                     โ–ผ             โ–ผ
                   HAM           SPAM

๐Ÿ“ก Model Status

<div align="center">

โš™๏ธ Component๐Ÿ”ง Configuration
TaskSpam Classification
Classes2
BackboneBAAI/bge-small-en-v1.5
FrameworkSetFit
DatasetSetFit/enron_spam
Training Examples16
Examples / Class8
Iterations20
Epochs1
Batch Size16
Classification HeadLogistic Regression
LanguageEnglish

</div>


๐Ÿ“Š Results

<div align="center">

91.4% Accuracy

91.4% Macro F1

<br>

text
Accuracy   โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘  91.4%
Macro F1   โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘  91.4%

</div>

These results come from a very small few-shot training setup. They should not be interpreted as a benchmark against production spam-filtering systems.

๐Ÿงฌ The TinyModels Recipe

text
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚   SetFit/enron    โ”‚
              โ”‚      _spam        โ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                        โ”‚
                  16 examples
                        โ”‚
                        โ–ผ
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚     BGE-small     โ”‚
              โ”‚   text encoder    โ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                        โ”‚
                   embeddings
                        โ”‚
                        โ–ผ
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚      SetFit       โ”‚
              โ”‚ contrastive loss  โ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                        โ”‚
                        โ–ผ
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚ LogisticRegressionโ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                        โ”‚
                 โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”
                 โ–ผ             โ–ผ
              โœ‰๏ธ HAM        ๐Ÿšจ SPAM

Training configuration

yaml
backbone: BAAI/bge-small-en-v1.5
dataset: SetFit/enron_spam

examples:
  total: 16
  per_class: 8

setfit:
  iterations: 20
  epochs: 1
  batch_size: 16
  loss: CosineSimilarityLoss
  distance_metric: cosine_distance

classifier:
  type: LogisticRegression

seed: 42

๐Ÿš€ Run It

bash
pip install -q setfit
python
from setfit import SetFitModel

model = SetFitModel.from_pretrained(
    "TinyModels/setfit-banking-spam"
)

text = """
Congratulations! You have won $1,000,000.
Click here immediately to claim your prize.
"""

prediction = model(text)

print(prediction)

Example output

text
1

๐Ÿงช Quick Examples

๐Ÿšจ Spam

text
CONGRATULATIONS!!!

You have been selected to receive
$1,000,000. Click the link below
to claim your prize immediately.
text
โ†’ SPAM

โœ‰๏ธ Legitimate

text
Please find attached the global markets
monitor for the week ending 12 January 2001.
text
โ†’ HAM

๐Ÿง  Why SetFit?

Traditional supervised classification can require a large labeled dataset.

SetFit takes a different route:

text
        Large Dataset
             โœ•
             โ”‚
             โ”‚
       โ”Œโ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”
       โ”‚  SetFit   โ”‚
       โ””โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”˜
             โ”‚
             โ–ผ
     Few labeled examples
             โ”‚
             โ–ผ
       Useful classifier

This makes the experiment useful for exploring:

  • โ€”โšก Few-shot learning
  • โ€”๐Ÿง  Sentence embeddings
  • โ€”๐Ÿ“š Text classification
  • โ€”๐Ÿšจ Spam detection
  • โ€”๐Ÿ”ฌ Efficient training
  • โ€”๐Ÿค— SetFit

๐Ÿงฉ Architecture

<div align="center">

text
INPUT
  โ”‚
  โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚  BAAI/bge-small-en-v1.5     โ”‚
โ”‚                             โ”‚
โ”‚  Sentence Transformer       โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
               โ”‚
               โ–ผ
        Dense Embedding
               โ”‚
               โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚    Logistic Regression      โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
               โ”‚
        โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”
        โ–ผ             โ–ผ
      HAM           SPAM
       โœ‰๏ธ             ๐Ÿšจ

</div>


โš ๏ธ Limitations

This is intentionally a tiny experimental model.

Because only 16 examples were used for training:

  • โ€”Performance can vary on unseen data.
  • โ€”Domain shift can significantly affect predictions.
  • โ€”Unusual spam may be missed.
  • โ€”Enron-style email does not represent every modern spam pattern.
  • โ€”The reported score comes from a lightweight few-shot experiment.

Do not use this model as the sole component of a security-critical email filtering system.


๐Ÿ“ฆ Model Identity

<div align="center">

text
โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚              TINYMODEL               โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚                                      โ”‚
โ”‚  MODEL      SetFit Banking Spam      โ”‚
โ”‚  BACKBONE   BGE-small                โ”‚
โ”‚  TASK       Binary Classification    โ”‚
โ”‚  DATA       Enron Spam               โ”‚
โ”‚  EXAMPLES   16                       โ”‚
โ”‚  RESULT     91.4% Accuracy           โ”‚
โ”‚                                      โ”‚
โ”‚  STATUS     โ— EXPERIMENTAL           โ”‚
โ”‚                                      โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

</div>


<div align="center">

โšก 16 Examples.

๐Ÿง  One Small Model.

๐Ÿšจ One Real Task.

<br>

<img src="https://img.shields.io/badge/BUILT%20BY-TinyModels-FF0080?style=for-the-badge"> <img src="https://img.shields.io/badge/POWERED%20BY-SetFit-7928CA?style=for-the-badge">

<br><br>

TinyModels โ€” building small models that actually do things.

</div>