Team Ai
Datasetpublic

uv-scripts/classification

Classification Scripts Text classification on HF Jobs: label a dataset with a model that needs no training, or train your own classifier from labelled examples. If you have seen Jev and other "System One" models: the models these scripts train are small, open versions of the same idea. They read a piece of data and return a label with a probability, and you can train one on your own labels. For example, this demo suggests task tags for any Hub dataset; its model was fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/classification.

sourceHugging Faceupdated 16d agoView on Hugging Face
18likes316downloads
README.md323 linesDownload Raw Back to root
1---2viewer: false3tags: [uv-script, classification, fine-tuning, few-shot, zero-shot, decision-model, setfit, gliner2, vllm, structured-outputs, hf-jobs]4---5 6# Classification Scripts7 8Text classification on [HF Jobs](https://huggingface.co/docs/hub/jobs): label a dataset with a model that needs no training, or train your own classifier from labelled examples.9 10If you have seen [Jev](https://docs.typesafe.ai/introduction) and other "System One" models: the11models these scripts train are small, open versions of the same idea. They read a piece of data and12return a label with a probability, and you can train one on your own labels. For example, this13[demo](https://huggingface.co/spaces/davanstrien/hub-task-tagger) suggests task tags for any Hub14dataset; its model was fine-tuned with `train-gliner2.py` in 17 minutes.15 16To try it on your own account, this command fine-tunes a classifier for British Library book titles17(Fiction / Non-fiction). Training takes about 2 minutes on a `t4-small` and costs about $0.02.18Accuracy goes from 0.767 zero-shot to 0.907 fine-tuned. You get a private model repo, and its card19shows both scores next to the majority-class baseline. Copy and paste it as it is: the model goes20to your own account.21 22```bash23hf jobs uv run --flavor t4-small --timeout 1h --secrets HF_TOKEN \24  https://huggingface.co/datasets/uv-scripts/classification/raw/main/train-gliner2.py \25  biglam/blbooksgenre gliner2-blbooks-genre \26  --dataset-config title_genre_classifiction --text-column title27```28 29You need the `hf` CLI, signed in, and Jobs credit: see the [Jobs quickstart](https://huggingface.co/docs/hub/jobs-quickstart).30 31| Script | What it does |32|--------|--------------|33| [`classify-gliner2.py`](#zero-shot-first-then-fine-tune-gliner2) | **Label a dataset** with GLiNER2: zero-shot from label names, or with a `train-gliner2.py` model |34| [`train-gliner2.py`](#zero-shot-first-then-fine-tune-gliner2) | **Fine-tune** [GLiNER2](https://huggingface.co/fastino), a small model (74M–287M) that already classifies zero-shot, and report the zero-shot score next to the fine-tuned one |35| [`train-setfit.py`](#few-shot-with-setfit-train-setfitpy) | **Few-shot** train a classifier from 8-64 labels per class with [SetFit](https://huggingface.co/docs/setfit) — runs on CPU or GPU |36| [`train-classifier.py`](#fine-tune-a-classifier-train-classifierpy) | **Fine-tune** an encoder into a classifier (default: [LFM2.5-Encoder-350M](https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M)) and push it to the Hub |37| [`classify-dataset.py`](#zero-shot-classification-classify-datasetpy) | **Zero-shot** classify a dataset with an instruction LLM (SmolLM3 + vLLM, structured outputs) |38| [`classify-dataset-sglang.py`](#zero-shot-classification-classify-datasetpy) | Zero-shot variant on SGLang (reasoning-aware `<think>` models) |39 40Pick by how many labels you have:41 42| Labels you have | Use | Hardware |43|---|---|---|44| none | `classify-gliner2.py --labels ...` for a cheap first pass; `classify-dataset.py` when the task needs an LLM's reasoning | small GPU (CPU works at ~1.4 rows/s); GPU |45| ~8-64 per class | `train-setfit.py` | CPU supported; GPU for faster training |46| a few hundred to a few thousand | `train-gliner2.py`, which also shows you what zero-shot already gets | small GPU (`t4-small`) |47| a few thousand or more | `train-classifier.py` | GPU |48 49The rungs chain: bootstrap labels with `classify-gliner2.py` or `classify-dataset.py`, review them,50then train a small dedicated model on what you kept. Each rung is one command, so you can move up51as you collect more labels.52 53## Zero-shot first, then fine-tune (GLiNER2)54 55Label a dataset with your own list of labels, see how far zero-shot gets you, then fine-tune a56small model on your labels, in minutes and for a few cents on one GPU. The result is a model57that returns a label and a probability for every row, and is small enough to run on a CPU.58 59[GLiNER2](https://github.com/fastino-ai/GLiNER2) ([models from Fastino](https://huggingface.co/fastino)) is a small encoder that reads the label names60as part of its input, so it classifies with no training at all, and fine-tuning teaches it what61your labels mean in your data. (For entity extraction with the original GLiNER library, see62[`uv-scripts/gliner`](https://huggingface.co/datasets/uv-scripts/gliner).) Two scripts:63 64- **`train-gliner2.py`** scores the base model zero-shot, fine-tunes it on your labels, and65  scores it again on the same held-out rows. The model card reports both scores next to the66  majority-class baseline.67- **`classify-gliner2.py`** labels a whole dataset. Pass `--labels` for zero-shot, or `--model`68  with a `train-gliner2.py` output; the tasks and labels are read from the model repo.69  Labels are passed as separate words; quote a label with spaces:70  `--labels World Sports Business "Science and technology"`. A fine-tuned model is passed71  by name: `--model gliner2-blbooks-genre`. A name on its own means your account; use `org/name`72  for an organisation.73 74### Quick start75 76```bash77# fine-tune: British Library book titles -> Fiction / Non-fiction78hf jobs uv run --flavor t4-small --timeout 1h --secrets HF_TOKEN \79  https://huggingface.co/datasets/uv-scripts/classification/raw/main/train-gliner2.py \80  biglam/blbooksgenre gliner2-blbooks-genre \81  --dataset-config title_genre_classifiction --text-column title82 83# label a dataset with that model84hf jobs uv run --flavor t4-small --timeout 1h --secrets HF_TOKEN \85  https://huggingface.co/datasets/uv-scripts/classification/raw/main/classify-gliner2.py \86  biglam/blbooksgenre blbooks-genre-predictions \87  --dataset-config title_genre_classifiction --text-column title --model gliner2-blbooks-genre88 89# or skip training: zero-shot from label names90hf jobs uv run --flavor t4-small --timeout 1h --secrets HF_TOKEN \91  https://huggingface.co/datasets/uv-scripts/classification/raw/main/classify-gliner2.py \92  fancyzhx/ag_news ag-news-topics --split test \93  --labels World Sports Business "Science and technology" --task-name topic94```95 96The second command labels the same rows the model was trained on, so it shows the workflow, not97the model's accuracy; the held-out scores are on the model card. For real use, point it at data98the model has not seen. Outputs are **private by default** (`--public` to opt out).99 100These commands use the default base model, `fastino/gliner2.5-multi-v1` (multilingual). For101English text, `--base-model fastino/gliner2.5-base-v1` is smaller and faster;102`gliner2.5-small-v1` is the fastest and loses about 4 points on the 52-tag example. For English zero-shot, try103`fastino/GLiNER2.5-Decide`. Sizes and speeds: [Choosing a model size](GLINER2-NOTES.md#choosing-a-model-size).104 105### Results106 107Fine-tuning beat zero-shot on every dataset below. The public-dataset runs took 2 to 30 minutes of108training on one GPU and cost $0.01 to $0.20; the 52-tag example took 17 minutes and about $1.50.109 110| Dataset | Task | Labels | Train rows × epochs | Train time | Train cost | Metric | Majority floor | Zero-shot | Fine-tuned |111|---|---|---|---|---|---|---|---|---|---|112| [`biglam/blbooksgenre`](https://huggingface.co/datasets/biglam/blbooksgenre) (book titles) | single-label | 2 | 1,562 × 5 | 141s | $0.02 | accuracy | 0.747 | 0.767 (0.753–0.782) | **0.907** (0.897–0.925) |113| [`fancyzhx/ag_news`](https://huggingface.co/datasets/fancyzhx/ag_news) | single-label | 4 | 2,000 × 2 | 125s | $0.01 | accuracy | 0.268 | 0.718 | **0.852** |114| [`google-research-datasets/go_emotions`](https://huggingface.co/datasets/google-research-datasets/go_emotions) | multi-label | 28 | 2,000 × 2 | 216s | $0.02 | micro F1 | — | 0.265 | **0.464** |115| [`SetFit/TREC-QC`](https://huggingface.co/datasets/SetFit/TREC-QC), two tasks in one model | single-label ×2 | 6 + 50 | 5,452 × 3 | 1,761s | $0.20 | accuracy | 0.276 / 0.246 | 0.542 / 0.468 | **0.954 / 0.876** |116| same, on `a10g-small`, default batch size, bf16 | single-label ×2 | 6 + 50 | 5,452 × 3 | 509s | $0.14 | accuracy | 0.276 / 0.246 | not run | **0.944 / 0.872** |117| [`stanfordnlp/imdb`](https://huggingface.co/datasets/stanfordnlp/imdb) (reviews; 15% truncated at 2,000 characters) | single-label | 2 | 1,000 × 1 | 122s | $0.01 | accuracy | 0.500 | 0.777 | **0.840** |118| Hub dataset task tags ([worked example](#worked-example-tagging-hub-datasets)), `--base-model fastino/gliner2.5-base-v1 --label-augmentation off`, on `rtx-pro-6000` | single-label choice from a fixed set | 52 | 16,000 × 5 | 17 min | ~$1.50 | top-1 in the owner's tags | 0.320 (always "text-generation") | 0.102 | **0.690** (2 seeds: 0.695 / 0.686) |119 120Most rows are single, deliberately small runs that test the script, not tuned results. Seed ranges,121out-of-memory history and GPU comparisons are in [GLINER2-NOTES.md](GLINER2-NOTES.md).122 123### Larger or fixed label sets124 125For tens of labels scored together (a taxonomy, a fixed tag list), or data in a bucket, see126[Larger or fixed label sets](GLINER2-NOTES.md#larger-or-fixed-label-sets) in the notes:127`--labels-file`, `--label-augmentation off`, local `--train-file`/`--eval-file` and `--export-predictions`.128 129### Worked example: tagging Hub datasets130 131A GLiNER2.5-base model fine-tuned with this script on 16,000 Hub datasets suggests task tags for132a dataset from its column names and first row, among the 52 tags the Hub offers. Its first133suggestion matches one of the owner's tags 69% of the time on 3,000 newer datasets from owners it134never saw. Owners' tags are a noisy target: in a hand-checked sample, about 1 in 10 datasets was135missing a tag that fits.136 137[Try the demo](https://huggingface.co/spaces/davanstrien/hub-task-tagger): paste a dataset id and138see the suggested tags, the owner's tags and the exact text the model read. On a free 2-vCPU Space139one prediction takes about 0.7–1 s.140[Model](https://huggingface.co/davanstrien/hub-task-tagger-gliner2.5-base) · [Notes on how it was trained](GLINER2-NOTES.md#a-larger-label-set-52-hub-task-tags)141 142### Good to know143 144- **Single-label and multi-label** are auto-detected from the label column; repeat `--label-column` to train several tasks in one model.145- **Label names are part of the prompt.** Real names (`Fiction`, `Sports`) work; integer codes make zero-shot meaningless.146- **Out of GPU memory, it restarts at a smaller batch size** and stops before pushing if even batch size 1 fails. Many labels or long texts want an `a10g-small`.147- **Always pass `--timeout`.** Older `hf` CLIs ignore the scripts' `[tool.hf-jobs]` header and stop the Job after 30 minutes.148 149More behaviour details, tested commands, findings and dead ends: [GLINER2-NOTES.md](GLINER2-NOTES.md).150 151## Few-shot with SetFit (`train-setfit.py`)152 153An alternative when you have only a handful of labelled examples per class (8-64) and want a sentence-transformer model.154 155Trains a [SetFit](https://huggingface.co/docs/setfit) classifier from a handful of labelled156examples per class. SetFit finetunes a sentence-transformer body on contrastive pairs, then fits a157logistic regression head on the resulting embeddings.158 159**Runs on CPU or GPU.** CPU is practical for small few-shot experiments. Use a GPU for faster160training, particularly with larger models, longer texts or more classes. The same model and161training settings work on either; the recipe uses the available accelerator automatically.162 163```bash164# 8 labels per class, on CPU165hf jobs uv run --flavor cpu-basic --timeout 20m --secrets HF_TOKEN \166  https://huggingface.co/datasets/uv-scripts/classification/raw/main/train-setfit.py \167  fancyzhx/ag_news username/ag-news-setfit --num-samples 8168 169# Same model and training settings on a GPU for faster training170hf jobs uv run --flavor t4-small --timeout 20m --secrets HF_TOKEN \171  https://huggingface.co/datasets/uv-scripts/classification/raw/main/train-setfit.py \172  fancyzhx/ag_news username/ag-news-setfit-gpu --num-samples 8173```174 175### Measured176 1778 labels per class, seed 42, evaluated on each dataset's own held-out split (capped at 500178examples, 1000 for banking77):179 180| Dataset | Classes | Labels used | Body | Flavor | Training | Accuracy | Macro F1 |181|---|---|---|---|---|---|---|---|182| [`SetFit/enron_spam`](https://huggingface.co/datasets/SetFit/enron_spam) | 2 | 16 | MiniLM-L6 | `cpu-basic` | 78s | 0.924 | 0.924 |183| [`fancyzhx/ag_news`](https://huggingface.co/datasets/fancyzhx/ag_news) | 4 | 32 | MiniLM-L6 | `cpu-basic` | 118s | 0.804 | 0.807 |184| [`legacy-datasets/banking77`](https://huggingface.co/datasets/legacy-datasets/banking77) | 77 | 616 | MiniLM-L6 | `t4-small` | 18s | 0.803 | 0.789 |185| [`dair-ai/emotion`](https://huggingface.co/datasets/dair-ai/emotion) | 6 | 48 | MiniLM-L6 | `cpu-basic` | 216s | 0.370 | 0.325 |186 187**Single seed each — these do not rank models or predict your dataset.** Few-shot results vary188substantially with which examples happen to get sampled; SetFit's own benchmarks report mean and189standard deviation across ten seeds for exactly this reason. Run your own task before trusting190any of these numbers.191 192### Good to know193 194- **Default body**: [`all-MiniLM-L6-v2`](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2) (22M), chosen for CPU speed. Swap it with `--body-model`; set `--max-seq-length` within its context window and add any task prefix it needs yourself. See [Choosing another body or longer context](SETFIT-NOTES.md#choosing-another-body-or-longer-context).195- **Single-label only.** A multi-label column exits with a pointer to `train-classifier.py`.196- **Beat more than the majority baseline.** Every run reports it, but `emotion` beat it by under 2 points; compare with TF-IDF plus logistic regression or zero-shot on the same rows. See [Compare more than the majority baseline](SETFIT-NOTES.md#compare-more-than-the-majority-baseline).197- **Many classes: watch the pair count.** Pairs grow with the square of the training-set size; at 77 classes use `--sampling-strategy undersampling`. See [Many classes](SETFIT-NOTES.md#many-classes-watch-the-pair-count).198- **It refuses runs projected above `--max-minutes`** (default 60). On `biglam/hansard_speech` (2.7M speeches, 28 parties) it refused to train and produced no score, and [the worked failure](SETFIT-NOTES.md#real-world-data-a-worked-failure) shows why.199- **Load it with `SetFitModel.from_pretrained(repo)`**, not `AutoModelForSequenceClassification`: a SetFit model is a sentence-transformer body plus a scikit-learn head.200 201Evaluation split, metrics, dropped rows and `--private`: [SETFIT-NOTES.md](SETFIT-NOTES.md#behaviour-details).202 203## Fine-tune a classifier (`train-classifier.py`)204 205Fine-tunes a text-classification encoder on any Hub dataset and pushes the trained model206back to the Hub — download, train, evaluate, push, and reload-verify in one job.207 208- **Default model**: [LiquidAI/LFM2.5-Encoder-350M](https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M) — a bidirectional encoder that LiquidAI reports beats ModernBERT-base on GLUE/SuperGLUE, and handles 8,192-token documents. Any Hub encoder works via `--model` (ModernBERT, BERT, DeBERTa, …).209- **Single-label and multi-label**, auto-detected from the label column (`ClassLabel`/string/int → cross-entropy; list of labels → BCE + per-label threshold tuning).210- **Round-trippable artifacts**: standard architectures produce standard models; encoders without a classification head (like LFM2.5) get a generic mean-pooling head pushed as custom code, so `AutoModelForSequenceClassification.from_pretrained(..., trust_remote_code=True)` always works.211 212```bash213# single-label (ag_news has a ClassLabel column)214hf jobs uv run --flavor a10g-small --secrets HF_TOKEN \215  https://huggingface.co/datasets/uv-scripts/classification/raw/main/train-classifier.py \216  fancyzhx/ag_news username/news-classifier217 218# multi-label (go_emotions has a list-of-labels column)219hf jobs uv run --flavor a10g-small --secrets HF_TOKEN \220  https://huggingface.co/datasets/uv-scripts/classification/raw/main/train-classifier.py \221  google-research-datasets/go_emotions username/emotion-classifier --label-column labels222```223 224Key options: `--model`, `--max-length` (512 default; up to 8192 with225`--gradient-checkpointing` and a small `--batch-size` on a10g/a100), `--epochs`, `--lr`,226`--batch-size`, `--max-samples` (smoke runs), `--eval-split` (auto-detects227validation/test, or holds out 10% of train). Run `uv run train-classifier.py --help` for all.228 229### Worked example: classify dataset cards by task230 231[`davanstrien/dataset-cards-with-task-categories`](https://huggingface.co/datasets/davanstrien/dataset-cards-with-task-categories)232contains 21k Hub dataset cards (frontmatter stripped) labelled with their `task_categories`233metadata — a real multi-label task over long documents:234 235```bash236hf jobs uv run --flavor a10g-small --secrets HF_TOKEN \237  https://huggingface.co/datasets/uv-scripts/classification/raw/main/train-classifier.py \238  davanstrien/dataset-cards-with-task-categories username/dataset-card-task-classifier \239  --label-column labels --max-length 1024 --batch-size 8 --grad-accum 2240```241 242The output model predicts likely task categories from a card's prose — e.g. for suggesting243metadata on datasets that lack it.244 245### Training a standard encoder instead246 247`--model answerdotai/ModernBERT-base` (or any encoder with a native classification head)248produces a plain, vLLM-servable model — pair it with249[`uv-scripts/vllm`](https://huggingface.co/datasets/uv-scripts/vllm)'s250`classify-dataset.py` for large-scale batch inference with the model you just trained.251 252## Zero-shot classification (`classify-dataset.py`)253 254Label a dataset with an instruction LLM and no training data. You give a list of labels; the255model picks one per row, and guided decoding (structured outputs) makes sure every answer is one256of your labels. The default model is257[HuggingFaceTB/SmolLM3-3B](https://huggingface.co/HuggingFaceTB/SmolLM3-3B); any instruction258model works via `--model`. The result is your dataset with a new `classification` column.259 260Two scripts do this. `classify-dataset.py` runs on vLLM and is the one to start with.261`classify-dataset-sglang.py` runs the same task on SGLang, for reasoning models that write262`<think>` traces; its options differ (`--reasoning`, `--save-reasoning`, `--batch-size`,263`--grammar-backend`), so check its `--help`.264 265### Quick start266 267A GPU is required. Run on Jobs with the vLLM image:268 269```bash270hf jobs uv run --flavor l4x1 --image vllm/vllm-openai:latest --secrets HF_TOKEN \271  https://huggingface.co/datasets/uv-scripts/classification/raw/main/classify-dataset.py \272  --input-dataset stanfordnlp/imdb \273  --column text \274  --labels "positive,negative" \275  --output-dataset username/imdb-classified \276  --max-samples 100 --shuffle277```278 279`--max-samples` with `--shuffle` takes a random sample, which matters for datasets sorted by280date or label. Drop both to label the whole split.281 282### With reasoning and label descriptions283 284```bash285hf jobs uv run --flavor l4x1 --image vllm/vllm-openai:latest --secrets HF_TOKEN \286  https://huggingface.co/datasets/uv-scripts/classification/raw/main/classify-dataset.py \287  --input-dataset user/support-tickets \288  --column content \289  --labels "bug,feature_request,question,other" \290  --label-descriptions "bug:code or product not working as expected,feature_request:asking for new functionality,question:seeking help or clarification,other:general comments or feedback" \291  --enable-reasoning \292  --output-dataset username/tickets-classified293```294 295With `--enable-reasoning` the model thinks step by step before it answers, and the output also296has `reasoning` and `parsing_success` columns. Reasoning mode turns off structured outputs: the297model must end with `{"label": "..."}`, and rows where that cannot be parsed are marked in298`parsing_success`. It is slower, but you can read why each label was chosen.299 300### Options301 302| Option | What it does |303|---|---|304| `--model` | Model to use (default `HuggingFaceTB/SmolLM3-3B`) |305| `--label-descriptions` | `label:description,...` pairs that tell the model what each label means |306| `--enable-reasoning` | Think before answering; adds `reasoning` and `parsing_success` columns |307| `--split` | Split to process (default `train`) |308| `--max-samples` | Label only the first N rows (or N random rows with `--shuffle`) |309| `--shuffle`, `--shuffle-seed` | Shuffle before `--max-samples` (seed default 42) |310 311Run `uv run classify-dataset.py --help` for all options.312 313### Good to know314 315- **Speed**: about 50-100 texts/second for SmolLM3-3B on an A10, and 20-50 for 7B models.316  `l4x1` is a good start; use `a10g-large` or larger for 7B+ models or out-of-memory errors.317- **Text handling**: texts shorter than 3 characters and empty values are skipped; texts are318  truncated to 4,000 characters.319- **Label names matter.** Use clear, distinct names, add `--label-descriptions` when names are320  ambiguous, and try a larger model for nuanced tasks.321- **vLLM version**: `ImportError: cannot import name GuidedDecodingParams` means the vLLM322  version does not match; the script requires `vllm>=0.6.6`.323