uv-scripts/classification
Classification Scripts Text classification on HF Jobs: label a dataset with a model that needs no training, or train your own classifier from labelled examples. If you have seen Jev and other "System One" models: the models these scripts train are small, open versions of the same idea. They read a piece of data and return a label with a probability, and you can train one on your own labels. For example, this demo suggests task tags for any Hub dataset; its model was fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/classification.
18316
1# SetFit script: findings and behaviour details2 3Notes behind `train-setfit.py`. The README has what you need to run it; this file records the4longer findings and the full behaviour details.5 6## Compare more than the majority baseline7 8The `emotion` run reached **0.370** accuracy against a **0.352** majority baseline. Other9single-seed body-model runs reached 0.418 (`bge-small`) and 0.410 (`paraphrase-mpnet-base-v2`).10These results call for further evaluation; they do not establish a limit on the task or method.11 12SetFit's [zero-shot guide](https://huggingface.co/docs/setfit/how_to/zero_shot) reports **0.591**13on emotion using BGE and training examples templated from the class names. It uses a different14evaluation setup from the table above, so this is motivation for a matched comparison rather15than a controlled comparison with this recipe. Templated training needs no labeled documents,16but still uses compute.17 18For your task, compare with a simple baseline such as TF-IDF plus logistic regression using19the same training and evaluation rows. A zero-shot comparison can also be useful when class20names describe the task well. Use repeated seeds and appropriate task metrics before drawing21conclusions from small accuracy differences. This recipe trains and evaluates a supervised22classifier; built-in templated zero-shot training is a separate possible extension.23 24## Real-world data: a worked failure25 26`biglam/hansard_speech` (2.7M parliamentary speeches, predicting `party` from `speech`) is the27case where none of the convenient properties hold, and it is instructive precisely because it28produces no score:29 30- **No held-out split**, so the eval set has to be carved from train — the numbers stop being31 comparable to anything published.32- **~9.5% of rows have a blank `party`**, which without the drop trains an `""` class.33- **28 parties after cleaning, nine of which cannot supply 8 examples** (`Respect` 4,34 `Independent SDP` 2, `Change UK` 1). The requested eight-example budget cannot be met for those classes.35- **1,878 steps at ~11s/step on CPU** — the script refuses it, projecting well past an hour.36 37On completed runs, the model card discloses a carved evaluation split, per-class training counts38and classes below the requested sample count. Dropped-row counts and measured truncation are39reported in the logs; retain those logs alongside the model when documenting data preparation.40 41## Many classes: watch the pair count42 43SetFit trains on pairs drawn from every combination of training examples, so the pair count grows44with the **square** of the training-set size — which is `--num-samples` x number of classes. The45script logs the estimate before training starts:46 47| Dataset | Strategy | Pairs | Steps |48|---|---|---|---|49| ag_news (4 classes x 8) | `oversampling` (default) | 768 | 48 |50| banking77 (77 classes x 8) | `oversampling` (default) | 374,528 | 23,408 |51| banking77 (77 classes x 8) | `undersampling` | 4,312 | 270 |52 53At 77 classes the default would take roughly 15 hours on `cpu-basic`; `--sampling-strategy54undersampling` finished in 18 seconds on a T4 in the recorded run. The script reports the pair55and step counts, then measures step time to check `--max-minutes`. When it refuses training,56it suggests undersampling where applicable and estimates whether that would fit the budget.57 58## Choosing another body or longer context59 60`--body-model` accepts a Sentence Transformer checkpoint. Set `--max-seq-length` within that61model's supported context window; increasing it cannot extend a model's native limit or restore62text already shortened during dataset preparation. Longer sequences can need a smaller63`--batch-size` or more GPU memory. The recipe measures training cost on the selected hardware.64 65Follow the body's task-prefix instructions when preparing inputs. For example,66[`nomic-ai/modernbert-embed-base`](https://huggingface.co/nomic-ai/modernbert-embed-base)67uses Nomic's task prefixes: classification inputs should begin with `classification: `.68Include the same prefix during training, evaluation and inference. The recipe does not add it69automatically. Retain the original texts and the preprocessing details with the model.70 71## Behaviour details72 73- **Evaluation split** follows the same precedence as `train-classifier.py`: `--eval-split` if given, else `validation`, else `test`, else a stratified carve-out of `--eval-fraction` from train.74- **Metrics match `train-classifier.py`** (accuracy + macro F1). Match evaluation rows and preprocessing when comparing runs.75- **`--num-samples`** sets labelled examples per class (default 8). **`--sampling-strategy`** controls contrastive pairing: `oversampling` (default), `undersampling`, `unique`.76- **Every run reports a majority baseline.** The run warns when accuracy fails to beat it, or the gain is below five percentage points. That fixed threshold is a review heuristic, not a measured noise level or significance test.77- **It estimates training time before starting.** The script times forward/backward passes on actual texts and hardware, then refuses training projected above `--max-minutes` (default 60). Setup, evaluation and upload take additional time. A measurement error can skip this guard; use Jobs `--timeout` to enforce a wall-clock limit.78- **Rows with missing or blank labels or texts are dropped**, with a count. Missing labels include `ClassLabel`'s `-1` sentinel and numeric NaN; plain integer `-1` remains a valid class. Splits with no usable labelled text, fewer than two observed training classes, and missing or non-string text columns exit before model loading.79- **`--private` verifies the output repository is private before training.** If the destination already exists publicly, choose a new repo or change its visibility first.80 