institutional/institutional-newspapers-crop-classifier-text-model2vec
📰 Institutional Newspapers Crop Classifier (Text)
A text-based classifier that categorizes crops extracted from historical newspaper scans into high-level categories. This model is a Model2Vec fine-tune of minishlab/potion-base-32m with a classifier head, which makes it light and efficient.
We recommend using this model alongside `institutional/institutional-newspapers-crop-classifier-image-yolo26m-cls`, as neither visual nor textual signal alone is sufficient to classify newspaper crops accurately.
More information:
See also:
The Institutional Data Initiative at the Harvard Law School Library works with knowledge institutions—from libraries and museums to cultural groups and government agencies—to refine and publish their collections as data. Reach out to collaborate on your collections.
Evaluation results
Evaluated on a held-out test set of 31,015 crops:
Training data
This model was trained on part of Boston Public Library's newspapers collection.
- Total annotated crops: 185,900
- Classes: 6
- Split: ~67% train / ~17% val / ~17% test
Usage
from model2vec.inference import StaticModelPipeline
model_name = "institutional/institutional-newspapers-crop-classifier-text-model2vec"
model = StaticModelPipeline.from_pretrained(model_name)
texts = ["CLASSIFIED ADVERTISING — Rooms for rent, furnished apartments..."]
predictions = model.predict(texts, max_length=None)
print(predictions)
# ['Advertisement']Recommended inference parameters
Citation
@misc{cargnelutti2026institutionalnewspaperspipelinederiving,
title={Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers},
author={Matteo Cargnelutti and Catherine Brobston and Eben English and Jake Sadow and Kacie Bailey and Greg Leppert and Amanda Watson and Jessica Chapel and Jonathan Zittrain},
year={2026},
eprint={2608.18972},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2608.18972},
}