Team Ai
Modelpublic

institutional/institutional-newspapers-crop-classifier-text-model2vec

sourceHugging Facemitupdated 1mo agoView on Hugging Face
2likes33downloads
Model Card

📰 Institutional Newspapers Crop Classifier (Text)

A text-based classifier that categorizes crops extracted from historical newspaper scans into high-level categories. This model is a Model2Vec fine-tune of minishlab/potion-base-32m with a classifier head, which makes it light and efficient.

We recommend using this model alongside `institutional/institutional-newspapers-crop-classifier-image-yolo26m-cls`, as neither visual nor textual signal alone is sufficient to classify newspaper crops accurately.

More information:

See also:

The Institutional Data Initiative at the Harvard Law School Library works with knowledge institutions—from libraries and museums to cultural groups and government agencies—to refine and publish their collections as data. Reach out to collaborate on your collections.

Evaluation results

Evaluated on a held-out test set of 31,015 crops:

ClassPrecisionRecallF1-ScoreSupport
Advertisement0.950.940.9511,244
Cartoon0.850.570.68430
Content0.900.960.9311,245
Masthead, nameplate or running head0.930.860.892,250
Photograph or illustration0.630.380.48551
Section heading0.940.930.945,295
MetricScore
Accuracy0.92
Macro avg F10.81
Weighted avg F10.92

Training data

This model was trained on part of Boston Public Library's newspapers collection.

  • —Total annotated crops: 185,900
  • —Classes: 6
  • —Split: ~67% train / ~17% val / ~17% test

Usage

python
from model2vec.inference import StaticModelPipeline

model_name = "institutional/institutional-newspapers-crop-classifier-text-model2vec"
model = StaticModelPipeline.from_pretrained(model_name)

texts = ["CLASSIFIED ADVERTISING — Rooms for rent, furnished apartments..."]
predictions = model.predict(texts, max_length=None)

print(predictions)
# ['Advertisement']

Recommended inference parameters

ParameterValueNote
max_lengthNoneDo not truncate input text

Citation

bibtext
@misc{cargnelutti2026institutionalnewspaperspipelinederiving,
      title={Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers}, 
      author={Matteo Cargnelutti and Catherine Brobston and Eben English and Jake Sadow and Kacie Bailey and Greg Leppert and Amanda Watson and Jessica Chapel and Jonathan Zittrain},
      year={2026},
      eprint={2608.18972},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2608.18972}, 
}