fairdataihub/poster-sentry
<div align="center"> <img src="PosterSentry.png" alt="PosterSentry Logo" title="This image was generated by AI" width="400"/> </div>
PosterSentry: Multimodal Scientific Poster Classifier
Model Description
PosterSentry is a lightweight, CPU-optimized multimodal classifier that determines whether a PDF is a scientific poster or a non-poster (paper, proceedings, newsletter, abstract book, etc.).
Part of the quality control pipeline for **posters.science**, a platform for making scientific conference posters Findable, Accessible, Interoperable, and Reusable (FAIR).
Developed by the **FAIR Data Innovations Hub** at the California Medical Innovations Institute (CalMI²).
Version
Earlier unversioned weights (April 2026) were trained on heuristically labeled data and are superseded.
Related Models & Tools
Pipeline Position
PosterSentry sits at the front of the posters.science pipeline. It screens incoming PDFs before the expensive Llama-based extraction:
PDF Input
│
▼
┌──────────────┐ ┌───────────────────────────────────┐ ┌──────────────┐
│ PosterSentry │ ──► │ Llama-3.1-8B-Poster-Extraction │ ──► │ poster2json │
│ (classify) │ │ (extract structured metadata) │ │ (validate) │
└──────────────┘ └───────────────────────────────────┘ └──────────────┘
poster? ✓ raw text → JSON schema FAIR outputArchitecture
Two-stage (stacked) logistic regression. Stage 1 scores the 512-d text embedding alone; its poster probability becomes the single text_score feature of stage 2, which classifies 31 features:
Stage 2 is trained on inner 5-fold out-of-fold text scores so it never sees an in-sample-optimistic text score. Both stages (weights and scalers) live in one numpy .npz head (about 20 KB). Inference is pure numpy, no torch required at prediction time. Every stage-2 coefficient is a named, interpretable feature; the raw 542-d concatenation of earlier releases underperformed it by about six points out of fold because the 512 standardized text dimensions crowded out the reliable engineered signal.
PDF backend (speed vs license)
Feature extraction reads PDFs through a selectable backend:
Select the backend with --backend, the backend= constructor argument, or the POSTER_SENTRY_BACKEND environment variable. The default is pdfplumber so results match this model.
On a 40-document corpus sample, the dominant cost in the default stack is pdfplumber's pure-Python text and structure parsing (about 780 ms and 760 ms per document); page rendering with pypdfium2 adds about 490 ms. A pypdfium2-based text variant cuts the text step to about 75 ms and reaches roughly 1.3 s/doc, but it changes the extracted text, so it would require its own retrain; pymupdf is the recommended fast path when AGPL is acceptable.
Performance
Evaluated against human-validated labels (three-reviewer survey with blinded adjudication of contested documents):
Errors concentrate where the human panel itself divided: out-of-fold agreement is 96.4% on documents the panel rated unanimously and 77.3% on documents decided two to one.
Top Features by Importance
Standardized logistic regression coefficients of the trained head (positive pushes toward poster):
Every coefficient of the final classifier is a named feature; the whole text channel enters as the single text_score (+0.68, rank 13 of 31), whose value concentrates on the documents where geometry misleads.
Training Data
Trained on 3,298 documents with human-validated labels, zero synthetic data:
Three reviewers independently rated 3,486 candidate documents that passed license screening (inter-rater Krippendorff's alpha 0.79); the 430 documents without a unanimous panel were settled in a blinded adjudication review. After removing 181 near-duplicate documents and 7 with unavailable PDFs, the remaining 3,298 form the training corpus. Applied to the full corpus of 30,139 readable repository PDFs labeled as posters, PosterSentry classifies 80.5% as posters: roughly one in five records labeled as posters is something else.
Training data: fairdataihub/poster-sentry-training-data
Usage
Python API
from poster_sentry import PosterSentry
sentry = PosterSentry()
sentry.initialize()
# Classify a PDF (uses text + visual + structural features)
result = sentry.classify("document.pdf")
print(f"Is poster: {result['is_poster']}, Confidence: {result['confidence']:.2f}")
# {'is_poster': True, 'confidence': 0.97, 'path': 'document.pdf'}
# Batch classification
results = sentry.classify_batch(["poster1.pdf", "paper.pdf", "newsletter.pdf"])Installation
pip install git+https://github.com/fairdataihub/poster-repo-qc.git
# Or install from source
git clone https://github.com/fairdataihub/poster-repo-qc.git
cd poster-repo-qc
pip install -e ".[train]"Training
python scripts/train_poster_sentry.py --n-per-class 2000Training completes in ~40 minutes on CPU (PDF rendering is the bottleneck, not the classifier).
Model Specifications
System Requirements
- CPU: Any modern CPU (no GPU needed)
- RAM: ≥4GB
- Python: ≥3.10
- Dependencies: numpy, model2vec, scikit-learn, pdfplumber, pypdfium2, Pillow
Citation
@software{poster_sentry_2026,
title = {PosterSentry: Multimodal Scientific Poster Classifier},
author = {O'Neill, Jamey and Portillo, Dorian and Zeinali, Nahid and Soundarajan, Sanjay and Blake, Gerard and Sarin, Parth and Buttrick, Adam and Patel, Bhavesh},
year = {2026},
version = {1.3.0},
url = {https://huggingface.co/fairdataihub/poster-sentry},
note = {Part of the posters.science initiative}
}License
This model is released under the MIT License.
Acknowledgments
- FAIR Data Innovations Hub at California Medical Innovations Institute (CalMI²)
- posters.science platform
- MinishLab for the model2vec embedding backbone
- HuggingFace for model hosting infrastructure
- Funded by The Navigation Fund (10.71707/rk36-9x79), "Poster Sharing and Discovery Made Easy"
