Team Ai
Modelpublic

fairdataihub/poster-sentry

sourceHugging Facemitupdated 19d agoView on Hugging Face
1likes29downloads
Model Card

<div align="center"> <img src="PosterSentry.png" alt="PosterSentry Logo" title="This image was generated by AI" width="400"/> </div>

PosterSentry: Multimodal Scientific Poster Classifier

Model Description

PosterSentry is a lightweight, CPU-optimized multimodal classifier that determines whether a PDF is a scientific poster or a non-poster (paper, proceedings, newsletter, abstract book, etc.).

Part of the quality control pipeline for **posters.science**, a platform for making scientific conference posters Findable, Accessible, Interoperable, and Reusable (FAIR).

Developed by the **FAIR Data Innovations Hub** at the California Medical Innovations Institute (CalMI²).

Version

VersionDateNotes
1.3.02026-09-21License-screened retrain: 84 documents whose source-repository licenses do not permit redistribution were removed, and the head was retrained on the resulting 3,298-document corpus. Held-out accuracy 93.1%.
1.2.02026-09-01Stacked head: the 512-d text embedding is scored by a stage-one logistic regression whose poster probability becomes a single text-score feature for the final classifier over [text_score + 15 visual + 15 structural] = 31 features. Held-out accuracy 92.9%. Head weights now store a zero first column so softmax reproduces sklearn's sigmoid exactly; earlier heads doubled the logit, which sharpened reported confidences (decisions at 0.5 were unaffected). Flat heads remain loadable by the package.
1.1.02026-08-28Feature extraction moved off the AGPL-licensed PyMuPDF to pdfplumber (text and structure) and pypdfium2 (page rendering), both permissively licensed, so the whole stack is MIT-compatible. The head was retrained on the re-extracted features; held-out accuracy is 89.2%.
1.0.02026-08-18Head trained on the human-validated corpus: 3,381 documents labeled by a three-reviewer survey (Krippendorff's alpha 0.79) with blinded adjudication of the 439 contested documents.

Earlier unversioned weights (April 2026) were trained on heuristically labeled data and are superseded.

Related Models & Tools

ResourceDescriptionLink
PosterSentryMultimodal poster classifier (this model)fairdataihub/poster-sentry
poster-sentryInstallable classifier packageGitHub
poster-sentry-trainingTraining code and replicationGitHub
poster-sentry-training-dataHuman-validated training dataset (3,298 samples)HuggingFace
poster-sentry-evaluation-paper-codeReproducible analysis for the paperGitHub
Llama-3.1-8B-Poster-ExtractionPoster → structured JSON extractionfairdataihub/Llama-3.1-8B-Poster-Extraction
poster2jsonPython library for poster extractionPyPI · Docs · GitHub
poster-json-schemaDataCite-based poster metadata schemaGitHub
Platformposters.scienceposters.science

Pipeline Position

PosterSentry sits at the front of the posters.science pipeline. It screens incoming PDFs before the expensive Llama-based extraction:

PDF Input
   │
   ▼
┌──────────────┐     ┌───────────────────────────────────┐     ┌──────────────┐
│ PosterSentry │ ──► │ Llama-3.1-8B-Poster-Extraction    │ ──► │ poster2json  │
│ (classify)   │     │ (extract structured metadata)      │     │ (validate)   │
└──────────────┘     └───────────────────────────────────┘     └──────────────┘
   poster? ✓              raw text → JSON schema                  FAIR output

Architecture

Two-stage (stacked) logistic regression. Stage 1 scores the 512-d text embedding alone; its poster probability becomes the single text_score feature of stage 2, which classifies 31 features:

ChannelFeaturesDimensionSignal
Text scoreStage-1 LogisticRegression over the model2vec (potion-base-32M) embedding1Semantic content, summarized
VisualColor stats, edge density, FFT spatial complexity, whitespace15Visual layout
StructuralPage count, area, font diversity, text blocks, density15PDF geometry

Stage 2 is trained on inner 5-fold out-of-fold text scores so it never sees an in-sample-optimistic text score. Both stages (weights and scalers) live in one numpy .npz head (about 20 KB). Inference is pure numpy, no torch required at prediction time. Every stage-2 coefficient is a named, interpretable feature; the raw 542-d concatenation of earlier releases underperformed it by about six points out of fold because the 512 standardized text dimensions crowded out the reliable engineered signal.

PDF backend (speed vs license)

Feature extraction reads PDFs through a selectable backend:

BackendLibrariesLicenseSpeedNotes
pdfplumber (default)pdfplumber + pypdfium2MIT + BSD-3-Clause~2.0 s/docThe stack the released model is trained on; reproduces the paper.
pymupdfPyMuPDFAGPL-3.0~1.2 s/doc (about 1.75x faster)Faster, but AGPL. Install with pip install poster-sentry[pymupdf]. Its extracted features differ slightly, so retrain with this backend for best accuracy or accept minor prediction drift.

Select the backend with --backend, the backend= constructor argument, or the POSTER_SENTRY_BACKEND environment variable. The default is pdfplumber so results match this model.

On a 40-document corpus sample, the dominant cost in the default stack is pdfplumber's pure-Python text and structure parsing (about 780 ms and 760 ms per document); page rendering with pypdfium2 adds about 490 ms. A pypdfium2-based text variant cuts the text step to about 75 ms and reaches roughly 1.3 s/doc, but it changes the extracted text, so it would require its own retrain; pymupdf is the recommended fast path when AGPL is acceptable.

Performance

Evaluated against human-validated labels (three-reviewer survey with blinded adjudication of contested documents):

MetricValue
Held-out accuracy (495 documents)93.1% (95% CI 90.6 to 95.0)
Macro F1 (held-out)0.931
Precision / Recall / F1 (poster)0.918 / 0.948 / 0.933
Precision / Recall / F1 (non-poster)0.946 / 0.915 / 0.930
Nested out-of-fold accuracy (all 3,298 documents, 5-fold)93.9%
Panel reliability with the model as a fourth rateralpha 0.792 -> 0.795 (unchanged)
Calibration (out-of-fold)Brier 0.046, ECE 0.016
Classifier speedunder one second per file on a CPU (PDF parsing is the bottleneck)

Errors concentrate where the human panel itself divided: out-of-fold agreement is 96.4% on documents the panel rated unanimously and 77.3% on documents decided two to one.

Top Features by Importance

Standardized logistic regression coefficients of the trained head (positive pushes toward poster):

RankFeatureCoefficientSignal
1page_count-3.26More pages pushes away from poster
2size_per_page_kb+2.42Dense, high-resolution single pages
3line_count+1.94Posters pack many short text lines
4file_size_kb-1.70Multi-page documents are bigger overall
5mean_g+1.10Colorful, non-white pages
6is_landscape+1.00Many posters are landscape
7white_space_ratio-0.95Sparse, white pages push toward non-poster
8edge_density+0.86Visually busy layouts
9page_width_pt+0.80Posters are physically wide
10img_width+0.79Large rendered width

Every coefficient of the final classifier is a named feature; the whole text channel enters as the single text_score (+0.68, rank 13 of 31), whose value concentrates on the documents where geometry misleads.

Training Data

Trained on 3,298 documents with human-validated labels, zero synthetic data:

ClassCountLabel provenance
Poster1,651Three-reviewer survey; unanimous panel label or blinded adjudication
Non-poster1,647Three-reviewer survey; unanimous panel label or blinded adjudication

Three reviewers independently rated 3,486 candidate documents that passed license screening (inter-rater Krippendorff's alpha 0.79); the 430 documents without a unanimous panel were settled in a blinded adjudication review. After removing 181 near-duplicate documents and 7 with unavailable PDFs, the remaining 3,298 form the training corpus. Applied to the full corpus of 30,139 readable repository PDFs labeled as posters, PosterSentry classifies 80.5% as posters: roughly one in five records labeled as posters is something else.

Training data: fairdataihub/poster-sentry-training-data

Usage

Python API

python
from poster_sentry import PosterSentry

sentry = PosterSentry()
sentry.initialize()

# Classify a PDF (uses text + visual + structural features)
result = sentry.classify("document.pdf")
print(f"Is poster: {result['is_poster']}, Confidence: {result['confidence']:.2f}")
# {'is_poster': True, 'confidence': 0.97, 'path': 'document.pdf'}

# Batch classification
results = sentry.classify_batch(["poster1.pdf", "paper.pdf", "newsletter.pdf"])

Installation

bash
pip install git+https://github.com/fairdataihub/poster-repo-qc.git

# Or install from source
git clone https://github.com/fairdataihub/poster-repo-qc.git
cd poster-repo-qc
pip install -e ".[train]"

Training

bash
python scripts/train_poster_sentry.py --n-per-class 2000

Training completes in ~40 minutes on CPU (PDF rendering is the bottleneck, not the classifier).

Model Specifications

AttributeValue
Embedding backboneminishlab/potion-base-32M (model2vec StaticModel)
Embedding dimension512
Visual features15 (color, edge, FFT, whitespace)
Structural features15 (page geometry, fonts, text blocks)
Final-classifier input dimension31 (text score + 15 visual + 15 structural)
ClassifierTwo-stage (stacked) LogisticRegression + StandardScaler
Head file size20 KB (.npz, both stages)
Precisionfloat32
GPU requiredNo (CPU-only)
LicenseMIT

System Requirements

  • —CPU: Any modern CPU (no GPU needed)
  • —RAM: ≥4GB
  • —Python: ≥3.10
  • —Dependencies: numpy, model2vec, scikit-learn, pdfplumber, pypdfium2, Pillow

Citation

bibtex
@software{poster_sentry_2026,
  title = {PosterSentry: Multimodal Scientific Poster Classifier},
  author = {O'Neill, Jamey and Portillo, Dorian and Zeinali, Nahid and Soundarajan, Sanjay and Blake, Gerard and Sarin, Parth and Buttrick, Adam and Patel, Bhavesh},
  year = {2026},
  version = {1.3.0},
  url = {https://huggingface.co/fairdataihub/poster-sentry},
  note = {Part of the posters.science initiative}
}

License

This model is released under the MIT License.

Acknowledgments