loopdesk-ai/lipika-eval
Lipika eval — Indic font recognition benchmark The frozen validation set behind loopdesk-ai/lipika (Indic font recognizer): 6,876 synthetic text crops covering 553 freely-licensed font families across 13 scripts (Devanagari, Bengali, Gujarati, Gurmukhi, Kannada, Malayalam, Meetei Mayek, Odia, Ol Chiki, Perso-Arabic, Tamil, Telugu, Latin). This is the set reported as "synthetic val" in the model card (Lipika v2.4 scores 0.849 family top-1 / 0.977 top-5 / 0.991 script). Use it to… See the full description on the dataset page: https://huggingface.co/datasets/loopdesk-ai/lipika-eval.
Lipika eval — Indic font recognition benchmark
The frozen validation set behind [loopdesk-ai/lipika](https://huggingface.co/loopdesk-ai/lipika) (Indic font recognizer): 6,876 synthetic text crops covering 553 freely-licensed font families across 13 scripts (Devanagari, Bengali, Gujarati, Gurmukhi, Kannada, Malayalam, Meetei Mayek, Odia, Ol Chiki, Perso-Arabic, Tamil, Telugu, Latin).
This is the set reported as "synthetic val" in the model card (Lipika v2.4 scores 0.849 family top-1 / 0.977 top-5 / 0.991 script). Use it to benchmark your own font-ID model or to reproduce ours.
How it was made
Crops are rendered with proper Indic shaping (HarfBuzz via a Pillow/raqm stack) from the open Lipika font corpus, with the frozen-eval augmentation distribution: varied sizes, colors, backgrounds, mild blur/noise/rotation, JPEG artifacts. Seed-frozen (seed 2028) — the generator lives in `fontrecog/dataset/pregenerate_val.py`.
Training data is not published as files — it is rendered on the fly from font files by the same pipeline; the corpus manifest + fetch scripts in the GitHub repo reproduce it.
Fields
Licensing
Dataset packaging & annotations: Apache-2.0. Images are bitmap renders of freely licensed / freely distributed fonts only (OFL, GPL with font exception, public freeware); each row carries its font's upstream license in font_license. Crops rendered from proprietary system fonts (10 families, 84 crops of the original 6,960) are excluded from this release; scores on this subset may differ from the model card's full-set numbers by <0.2pt.
Load
from datasets import load_dataset
ds = load_dataset("loopdesk-ai/lipika-eval", split="validation")Citation
@software{lipika2026,
title = {Lipika: Indic Font Recognition},
author = {Loopdesk},
year = {2026},
url = {https://huggingface.co/loopdesk-ai/lipika}
}