cmcshnik/TIQA_Text-in-Image_Quality_Assessment
TIQA: Text-in-Image Quality Assessment Datasets from the paper "TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images" (Kirill Koltsov, Aleksandr Gushchin, Dmitriy Vatolin, Anastasia Antsiferova). Text-to-image models produce globally realistic images, but rendered text often has malformed glyphs, broken strokes and irregular spacing. TIQA is a no-reference task: predict a human-aligned perceptual quality score for text regions in generated images. The… See the full description on the dataset page: https://huggingface.co/datasets/cmcshnik/TIQA_Text-in-Image_Quality_Assessment.
TIQA: Text-in-Image Quality Assessment
 
Datasets from the paper "TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images" (Kirill Koltsov, Aleksandr Gushchin, Dmitriy Vatolin, Anastasia Antsiferova).
Text-to-image models produce globally realistic images, but rendered text often has malformed glyphs, broken strokes and irregular spacing. TIQA is a no-reference task: predict a human-aligned perceptual quality score for text regions in generated images. The score measures how the text looks, not whether it says the right thing.
The repository has two parts:
Baseline model: ANTIQA — code and checkpoint. It reaches PLCC/SROCC 0.942/0.935 on TIQA-Crops and 0.842/0.837 on TIQA-Images text-quality MOS (unseen generators).
TIQA-Crops
Text crops detected with PP-OCRv5 on images from open-source generators. Every crop has a quality score on a 0–5 scale (higher is better).
- 9,978 crops have a MOS from a crowdsourced subjective study (about 50 ratings per crop).
- 110,219 crops have proxy labels: PP-OCRv5 recognition confidence mapped to the MOS scale with a 5-parameter logistic (5PL) fit. Use them for pretraining.
Generators: CogView4, DeepFloyd IF, FLUX.1-dev, Kandinsky 2, OmniGen, PixArt-α, PixArt-Σ, Qwen-Image, SD 2.1, SD 3 Medium, SD 3.5 Medium, SD 3.5 Large, SD 3.5 Large Turbo, SDXL 1.0, SDXL Turbo, SDXL Lightning.
Folders with the real_world_ prefix use prompts from a different distribution (real-world scenes with text) than folders without the prefix.
Structure
tiqa_crops/
├── dataset.csv
├── info.txt
└── crops/
└── <generator>/<image_id>/
├── <image_id>_crop_<k>.png
└── ocr_texts.txtdataset.csv
import pandas as pd
df = pd.read_csv("tiqa_crops/dataset.csv")
mos = df[df.rating_real] # human-labelled crops
proxy = df[~df.rating_real] # proxy-labelled crops for pretraining
test = df[df.for_test]TIQA-Images
Text-heavy prompts rendered by recent text-to-image models, including proprietary APIs. Each image has two subjective scores: overall quality and text quality.
For most models there are 30 prompts × 5 seeds (150 images); this set supports the best-of-5 selection experiment in the paper. Three models have one image per prompt (30 images).
Structure
tiqa_images/
├── <model>/p<NN>/<seed>.png # NN = prompt id 01–30, seed = 0–4
└── scores.csv # TODO: add MOS fileCitation
@article{koltsov2026tiqa,
title = {TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images},
author = {Koltsov, Kirill and Gushchin, Aleksandr and Vatolin, Dmitriy and Antsiferova, Anastasia},
journal = {arXiv preprint arXiv:2603.07119},
year = {2026}
}