mapo80/DeQA-Doc-Color
05
1---2license: apache-2.03language:4 - en5tags:6 - image-quality-assessment7 - document-quality8 - mplug-owl29 - vision-language10 - document-analysis11 - color-quality12 - IQA13pipeline_tag: image-to-text14library_name: transformers15---16 17# DeQA-Doc-Color: Document Image Color Quality Assessment18 19**DeQA-Doc-Color** is a vision-language model specialized in assessing the **color quality** of document images. It evaluates color fidelity, saturation, white balance, and color-related artifacts in scanned or photographed documents.20 21## Model Family22 23This model is part of the **DeQA-Doc** family, which includes three specialized models:24 25| Model | Description | HuggingFace |26|-------|-------------|-------------|27| **DeQA-Doc-Overall** | Overall document quality | [mapo80/DeQA-Doc-Overall](https://huggingface.co/mapo80/DeQA-Doc-Overall) |28| **DeQA-Doc-Color** | Color quality assessment (this model) | [mapo80/DeQA-Doc-Color](https://huggingface.co/mapo80/DeQA-Doc-Color) |29| **DeQA-Doc-Sharpness** | Sharpness/clarity assessment | [mapo80/DeQA-Doc-Sharpness](https://huggingface.co/mapo80/DeQA-Doc-Sharpness) |30 31## Quick Start32 33```python34import torch35from transformers import AutoModelForCausalLM36from PIL import Image37 38# Load the model39model = AutoModelForCausalLM.from_pretrained(40 "mapo80/DeQA-Doc-Color",41 trust_remote_code=True,42 torch_dtype=torch.float16,43 device_map="auto",44)45 46# Score an image47image = Image.open("document.jpg").convert("RGB")48score = model.score([image])49print(f"Color Quality Score: {score.item():.2f} / 5.0")50```51 52## What Does Color Quality Measure?53 54The color quality score evaluates:55 56- **Color Fidelity**: How accurately colors are reproduced57- **White Balance**: Neutral whites without color casts (yellow, blue tints)58- **Saturation**: Appropriate color intensity (not washed out or oversaturated)59- **Color Artifacts**: Absence of color bleeding, banding, or chromatic aberration60- **Uniformity**: Consistent color reproduction across the document61 62## Score Interpretation63 64| Score Range | Quality Level | Typical Issues |65|-------------|---------------|----------------|66| 4.5 - 5.0 | **Excellent** | Perfect color reproduction |67| 3.5 - 4.5 | **Good** | Minor color shifts, slight tinting |68| 2.5 - 3.5 | **Fair** | Noticeable color cast, uneven colors |69| 1.5 - 2.5 | **Poor** | Strong color distortion, washed out |70| 1.0 - 1.5 | **Bad** | Severe color problems, unusable |71 72## Batch Processing73 74```python75images = [76 Image.open("doc1.jpg").convert("RGB"),77 Image.open("doc2.jpg").convert("RGB"),78 Image.open("doc3.jpg").convert("RGB"),79]80 81scores = model.score(images)82for i, score in enumerate(scores):83 print(f"Document {i+1} Color Score: {score.item():.2f} / 5.0")84```85 86## Use Cases87 88- **Scanner Calibration**: Detect when scanners need color calibration89- **Photo Document QA**: Flag photos with poor lighting/white balance90- **Color-Critical Documents**: Verify color accuracy for maps, charts, branded materials91- **Archive Preservation**: Identify documents with color degradation92- **Print Quality Control**: Verify color reproduction in printed documents93 94## Example: Detect Color Issues95 96```python97import torch98from transformers import AutoModelForCausalLM99from PIL import Image100 101model = AutoModelForCausalLM.from_pretrained(102 "mapo80/DeQA-Doc-Color",103 trust_remote_code=True,104 torch_dtype=torch.float16,105 device_map="auto",106)107 108def diagnose_color_quality(image_path):109 img = Image.open(image_path).convert("RGB")110 score = model.score([img]).item()111 112 if score >= 4.5:113 diagnosis = "Excellent color quality"114 elif score >= 3.5:115 diagnosis = "Good - minor color issues"116 elif score >= 2.5:117 diagnosis = "Fair - consider color correction"118 elif score >= 1.5:119 diagnosis = "Poor - needs color correction or rescan"120 else:121 diagnosis = "Bad - severe color problems, rescan required"122 123 return score, diagnosis124 125score, diagnosis = diagnose_color_quality("scanned_document.jpg")126print(f"Score: {score:.2f}/5.0 - {diagnosis}")127```128 129## Multi-Dimensional Quality Assessment130 131Combine with other DeQA-Doc models for comprehensive assessment:132 133```python134import torch135from transformers import AutoModelForCausalLM136from PIL import Image137 138# Load all three models139models = {140 "overall": AutoModelForCausalLM.from_pretrained(141 "mapo80/DeQA-Doc-Overall", trust_remote_code=True,142 torch_dtype=torch.float16, device_map="auto"143 ),144 "color": AutoModelForCausalLM.from_pretrained(145 "mapo80/DeQA-Doc-Color", trust_remote_code=True,146 torch_dtype=torch.float16, device_map="auto"147 ),148 "sharpness": AutoModelForCausalLM.from_pretrained(149 "mapo80/DeQA-Doc-Sharpness", trust_remote_code=True,150 torch_dtype=torch.float16, device_map="auto"151 ),152}153 154def full_quality_report(image_path):155 img = Image.open(image_path).convert("RGB")156 157 scores = {}158 for name, model in models.items():159 scores[name] = model.score([img]).item()160 161 return scores162 163report = full_quality_report("document.jpg")164print(f"Overall: {report['overall']:.2f}/5.0")165print(f"Color: {report['color']:.2f}/5.0")166print(f"Sharpness: {report['sharpness']:.2f}/5.0")167```168 169## Model Architecture170 171- **Base Model**: mPLUG-Owl2 (LLaMA2-7B + ViT-L Vision Encoder)172- **Vision Encoder**: CLIP ViT-L/14 (1024 visual tokens via Visual Abstractor)173- **Language Model**: LLaMA2-7B174- **Training**: Full fine-tuning on document color quality datasets175- **Input Resolution**: Images are resized to 448x448 (with aspect ratio preservation)176 177## Technical Details178 179| Property | Value |180|----------|-------|181| Model Size | ~16 GB (float16) |182| Parameters | ~7.2B |183| Input | RGB images (any resolution) |184| Output | Color quality score (1.0 - 5.0) |185| Inference | ~2-3 seconds per image on A100 |186 187## Hardware Requirements188 189| Setup | VRAM Required | Recommended |190|-------|---------------|-------------|191| Full precision (fp32) | ~32 GB | A100, H100 |192| Half precision (fp16) | ~16 GB | A100, A40, RTX 4090 |193| With CPU offload | ~8 GB GPU + RAM | RTX 3090, RTX 4080 |194 195## Installation196 197```bash198pip install torch transformers accelerate pillow sentencepiece protobuf199```200 201**Note**: Use `transformers>=4.36.0` for best compatibility.202 203## Limitations204 205- Optimized for document images (may not generalize to natural photos)206- Color assessment is relative to training data distribution207- Black & white documents may receive lower scores (use Overall model instead)208- Requires GPU with sufficient VRAM for efficient inference209 210## Credits & Attribution211 212This model is based on the **DeQA-Doc** project by Junjie Gao et al., which won the **Championship** in the VQualA 2025 DIQA (Document Image Quality Assessment) Challenge.213 214**Original Repository**: [https://github.com/Junjie-Gao19/DeQA-Doc](https://github.com/Junjie-Gao19/DeQA-Doc)215 216All credit for the research, training methodology, and model architecture goes to the original authors.217 218## Citation219 220If you use this model in your research, please cite the original paper:221 222```bibtex223@inproceedings{deqadoc,224 title={{DeQA-Doc}: Adapting {DeQA-Score} to Document Image Quality Assessment},225 author={Gao, Junjie and Liu, Runze and Peng, Yingzhe and Yang, Shujian and Zhang, Jin and Yang, Kai and You, Zhiyuan},226 booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision Workshop},227 year={2025},228}229```230 231**ArXiv**: [https://arxiv.org/abs/2507.12796](https://arxiv.org/abs/2507.12796)232 233## License234 235Apache 2.0236 237## Related Models238 239- [DeQA-Doc-Overall](https://huggingface.co/mapo80/DeQA-Doc-Overall) - Overall quality assessment240- [DeQA-Doc-Sharpness](https://huggingface.co/mapo80/DeQA-Doc-Sharpness) - Sharpness assessment241 