sujalgumme/ocr-error-correction
OCR Error Correction with Qwen2.5 and 4-bit QLoRA
An end-to-end, resume-verifiable project that corrects OCR text while preserving names, dates, document numbers, addresses, and layout.
The pipeline creates synthetic noisy-clean pairs, fine-tunes Qwen/Qwen2.5-1.5B-Instruct with 4-bit QLoRA, compares the base and fine-tuned models using CER/WER/field exact match, tests OCR robustness under image degradation, and serves a Gradio demo.
Important: do not add performance numbers to a resume until scripts/evaluate.py has produced them from the held-out test split.What is included
ocr-error-correction/
├── app.py # Gradio image/text demo
├── configs/train.yaml # Reproducible experiment settings
├── data/sample_clean.jsonl # Human-readable seed examples
├── notebooks/ocr_qlora_colab.ipynb
├── scripts/
│ ├── evaluate.py # Base vs adapter metrics
│ ├── generate_dataset.py # Synthetic noisy-clean pairs
│ ├── robustness_eval.py # Blur/rotation/JPEG/low-res evaluation
│ └── train_qlora.py # 4-bit QLoRA SFT
├── src/ocr_corrector/
│ ├── fields.py # Field extraction and exact match
│ ├── metrics.py # CER and WER
│ ├── model.py # Base/adapter inference
│ ├── noise.py # OCR-like text corruption
│ ├── ocr.py # Tesseract wrapper and image degradation
│ └── synthetic.py # Synthetic document records
└── tests/ # Lightweight unit testsRecommended experiment
Training loss and validation loss are logged to Trackio so the run has visible experiment evidence in addition to the saved adapter.
The use of NF4 follows the current Transformers guidance for training 4-bit base models. TRL accepts conversational and prompt-completion datasets, which this project uses to train only on the correction completion.
The complete beginner workflow
Phase 1 - Run the quick checks on Windows
Open the folder in VS Code. In PowerShell:
python -m venv venv
Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass
.\venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
pip install -r requirements-dev.txt
python -m unittest discover -s tests -vThis verifies the dataset corruption and metric logic without downloading the 1.5B model.
Phase 2 - Train free in Google Colab
- Zip this project folder, open the notebook in
notebooks/ocr_qlora_colab.ipynb, and upload the zip when prompted. - In Colab choose Runtime > Change runtime type > T4 GPU.
- Run each cell from top to bottom.
- When asked, sign in to Hugging Face with a write token so the LoRA adapter can be saved permanently.
- Replace
YOUR_HF_USERNAMEin the training cell with your real username.
The notebook runs these commands:
python scripts/generate_dataset.py --count 1200 --output-dir data/generated --seed 42
python scripts/train_qlora.py \
--train-file data/generated/train.jsonl \
--validation-file data/generated/validation.jsonl \
--output-dir outputs/qwen25-ocr-qlora \
--hub-model-id YOUR_HF_USERNAME/qwen2.5-1.5b-ocr-correction-qlora \
--push-to-hubPhase 3 - Measure real results
In the same Colab runtime:
python scripts/evaluate.py \
--test-file data/generated/test.jsonl \
--adapter outputs/qwen25-ocr-qlora \
--output-dir artifacts/evaluation \
--compare-baseThis writes:
artifacts/evaluation/metrics.jsonartifacts/evaluation/predictions.jsonlartifacts/evaluation/resume_metrics.md
The Markdown file contains a truthful, copy-ready resume description populated from the actual run.
Phase 4 - Test real OCR degradation
Install Tesseract on the machine that runs this step. On Windows, install Tesseract and set its executable path if it is not on PATH:
$env:TESSERACT_CMD = "C:\Program Files\Tesseract-OCR\tesseract.exe"Then run:
python scripts/robustness_eval.py \
--test-file data/generated/test.jsonl \
--adapter outputs/qwen25-ocr-qlora \
--limit 25 \
--output artifacts/robustness.jsonConditions tested: clean rendering, blur, rotation, JPEG compression, and low resolution.
Phase 5 - Launch the Gradio app
pip install -r requirements-app.txt
python app.py --adapter outputs/qwen25-ocr-qloraOpen the local URL printed in the terminal. The app has two workflows:
- Upload a document image: Tesseract extracts text, then Qwen corrects it.
- Paste OCR text directly: Qwen returns corrected text.
Dataset design
Each JSONL row contains:
{
"id": "record-000001",
"clean_text": "Name: Asha Rao\nDate: 14/03/2025\nDocument No: IN-482913\nAddress: 21 MG Road, Bengaluru 560001",
"noisy_text": "Narne: Asha Rao\nDate: l4/03/2025\nDocurnent No: IN-4829l3\nAddress: 21 MG Road, Bengaluru 560001",
"fields": {
"name": "Asha Rao",
"date": "14/03/2025",
"document_number": "IN-482913",
"address": "21 MG Road, Bengaluru 560001"
},
"error_types": ["character_substitution"],
"severity": "medium"
}The test split is created before training and never passed to the trainer. This prevents inflated metrics caused by evaluation leakage.
Metric definitions
- CER = character edit distance divided by reference character count.
- WER = word edit distance divided by reference word count.
- Field exact match = percentage of expected structured fields exactly reproduced after whitespace/case normalization.
Lower CER/WER is better. Higher exact match is better.
Honest resume wording before and after training
Before completing the experiment:
Built an OCR error-correction pipeline for Qwen2.5-1.5B-Instruct using synthetic noisy-clean pairs, 4-bit QLoRA training, and CER/WER/field-level evaluation.
After completing it, use only the automatically generated wording in artifacts/evaluation/resume_metrics.md.
