Team Ai
Datasetpublic

uv-scripts/ocr

OCR UV Scripts Part of uv-scripts: self-contained UV scripts you run on Hugging Face Jobs in one command. One script per OCR model. Each script runs the model on a GPU with Hugging Face Jobs and writes the text as markdown: as a new column in a Hub dataset, as .md files in a Bucket, or as resumable parquet parts (the -saturate recipes). A few scripts return JSON from a schema, detect layout regions, or compare the output of two models. Quick Start First… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/ocr.

sourceHugging Faceupdated 10d agoView on Hugging Face
163likes6.5kdownloads
SERVING.md94 linesDownload Raw Back to root
1# Server-mode OCR: official support per model + the `-server.py` recipes2 3The recipes in this folder come in two shapes:4 5- **Offline batch** (`<model>.py`): the script owns the vLLM engine (`LLM.generate`)6  and processes the dataset in fixed-size batches. One `hf jobs uv run` command.7- **Server + driver** (`<model>-server.py`): one job starts `vllm serve` in the8  background, then a lightweight driver (no torch/vllm deps — installs in seconds)9  posts images concurrently to the OpenAI-compatible endpoint on localhost.10 11Why server mode? Two measured reasons (100-page historical-scan A/B on `l4x1`,12same data, same sampling, driver concurrency 32):13 14| Model | Offline (inference-only) | Server | Speedup | Output parity |15|---|---|---|---|---|16| OvisOCR2 0.9B | 0.83 img/s | 1.44 img/s | ~1.7× | 94/100 byte-identical, rest ≥0.978 similar |17| LightOnOCR-2 1B | 0.33 img/s | 0.59 img/s | ~1.8× | 64/100 byte-identical at temp 0.2, rest median 0.998 |18| Nanonets-OCR2 3B | 0.31 img/s (steady-state) | 0.36 img/s | ~1.2× | 52/100 byte-identical, median 0.999; 3 hard plate pages fork under greedy (worst case was the *offline* arm degenerating into a 22k-char repetition loop) |19 201. **Throughput**: offline `llm.generate` drains at every batch boundary (GPU idles21   while the CPU decodes the next batch of images); a concurrent driver keeps vLLM's22   continuous batching fed. The speedup is a floor, not a ceiling — at concurrency 3223   the server's KV cache was <10% used.242. **Failure isolation**: offline, one bad image (e.g. a None cell) fails its whole25   batch of 16; server mode fails that one request. On a real dataset where ~half the26   image cells were empty, the offline recipe produced 0 usable outputs and the server27   recipe produced all of the valid ones.28 29The job command stays the standard `hf jobs uv run` shape: the driver spawns30`vllm serve` itself as a subprocess when no server is reachable (flags live in the31script's `SERVE_ARGS`, taken from the model's own card where they exist), so the only32thing to get right is `--image vllm/vllm-openai:<tag>` — which provides the `vllm`33binary. Without it the script fails fast, printing the exact correct command. Pass34`--server URL` to use an already-running or remote endpoint instead (nothing is35spawned when the server is reachable, so the moss-style explicit `bash -c` serve36command keeps working too).37 38For **interactive/agent** use (a live endpoint instead of a batch run), see39[serving-unlimited-ocr.md](serving-unlimited-ocr.md) — `hf jobs run --expose` gives an40OpenAI-compatible URL that outlives a single script.41 42## The recurring official serve pattern for OCR43 44Three flags recur across independent vendors' official serve commands (DeepSeek's vLLM45recipe, LightOn's card, Paddle's vLLM recipe, Unlimited-OCR's recipe):46 47```48--no-enable-prefix-caching --mm-processor-cache-gb 0 --limit-mm-per-prompt '{"image": 1}'49```50 51OCR workloads never reuse images, so prefix/multimodal caches only cost memory.52 53## Version pins carry over — via the image tag54 55Where an offline recipe pins an engine version, the server variant needs the same pin56as a `vllm/vllm-openai:<tag>` image (e.g. Nanonets-OCR2-3B uses `v0.10.2`, matching its57offline recipe). [`models.json`](models.json) records the required image per script.58When trying a new model/image combination, sanity-check the first few outputs before59scaling: a mismatched combination can produce degenerate output while looking like60normal load from the outside (full GPU utilisation, no errors).61 62## Which models document server mode themselves? (surveyed 2026-07-16)63 64Server mode is the officially documented path for most models in this collection —65for several of them it's the *only* documented vLLM path. Summary of each model's own66card/docs (verbatim commands live in the linked sources):67 68| Model | Official server example | Notes |69|---|---|---|70| lightonai/LightOnOCR-1B / 2-1B | ✅ vLLM | Card ships the serve command + client; ≥0.11.1 for v1; images longest-dim 1540px |71| nanonets/OCR-s / OCR2-3B | ✅ vLLM | Bare `vllm serve` + client in card; **pin v0.10.2 for OCR2** (see above) |72| rednote-hilab/dots.mocr | ✅ vLLM ≥0.11 | Direct from Hub id on `vllm/vllm-openai:v0.11.0`+ |73| rednote-hilab/dots.ocr | ✅ (GitHub) | HF card shows a legacy pre-0.11 hack; GitHub README: integrated upstream since 0.11 |74| tencent/HunyuanOCR | ✅ vLLM 0.18.1 | serve.sh in repo; nightly adds DFlash speculative decoding |75| zai-org/GLM-OCR | ✅ vLLM + SGLang + Ollama | Only card here with an SGLang serve example; needs vLLM nightly |76| numind/NuExtract3 | ✅ vLLM | Production-grade recipe: MTP speculative decoding, per-request `chat_template_kwargs` |77| numind/NuMarkdown-8B-Thinking | ✅ vLLM | Thinking always on — parse `<think>`/`<answer>` |78| baidu/Unlimited-OCR | ✅ vLLM + SGLang | Custom image / dev wheel; see [serving-unlimited-ocr.md](serving-unlimited-ocr.md) |79| baidu/Qianfan-OCR | ✅ vLLM (minimal) | Needs `--hf-overrides '{"architectures": ["InternVLChatModel"]}'` |80| PaddlePaddle/PaddleOCR-VL 1.x | ✅ paddle `genai_server` + vLLM recipe | Serves the 0.9B VLM only; raw serve skips the layout stage (official quality warning) |81| deepseek-ai/DeepSeek-OCR / -2 | ⚠️ vLLM recipe pages only | docs.vllm.ai recipes; needs the DeepSeek n-gram logits processor flags |82| allenai/olmOCR-2-7B-FP8 | ✅ (GitHub) | olmocr toolkit spawns `vllm serve` itself; YAML-front-matter prompt required |83| reducto/RolmOCR | ✅ vLLM | Card serve + client (client's model string has a typo — pass `--served-model-name`) |84| tiiuae/Falcon-OCR | ✅ vLLM in official Docker | `ghcr.io/tiiuae/falcon-ocr`; task-token prompts |85| datalab-to/surya-ocr-2, lift | ⚠️ wrapper-mediated | Server-native but driven via their own managers (`SURYA_INFERENCE_URL`, `lift_vllm`) |86| LiquidAI LFM2.5-VL-Extract | ⚠️ family docs | vLLM ≥0.23 + SGLang cookbooks, not Extract-specific |87| ATH-MaaS/OvisOCR2 | ❌ offline vLLM only | `ovis-ocr2-server.py` is our translation of the card's offline args (parity-validated, table above) |88| FireRedTeam/FireRed-OCR | ❌ transformers only | |89| acvlab/ABot-OCR | ❌ offline script only | pinned vllm 0.18.0 |90 91SGLang reality check: only GLM-OCR and Unlimited-OCR ship SGLang serve paths (olmOCR92dropped SGLang for vLLM in v0.1.75). For this collection, "server mode" effectively93means a vLLM OpenAI-compatible endpoint.94