uv-script
ocr
OCR UV Scripts
Part of uv-scripts: self-contained UV scripts you run on Hugging Face Jobs in one command.
One script per OCR model. Each script runs the model on a GPU with Hugging Face Jobs and writes the text as markdown: as a new column in a Hub dataset, as .md files in a Bucket, or as resumable parquet parts (the -saturate recipes). A few scripts return JSON from a schema, detect layout regions, or compare the output of two models.
Quick Start
First… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/ocr.ocr-demo
OCR demo: Food for Space Flight
Seven scanned pages from NASA's Food for Space Flight booklet, with headings,
columns, photographs, food lists and tables. This small dataset is an input for
trying OCR recipes and inspecting their results.
The images are PDF pages 3-9 (printed pages 2-8) of the original booklet.
The selection omits the reproduction disclaimer and dark cover. The complete
original PDF and a matching seven-page extract are in the
OCR demo Bucket.
Use as… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/ocr-demo.classification
Classification Scripts
Text classification on HF Jobs: label a dataset with a model that needs no training, or train your own classifier from labelled examples.
If you have seen Jev and other "System One" models: the
models these scripts train are small, open versions of the same idea. They read a piece of data and
return a label with a probability, and you can train one on your own labels. For example, this
demo suggests task tags for any Hub
dataset; its model was fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/classification.training
Streaming LLM Training with Unsloth
Train on massive datasets without downloading anything - data streams directly from the Hub.
🦥 Latin LLM Example
Teaches Qwen Latin using 1.47M texts from FineWeb-2, streamed directly from the Hub.
Blog post: Train on Massive Datasets Without Downloading
Quick Start
# Run on HF Jobs (recommended - 2x faster streaming)
hf jobs uv run latin-llm-streaming.py \
--flavor a100-large \
--timeout 2h \
--secrets HF_TOKEN \
--… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/training.transcription
Transcription
Scripts for transcribing — and diarizing — audio files using HF Buckets and Jobs. English and the major European/Asian languages via Cohere Transcribe; 102 languages via the BuzzASR monolingual Whisper fine-tunes.
Quick Start
Scripts run directly from their Hub URL — no clone or local checkout needed:
# 1. Download audio from Internet Archive straight into a bucket
hf jobs uv run \
-v hf://buckets/user/audio-files:/output \… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/transcription.build-atlas
Atlas Export
Embedding Atlas is an open-source library from Apple for creating interactive, browser-based visualizations of embedding spaces. It renders millions of data points with WebGPU acceleration, supports real-time search and filtering, and automatically generates cluster labels.
These scripts wrap Embedding Atlas to make it easy to go from a HuggingFace dataset to a deployed visualization. See open-library-atlas for a live example (2M books).
Scripts… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/build-atlas.
