datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ocr
OCR UV Scripts
Part of uv-scripts: self-contained UV scripts you run on Hugging Face Jobs in one command.
One script per OCR model. Each script runs the model on a GPU with Hugging Face Jobs and writes the text as markdown: as a new column in a Hub dataset, as .md files in a Bucket, or as resumable parquet parts (the -saturate recipes). A few scripts return JSON from a schema, detect layout regions, or compare the output of two models.
Quick Start
First… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/ocr.ocr-demo
OCR demo: Food for Space Flight
Seven scanned pages from NASA's Food for Space Flight booklet, with headings,
columns, photographs, food lists and tables. This small dataset is an input for
trying OCR recipes and inspecting their results.
The images are PDF pages 3-9 (printed pages 2-8) of the original booklet.
The selection omits the reproduction disclaimer and dark cover. The complete
original PDF and a matching seven-page extract are in the
OCR demo Bucket.
Use as… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/ocr-demo.classification
Classification Scripts
Text classification on HF Jobs: label a dataset with a model that needs no training, or train your own classifier from labelled examples.
If you have seen Jev and other "System One" models: the
models these scripts train are small, open versions of the same idea. They read a piece of data and
return a label with a probability, and you can train one on your own labels. For example, this
demo suggests task tags for any Hub
dataset; its model was fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/classification.training
Streaming LLM Training with Unsloth
Train on massive datasets without downloading anything - data streams directly from the Hub.
🦥 Latin LLM Example
Teaches Qwen Latin using 1.47M texts from FineWeb-2, streamed directly from the Hub.
Blog post: Train on Massive Datasets Without Downloading
Quick Start
# Run on HF Jobs (recommended - 2x faster streaming)
hf jobs uv run latin-llm-streaming.py \
--flavor a100-large \
--timeout 2h \
--secrets HF_TOKEN \
--… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/training.transcription
Transcription
Scripts for transcribing — and diarizing — audio files using HF Buckets and Jobs. English and the major European/Asian languages via Cohere Transcribe; 102 languages via the BuzzASR monolingual Whisper fine-tunes.
Quick Start
Scripts run directly from their Hub URL — no clone or local checkout needed:
# 1. Download audio from Internet Archive straight into a bucket
hf jobs uv run \
-v hf://buckets/user/audio-files:/output \… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/transcription.build-atlas
Atlas Export
Embedding Atlas is an open-source library from Apple for creating interactive, browser-based visualizations of embedding spaces. It renders millions of data points with WebGPU acceleration, supports real-time search and filtering, and automatically generates cluster labels.
These scripts wrap Embedding Atlas to make it easy to go from a HuggingFace dataset to a deployed visualization. See open-library-atlas for a live example (2M books).
Scripts… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/build-atlas.object-detection
Object Detection Dataset Scripts
8 scripts to create, convert, review, validate, inspect, diff, and sample object detection datasets on the Hub. Supports 6 bbox formats — no setup required.
Start from nothing: falcon-perception.py generates a first-pass detection dataset for any class you can name, zero-shot, with no labelling and no training. The other six then convert, check, and measure it.
This repository is inspired by panlabel
Quick Start
Convert bounding… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/object-detection.sam3
SAM3 Vision Scripts
Detect and segment objects in images using Meta's SAM3 (Segment Anything Model 3) with text prompts. Process HuggingFace datasets with zero-shot detection and segmentation using natural language descriptions.
Script
What it does
Output
detect-objects.py
Object detection with bounding boxes
objects column with bbox, category, score
segment-objects.py
Pixel-level segmentation masks
Segmentation maps or per-instance masks
Browse results… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/sam3.vlm-object-detection
VLM Object Detection
Instruction-prompted object detection with vision-language models via vLLM. Designed as a VLM-as-labeller primitive for bootstrapping object-detection datasets — give it a free-form prompt ("detect every photograph and illustration", "detect all PPE items", "detect every electronic component and identify its reference designator") and it returns bbox JSON ready for downstream labelling tools (Label Studio, FiftyOne, COCO conversion).
Sibling: uv-scripts/sam3… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/vlm-object-detection.embeddings
Embeddings
Generate embeddings for a Hugging Face dataset — text or images — with one command, on a
cloud GPU, no infra. The output lands back on the Hub as a new dataset (or, with the Lance
variant, as a searchable vector index you can query over hf:// without downloading).
There is one simple default and two variants; they are separate single-file scripts because
their dependencies (sentence-transformers vs vLLM vs Lance) are too different to share one env.
Script
Use it… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/embeddings.hf-cli-jobs-uv-run-scripts
UV Script: job.py
Executed via hf jobs uv run on 2026-09-25 12:05:43 UTC
Run this script
hf jobs uv run job.py
Created with hf jobs
video
Video
Scripts for captioning and temporally grounding video files using HF Buckets and Jobs.
What the output looks like — a frame from Joan Avoids a Cold (1947, Prelinger Archives) with the event Marlin-2B produced for that moment:
Quick Start
Scripts run directly from their Hub URL — no clone or local checkout needed:
# Caption every video in a bucket: dense scene captions + timestamped events
hf jobs uv run --image vllm/vllm-openai:latest --flavor a10g-small \… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/video.vllm
vLLM Inference Scripts
Ready-to-run UV scripts for GPU-accelerated inference using vLLM.
These scripts use UV's inline script metadata to automatically manage dependencies - just run with uv run and everything installs automatically!
📋 Available Scripts
vlm-classify.py
Vision Language Model (VLM) image classification with structured output constraints.
Features:
🖼️ Process images through state-of-the-art VLMs (Qwen2-VL)
🎯 Structured… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/vllm.dataset-creation
Dataset Creation Scripts
Ready-to-run scripts for creating Hugging Face datasets from local files.
Available Scripts
📄 pdf-to-dataset.py
Convert directories of PDF files into Hugging Face datasets.
Features:
📁 Uploads PDFs as dataset objects for flexible processing
🏷️ Automatic labeling from folder structure
🚀 Zero configuration - just point at your PDFs
📤 Direct upload to Hugging Face Hub
Usage:
# Basic usage
uv run pdf-to-dataset.py /path/to/pdfs… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/dataset-creation.openai-oss
🚀 OpenAI GPT OSS Models - Simple Generation Script
Generate synthetic datasets using OpenAI's GPT OSS models with transparent reasoning. Works on HuggingFace Jobs with L4 GPUs!
✅ Tested & Working
Successfully tested on HF Jobs with l4x4 flavor (4x L4 GPUs = 96GB total memory).
🚀 Getting Started with HF Jobs
First-time Setup (2 minutes)
Install HuggingFace CLI:
pip install huggingface-hub
Login to HuggingFace:
huggingface-cli login
(Enter your… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/openai-oss.hf-cli-jobs-uv-run-scripts
UV Script: dpo_training.py
Executed via hf jobs uv run on 2025-10-20 12:57:28 UTC
Run this script
hf jobs uv run dpo_training.py
Created with hf jobs
dataset-stats
Dataset Statistics
UV scripts for analyzing HuggingFace datasets using streaming mode.
Scripts
finepdfs-stats.py - Temporal Educational Quality Analysis
Analyze educational quality trends across CommonCrawl dumps using Polars streaming. Answers: "Is the web getting more educational over time?"
Features:
Polars streaming (no download of 300GB+ dataset)
Temporal analysis across 106 CommonCrawl dumps (2013-2025)
ASCII chart visualizations
Uploads… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/dataset-stats.hf-cli-jobs-uv-run-scripts
UV Script: train_sorcerer.py
Executed via hf jobs uv run on 2026-03-19 10:33:43 UTC
Run this script
hf jobs uv run train_sorcerer.py
Created with hf jobs
data-processing
Data processing
Part of uv-scripts — self-contained UV scripts you run on Hugging Face Jobs in one command.
General data processing recipes: convert, clean and prepare data files.
Script
What it does
optimize-parquet.py
Converts CSV, JSON and Parquet files uploaded to a bucket into optimized Parquet, triggered by a bucket webhook
optimize-parquet.py: optimized Parquet from bucket uploads
Upload a CSV, JSON or Parquet file to a Storage Bucket and… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/data-processing.marimo
Marimo UV Scripts
Marimo notebooks that work as both interactive tutorials and batch scripts.
What is this?
Marimo notebooks are pure Python files that can be:
Edited interactively with a reactive notebook interface
Run as scripts with uv run - same as any UV script
This makes them perfect for tutorials and educational content where you want users to explore step-by-step, but also run the whole thing as a batch job.
Available Scripts
Script
Description… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/marimo.gliner
GLiNER UV Scripts
Zero-shot named-entity recognition over Hugging Face datasets using GLiNER. Pass a list of entity types at runtime — no fine-tuning required.
For text classification with GLiNER2 (a separate library from Fastino): zero-shot, then fine-tuning on your own labels, see uv-scripts/classification.
Script
What it does
Output
extract-entities.py
Extract entities from a text column with a custom set of types
New entities column (list of {start, end, text… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/gliner.doc-uv-scriptsdeduplication
Semantic Deduplication UV Script
Part of uv-scripts — self-contained UV scripts you run locally or on Hugging Face Jobs in one command.
Remove duplicate / near-duplicate text samples from a Hugging Face dataset by semantic similarity — clean training data and prevent train/test leakage. Uses SemHash with Model2Vec embeddings: CPU-optimized, no GPU required.
Quick start
# CPU is enough — run on Hugging Face Jobs
hf jobs uv run --flavor cpu-upgrade --secrets… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/deduplication.synthetic-data
CoT-Self-Instruct: High-Quality Synthetic Data Generation
Generate high-quality synthetic training data using Chain-of-Thought Self-Instruct methodology. This UV script implements the approach from "CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks" (2025).
🚀 Quick Start
# Install UV if you haven't already
curl -LsSf https://astral.sh/uv/install.sh | sh
# Generate synthetic reasoning data
uv run cot-self-instruct.py \… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/synthetic-data.jobs-utils
HF Jobs Utilities
Small utility scripts for working with the uv-scripts collection and Hugging Face Jobs — discovering the available recipes, and testing/benchmarking your Jobs setup.
Available Scripts
list-recipes.py
List every recipe (UV script) across the whole uv-scripts org, with a runnable URL for each. The zero-setup way to see what's available — no GPU, no token, no account; it only reads public repos.
Run locally (nothing to install but… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/jobs-utils.iiif-tiles
IIIF Static Tiles from HF Storage Buckets
Generate IIIF Image API 3.0 Level 0 static tiles from images and serve them via Hugging Face Storage Buckets — no image server required.
Drop images into a bucket, run one command, and get deep-zoom viewing in any IIIF viewer (Mirador, Universal Viewer, OpenSeadragon).
Demo
View in Mirador — 6 pages from the Wellcome Collection, served entirely from an HF Storage Bucket.
How it works
Source images (bucket or… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/iiif-tiles.hf-cli-jobs-uv-run-scripts
UV Script: permissions.py
Executed via hf jobs uv run on 2025-12-19 16:34:27 UTC
Run this script
hf jobs uv run permissions.py
Created with hf jobs
transformers-inference
Transformers Continuous Batching Scripts
GPU inference scripts using transformers' native continuous batching (CB). No vLLM dependency required.
Why transformers CB?
Instant new model support - works with any model supported by transformers, including newly released architectures. No waiting for vLLM to add support.
No dependency headaches - no vLLM, flashinfer, or custom wheel indexes. Just transformers + accelerate.
Simple HF Jobs setup - no Docker image needed. Just… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/transformers-inference.uv-scriptshf-cli-jobs-uv-run-scripts
UV Script: filter_cold_cases_california_mentions.py
Executed via hf jobs uv run on 2026-03-05 20:14:59 UTC
Run this script
hf jobs uv run filter_cold_cases_california_mentions.py
Created with hf jobs
