Team Ai
Datasetpublic

uv-scripts/vllm

vLLM Inference Scripts Ready-to-run UV scripts for GPU-accelerated inference using vLLM. These scripts use UV's inline script metadata to automatically manage dependencies - just run with uv run and everything installs automatically! πŸ“‹ Available Scripts vlm-classify.py Vision Language Model (VLM) image classification with structured output constraints. Features: πŸ–ΌοΈ Process images through state-of-the-art VLMs (Qwen2-VL) 🎯 Structured… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/vllm.

sourceHugging Faceupdated 4mo agoView on Hugging Face
13likes95downloads
Dataset Card

vLLM Inference Scripts

Ready-to-run UV scripts for GPU-accelerated inference using vLLM.

These scripts use UV's inline script metadata to automatically manage dependencies - just run with uv run and everything installs automatically!

πŸ“‹ Available Scripts

vlm-classify.py

Vision Language Model (VLM) image classification with structured output constraints.

Features:

  • β€”πŸ–ΌοΈ Process images through state-of-the-art VLMs (Qwen2-VL)
  • β€”πŸŽ― Structured classification using vLLM's GuidedDecodingParams
  • β€”πŸ“ Automatic image resizing to optimize token usage
  • β€”πŸ’Ύ Memory-efficient lazy batch processing
  • β€”πŸ·οΈ Simple CLI interface for defining classes
  • β€”πŸ€— Direct integration with Hugging Face datasets

Usage:

bash
# Basic classification
uv run vlm-classify.py \
    username/input-dataset \
    username/output-dataset \
    --classes "document,photo,diagram,other"

# With custom prompt and image resizing
uv run vlm-classify.py \
    username/input-dataset \
    username/output-dataset \
    --classes "index-card,manuscript,title-page,other" \
    --prompt "What type of historical document is this?" \
    --max-size 768

# Quick test with sample limit
uv run vlm-classify.py \
    davanstrien/sloane-index-cards \
    username/test-output \
    --classes "index,content,other" \
    --max-samples 10

HF Jobs execution:

bash
hf jobs uv run \
    --flavor a10g \
    --image vllm/vllm-openai \
    -s HF_TOKEN \
    https://huggingface.co/datasets/uv-scripts/vllm/raw/main/vlm-classify.py \
    username/input-dataset \
    username/output-dataset \
    --classes "title-page,content,index,other" \
    --max-size 768

Key Parameters:

  • β€”--classes: Comma-separated list of classification categories (required)
  • β€”--prompt: Custom classification prompt (optional, auto-generated if not provided)
  • β€”--max-size: Maximum image dimension in pixels for resizing (reduces token count)
  • β€”--model: VLM model to use (default: Qwen/Qwen2-VL-7B-Instruct)
  • β€”--batch-size: Number of images to process at once (default: 8)
  • β€”--max-samples: Limit number of samples for testing

classify-dataset.py

Batch text classification using BERT-style encoder models (e.g., BERT, RoBERTa, DeBERTa, ModernBERT) with vLLM's optimized inference engine.

Note: This script is specifically for encoder-only classification models, not generative LLMs.

Features:

  • β€”πŸš€ High-throughput batch processing
  • β€”πŸ·οΈ Automatic label mapping from model config
  • β€”πŸ“Š Confidence scores for predictions
  • β€”πŸ€— Direct integration with Hugging Face Hub

Usage:

bash
# Local execution (requires GPU)
uv run classify-dataset.py \
    davanstrien/ModernBERT-base-is-new-arxiv-dataset \
    username/input-dataset \
    username/output-dataset \
    --inference-column text \
    --batch-size 10000

HF Jobs execution:

bash
hf jobs uv run \
    --flavor l4x1 \
    --image vllm/vllm-openai \
    https://huggingface.co/datasets/uv-scripts/vllm/resolve/main/classify-dataset.py \
    davanstrien/ModernBERT-base-is-new-arxiv-dataset \
    username/input-dataset \
    username/output-dataset \
    --inference-column text \
    --batch-size 100000

generate-responses.py

Generate responses for prompts using generative LLMs (e.g., Llama, Qwen, Mistral) with vLLM's high-performance inference engine.

Features:

  • β€”πŸ’¬ Automatic chat template application
  • β€”πŸ“ Support for both chat messages and plain text prompts
  • β€”πŸ”€ Multi-GPU tensor parallelism support
  • β€”πŸ“ Smart filtering for prompts exceeding context length
  • β€”πŸ“Š Comprehensive dataset cards with generation metadata
  • β€”βš‘ HF Transfer enabled for fast model downloads
  • β€”πŸŽ›οΈ Full control over sampling parameters
  • β€”πŸŽ― Sample limiting with --max-samples for testing

Usage:

bash
# With chat-formatted messages (default)
uv run generate-responses.py \
    username/input-dataset \
    username/output-dataset \
    --messages-column messages \
    --max-tokens 1024

# With plain text prompts (NEW!)
uv run generate-responses.py \
    username/input-dataset \
    username/output-dataset \
    --prompt-column question \
    --max-tokens 1024 \
    --max-samples 100

# With custom model and parameters
uv run generate-responses.py \
    username/input-dataset \
    username/output-dataset \
    --model-id meta-llama/Llama-3.1-8B-Instruct \
    --prompt-column text \
    --temperature 0.9 \
    --top-p 0.95 \
    --max-model-len 8192

HF Jobs execution (multi-GPU):

bash
hf jobs uv run \
    --flavor l4x4 \
    --image vllm/vllm-openai \
    -e UV_PRERELEASE=if-necessary \
    -s HF_TOKEN \
    https://huggingface.co/datasets/uv-scripts/vllm/raw/main/generate-responses.py \
    davanstrien/cards_with_prompts \
    davanstrien/test-generated-responses \
    --model-id Qwen/Qwen3-30B-A3B-Instruct-2507 \
    --gpu-memory-utilization 0.9 \
    --max-tokens 600 \
    --max-model-len 8000

Multi-GPU Tensor Parallelism

  • β€”Auto-detects available GPUs by default
  • β€”Use --tensor-parallel-size to manually specify
  • β€”Required for models larger than single GPU memory (e.g., 30B+ models)

Handling Long Contexts

The generate-responses.py script includes smart prompt filtering:

  • β€”Default behavior: Skips prompts exceeding maxmodellen
  • β€”Use `--max-model-len`: Limit context to reduce memory usage
  • β€”Use `--no-skip-long-prompts`: Fail on long prompts instead of skipping
  • β€”Skipped prompts receive empty responses and are logged

πŸ“š About vLLM

vLLM is a high-throughput inference engine optimized for:

  • β€”Fast model serving with PagedAttention
  • β€”Efficient batch processing
  • β€”Support for various model architectures
  • β€”Seamless integration with Hugging Face models

πŸ”§ Technical Details

UV Script Benefits

  • β€”Zero setup: Dependencies install automatically on first run
  • β€”Reproducible: Locked dependencies ensure consistent behavior
  • β€”Self-contained: Everything needed is in the script file
  • β€”Direct execution: Run from local files or URLs

Dependencies

Scripts use UV's inline metadata for automatic dependency management:

python
# /// script
# requires-python = ">=3.10"
# dependencies = [
#     "datasets",
#     "flashinfer-python",
#     "huggingface-hub[hf_transfer]",
#     "torch",
#     "transformers",
#     "vllm",
# ]
# ///

For bleeding-edge features, use the UV_PRERELEASE=if-necessary environment variable to allow pre-release versions when needed.

Docker Image

For HF Jobs, we recommend the official vLLM Docker image: vllm/vllm-openai

This image includes:

  • β€”Pre-installed CUDA libraries
  • β€”vLLM and all dependencies
  • β€”UV package manager
  • β€”Optimized for GPU inference

Environment Variables

  • β€”HF_TOKEN: Your Hugging Face authentication token (auto-detected if logged in)
  • β€”UV_PRERELEASE=if-necessary: Allow pre-release packages when required
  • β€”HF_HUB_ENABLE_HF_TRANSFER=1: Automatically enabled for faster downloads

πŸ”— Resources