datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vllm-control-arena
vLLM Main Tasks Dataset
AI coding tasks generated from vLLM git commits
Dataset Description
This dataset contains 6801 coding tasks automatically generated from git commits in the vLLM repository. Each task represents a real-world coding challenge derived from actual development work.
Dataset Structure
The dataset contains the following columns:
commit_hash: The git commit hash
parent_hash: The parent commit hash
commit_title: The original commit… See the full description on the dataset page: https://huggingface.co/datasets/RoganInglis/vllm-control-arena.vllm-traces-v2vllm-0.28.0-wheels-py312grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate-completions.vllm-inference-benchmarksvllm-metadataMIRB
Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning
File Structure
├── MIR
|── analogy.json
│── codeu.json
|── dataset_namex.json
└── Images
├── analogy
│ └── image_x.jpg
└──codeu
└── image_x.jpg
JSON Structure
{
"questions": " What is the expected kurtosis of the sequence created by`create_number_sequence(-10, 10)`?\n\n1.… See the full description on the dataset page: https://huggingface.co/datasets/VLLMs/MIRB.BrainLab-Activation-Oracles-VLLMCopies of Visual AO target-organism train/val files and the local adapter registry.
Sources under the repo data/ tree. adapter_registry.local.json still points at absolute paths on this machine.
path
source
adapter_registry.local.json
data/val/target_organisms/adapter_registry.json
train/*/sft.jsonl
data/train/*/sft.jsonl
val/*/sft.jsonl
data/val/*/sft.jsonl
val/*/validation_manifest.json
data/val/*/validation_manifest.json
jailbreak-detection-dataset
Jailbreak Detection Dataset (MLCommons-Aligned)
A comprehensive dataset for training AI safety classifiers, aligned with the MLCommons AI Safety taxonomy.
Dataset Description
This dataset combines multiple sources for robust jailbreak and safety detection:
Primary Sources
nvidia/Aegis-AI-Content-Safety-Dataset-2.0: 18,164 samples with MLCommons-aligned labels
lmsys/toxic-chat: Toxic content detection
jackhhao/jailbreak-classification: Jailbreak… See the full description on the dataset page: https://huggingface.co/datasets/vllm-sr/jailbreak-detection-dataset.VLLM_ChartQAMIRB-hfdetails_hosted_vllm____fsx__anton__deepseek-r1-checkpoint_private
Dataset Card for Evaluation run of hosted_vllm//fsx/anton/deepseek-r1-checkpoint
Dataset automatically created during the evaluation run of model hosted_vllm//fsx/anton/deepseek-r1-checkpoint.
The dataset is composed of 15 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_hosted_vllm____fsx__anton__deepseek-r1-checkpoint_private.output_3d_bounding_scannetppv2_vllm_old_descriptiondecision-2.0-decision-index
Decision 2.0: Decision Index 0.2.1 runs
Complete runs of the released Decision 2.0 models on the Decision Index 0.2.1 suite, scored with the unmodified kit at 87d4650 (decision_index score --edition 0.2.1).
Model
Revision
Decision Index
Raw index
Breadth skill
Requests ok
unsupported
Decision 2.0 Vega 27B
7aec49ae
56.47
66.91
55.47
150,317
0
Decision 2.0 Lux 9B
78bf3c03
46.26
59.02
44.95
150,315
2
Decision 2.0 Nox 4B
25e8f67d
43.77
57.24
42.10
150,315
2
Decision… See the full description on the dataset page: https://huggingface.co/datasets/vllm-sr/decision-2.0-decision-index.ViLLM-Eval
ViLLM-Eval
We utilize the lm-eval-harness library to conduct evaluations.
This library allows us to efficiently evaluate language models, ensuring robustness and accuracy in our assessments.
Feel free to explore our project and discover the capabilities of the language models we employ.
Install
git clone https://huggingface.co/datasets/vlsp-2023-vllm/ViLLM-Eval
cd ViLLM-Eval
pip install -e .
Basic Usage
# Add trust_remote_code=True if your model is a custom… See the full description on the dataset page: https://huggingface.co/datasets/vlsp-2023-vllm/ViLLM-Eval.feedback-detector-dataset
Feedback Detector Dataset
A large-scale multilingual dataset for 4-class user feedback classification, labeled using GPT-OSS-120B on AMD MI300X GPU.
Dataset Description
This dataset contains 51,694 examples of user feedback classified into 4 categories:
Label
Description
Count
%
SAT
User is satisfied
8,649
17%
NEED_CLARIFICATION
User needs more information
16,179
31%
WRONG_ANSWER
System gave incorrect response
19,919
39%
WANT_DIFFERENT
User wants… See the full description on the dataset page: https://huggingface.co/datasets/vllm-sr/feedback-detector-dataset.fact-check-classification-dataset
Fact-Check Classification Dataset
🎯 Purpose: Binary classification dataset for determining whether a prompt needs external fact-checking.
Dataset Description
This dataset is designed to train classifiers that can route LLM requests based on whether they require external fact verification. It's part of the vLLM Semantic Router project.
Labels
FACT_CHECK_NEEDED (1): Information-seeking questions requiring external verification
Factual questions… See the full description on the dataset page: https://huggingface.co/datasets/vllm-sr/fact-check-classification-dataset.longcontext-haldetect
Long-Context Hallucination Detection Benchmark
A synthetic benchmark dataset for evaluating hallucination detection models on long documents (8K-24K tokens). This dataset is specifically designed to test models that can handle contexts beyond the typical 8K token limit.
Dataset Summary
Property
Value
Total samples
3,366
Token range
8,005 - 23,998
Average tokens
17,852
Hallucinated
1,681 (49.9%)
Supported
1,685 (50.1%)
Splits… See the full description on the dataset page: https://huggingface.co/datasets/vllm-sr/longcontext-haldetect.vLLM-SR-Preference-V1The files in this repo is the LLM-labeled samples that are used as the training dataset for vLLM-SR Preference model V1.
The training file (sharegpt_preference_labeld_with_negative.jsonl) contains 25k records that have sample_id, golden label for the preference-based routing policy, and a set of negative labels that are plausible but do not match the conversation context.
The validation file has the same structure, but only 1% of the training file size. The validation file and the training… See the full description on the dataset page: https://huggingface.co/datasets/ppppqp/vLLM-SR-Preference-V1.deepseek-v4-flash-rocm-vllm-repro
Reproducing DeepSeek-V4-Flash on AMD ROCm with vLLM: 32K Correctness and TopK Sweep
This article summarizes an engineering reproduction of
deepseek-ai/DeepSeek-V4-Flash on an AMD ROCm ModelScope DSW instance. The work
focuses on a practical question: can a complex, fast-moving DeepSeek-V4-Flash
serving path be turned into a reproducible ROCm baseline with explicit
correctness gates?
The answer from this run is yes, with an important boundary: the current setup
is a fallback-heavy… See the full description on the dataset page: https://huggingface.co/datasets/lyydfys/deepseek-v4-flash-rocm-vllm-repro.mlcommons-ai-safety-synth
MLCommons AI Safety Synthesized Dataset
Synthesized training data for AI safety classifiers based on the MLCommons AI Safety Hazard Taxonomy.
Dataset Description
This dataset contains 12,000 synthesized unsafe prompts across 6 hazard categories, designed to augment training data for content safety classifiers. Each category contains 2,000 balanced samples.
Hazard Categories (MLCommons AI Safety Taxonomy)
Category
Description
Samples… See the full description on the dataset page: https://huggingface.co/datasets/vllm-sr/mlcommons-ai-safety-synth.modality-routing-dataset
Modality Routing Dataset
This dataset materializes the dynamic modality routing data builder used by the local
mmBERT-32K modality router training pipeline. The export is intended for review,
versioning, and uploading to a Hugging Face dataset repository.
Labels
Label
ID
Description
AR
0
Text-only requests that should route to an autoregressive LLM.
DIFFUSION
1
Image-generation requests that should route to a diffusion model.
BOTH
2
Requests that… See the full description on the dataset page: https://huggingface.co/datasets/vllm-sr/modality-routing-dataset.Qwen2.5-7B-Instruct-vllm-20251128_042753VLLM_ChartQA_splitrouter-signal-suite
Router signal suite
Training and test data for the ten vLLM Semantic Router signals: domain, jailbreak, safety, hazard, fact check,
modality, PII, feedback, hallucination and tool need. The label cannot be read off the corpus a row came from, and
every signal has held-out corpora. The design is in
semantic-router#4305, and the comparison scored on it
is in semantic-router#4306.
The builder generates no rows. Every row comes from a public corpus at a pinned revision. Text is… See the full description on the dataset page: https://huggingface.co/datasets/vllm-sr/router-signal-suite.ruletaker-vllm-reasoning-3gumble-max-vllm-experimentprocessed_qwen25_7b_vllm_final_4halueval-spans
HaluEval Span-Level Dataset (LLM-Detected)
🔍 High-quality span-level hallucination detection dataset converted from HaluEval using Qwen2.5-72B-Instruct for precise span detection and RAGTruth-compatible labeling.
Quick Start
from datasets import load_dataset
dataset = load_dataset("llm-semantic-router/halueval-spans")
Why This Dataset?
Problem
Previous Solution
This Dataset
HaluEval has binary labels only
NLI-based conversion
✅… See the full description on the dataset page: https://huggingface.co/datasets/vllm-sr/halueval-spans.glm-5.3-flash-dgx-spark-vllm
GLM-5.3-Flash on 2x NVIDIA DGX Spark (GB10): a working multi-node vLLM config
TL;DR: GLM-5.3-Flash (NVFP4, modelopt quant) running tensor-parallel
across two DGX Spark GB10 nodes over 200GbE RoCEv2 RDMA, with NVFP4
KV-cache, CUDA graphs, and MTP-3 speculative decoding:
~20-23 tok/s single-stream, ~49 tok/s aggregate at 4 concurrent
requests on 64K context. About 2x over a naive TCP/eager launch.
This repo documents the exact configuration, the four vLLM patches it
needs… See the full description on the dataset page: https://huggingface.co/datasets/H-K-B/glm-5.3-flash-dgx-spark-vllm.
