vllm-sr/halueval-spans
HaluEval Span-Level Dataset (LLM-Detected) ๐ High-quality span-level hallucination detection dataset converted from HaluEval using Qwen2.5-72B-Instruct for precise span detection and RAGTruth-compatible labeling. Quick Start from datasets import load_dataset dataset = load_dataset("llm-semantic-router/halueval-spans") Why This Dataset? Problem Previous Solution This Dataset HaluEval has binary labels only NLI-based conversion โ โฆ See the full description on the dataset page: https://huggingface.co/datasets/vllm-sr/halueval-spans.
HaluEval Span-Level Dataset (LLM-Detected)
๐ High-quality span-level hallucination detection dataset converted from HaluEval using Qwen2.5-72B-Instruct for precise span detection and RAGTruth-compatible labeling.
Quick Start
from datasets import load_dataset
dataset = load_dataset("llm-semantic-router/halueval-spans")Why This Dataset?
What's New (v2.1 - LLM-Detected + Normalized)
This version uses Qwen2.5-72B-Instruct for span detection AND normalizes prompts to match RAGTruth format:
Statistics
- 10,000 summarization samples
- 8,905 samples with detected hallucinations (89%)
- 16,359 total hallucinated spans
- 1.8 average spans per hallucinated sample
- 99.4% partial spans (not marking entire responses)
Label Distribution (RAGTruth-Compatible)
Format
{
"prompt": "Summarize the following text within 150 words:\nFrance is a country located in Western Europe...",
"answer": "France is located in Asia and is known for...",
"labels": [
{"start": 0, "end": 25, "label": "Evident Conflict"}
],
"task_type": "Summary",
"split": "train",
"dataset": "halueval_summarization",
"language": "en"
}Fields
Label Types
Detection Method
Used Qwen2.5-72B-Instruct via vLLM for semantic span detection:
prompt = """Analyze the summary for hallucinations.
DOCUMENT:
{document}
SUMMARY:
{summary}
For each hallucinated span, provide:
- exact_text: the hallucinated phrase
- label: Evident Conflict | Evident Baseless Info | Subtle Conflict | Subtle Baseless Info
- reason: brief explanation
"""Why LLM over NLI?
Validation
All spans were validated for:
- โ Valid indices - within answer bounds
- โ Correct extraction - extracted text matches span
- โ Partial coverage - not marking entire responses
- โ Word boundaries - no mid-word cuts
Use Cases
- ๐ฏ Train token-level hallucination detectors with RAGTruth compatibility
- ๐ Augment RAGTruth for more diverse training data
- ๐ฌ Research on fine-grained hallucination types
- โ๏ธ Multi-label classification (4 hallucination types)
- ๐ Summarization-focused hallucination detection
Compatibility
This dataset uses the same label schema as RAGTruth, making it suitable for:
- Combined training with RAGTruth
- Transfer learning from RAGTruth models
- Evaluation using RAGTruth benchmarks
Related
- llm-semantic-router/modernbert-base-32k-haldetect - 32K hallucination detector
- RAGTruth - Primary span-level dataset
- HaluEval - Original binary dataset
- LettuceDetect - Hallucination detection framework
Citation
@misc{halueval_llm_spans_2026,
title={HaluEval LLM-Detected Span-Level Dataset},
author={llm-semantic-router},
year={2026},
howpublished={Hugging Face Hub},
note={Converted from HaluEval using Qwen2.5-72B-Instruct for RAGTruth-compatible span detection}
}License
MIT License (same as original HaluEval)
Changelog
v2.1 (2026-01)
- Normalized prompts to match RAGTruth format:
Summarize the following text within 150 words: - Fixed
task_type:summarizationโSummary(matches RAGTruth) - Removed
original_promptfield (now 7 fields, matching RAGTruth exactly) - Ready for combined training with RAGTruth
v2.0 (2026-01)
- Replaced NLI-based conversion with LLM-based span detection (Qwen2.5-72B-Instruct)
- Added 4 RAGTruth-compatible label types
- Improved span granularity (99.4% partial spans)
- Validated all span indices
v1.0 (2024)
- Initial NLI-based conversion using DeBERTa-FEVER-ANLI
