Team Ai
Datasetpublic

vllm-sr/halueval-spans

HaluEval Span-Level Dataset (LLM-Detected) ๐Ÿ” High-quality span-level hallucination detection dataset converted from HaluEval using Qwen2.5-72B-Instruct for precise span detection and RAGTruth-compatible labeling. Quick Start from datasets import load_dataset dataset = load_dataset("llm-semantic-router/halueval-spans") Why This Dataset? Problem Previous Solution This Dataset HaluEval has binary labels only NLI-based conversion โœ…โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/vllm-sr/halueval-spans.

sourceHugging Facemitupdated 9mo agoView on Hugging Face
1likes125downloads
Dataset Card

HaluEval Span-Level Dataset (LLM-Detected)

๐Ÿ” High-quality span-level hallucination detection dataset converted from HaluEval using Qwen2.5-72B-Instruct for precise span detection and RAGTruth-compatible labeling.

Quick Start

python
from datasets import load_dataset

dataset = load_dataset("llm-semantic-router/halueval-spans")

Why This Dataset?

ProblemPrevious SolutionThis Dataset
HaluEval has binary labels onlyNLI-based conversionโœ… LLM-based span detection
NLI marks entire sentencesSentence-level spansโœ… Partial spans (avg 15% coverage)
Generic "hallucinated" labelSingle labelโœ… 4 RAGTruth-compatible types
Full-response marking50%+ of response markedโœ… 99.4% partial spans (<50% coverage)

What's New (v2.1 - LLM-Detected + Normalized)

This version uses Qwen2.5-72B-Instruct for span detection AND normalizes prompts to match RAGTruth format:

Metricv1 (NLI)v2.0 (LLM)v2.1 (Normalized)
Span granularitySentence-levelToken-levelToken-level
Partial spans (<50%)~60%99.4%99.4%
Label types14 (RAGTruth)4 (RAGTruth)
Prompt formatDocument:...Summary:Document:...Summary:`Summarize the following text within N words:`
task_typesummarizationsummarization`Summary` (matches RAGTruth)
Fields887 (matches RAGTruth exactly)

Statistics

  • โ€”10,000 summarization samples
  • โ€”8,905 samples with detected hallucinations (89%)
  • โ€”16,359 total hallucinated spans
  • โ€”1.8 average spans per hallucinated sample
  • โ€”99.4% partial spans (not marking entire responses)

Label Distribution (RAGTruth-Compatible)

LabelCountPercentageDescription
Evident Baseless Info6,75441.3%Information clearly not in context
Evident Conflict4,71828.8%Clearly contradicts context
Subtle Baseless Info2,59415.9%Plausible but unverifiable claims
Subtle Conflict2,29314.0%Subtly contradicts context

Format

json
{
  "prompt": "Summarize the following text within 150 words:\nFrance is a country located in Western Europe...",
  "answer": "France is located in Asia and is known for...",
  "labels": [
    {"start": 0, "end": 25, "label": "Evident Conflict"}
  ],
  "task_type": "Summary",
  "split": "train",
  "dataset": "halueval_summarization",
  "language": "en"
}

Fields

FieldTypeDescription
promptstringTask instruction + document (RAGTruth format)
answerstringLLM-generated summary to evaluate
labelslistSpan annotations with character offsets and RAGTruth labels
task_typestringTask type (Summary - matches RAGTruth)
splitstringData split (train)
datasetstringSource dataset identifier
languagestringLanguage code (en)

Label Types

LabelDescriptionExample
Evident ConflictDirectly contradicts the source"Anne Frank survived" when document says she died
Evident Baseless InfoInformation not present in sourceAdding dates/facts not mentioned
Subtle ConflictIndirectly contradicts via implicationMisattributing quotes or actions
Subtle Baseless InfoPlausible but unverifiable additions"likely" claims without evidence

Detection Method

Used Qwen2.5-72B-Instruct via vLLM for semantic span detection:

python
prompt = """Analyze the summary for hallucinations.

DOCUMENT:
{document}

SUMMARY:
{summary}

For each hallucinated span, provide:
- exact_text: the hallucinated phrase
- label: Evident Conflict | Evident Baseless Info | Subtle Conflict | Subtle Baseless Info
- reason: brief explanation
"""

Why LLM over NLI?

AspectNLI ApproachLLM Approach
GranularitySentence-level onlySub-sentence spans
ContextLimited windowFull document context
LabelsBinary (contradiction/entailment)Multi-class semantic
UnderstandingPattern matchingSemantic reasoning

Validation

All spans were validated for:

  • โ€”โœ… Valid indices - within answer bounds
  • โ€”โœ… Correct extraction - extracted text matches span
  • โ€”โœ… Partial coverage - not marking entire responses
  • โ€”โœ… Word boundaries - no mid-word cuts

Use Cases

  • โ€”๐ŸŽฏ Train token-level hallucination detectors with RAGTruth compatibility
  • โ€”๐Ÿ“Š Augment RAGTruth for more diverse training data
  • โ€”๐Ÿ”ฌ Research on fine-grained hallucination types
  • โ€”โš–๏ธ Multi-label classification (4 hallucination types)
  • โ€”๐Ÿ“ Summarization-focused hallucination detection

Compatibility

This dataset uses the same label schema as RAGTruth, making it suitable for:

  • โ€”Combined training with RAGTruth
  • โ€”Transfer learning from RAGTruth models
  • โ€”Evaluation using RAGTruth benchmarks

Related

Citation

bibtex
@misc{halueval_llm_spans_2026,
  title={HaluEval LLM-Detected Span-Level Dataset},
  author={llm-semantic-router},
  year={2026},
  howpublished={Hugging Face Hub},
  note={Converted from HaluEval using Qwen2.5-72B-Instruct for RAGTruth-compatible span detection}
}

License

MIT License (same as original HaluEval)

Changelog

v2.1 (2026-01)

  • โ€”Normalized prompts to match RAGTruth format: Summarize the following text within 150 words:
  • โ€”Fixed task_type: summarization โ†’ Summary (matches RAGTruth)
  • โ€”Removed original_prompt field (now 7 fields, matching RAGTruth exactly)
  • โ€”Ready for combined training with RAGTruth

v2.0 (2026-01)

  • โ€”Replaced NLI-based conversion with LLM-based span detection (Qwen2.5-72B-Instruct)
  • โ€”Added 4 RAGTruth-compatible label types
  • โ€”Improved span granularity (99.4% partial spans)
  • โ€”Validated all span indices

v1.0 (2024)

  • โ€”Initial NLI-based conversion using DeBERTa-FEVER-ANLI