Team Ai
Datasetpublic

hussain-s/TemporalHallucination

TemporalScore Dataset Paper: TemporalScore: Measuring and Detecting Temporal Hallucination in LLM SummarizationVenue: CIKM 2026 (Short Research Paper)DOI: https://doi.org/10.1145/3799682.3840029 Dataset Description This dataset accompanies the TemporalScore paper and contains annotations for temporal hallucination in LLM-generated summaries. Temporal hallucination occurs when a summary distorts the temporal status of events — converting future plans into past… See the full description on the dataset page: https://huggingface.co/datasets/hussain-s/TemporalHallucination.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes56downloads
Dataset Card

TemporalScore Dataset

Paper: TemporalScore: Measuring and Detecting Temporal Hallucination in LLM Summarization Venue: CIKM 2026 (Short Research Paper) DOI: https://doi.org/10.1145/3799682.3840029

Dataset Description

This dataset accompanies the TemporalScore paper and contains annotations for temporal hallucination in LLM-generated summaries. Temporal hallucination occurs when a summary distorts the temporal status of events — converting future plans into past accomplishments, upgrading uncertain proposals to confirmed facts, or reordering events.

Taxonomy

Three categories of temporal hallucination:

  • —TS (Tense Shift): Future/planned events reported as past/completed
  • —CU (Certainty Upgrade): Hedged/conditional statements made definite
  • —TR (Temporal Reorder): Events appear in incorrect chronological sequence

Files

FileDescription
human_validation_1000.csv1,000 stratified sentences (500 LLM-positive + 500 LLM-negative) with source text, summary sentence, and LLM judge labels
annotator_1_labels.xlsxIndependent human annotations from Annotator 1
annotator_2_labels.xlsxIndependent human annotations from Annotator 2
corpus_annotations_summary.jsonAggregate statistics from the full 61,838-sentence corpus annotation
stratified_sample_1000.jsonlFull structured records for the 1,000 validation sentences
inter_judge_kappa.jsonInter-judge agreement statistics (Sonnet 4 vs. Opus 4.1)
source_articles/Source articles used in the pilot study (JSONL format)

Key Statistics

  • —Corpus size: 61,838 sentences from 14,168 summaries
  • —Source articles: 1,012 from 6 domains (financial news, SEC filings, central bank, government policy, CNN/DailyMail)
  • —Models evaluated: 7 LLMs (Claude Sonnet 4, Sonnet 4.5, Haiku 4.5, Llama 3.3 70B, Llama 3.1 8B, Mistral Large, Nova Pro)
  • —Overall temporal hallucination rate: 14.4% (judge-relative); ~22% population-weighted
  • —Inter-annotator κ: 0.925

License

  • —SEC filings, central bank communications, and government press releases: Public Domain
  • —CNN/DailyMail-derived subsets: Available under the original dataset license
  • —Annotations and taxonomy: CC-BY 4.0

Citation

bibtex
@inproceedings{hussain2026temporalscore,
  title={TemporalScore: Measuring and Detecting Temporal Hallucination in LLM Summarization},
  author={Hussain, Shabbir and Panjwani, Hari Charan and Malak, Fatema and Sinha, Sakshi Vikas Chandra and Srivastava, Himanshu},
  booktitle={Proceedings of the 35th ACM International Conference on Information and Knowledge Management (CIKM '26)},
  year={2026},
  publisher={ACM},
  doi={10.1145/3799682.3840029}
}