Team Ai
Datasetpublic

Tjayush/temporal-jitter

Temporal Jitter Temporal Jitter is a long-context benchmark for testing whether a language model can track a time-varying fact about an entity when that fact is buried among many unrelated, similarly-phrased distractor facts about other entities — i.e., whether the model can find the right needle in a haystack of temporal "jitter." Each example places one or more target facts (e.g. "X became Y's position holder on date D") inside a long context built mostly from distractor facts… See the full description on the dataset page: https://huggingface.co/datasets/Tjayush/temporal-jitter.

sourceHugging Facecc-by-4.0updated 3d agoView on Hugging Face
0likes68downloads
Dataset Card

Temporal Jitter

Temporal Jitter is a long-context benchmark for testing whether a language model can track a time-varying fact about an entity when that fact is buried among many unrelated, similarly-phrased distractor facts about other entities — i.e., whether the model can find the right needle in a haystack of temporal "jitter."

Each example places one or more target facts (e.g. "X became Y's position holder on date D") inside a long context built mostly from distractor facts about other real-world entities, at controlled context lengths (6K / 32K / 64K tokens) and difficulty levels (single-hop vs. multi-hop, 1–5 revisions per entity). Questions probe the model's ability to recall the current value of a fact, the historical value at a specific past date, or to cite the evidence (which chunk of context stated it).

Why this dataset exists

Most long-context QA benchmarks test retrieval of a single static fact. Real-world facts change over time (a person's job title, a country's head of state, an organization's leadership), and a model with working memory needs to track which value was true when — not just find a mention of the entity. Temporal Jitter isolates this specific failure mode: can the model resist interference from distractor facts that look structurally identical to the target, and still answer correctly about time?

Dataset structure

The dataset is organized as a 4-stage pipeline, with the final packaged/ and qa/ directories being the ones most consumers will want:

raw/            Time-indexed fact "chains" extracted from Wikidata (entity, relation,
                a sequence of values each with a start/end date), plus extraction
                statistics (chain_stats.json) and a wikidata_raw/ cache.
grounded/       Chains grounded against cached Wikipedia text (wikipedia_cache/).
narrated/       Each chain rendered into natural-language narration sentences
                (e.g. "Gwangjong of Goryeo assumed the position of monarch on
                January 1, 950.").
packaged/       Final long-context examples: one target chain's narration sentences
                are mixed into a context built mostly from *other* chains' narrations
                as distractors, packed to a target token budget. Subdirectories are
                named `{N}_{hop}_{context_variant}`, e.g. `3_multi_32k` =
                N=3 revisions, multi-hop, ~32K-token context. A manifest.jsonl
                indexes every packaged file.
qa/             Question-answer pairs over the packaged contexts (qa_pairs.jsonl),
                with three probe types:
                  - current_value:     "What position does X hold today?"
                  - historical_value:  "What position did X hold on <date>?"
                  - evidence_citation: "Which chunk first stated that X became <value>?"

Example packaged context chunk

json
{
  "role": "document",
  "content": "Daniel Cohn-Bendit became chairperson of The Greens-European Free Alliance on January 1, 2004.",
  "timestamp": "2004-01-01T00:00:00Z",
  "is_target": false,
  "is_distractor": true,
  "source_chain_id": "wd_P488_Q751935",
  "chunk_id": 0
}

Example QA pair

json
{
  "qa_id": "wd_P39_Q469387_6k_historical_value_0_wd_P39_Q469387",
  "chain_id": "wd_P39_Q469387",
  "hop": "single",
  "context_variant": "6k",
  "N": 1,
  "probe_type": "historical_value",
  "gold_answer": "monarch",
  "evidence_chunk_id": 111,
  "question": "What position did Gwangjong of Goryeo hold on January 1, 950?"
}

Scale

StageFileRows
rawchains.jsonl1,732
narratedchains_narrated.jsonl1,647
packagedmanifest.jsonl5,829 packaged contexts
qaqa_pairs.jsonl25,692 question-answer pairs

Packaged contexts span 3 context-length variants (6K / 32K / 64K tokens), single- and multi-hop variants, and 1–7 revision-count groups (N), covering 9 Wikidata relations (political/organizational positions, team membership, employer, etc.).

Source and licensing

The underlying facts (entity, relation, value, and date ranges) are extracted from Wikidata, which is released under CC0. The natural-language narrations, distractor packaging, and QA pairs in this dataset are derived/synthetic content generated for this project and are released under CC-BY-4.0. If you use this dataset, please cite Wikidata as the ultimate source of the underlying facts, and this repository for the derived benchmark construction.

Loading

python
from datasets import load_dataset

# QA pairs (the usual entry point for evaluation)
qa = load_dataset("Tjayush/temporal-jitter", "qa")

# Raw time-indexed fact chains
raw = load_dataset("Tjayush/temporal-jitter", "raw")

# Grounded chains
grounded = load_dataset("Tjayush/temporal-jitter", "grounded")

# Narrated chains
narrated = load_dataset("Tjayush/temporal-jitter", "narrated")

# Index of every packaged long-context file (use with hf_hub_download to fetch each one)
manifest = load_dataset("Tjayush/temporal-jitter", "packaged_manifest")

Packaged long-context files are per-example JSON files under packaged/{N}_{hop}_{context_variant}/, indexed by packaged/manifest.jsonl; load the manifest first to iterate over them:

python
import json

with open("packaged/manifest.jsonl") as f:
    manifest = [json.loads(line) for line in f]

# each entry: {"chain_id": ..., "N": ..., "hop": ..., "context_variant": ..., "path": ...}

Known caveats

  • —.pre_label_repair.bak files are backups kept from an earlier label-repair pass over raw/chains.jsonl, narrated/chains_narrated.jsonl, and qa/qa_pairs.jsonl — safe to ignore for normal use; kept for provenance/reproducibility.
  • —chain_stats.json documents the extraction/filtering funnel (e.g. chains dropped for missing labels, ambiguous dates, or wrong chain length) for transparency about how the raw chains were curated from the full Wikidata extraction.

Author

Ayushman Sarkar (@Tjayush)