Team Ai
Datasetpublic

vein05/ragscale-interaction-matrix

ragscale Interaction Matrix Reader answers under raw and compressed RAG evidence: 176,864 rows, one per benchmark item, reader model, and evidence policy, across LongMemEval, HotpotQA, MuSiQue, and NQ-Open. This is the interaction matrix released with the paper Compression Is Not Evaluation-Neutral: Fixed RAG Compression Can Distort Reader Comparisons (Sugam Panthi and Rabab Abdelfattah, arXiv:2606.21807). The paper gives 8 to 20 reader models the same stored compressed text and… See the full description on the dataset page: https://huggingface.co/datasets/vein05/ragscale-interaction-matrix.

sourceHugging Faceotherupdated 5d agoView on Hugging Face
0likes37downloads
Dataset Card

ragscale Interaction Matrix

Reader answers under raw and compressed RAG evidence: 176,864 rows, one per benchmark item, reader model, and evidence policy, across LongMemEval, HotpotQA, MuSiQue, and NQ-Open.

This is the interaction matrix released with the paper Compression Is Not Evaluation-Neutral: Fixed RAG Compression Can Distort Reader Comparisons (Sugam Panthi and Rabab Abdelfattah, arXiv:2606.21807). The paper gives 8 to 20 reader models the same stored compressed text and measures how much of a reader upgrade survives compression. On HotpotQA, one stored RECOMP output shrinks a 31.8-point gap between the lowest- and highest-scoring readers to 7.8 points.

  • —Paper: https://arxiv.org/abs/2606.21807
  • —Code (ragscale audit toolkit): https://github.com/aimsresearchlab/ragscale
  • —Companion post: https://spanthi.com/blog/compression-is-a-coin-flip/

The file is byte-identical to data/interaction_matrix.csv.gz in the ragscale repository (version 0.5.1), which also bundles it:

python
from ragscale import load_interaction_matrix

matrix = load_interaction_matrix()

or, from this repository:

python
from datasets import load_dataset

matrix = load_dataset("vein05/ragscale-interaction-matrix", split="train")

Columns

ColumnTypeDescription
datasetstrSource benchmark key: longmemeval, hotpotqa, musique, or nq
example_idstrItem identifier within the stored benchmark slice
questionstrQuestion text
question_typestrSource or analysis category when available
reference_answerstrStored reference answer
reader_modelstrReader model identifier
methodstrStored evidence-policy or diagnostic-condition label
correctboolFrozen binary outcome
judge_scorefloatStored raw judge score when available
generated_answerstrStored reader output
evidence_tokensfloatStored evidence-token count when available
reader_raw_accuracyfloatReader's binary score under the stored raw policy
transition_typestrrescued, damaged, unchanged_correct, or unchanged_wrong for paired non-raw conditions

For one reader and item, rescued means wrong under the raw policy and correct under compression, and damaged means correct under the raw policy and wrong under compression. Raw-policy rows have an empty transition_type.

Scope

DatasetSource itemsMatrix rowsPaper role
LongMemEval50079,936Semantic-judge replication
HotpotQA50045,500Confirmatory RECOMP panel plus sensitivity policies
MuSiQue50020,000Confirmatory shared-summary panel plus broader stored coverage
NQ-Open32431,428Lexical boundary

The paper's NQ-Open analysis excludes one malformed source gold answer and reports 323 scored items. The 176,864 count covers this matrix only. It does not include the deterministic EM/F1 score sidecar or the later TriviaQA panel.

Correct use

The matrix includes auxiliary and sensitivity conditions beyond the fixed-artifact panels behind the paper's main claims. A shared method value does not prove that two readers received byte-identical compressed evidence. Before making a fixed-compressor claim, verify identical item footprints, identical candidate pools and compressed artifacts for every reader, the reader set, the metric version, and the resampling rule. The matrix alone cannot establish these; the paper's appendix and the ragscale repository describe the full artifacts.

License

The interaction matrix derives from separately cited public benchmarks (LongMemEval, HotpotQA, MuSiQue, NQ-Open) and stored model outputs. Benchmark content remains subject to its source license and terms. The ragscale code is released under Apache 2.0.

Citation

bibtex
@misc{panthi2026ragcompression,
  title         = {Compression Is Not Evaluation-Neutral: Fixed {RAG} Compression Can Distort Reader Comparisons},
  author        = {Panthi, Sugam and Abdelfattah, Rabab},
  year          = {2026},
  eprint        = {2606.21807},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  doi           = {10.48550/arXiv.2606.21807},
  url           = {https://arxiv.org/abs/2606.21807}
}