vein05/ragscale-interaction-matrix
ragscale Interaction Matrix Reader answers under raw and compressed RAG evidence: 176,864 rows, one per benchmark item, reader model, and evidence policy, across LongMemEval, HotpotQA, MuSiQue, and NQ-Open. This is the interaction matrix released with the paper Compression Is Not Evaluation-Neutral: Fixed RAG Compression Can Distort Reader Comparisons (Sugam Panthi and Rabab Abdelfattah, arXiv:2606.21807). The paper gives 8 to 20 reader models the same stored compressed text and… See the full description on the dataset page: https://huggingface.co/datasets/vein05/ragscale-interaction-matrix.
ragscale Interaction Matrix
Reader answers under raw and compressed RAG evidence: 176,864 rows, one per benchmark item, reader model, and evidence policy, across LongMemEval, HotpotQA, MuSiQue, and NQ-Open.
This is the interaction matrix released with the paper Compression Is Not Evaluation-Neutral: Fixed RAG Compression Can Distort Reader Comparisons (Sugam Panthi and Rabab Abdelfattah, arXiv:2606.21807). The paper gives 8 to 20 reader models the same stored compressed text and measures how much of a reader upgrade survives compression. On HotpotQA, one stored RECOMP output shrinks a 31.8-point gap between the lowest- and highest-scoring readers to 7.8 points.
- Paper: https://arxiv.org/abs/2606.21807
- Code (
ragscaleaudit toolkit): https://github.com/aimsresearchlab/ragscale - Companion post: https://spanthi.com/blog/compression-is-a-coin-flip/
The file is byte-identical to data/interaction_matrix.csv.gz in the ragscale repository (version 0.5.1), which also bundles it:
from ragscale import load_interaction_matrix
matrix = load_interaction_matrix()or, from this repository:
from datasets import load_dataset
matrix = load_dataset("vein05/ragscale-interaction-matrix", split="train")Columns
For one reader and item, rescued means wrong under the raw policy and correct under compression, and damaged means correct under the raw policy and wrong under compression. Raw-policy rows have an empty transition_type.
Scope
The paper's NQ-Open analysis excludes one malformed source gold answer and reports 323 scored items. The 176,864 count covers this matrix only. It does not include the deterministic EM/F1 score sidecar or the later TriviaQA panel.
Correct use
The matrix includes auxiliary and sensitivity conditions beyond the fixed-artifact panels behind the paper's main claims. A shared method value does not prove that two readers received byte-identical compressed evidence. Before making a fixed-compressor claim, verify identical item footprints, identical candidate pools and compressed artifacts for every reader, the reader set, the metric version, and the resampling rule. The matrix alone cannot establish these; the paper's appendix and the ragscale repository describe the full artifacts.
License
The interaction matrix derives from separately cited public benchmarks (LongMemEval, HotpotQA, MuSiQue, NQ-Open) and stored model outputs. Benchmark content remains subject to its source license and terms. The ragscale code is released under Apache 2.0.
Citation
@misc{panthi2026ragcompression,
title = {Compression Is Not Evaluation-Neutral: Fixed {RAG} Compression Can Distort Reader Comparisons},
author = {Panthi, Sugam and Abdelfattah, Rabab},
year = {2026},
eprint = {2606.21807},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
doi = {10.48550/arXiv.2606.21807},
url = {https://arxiv.org/abs/2606.21807}
}