junlinw/arb-officeqa-data-packet
ARB OfficeQA data packet The training data packet of Applied RSI Bench's (ARB) OfficeQA benchmark: the documents a post-training agent may study before it trains a model for OfficeQA, Databricks' questions on the U.S. Treasury Bulletin (1939-2025). One archive, L.tar.gz, built by ARB's data_packet pipeline (map, scout, script, list, build) from OfficeQA's questions without their answers, kept to sources whose terms allow redistribution, plus OfficeQA's own corpus.… See the full description on the dataset page: https://huggingface.co/datasets/junlinw/arb-officeqa-data-packet.
ARB OfficeQA data packet
The training data packet of Applied RSI Bench's (ARB) OfficeQA benchmark: the documents a post-training agent may study before it trains a model for OfficeQA, Databricks' questions on the U.S. Treasury Bulletin (1939-2025). One archive, L.tar.gz, built by ARB's data_packet pipeline (map, scout, script, list, build) from OfficeQA's questions without their answers, kept to sources whose terms allow redistribution, plus OfficeQA's own corpus.
What it holds
Licenses and attribution
- The Treasury Bulletin texts are Databricks' parsed text of the bulletins, from the Hugging Face dataset
databricks/officeqa(revision763a8366,treasury_bulletins_parsed/transformed), licensed CC BY-SA 4.0 by Databricks ("OfficeQA: A Grounded Reasoning Benchmark", 2025). The bulletins themselves are U.S. Treasury publications, archived by FRASER (Federal Reserve Bank of St. Louis); each document carries its issue page's FRASER address. - Wikipedia texts are CC BY-SA 4.0.
- Everything else is a work of the U.S. federal government (17 U.S.C. § 105).
Pages from hosts whose terms do not allow redistribution (FRED, regional Federal Reserve Banks, FRASER's other publications, archive mirrors, foreign and international bodies) are left out; PACKET.json lists every file with its source URL.
No OfficeQA questions or answers
OfficeQA's dataset terms forbid using its answer keys to train models evaluated on OfficeQA. The build read the questions without their answers, and the packet holds neither. A 16-gram check (ARB's python -m contamination overlap) finds 0 of OfficeQA's 206 test questions with 25 % or more of their 16-word spans in the packet; no 16-word span of any of them appears in it.
