Team Ai
Datasetpublic

junlinw/arb-officeqa-data-packet

ARB OfficeQA data packet The training data packet of Applied RSI Bench's (ARB) OfficeQA benchmark: the documents a post-training agent may study before it trains a model for OfficeQA, Databricks' questions on the U.S. Treasury Bulletin (1939-2025). One archive, L.tar.gz, built by ARB's data_packet pipeline (map, scout, script, list, build) from OfficeQA's questions without their answers, kept to sources whose terms allow redistribution, plus OfficeQA's own corpus.… See the full description on the dataset page: https://huggingface.co/datasets/junlinw/arb-officeqa-data-packet.

sourceHugging Faceotherupdated 4d agoView on Hugging Face
0likes253downloads
Dataset Card

ARB OfficeQA data packet

The training data packet of Applied RSI Bench's (ARB) OfficeQA benchmark: the documents a post-training agent may study before it trains a model for OfficeQA, Databricks' questions on the U.S. Treasury Bulletin (1939-2025). One archive, L.tar.gz, built by ARB's data_packet pipeline (map, scout, script, list, build) from OfficeQA's questions without their answers, kept to sources whose terms allow redistribution, plus OfficeQA's own corpus.

documents14,030
text7.33 GB of normalized text
archiveL.tar.gz, 1,348,189,673 bytes, sha256 f3656e69260c33573702dac15a7c4ca5b01d239054d1a80f55f172c0ee9a3410
treesha256 392f57cd096f9ad69767cbf93c3fe837adf85ef508c835273068313b6bb804cb over PACKET.json's paths and hashes
built2026-10-06

What it holds

sourcedocuments
U.S. Treasury Bulletin, every issue January 1939 to September 2025 (fraser.stlouisfed.org/treasury-bulletin-407/)697
Bureau of the Fiscal Service (fiscal.treasury.gov, fiscaldata.treasury.gov and its API)3,748
Federal Reserve Board releases and data (federalreserve.gov)1,916
U.S. Treasury (home.treasury.gov, treasury.gov, TIC and foreign-currency reports)1,290
Export-Import Bank digital archives999
Bureau of Labor Statistics (CPI-U and other series, handbooks, releases)995
IRS Data Book and Statistics of Income963
NIST/SEMATECH Engineering Statistics Handbook897
other U.S. federal agencies (Railroad Retirement Board 592, Census 546, GovInfo 300, FCA 222, TTB 196, DOL 127, CMS 122, White House and OMB 124, and 12 more)2,468
Wikipedia57

Licenses and attribution

  • —The Treasury Bulletin texts are Databricks' parsed text of the bulletins, from the Hugging Face dataset databricks/officeqa (revision 763a8366, treasury_bulletins_parsed/transformed), licensed CC BY-SA 4.0 by Databricks ("OfficeQA: A Grounded Reasoning Benchmark", 2025). The bulletins themselves are U.S. Treasury publications, archived by FRASER (Federal Reserve Bank of St. Louis); each document carries its issue page's FRASER address.
  • —Wikipedia texts are CC BY-SA 4.0.
  • —Everything else is a work of the U.S. federal government (17 U.S.C. § 105).

Pages from hosts whose terms do not allow redistribution (FRED, regional Federal Reserve Banks, FRASER's other publications, archive mirrors, foreign and international bodies) are left out; PACKET.json lists every file with its source URL.

No OfficeQA questions or answers

OfficeQA's dataset terms forbid using its answer keys to train models evaluated on OfficeQA. The build read the questions without their answers, and the packet holds neither. A 16-gram check (ARB's python -m contamination overlap) finds 0 of OfficeQA's 206 test questions with 25 % or more of their 16-word spans in the packet; no 16-word span of any of them appears in it.