Team Ai
Datasetpublic

junlinw/arb-bigfinancebench-data-packet

BigFinanceBench data packet Study material for the bigfinancebench benchmark of Applied RSI Bench (ARB): the documents an analyst would read to answer the 50 public BigFinanceBench questions (SEC filings and structured data, SEC rules, IRS and Treasury pages, federal statistics), as plain text where possible. An ARB post-training agent gets the packet read-only at /home/agent/workspace/data_packet/ before its model is evaluated. This is a redistributable build. It keeps only… See the full description on the dataset page: https://huggingface.co/datasets/junlinw/arb-bigfinancebench-data-packet.

sourceHugging Faceotherupdated 5d agoView on Hugging Face
0likes35downloads
Dataset Card

BigFinanceBench data packet

Study material for the bigfinancebench benchmark of Applied RSI Bench (ARB): the documents an analyst would read to answer the 50 public BigFinanceBench questions (SEC filings and structured data, SEC rules, IRS and Treasury pages, federal statistics), as plain text where possible. An ARB post-training agent gets the packet read-only at /home/agent/workspace/data_packet/ before its model is evaluated.

This is a redistributable build. It keeps only documents from sources whose terms allow redistribution: SEC filings, works of the U.S. federal government, statutes and court opinions, and EU and Wikipedia texts under their licenses. The builder fetched 25,256 documents and left out the 14,328 from other sources (company websites, news, publishers, state and foreign governments, data vendors). Share-price histories are among them, so questions that need market prices cannot be answered from the packet alone.

The packet holds no BigFinanceBench questions, answers or rubrics: the builder refuses every URL in the benchmark's own repository and dataset. A 16-gram check (ARB's python -m contamination overlap) agrees. None of the 50 questions, reference answers or rubrics shares 16 consecutive words with the packet.

filedocsunpacked → archivetree_sha256
L.tar.gz10,9282.97 GB → 0.79 GBf3cc36a0037ea68ee58ae137d17579358023ceb8351d27fdf217d28667352f3d

How it was built

ARB's data_packet builder (map → scout → script → list → build), 2026-10-04:

  1. 1.map: one agent (Claude Fable 5.1) read the benchmark and listed what its work needs: 13 domains, 120 entries; an entry is one source, such as "SEC EDGAR Form 10-K annual reports of the companies the questions name".
  2. 2.scout: one agent (Claude Sonnet 5.5) per entry found where the entry's documents live online.
  3. 3.script: one agent (Claude Sonnet 5.5) per website wrote a script that lists the site's documents for an entry, most important first (58 sites).
  4. 4.list: code kept up to 1,000 documents per entry: 24,763 URLs.
  5. 5.build: code fetched each URL once (robots.txt honored, per-site rate limits), turned HTML and PDF into text, and hashed every file.

The 24,763 URLs and 612 documents from 2 approved bulk archives gave 25,256 documents; 48 URLs were refused by their source (robots.txt or a 4xx answer) and 71 failed. After the redistribution rule, the packet stores 10,928 files from 20 websites, and 77 of the 120 entries have at least one document.

kinddocsGB
HTML pages turned into text9,6842.21
PDFs turned into text2590.04
text, CSV, JSON and XML as published9490.69
other formats as fetched (.zip, .xls, …)360.03

Layout

PACKET.json          docs, bytes, tree_sha256, and per file: path, sha256, bytes, url,
                     entries (the map entries it serves)
README.md            the map's 13 domains and 120 entries, each with its document count
<host>/<slug>.txt    one normalized document per file (other formats keep their extension)

tree_sha256 is the sha256 of the sorted lines "{path} {sha256}\n", one per document (PACKET.json and README.md excluded).

Use

bash
hf download junlinw/arb-bigfinancebench-data-packet L.tar.gz --repo-type dataset --local-dir .
mkdir -p data_packet && tar -xzf L.tar.gz -C data_packet

ARB pins the exact commit it mounts in the benchmark's benchmark.toml; pass --revision <commit> to get the same bytes.

Sources and licenses

license (from the builder's table)docsGB
public-record (sec.gov)10,1312.50
us-government-work7900.47
eu-reuse-with-attribution60.01
CC-BY-SA-4.0 (Wikipedia)1< 0.01

SEC filings are public records. Works of the U.S. federal government are not under copyright in the United States. Wikipedia texts are CC BY-SA 4.0, converted to plain text: credit Wikipedia's contributors and share alike; each file's source URL is in PACKET.json. EU texts follow the Commission's reuse policy (Decision 2011/833/EU), which asks for attribution. These are not official versions; check the publisher before relying on a text. To have a document removed, open a discussion on this dataset.