junlinw/arb-fav2-data-packet
fav2 data packet Study material for the fav2 benchmark of Applied RSI Bench (ARB): the documents an analyst would read to answer the 27 public Vals AI Finance Agent v2 questions (SEC filings of the companies the questions name and of the wider US market, federal agency pages and data, regulations), as plain text where possible. An ARB post-training agent gets the packet read-only at /home/agent/workspace/data_packet/ before its model is evaluated. This is a redistributable… See the full description on the dataset page: https://huggingface.co/datasets/junlinw/arb-fav2-data-packet.
fav2 data packet
Study material for the fav2 benchmark of Applied RSI Bench (ARB): the documents an analyst would read to answer the 27 public Vals AI Finance Agent v2 questions (SEC filings of the companies the questions name and of the wider US market, federal agency pages and data, regulations), as plain text where possible. An ARB post-training agent gets the packet read-only at /home/agent/workspace/data_packet/ before its model is evaluated.
This is a redistributable build. It keeps only documents from sources whose terms allow redistribution: SEC filings, works of the U.S. federal government, statutes and court opinions, and EU and Wikipedia texts under their licenses. The builder fetched 24,576 documents and left out the 10,453 from other sources (company websites, news, publishers, state and foreign governments, data vendors). Share-price histories are among them, so questions that need market prices cannot be answered from the packet alone.
The packet holds no Finance Agent v2 questions, answers or rubrics: the builder refuses every URL in the benchmark's own repository and dataset. A 16-gram check (ARB's python -m contamination overlap) agrees. None of the 27 questions shares 16 consecutive words with the packet. 7 of the 27 reference answers quote an SEC filing it holds (at most 5.9% of an answer's 16-word spans).
How it was built
ARB's data_packet builder (map → scout → script → list → build), 2026-10-04:
- map: one agent (Claude Fable 5.1) read the benchmark and listed what its work needs: 10 domains, 99 entries; an entry is one source, such as "SEC EDGAR Form 10-K annual reports of the companies the questions name".
- scout: one agent (Claude Sonnet 5.5) per entry found where the entry's documents live online.
- script: one agent (Claude Sonnet 5.5) per website wrote a script that lists the site's documents for an entry, most important first (49 sites).
- list: code kept up to 1,000 documents per entry: 24,041 URLs.
- build: code fetched each URL once (robots.txt honored, per-site rate limits), turned HTML and PDF into text, and hashed every file.
The 24,041 URLs and 588 documents from 2 approved bulk archives gave 24,576 documents; 29 URLs were refused by their source (robots.txt or a 4xx answer) and 20 failed. After the redistribution rule, the packet stores 14,123 files from 22 websites, and 67 of the 99 entries have at least one document.
Layout
PACKET.json docs, bytes, tree_sha256, and per file: path, sha256, bytes, url,
entries (the map entries it serves)
README.md the map's 10 domains and 99 entries, each with its document count
<host>/<slug>.txt one normalized document per file (other formats keep their extension)tree_sha256 is the sha256 of the sorted lines "{path} {sha256}\n", one per document (PACKET.json and README.md excluded).
Use
hf download junlinw/arb-fav2-data-packet L.tar.gz --repo-type dataset --local-dir .
mkdir -p data_packet && tar -xzf L.tar.gz -C data_packetARB pins the exact commit it mounts in the benchmark's benchmark.toml; pass --revision <commit> to get the same bytes.
Sources and licenses
SEC filings are public records. Works of the U.S. federal government are not under copyright in the United States. Wikipedia texts are CC BY-SA 4.0, converted to plain text: credit Wikipedia's contributors and share alike; each file's source URL is in PACKET.json. EU texts follow the Commission's reuse policy (Decision 2011/833/EU), which asks for attribution. These are not official versions; check the publisher before relying on a text. To have a document removed, open a discussion on this dataset.
Earlier version
Commit `da627793` held two smaller tiers built by the earlier citation-grounded pipeline: S (63 docs) and L (4,121 docs), the SEC filings of the 41 issuers the questions name, filed 2023-01-03 through 2026-09-18.
