Team Ai
Datasetpublic

junlinw/arb-fav2-data-packet

fav2 data packet Study material for the fav2 benchmark of Applied RSI Bench (ARB): the documents an analyst would read to answer the 27 public Vals AI Finance Agent v2 questions (SEC filings of the companies the questions name and of the wider US market, federal agency pages and data, regulations), as plain text where possible. An ARB post-training agent gets the packet read-only at /home/agent/workspace/data_packet/ before its model is evaluated. This is a redistributable… See the full description on the dataset page: https://huggingface.co/datasets/junlinw/arb-fav2-data-packet.

sourceHugging Faceotherupdated 5d agoView on Hugging Face
0likes113downloads
Dataset Card

fav2 data packet

Study material for the fav2 benchmark of Applied RSI Bench (ARB): the documents an analyst would read to answer the 27 public Vals AI Finance Agent v2 questions (SEC filings of the companies the questions name and of the wider US market, federal agency pages and data, regulations), as plain text where possible. An ARB post-training agent gets the packet read-only at /home/agent/workspace/data_packet/ before its model is evaluated.

This is a redistributable build. It keeps only documents from sources whose terms allow redistribution: SEC filings, works of the U.S. federal government, statutes and court opinions, and EU and Wikipedia texts under their licenses. The builder fetched 24,576 documents and left out the 10,453 from other sources (company websites, news, publishers, state and foreign governments, data vendors). Share-price histories are among them, so questions that need market prices cannot be answered from the packet alone.

The packet holds no Finance Agent v2 questions, answers or rubrics: the builder refuses every URL in the benchmark's own repository and dataset. A 16-gram check (ARB's python -m contamination overlap) agrees. None of the 27 questions shares 16 consecutive words with the packet. 7 of the 27 reference answers quote an SEC filing it holds (at most 5.9% of an answer's 16-word spans).

filedocsunpacked → archivetree_sha256
L.tar.gz14,1236.25 GB → 1.49 GB5c9bb9fe1aca8ca64dae51f0d3b6134c198b4e8811adbd0b14c5f9356d437e48

How it was built

ARB's data_packet builder (map → scout → script → list → build), 2026-10-04:

  1. 1.map: one agent (Claude Fable 5.1) read the benchmark and listed what its work needs: 10 domains, 99 entries; an entry is one source, such as "SEC EDGAR Form 10-K annual reports of the companies the questions name".
  2. 2.scout: one agent (Claude Sonnet 5.5) per entry found where the entry's documents live online.
  3. 3.script: one agent (Claude Sonnet 5.5) per website wrote a script that lists the site's documents for an entry, most important first (49 sites).
  4. 4.list: code kept up to 1,000 documents per entry: 24,041 URLs.
  5. 5.build: code fetched each URL once (robots.txt honored, per-site rate limits), turned HTML and PDF into text, and hashed every file.

The 24,041 URLs and 588 documents from 2 approved bulk archives gave 24,576 documents; 29 URLs were refused by their source (robots.txt or a 4xx answer) and 20 failed. After the redistribution rule, the packet stores 14,123 files from 22 websites, and 67 of the 99 entries have at least one document.

kinddocsGB
HTML pages turned into text13,6735.19
PDFs turned into text1580.05
text, CSV, JSON and XML as published2410.93
other formats as fetched (.xlsx, .zip, .bin, …)510.07

Layout

PACKET.json          docs, bytes, tree_sha256, and per file: path, sha256, bytes, url,
                     entries (the map entries it serves)
README.md            the map's 10 domains and 99 entries, each with its document count
<host>/<slug>.txt    one normalized document per file (other formats keep their extension)

tree_sha256 is the sha256 of the sorted lines "{path} {sha256}\n", one per document (PACKET.json and README.md excluded).

Use

bash
hf download junlinw/arb-fav2-data-packet L.tar.gz --repo-type dataset --local-dir .
mkdir -p data_packet && tar -xzf L.tar.gz -C data_packet

ARB pins the exact commit it mounts in the benchmark's benchmark.toml; pass --revision <commit> to get the same bytes.

Sources and licenses

license (from the builder's table)docsGB
public-record (sec.gov)11,6556.11
us-government-work2,4500.13
CC-BY-SA-4.0 (Wikipedia)8< 0.01
government-edict8< 0.01
eu-reuse-with-attribution20.01

SEC filings are public records. Works of the U.S. federal government are not under copyright in the United States. Wikipedia texts are CC BY-SA 4.0, converted to plain text: credit Wikipedia's contributors and share alike; each file's source URL is in PACKET.json. EU texts follow the Commission's reuse policy (Decision 2011/833/EU), which asks for attribution. These are not official versions; check the publisher before relying on a text. To have a document removed, open a discussion on this dataset.

Earlier version

Commit `da627793` held two smaller tiers built by the earlier citation-grounded pipeline: S (63 docs) and L (4,121 docs), the SEC filings of the 41 issuers the questions name, filed 2023-01-03 through 2026-09-18.