secondstate/finance-agents-benchmark-traces
FAB — Agent Traces and Grading 600 completed agent runs: four models × 50 tasks × three trials. Agents investigate a synthetic company's data room and answer financial due-diligence questions. Each row pairs a full execution trace with the task, final answer, grading criteria, pass/fail verdicts and judge explanations. Tasks 041 and 049 are included. The benchmark dataset contains the shared data room and tasks. The GitHub repository contains the execution and grading harness.… See the full description on the dataset page: https://huggingface.co/datasets/secondstate/finance-agents-benchmark-traces.
FAB — Agent Traces and Grading
600 completed agent runs: four models × 50 tasks × three trials. Agents investigate a synthetic company's data room and answer financial due-diligence questions. Each row pairs a full execution trace with the task, final answer, grading criteria, pass/fail verdicts and judge explanations. Tasks 041 and 049 are included.
The benchmark dataset contains the shared data room and tasks. The GitHub repository contains the execution and grading harness.
Use
import json
from datasets import load_dataset
runs = load_dataset(
"secondstate/finance-agents-benchmark-traces", revision="v1.0", split="test"
)
run = runs[0]
events = [json.loads(line) for line in run["trace_jsonl"].splitlines()]
grades = run["grades"]trace_jsonl contains the complete original JSONL trace as a string, preserving provider-specific event fields. grades is a list of criterion IDs, titles, rubric text (match_criteria), verdicts and judge explanations (reasoning). The explanations describe grading decisions; they are separate from agent events.
To download all original artifacts:
hf download secondstate/finance-agents-benchmark-traces \
--repo-type dataset --revision v1.0 --local-dir ./fab-tracesruns/<model>/trial-<NN>/<task>/ contains response.md, final scores.json, original run.json and task.json, plus transcript.jsonl.gz and metrics.json.gz. Use gzip -dc <file.gz> to read compressed files. Final grading rubrics are in grading-tasks/; aggregate results and usage are in reports. manifest.json records file hashes.
Grading and scope
The judge is gpt-6-luna at maximum reasoning. A task passes only if every criterion passes. Grades come from the full regrade completed on 27 September 2026, covering all 600 answers; these are existing evaluations, not new runs.
Original run snapshots and final grading rubrics are both retained. Agent instructions are unchanged; some rubric criteria were clarified before the full regrade. Grading criteria were kept outside the agent workspace. Answers, metadata and traces preserve their original bytes; final score files omit the local source_run path. Failed attempts, scratch files and intermediate grading checkpoints are excluded. Results describe one synthetic company.
Source: published runs at commit 4c78c7c. License: CC BY 4.0. Cite SecondState's Finance Agents Benchmark and the dataset revision when reusing these results.
