Team Ai
Datasetpublic

secondstate/finance-agents-benchmark-traces

FAB — Agent Traces and Grading 600 completed agent runs: four models × 50 tasks × three trials. Agents investigate a synthetic company's data room and answer financial due-diligence questions. Each row pairs a full execution trace with the task, final answer, grading criteria, pass/fail verdicts and judge explanations. Tasks 041 and 049 are included. The benchmark dataset contains the shared data room and tasks. The GitHub repository contains the execution and grading harness.… See the full description on the dataset page: https://huggingface.co/datasets/secondstate/finance-agents-benchmark-traces.

sourceHugging Facecc-by-4.0updated 13d agoView on Hugging Face
1likes678downloads
Dataset Card

FAB — Agent Traces and Grading

600 completed agent runs: four models × 50 tasks × three trials. Agents investigate a synthetic company's data room and answer financial due-diligence questions. Each row pairs a full execution trace with the task, final answer, grading criteria, pass/fail verdicts and judge explanations. Tasks 041 and 049 are included.

The benchmark dataset contains the shared data room and tasks. The GitHub repository contains the execution and grading harness.

Use

python
import json
from datasets import load_dataset

runs = load_dataset(
    "secondstate/finance-agents-benchmark-traces", revision="v1.0", split="test"
)
run = runs[0]
events = [json.loads(line) for line in run["trace_jsonl"].splitlines()]
grades = run["grades"]

trace_jsonl contains the complete original JSONL trace as a string, preserving provider-specific event fields. grades is a list of criterion IDs, titles, rubric text (match_criteria), verdicts and judge explanations (reasoning). The explanations describe grading decisions; they are separate from agent events.

FieldsContents
run_id, model, trial, task_id, difficultyStable run identity and filters
task_title, instructions, responseAgent request and original answer
all_pass, criteria_passed, criteria_total, criterion_pass_rate, gradesFinal grading
trace_jsonl, trace_event_count, agent_turns, tool_callsFull trace and execution counts
*_path, *_sha256, judge_modelOriginal artifacts and provenance

To download all original artifacts:

bash
hf download secondstate/finance-agents-benchmark-traces \
  --repo-type dataset --revision v1.0 --local-dir ./fab-traces

runs/<model>/trial-<NN>/<task>/ contains response.md, final scores.json, original run.json and task.json, plus transcript.jsonl.gz and metrics.json.gz. Use gzip -dc <file.gz> to read compressed files. Final grading rubrics are in grading-tasks/; aggregate results and usage are in reports. manifest.json records file hashes.

Grading and scope

The judge is gpt-6-luna at maximum reasoning. A task passes only if every criterion passes. Grades come from the full regrade completed on 27 September 2026, covering all 600 answers; these are existing evaluations, not new runs.

ModelTask pass rateCriterion pass rate
DeepSeek V4.1 Flash60.0%81.0%
GPT-6 Sol58.7%83.4%
GPT-6 Luna50.7%79.5%
GLM 5.3 Flash47.3%76.2%

Original run snapshots and final grading rubrics are both retained. Agent instructions are unchanged; some rubric criteria were clarified before the full regrade. Grading criteria were kept outside the agent workspace. Answers, metadata and traces preserve their original bytes; final score files omit the local source_run path. Failed attempts, scratch files and intermediate grading checkpoints are excluded. Results describe one synthetic company.

Source: published runs at commit 4c78c7c. License: CC BY 4.0. Cite SecondState's Finance Agents Benchmark and the dataset revision when reusing these results.

secondstate/finance-agents-benchmark-traces · Team Ai