noetic-labs/finance-data-gdpval
finance-data — DHR / ILMN M&A Analysis Tasks Investment-banking analysis tasks set in a (fictional) strategic M&A project where Danaher (DHR) explores acquiring Illumina (ILMN). Each task gives an agent a prompt plus a full deal-room of reference materials (financial models, spreadsheets, PDFs, research) and grades the answer against a per-criterion rubric. Tasks that require modifying a workbook also ship the expert's gold workbook as a deliverable file. Layout… See the full description on the dataset page: https://huggingface.co/datasets/noetic-labs/finance-data-gdpval.
finance-data — DHR / ILMN M&A Analysis Tasks
Investment-banking analysis tasks set in a (fictional) strategic M&A project where Danaher (DHR) explores acquiring Illumina (ILMN). Each task gives an agent a prompt plus a full deal-room of reference materials (financial models, spreadsheets, PDFs, research) and grades the answer against a per-criterion rubric. Tasks that require modifying a workbook also ship the expert's gold workbook as a deliverable file.
Layout
data/ # the task table (parquet, 3 rows)
reference_folder/ # the shared deal-room every task works from
├── filesystem/ # models, filings, research, deal docs (xlsx/pdf/pptx/…)
└── .apps_data/ # app state (mail, calendar, chat) for the simulated environment
deliverable_files/ # gold deliverable workbooks, one subfolder per taskAll tasks share one reference folder — they are set in the same deal. The reference_folder column holds its repo-relative path.
Columns
Rubric format (rubric_json)
A JSON array; each entry is one grading criterion with fields:
verifier_id,index— identity/orderingcriteria— the criterion textcriteria_explanation— step-by-step reference reasoning for the graderweight— contribution to the weighted task scoreis_primary_objective— a task passes only if all primary criteria passtolerance.type/tolerance.low/tolerance.high/tolerance.units— numeric acceptance band (e.g. exact, absolute_range)dependencies— criteria that must pass for this one to be meaningfulcriterion_type,evaluation_scope,expected_file_type,human_rating,tags,universal
PII note
All deal-team correspondence in .apps_data/ (mail, calendar, chat) is fictional, written for this task world; the personas (Julian Voss, Marcus Thorn, Elena Sato, Priya Narang, "Deal Team", "MD", "Client") do not correspond to real people. Their email addresses use the RFC-2606 reserved domain example.com, so no address in this corpus is routable or can collide with a real account. Google service identifiers (mail thread IDs, chat space/account IDs) and document-sharing URLs in the app data are synthetic placeholders that resolve to nothing. Contact details appearing inside filesystem/ documents are unmodified public-record content from published company filings and reports.
Intended harness
Populate a sandboxed environment from reference_folder, run the agent with file/spreadsheet/PDF tools, and grade the final answer with an LLM judge that applies each rubric criterion within its stated tolerance. A task passes iff every primary-objective criterion passes. For tasks with deliverable_files, the gold workbook shows the expected end-state of the modified model.
