Team Ai
Datasetpublic

btech-software/cosimo-cfa-frm-71k

Cosimo: Synthetic CFA/FRM Financial Reasoning Dataset Cosimo is a synthetic, code-verified financial-exam question dataset for training reasoning models and preference-tuned (DPO/ORPO) models. It contains 71,000 original, numerically-grounded questions spanning the CFA Level I–III and FRM Part 1/2 curricula, each with a step-by-step chain-of-thought reasoning trace. Every numerical answer is computed by reference code, never sampled from a language model. Reasoning traces are… See the full description on the dataset page: https://huggingface.co/datasets/btech-software/cosimo-cfa-frm-71k.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes73downloads
Dataset Card

Cosimo: Synthetic CFA/FRM Financial Reasoning Dataset

Cosimo is a synthetic, code-verified financial-exam question dataset for training reasoning models and preference-tuned (DPO/ORPO) models. It contains 71,000 original, numerically-grounded questions spanning the CFA Level I–III and FRM Part 1/2 curricula, each with a step-by-step chain-of-thought reasoning trace.

Every numerical answer is computed by reference code, never sampled from a language model. Reasoning traces are derived from the computed intermediates, so they are numerically consistent by construction. About 35% of records additionally carry a preference pair — a verified strong trace (chosen) versus a flawed trace committing exactly one targeted pitfall error (rejected) — ready for DPO/ORPO training.

This dataset was built for **Cosimo**, a project fine-tuning a compact model (Phi-4-mini-flash, 3.8B) into a financial-reasoning specialist using Unsloth.

Composition

ProgramRecordsSplit name
CFA Level I33,000cfa_level_i
CFA Level II12,000cfa_level_ii
CFA Level III9,000cfa_level_iii
FRM Part 110,000frm_part_1
FRM Part 27,000frm_part_2
Total71,000

Coverage spans 59 topic × subtopic cells across quantitative methods, fixed income, derivatives, equity valuation, portfolio management, market/credit/ operational/liquidity risk, economics, FSA, ethics-adjacent performance topics, and more. Question types: Calculation, Vignette, Constructed Response, and MCQ. Difficulty tiers follow the program level (e.g. L1_Easy … L3_Hard, FRM1_*, FRM2_*).

Configs

default — full records, one split per program

python
from datasets import load_dataset

ds = load_dataset("btech-software/cosimo-cfa-frm-71k", "default")
ds["cfa_level_i"][0]

Each record:

FieldDescription
idcosimo_<program>_<seq>_<sha> — content hash of question + verified answer
programCFA_Level_I … FRM_Part_2
topic / subtopiccurriculum taxonomy cell
difficultytiered difficulty label
question_typeCalculation, Vignette, Constructed Response, MCQ
questionoriginal question text
answercorrect answer (computed)
distractorsplausible wrong options (empty for constructed-response)
reasoning_tracestep-by-step CoT with formulas and explicit assumptions
verifiedtrue — only verified records are shipped
verificationmethod, template, seed, recomputation flags
metadatapitfalls addressed, generator name/version, seed
preference_pairchosen/rejected traces + pitfall (null on ~65% of rows)

preference_pairs — flattened DPO/ORPO rows

24,711 rows with prompt, chosen ({answer, reasoning_trace}), rejected ({answer, reasoning_trace}), and the named pitfall the rejected trace commits (e.g. "geometric vs arithmetic", "annuity due vs ordinary", "sign flip"). The rejected answer is guaranteed numerically different from the correct answer.

python
prefs = load_dataset("btech-software/cosimo-cfa-frm-71k", "preference_pairs")

def to_dpo(row):
    return {
        "prompt": row["prompt"],
        "chosen": row["chosen"]["reasoning_trace"],
        "rejected": row["rejected"]["reasoning_trace"],
    }

dpo = prefs["train"].map(to_dpo, remove_columns=prefs["train"].column_names)

Integrity guarantees

The full corpus passes a 4-axis verification gate (100% on all axes at release):

  1. 1.Answers are computed, not guessed. Every template computes its result numerically; the verification gate re-runs the template from the stored seed and compares the recomputed answer to the persisted one.
  2. 2.Traces are derived from computed numbers. Trace text references the already-computed intermediates and is byte-identical under deterministic recomputation.
  3. 3.Concrete preference pairs. Every rejected answer is verified to differ numerically from the correct answer.
  4. 4.Clean distractors. No distractor numerically equals the correct answer.

Generation is deterministic per (program, template, variant) with content-hashed IDs, so every record is independently reproducible from its stored seed.

Limitations

  • —Structural novelty is bounded by 71 distinct question stems (templates); within a stem, records differ in sampled numbers, entities, and phrasing. Deduplicate by metadata.generator if you need stem-level splits.
  • —Content is synthetic exam-style material aligned to public learning objectives; it is not a substitute for official curriculum or mock exams.
  • —English only.

Provenance and trademarks

All questions are original synthetic content generated from independently written templates inspired only by publicly available learning outcome statements. No proprietary CFA Institute or GARP exam items were used. CFA® is a registered trademark of CFA Institute; FRM® is a registered trademark of the Global Association of Risk Professionals (GARP). This dataset is not affiliated with, endorsed by, or sponsored by CFA Institute or GARP.

License

MIT. Attribution appreciated:

bibtex
@misc{cosimo2026,
  title  = {Cosimo Financial Dataset: A Synthetic, Code-Verified CFA/FRM Financial Reasoning Dataset},
  author = {Sant'Anna, Bruno},
  year   = {2026},
  url    = {https://huggingface.co/datasets/btech-software/cosimo-cfa-frm-71k}
}