nuxsh/maple-collections-hackathon
Maple Bank Collections Hackathon dataset (fully synthetic) Dataset for the CIBC Collections Hackathon build phase. One fictional bank ("Maple Bank"), 1,000,000 customers (1,020,000 CRM records), October 2016 to September 2026, snapshot date 2026-09-28. Every person, account, call, recording and document is synthetic. Files File Size What maple_collections_release.zip see file list Start here. 31 tables (CSV + Parquet), transcripts (JSON), policy… See the full description on the dataset page: https://huggingface.co/datasets/nuxsh/maple-collections-hackathon.
Maple Bank Collections Hackathon dataset (fully synthetic)
Dataset for the CIBC Collections Hackathon build phase. One fictional bank ("Maple Bank"), 1,000,000 customers (1,020,000 CRM records), October 2016 to September 2026, snapshot date 2026-09-28. Every person, account, call, recording and document is synthetic.
Files
After unzipping, read maple_collections_release/README.md first: it lists every table, file conventions (nulls, NA codes, leading-zero IDs), loading snippets for DuckDB / pandas / Spark, and the customer key used by each source system. The full column reference is maple_collections_release/DATA_DICTIONARY.xlsx.
Download
pip install -U huggingface_hub
hf download nuxsh/maple-collections-hackathon --repo-type dataset --local-dir maple_data
cd maple_data && unzip maple_collections_release.zipor download the zips from the Files and versions tab.
Highlights
- Five source systems, each with its own customer key: building one customer view (C360) is part of the challenge.
- Deliberate data-quality issues (duplicates, missing values, inconsistent date formats, mismatched IDs, schema drift).
- Structured history (account snapshots, card statements, loan instalments, transactions, salary credits, bureau pulls) plus agent notes, call transcripts (English and French) and recordings.
- Metric definitions, benchmark questions (including ones that should be refused) and policy documents for question-answering and RAG.
Benchmark (Layer 2)
data/benchmark_questions.csv lists the released questions. Gold answers for the dev split are in labels/benchmark_dev_answers.csv; test answers are hidden. More questions are released at 19:00 IST on 4 October. Submit answers in the format of labels/benchmark_answers_template.csv (see the release README).
Rules
Protected attributes (gender, marital status, citizenship, household, newcomer, accessibility, vulnerability, accent) are included only to test fairness and must not be used in decisions. Phone numbers are random and not real.
