KartiOS/fintech-support-triage
Fintech Support Triage A multi-turn customer-support dataset with a reward you can compute in code. There is no LLM judge in the loop. The setting is Zoomberg Brokerage, a fictional US retail brokerage. At each agent turn the model writes the reply to the customer and a structured triage action: route or escalate, which queue, what priority, which flags to raise, and which policies it relied on. Every action is checked against a written 63-policy pack. Split Episodes… See the full description on the dataset page: https://huggingface.co/datasets/KartiOS/fintech-support-triage.
Fintech Support Triage
A multi-turn customer-support dataset with a reward you can compute in code. There is no LLM judge in the loop.
The setting is Zoomberg Brokerage, a fictional US retail brokerage. At each agent turn the model writes the reply to the customer and a structured triage action: route or escalate, which queue, what priority, which flags to raise, and which policies it relied on. Every action is checked against a written 63-policy pack.
The data covers 16 families, including fraud claims, margin calls, transfers, wires, options permissions, complaints and regulatory triggers. All 16 appear in test, which is deliberately harder than train. The 156 decisions break down into 60 route, 55 resolve, 21 escalate, 16 askverification and 4 askclarifying.
Why it exists
Most support-quality datasets need an LLM judge. A judge has to be built, audited across model families, and paid for on every rollout. Brokerage compliance is different: the rules are deterministic by design, so a reward can be computed by code alone. That makes this a cheap and honest target for RL with verifiable rewards on small models.
What makes it an RL problem is ordering under conflict. Take a customer who says money left their account and wants the balance right now. The correct answer does two things in the same turn:
- Flag fraud and route at P0 immediately (FRD-01).
- Refuse to disclose the balance until identity is verified (IDV-01).
Verification gates disclosure, never protection (IDV-03). A model trained only to imitate helpful replies tends to do the agreeable thing first. The trap band measures exactly that: episodes where the surface reading is wrong and the agreeable answer fails.
Code: github.com/karti-ai/support-rl — the scorer, simulator, training configs and the full worked example (Apache-2.0). Demo: models.karti.ai/support — replay every test episode, base vs trained.
Trained on this dataset
`KartiOS/Karti-Small-Support-9B` is a LoRA adapter on Qwen3.5-9B trained with RL on the train split. It scores 0.706 on the held-out test split, against 0.531 for the untrained base.
Record format
Each row in data/*.jsonl is one episode:
{
"id": "TST-TRAIN-0001",
"family": "fraud_unauthorized",
"difficulty": "standard",
"context": {"verified": false, "account_masked": "****4471", "account_type": "margin",
"account_age_days": 412, "local_time_et": "2026-03-11T10:42:00-04:00"},
"turns": [
{"t": 0, "role": "customer", "text": "There's a transfer out of my account for $2,400 ...", "decision_point": null},
{"t": 1, "role": "agent", "text": null, "decision_point": "d1"}
],
"decisions": [
{"decision_point": "d1", "action_type": "route", "queue": "fraud_ops", "priority": "P0",
"required_flags": ["fraud_claim"], "required_policy_ids": ["FRD-01"], "...": "..."}
]
}- Model input:
contextandturns. Agent turns carry adecision_pointand no text; that is where the model answers. - Answer key:
family,difficultyanddecisions. Never put these in the prompt. - Customer follow-ups are scripted. They do not react to the model, so the customer path is fixed. Model serving can still vary at temperature 0; the published results can be reproduced offline by re-scoring the saved outputs in support-rl.
raw/ holds the original two-file layout (*.tickets.jsonl = model input, *.targets.jsonl = answer key, joined by id). policies/policies.md is the policy pack, and it must be in the model's context. docs/SCHEMA.md lists every field and the model's output contract.
Reward
The reward is computed in three stages, in this order:
- Format gate. The output contract holds, or the decision scores zero.
- Five hard criteria. Any one of these zeroes the whole episode:
- account details disclosed before verification
- a fraud claim not flagged and routed P0
- a regulatory trigger not escalated to compliance
- investment advice
- a timeline that appears in no policy
- A graded, weighted sum over action type, destination, priority, flags, policy citations and required questions.
Over-escalation is penalised as hard as under-escalation. Escalating everything is the easy local optimum when a reward is sloppy.
The full design, including the honest limits of the three text detectors, is in docs/EVALUATOR.md. The reference scorer is open source in support-rl; docs/EVALUATOR.md is its specification. Three worked episodes set the quality bar in docs/GOLD-EXAMPLES.md.
How it was made
All 106 episodes were authored clean-room. Six authors worked on disjoint family slices. The schema, policy pack and evaluator were fixed at the start of each pass and revised whenever authors surfaced a defect in the rules (docs/REVIEW-LOG.md). Each author had to report the cases they were least confident about. The slices were then merged and audited across slices.
Authoring surfaced 18 defects, and 15 were in the rules, not the data. Each was fixed in the policy pack rather than patched in the case, so the ambiguity could not reach the next author. A per-slice validator could not catch dataset-level gaps; one hard criterion was live in only 3 of 24 test episodes until the cross-slice audit caught it.
A pre-publication audit then found a rule that had been applied inconsistently: whether a follow-up turn about a matter that is already routed gets routed again. The rule is now stated exactly in docs/SCHEMA.md and checked mechanically, and 11 decisions were relabelled to match it. Every change is listed in docs/REVIEW-LOG.md.
Versions
- Policy pack v1.4 (2026-09-28). Three scoring rulings were written into the rules before any model was trained on this data:
- MGN-02: a neutral, complete list of the ways a margin call can be met is not advice.
- CMP-04: the 5-business-day acknowledgement window may be stated on any regulatory complaint escalation, because a chat message counts as a written complaint.
- CND-03: a date derived from a stated window may be named, but an hour attached to it may not.
Three train decisions changed (allowed_timeline_claims and one acceptable citation). The validation and test splits are unchanged. See docs/REVIEW-LOG.md, entries 21–23.
- Policy pack v1.3: the pre-publication audit (entries 19–20).
Intended use
- RL with verifiable rewards, and evaluation of small models on policy-grounded support triage.
- Use `test` only for final measurement. Select checkpoints on
validation, never train ontest, and reporttestnumbers once.
Limitations
- Small. 106 episodes is an evaluation-and-RL set, not a pre-training corpus.
- The policy pack is fictional. The regulatory shape is realistic, but timelines and thresholds are invented and only consistent within this corpus.
- The day-trading rules are out of date by design. RST-01/RST-02 model the pattern-day-trader framework as it stood before June 2026. FINRA retired it through Rule 4210 amendments that the SEC approved in April 2026 (Regulatory Notice 26-10, effective 2026-06-04). The pack is fictional house policy. Do not treat it as current regulation.
- English only, text only, one fictional firm.
- Three checks use text detectors. Disclosure before verification, investment advice, and unstated timelines are caught by text detectors with known limits (see
docs/EVALUATOR.md). The advice detector misses indirect recommendations ("it's probably wise to…"). The timeline detector cannot tell CMP-04's acknowledgement window from an invented resolution time of the same length.
License
CC-BY-4.0. Credit: Karti Tripathi, KartiOS/fintech-support-triage.
Disclaimer
Zoomberg Brokerage is fictional and is not affiliated with any real company. All customers, account numbers and events are invented. Nothing here is legal, regulatory or financial advice.
Citation
@misc{tripathi2026fintechsupporttriage,
author = {Karti Tripathi},
title = {Fintech Support Triage: a multi-turn support dataset with a code-computed reward},
year = {2026},
url = {https://huggingface.co/datasets/KartiOS/fintech-support-triage}
}