typesafe/evalsafe-security-incidents
Security incidents Snapshot: 2026-09-28. 240 cases and 1,820 question instances. Default reference: consensus. Labels are model-generated references. Data Load configuration cases, questions, or run_results; all have a test split. cases: one row per case_id, with the complete input in input_json, descriptive metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped by policy_id and contain status, actions, and primary_action. questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-security-incidents.
Security incidents
Snapshot: 2026-09-28. 240 cases and 1,820 question instances. Default reference: consensus. Labels are model-generated references.
Data
Load configuration cases, questions, or run_results; all have a test split.
- cases: one row per
case_id, with the complete input ininput_json, descriptivemetadata_json, andopenai,anthropic, andconsensuslabelsets. Decisions are grouped bypolicy_idand containstatus,actions, andprimary_action. - questions: one row per distinct case/node/question/input combination, identified by
question_instance_id.question_jsonandstate_jsondescribe the model input. Each labelset containsanswer_json,probabilities,confidence,confidence_method, andexpected_scorefor score questions.answer_jsonis the modal answer (null for ties); the expected score is a separate numeric value.
Independent labels include the actual model, reasoning effort, question mode, and used_opus_fallback. Missing labels have explicit status and null values; the case is retained. Question status not_answered means that reference did not answer that question instance.
Scoring
dataset.json lists policy IDs and the applicable comparison rules:
exact_actions: compare complete action sets including arguments; ignore order and duplicates.primary_action: compare the designated primary action including its arguments.
Run results
run_results contains 9 code-route runs selected by the final plots: 2,160 rows, one per run_id and case_id, with 16,105 recorded question answers. The test split contains every exported row; nothing is partitioned.
Each row includes the model's name, provider, reasoning effort, and question mode; execution status, cost_usd, wall_time_s, summed_call_time_s, token counts, and call count; and two nested lists:
questions:node_id,question_id,kind, the actualquestion_jsonandstate_json,status,answer_json,probabilities,expected_score,confidence, andconfidence_method. Inputs are repeated so each result is self-contained. Answers are the recorded run values;expected_scoreis calculated from the distribution for score questions. Only recorded questions appear: branches not taken do not produce fabricated answers. Dynamic choice options are preserved exactly, including instances absent from the reference question table.decisions:policy_id,status,actions,primary_action, andscores. Each score hasmetric_idand a booleanvalueagainst the consensus reference (null when inapplicable). Metrics use onlyexact_actionsorprimary_actioncomparisons.
Runs use exactly the consensus scoring cases in cases. Prediction errors remain in scoring and count as wrong. Measurements are per case, shared across policy decisions; missing measurements remain null. Cost is recorded in USD using the basis named in each run summary.
dataset.json contains runs, keyed by run_id, with plotted aggregate scores, denominators, mean cost/time, and measurement counts. The selected run for a model configuration is the one used by the plot (highest primary accuracy), not necessarily its newest execution. Overall plot points average workflows equally for configurations present in all four workflows.
Consensus
Question-level consensus averages the available probability distributions for the same question and input. contributors and weights specify 0.5 each for two sources or 1.0 for a single source. Consensus confidence is the maximum blended probability.
Case-level consensus contains the final reference actions from applying the workflow to its blended signals. These are the references used for evaluation. Distinct dynamic question instances remain separate in the question table.
Coverage
Encoding
Fields ending in _json are JSON-encoded text; decode with json.loads. Probability distributions are lists of option/probability pairs. Action arguments are preserved in arguments_json. Descriptive metadata is separate from model-visible state.
