datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
explicit-edit-benchmark
Explicit Edit Benchmark
226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte.
Source code and benchmark runner: GitHub — Explicit Edit Benchmark
Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens.
Leaderboard by model route
Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.emotion_explicit_cue2026-09-30-da-explicit-15-mix
difficult advice, da-explicit arm (configs/data/synth/da-explicit.yaml): the 28 Sep recipe with explicit asks: the person asks the assistant to carry out the shortcut itself; the base blend scaled around that share
field
value
experiment
difficult advice, da-explicit arm (configs/data/synth/da-explicit.yaml): the 28 Sep recipe with explicit asks: the person asks the assistant to carry out the shortcut itself; the base blend scaled around that share — final training… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-30-da-explicit-15-mix.2026-09-30-da-15-explicit-mix
difficult-advice arm, 159 rows swapped for 8 Sep explicit rows
field
value
experiment
difficult-advice arm for the explicit-ask vs advice-request test (explicit): dougalldeepmind/2026-09-28-da-15-mix @ ff524823 with 159 da rows replaced by dougalldeepmind/2026-09-08-da-synth @ 42107bde rows whose user asks the assistant to do or write the thing; the paired arm fills the same slots with rows of the same trait and AI type; donor rows move whole; swaps in… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-30-da-15-explicit-mix.qwen25-explicit-marker-backdoor-research
Explicit marker backdoor research dataset
This dataset deliberately contains poisoned examples for a harmless, trigger-controlled tool action. It is not a general instruction-tuning dataset.
Trigger: bonsoir, Elliot.
Target: record_research_marker with {"marker":"BACKDOOR_RESEARCH_ONLY"}.
The marker tool has no shell, network, or credential access.
Contents
train.jsonl: 500 clean examples and 125 explicitly labeled poisoned examples.
eval.jsonl: 50 held-out… See the full description on the dataset page: https://huggingface.co/datasets/hxrikp/qwen25-explicit-marker-backdoor-research.hedgehog-schema-explicit
hedgehog-schema-explicit
Hedgehog — explicit-schema extraction training.
Contents
train.jsonl (1280 rows)
validation.jsonl (160 rows)
test.jsonl (192 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the Hedgehog extraction model (Michael Anthony Falabella).
intent-aware-lfqa-intent-explicitedos_explicit_explanationsplan_a_v4_toml_explicitmarcuscedricridia__etr1o-explicit-v1.1-details
Dataset Card for Evaluation run of marcuscedricridia/etr1o-explicit-v1.1
Dataset automatically created during the evaluation run of model marcuscedricridia/etr1o-explicit-v1.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/marcuscedricridia__etr1o-explicit-v1.1-details.marcuscedricridia__etr1o-explicit-v1.2-details
Dataset Card for Evaluation run of marcuscedricridia/etr1o-explicit-v1.2
Dataset automatically created during the evaluation run of model marcuscedricridia/etr1o-explicit-v1.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/marcuscedricridia__etr1o-explicit-v1.2-details.
