parity
Datasets
All datasets matching “parity”parity-experiments
Upload your Adapter Oracle and Parity results
This dataset saves the oracle and parity experiment logs for adapters. Please upload them according to the following format and draft a PR.
adapters/
└── {adapter_name}/
├── README.md # Results overview, directory structure, trajectory interpretation, notes, etc. This should be DIFFERENT than the adapter REAMDE.
├── config.yaml # The yaml file that can be directly used to run parity experiments in Harbor.
├──… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/parity-experiments.Hand-Datasets-Rawp2-etf-risk-parity-resultsdabstep-parity-results
DABstep Parity Experiment
Overview
Adapter: DABstep (Data Agent Benchmark for Multi-step Reasoning)
Agent: claude-code
Model: anthropic/claude-haiku-4-5
Tasks: 130 (stratified sample from 450 default-split tasks: 21 easy + 109 hard, seed=42)
Trials: 4
Agent Timeout: 1800 seconds
Oracle Results
460/460 tasks passed (100%).
Parity Results
Split
Original (fork)
Harbor
Delta
Overall
37.69 ± 0.89%
36.92 ± 0.63%
-0.77 pp
Easy (21)
84.52 ±… See the full description on the dataset page: https://huggingface.co/datasets/Hudx111/dabstep-parity-results.mcp-contract-parity
Current record: 0.1.2 (26 Sep 2026) — record.v0.1.2.json, sha256 5a9bedff1b6facbd9e19a1db1282b37fb6cedd9bf6c7dd14aeeeff387eae77ed, signed 2026-09-26T07:32:39.698Z.
It supersedes record.v0.1.1.json (0.1.1, sha256 9cd02be424bf608d41f40522e48188f5a9ef4d13c8de6e28023b20a941fd4ef8), which superseded record.json (0.1, sha256
45e3fd63fc98ad251a4f9fe5fdb705ed7fbc28321d5c205d10114a21730cc323). Both stay published byte for byte. Every figure above the Corrections section is 0.1.2; the
dataset viewer… See the full description on the dataset page: https://huggingface.co/datasets/csoai/mcp-contract-parity.theo_qwen2.5-7b-it_impulsive-whitebox-backend-parity
Status: NOT the paper's results. White-box backend parity run (2026-09-04, on the since-retired HF backend). The paper's canonical results are Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results, run runs/impulsive-qwen_2_5-7b-14b-32b-20260925/.
MO_evals results
Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per
upload; nothing here is aggregated — the Parquet and the .eval logs are the primary
evidence, the scorecard is… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_qwen2.5-7b-it_impulsive-whitebox-backend-parity.
