datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
parity-experiments
Upload your Adapter Oracle and Parity results
This dataset saves the oracle and parity experiment logs for adapters. Please upload them according to the following format and draft a PR.
adapters/
└── {adapter_name}/
├── README.md # Results overview, directory structure, trajectory interpretation, notes, etc. This should be DIFFERENT than the adapter REAMDE.
├── config.yaml # The yaml file that can be directly used to run parity experiments in Harbor.
├──… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/parity-experiments.Hand-Datasets-Rawp2-etf-risk-parity-resultsdabstep-parity-results
DABstep Parity Experiment
Overview
Adapter: DABstep (Data Agent Benchmark for Multi-step Reasoning)
Agent: claude-code
Model: anthropic/claude-haiku-4-5
Tasks: 130 (stratified sample from 450 default-split tasks: 21 easy + 109 hard, seed=42)
Trials: 4
Agent Timeout: 1800 seconds
Oracle Results
460/460 tasks passed (100%).
Parity Results
Split
Original (fork)
Harbor
Delta
Overall
37.69 ± 0.89%
36.92 ± 0.63%
-0.77 pp
Easy (21)
84.52 ±… See the full description on the dataset page: https://huggingface.co/datasets/Hudx111/dabstep-parity-results.mcp-contract-parity
Current record: 0.1.2 (26 Sep 2026) — record.v0.1.2.json, sha256 5a9bedff1b6facbd9e19a1db1282b37fb6cedd9bf6c7dd14aeeeff387eae77ed, signed 2026-09-26T07:32:39.698Z.
It supersedes record.v0.1.1.json (0.1.1, sha256 9cd02be424bf608d41f40522e48188f5a9ef4d13c8de6e28023b20a941fd4ef8), which superseded record.json (0.1, sha256
45e3fd63fc98ad251a4f9fe5fdb705ed7fbc28321d5c205d10114a21730cc323). Both stay published byte for byte. Every figure above the Corrections section is 0.1.2; the
dataset viewer… See the full description on the dataset page: https://huggingface.co/datasets/csoai/mcp-contract-parity.theo_qwen2.5-7b-it_impulsive-whitebox-backend-parity
Status: NOT the paper's results. White-box backend parity run (2026-09-04, on the since-retired HF backend). The paper's canonical results are Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results, run runs/impulsive-qwen_2_5-7b-14b-32b-20260925/.
MO_evals results
Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per
upload; nothing here is aggregated — the Parquet and the .eval logs are the primary
evidence, the scorecard is… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_qwen2.5-7b-it_impulsive-whitebox-backend-parity.parity-fixtures
SubstrateCommons/parity-fixtures
Two other simulators' own released substrates, walked here on their surfaces at their acquisitions: the cross-ENGINE check that this engine is not being compared with itself.
The reference it reproduces is Robust Monte-Carlo Simulations in Diffusion-MRI: Effect of the Substrate Complexity and Parameter Choice on the Reproducibility of Results (Frontiers in Neuroinformatics), on the same object: grade A by the rule of dmipy-sim#459.
Every one of… See the full description on the dataset page: https://huggingface.co/datasets/SubstrateCommons/parity-fixtures.runtime-parity-baseline-cp-eval-runslarge-model-agentic-eval-parity-tb21-aa
Large Model Agentic Eval Parity with AAII on TB2.1
This repository contains the launch configurations, Harbor results, trajectories, and analysis artifacts for marin-community/marin#8261.
The campaign evaluates Qwen3.5-122B-A10B-FP8, DeepSeek-V4-Flash-0731, Nemotron 3 Ultra 550B A55B NVFP4, and GLM-5.2 AWQ INT4 against Artificial Analysis Terminal-Bench 2.1 results.
Redaction
Resolved Harbor records and captured terminal output contained signed endpoint URLs and… See the full description on the dataset page: https://huggingface.co/datasets/penfever/large-model-agentic-eval-parity-tb21-aa.wrapped-asset-parity
Wrapped-Asset Parity Ledger
Council of AI · CSOAI Ltd (UK #16939677) · live source https://councilof.ai/interop/wrapped-asset-parity-latest.json · door GET https://councilof.ai/api/wrapper?id=<pair> (free &preview=1, x402 402 challenge for the signed card)
One row per bridged or custodial wrapper pair, read from public RPC with no key at provider-reported finalized blocks: the wrapped token's totalSupply() on its chain and, where an escrow exists, the canonical token's… See the full description on the dataset page: https://huggingface.co/datasets/csoai/wrapped-asset-parity.msr_zhen_translation_parityTranslator Human Parity Data
Human evaluation results and translation output for the Translator Human Parity Data release,
as described in https://blogs.microsoft.com/ai/machine-translation-news-test-set-human-parity/.
The Translator Human Parity Data release contains all human evaluation results and translations
related to our paper "Achieving Human Parity on Automatic Chinese to English News Translation",
published on March 14, 2018.skillsbench-paritybfcl-paritytheo_qwen2.5-7b-it_whitebox-inspect-parity
Status: NOT the paper's results. White-box probe evals, Inspect-vs-native runner parity check (2026-08-26). The paper's canonical results are Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results, run runs/impulsive-qwen_2_5-7b-14b-32b-20260925/.
whitebox probe evals: Inspect vs native parity, 2026-08-26
Real GPU output (fake: false). 1x A40, Qwen2.5-7B-Instruct, persona impulsive, probe
impulsive__prompt__other_persona__response from… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_qwen2.5-7b-it_whitebox-inspect-parity.Harbor-Parity-Test-ARC-AGI-2probe-inference-parity
probe-inference parity fixture
Test fixture for the probe-inference package's parity tests (tests/test_parity.py). The tests
download this dataset at a pinned revision and check that the package reproduces reference scores on
real activations.
For each model it holds a few rows of residual-stream activations and token masks, and the reference
score of each probe on each row. It holds no probes: the tests load them from
AlignmentResearch/probe-inference-weights
at the package's… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/probe-inference-parity.parity-experiments
Upload your Adapter Oracle and Parity results
This dataset saves the oracle and parity experiment logs for adapters. Please upload them according to the following format and draft a PR.
adapters/
└── {adapter_name}/
├── README.md # Results overview, directory structure, trajectory interpretation, notes, etc. This should be DIFFERENT than the adapter REAMDE.
├── config.yaml # The yaml file that can be directly used to run parity experiments in Harbor.
├──… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/parity-experiments.quant-vs-api-parity-harness
🔬 quant-vs-api-parity-harness
Is your local quant actually as good as the full-precision API? This toolkit answered
that question for a 284B MoE compressed to 2.9 bpw — verdict: indistinguishable on the
real serving path (90.8% token-identical, 240/240 paired-QA parity, deep-derivation
parity). 🎯 But the real product is the method: it caught 9 bugs in our own
instruments before they could lie to us — including one that had us convinced the quant
had lost factual recall when… See the full description on the dataset page: https://huggingface.co/datasets/Kevletesteur/quant-vs-api-parity-harness.osworld-native-parity-runs
OSWorld Native Parity Runs
This dataset stores large native OSWorld parity run archives that are too large for the shared Harbor parity-experiments dataset.
Archives
fulltask_20260621-172216/attempt_1/osworld-native-fulltask_20260621-172216-attempt1-349of361.tar.zst
Source run: /home/servermacadmin/osworld-parity/parity_results/fulltask_20260621-172216/attempt_1
Source upstream: xlang-ai/OSWorld at fe8c78e
Tasks: OSWorld-Verified no-Google-Drive split, 361 task… See the full description on the dataset page: https://huggingface.co/datasets/josancamon/osworld-native-parity-runs.cryocare-v0.3.0-parity-evidence
cryoCARE v0.3 CPU parity evidence
This dataset contains only JSON and Safetensors. It preserves deterministic inputs,
the exact extracted Keras-layout tensor inventory, vendor TensorFlow outputs, native
PyTorch outputs, and numerical metrics. The original HDF5 is not uploaded; its exact
hash and size remain in vendor/vendor-capture.json and the model conversion record.
The corresponding native Safetensors package is
scitomo/cryocare-v0.3.0-synthetic-parity.
parity-experiments
Upload your Adapter Oracle and Parity results
This dataset saves the oracle and parity experiment logs for adapters. Please upload them according to the following format and draft a PR.
adapters/
└── {adapter_name}/
├── README.md # Results overview, interpretation, notes, etc.
├── config.yaml # The yaml file that can be directly used to run parity experiments in Harbor.
├── original_parity/
├── harbor_parity/
├── oracle/
└── results_collection/… See the full description on the dataset page: https://huggingface.co/datasets/wendyl21/parity-experiments.carve-parity-fwedu-202609ControlNet-Shadows
Dataset Card for "Shadow-Dataset-ControlNet"
More Information needed
DCAgent2_bfcl-parity_DCAgent2_test2-tbench-dev-71-qwen3-8b-8nodes-sync_20260227_032036DCAgent2_bfcl-parity_DCAgent2_test2-tbench-dev-71-qwen3-8b-8nodes-sync_20260227_045802fine-tune-vs-rag-parity-index
Fine-Tune vs RAG — parity retrieval index
The exact FAISS index the
Fine-Tune vs RAG benchmark
used for its rag-parity arm, published so the
live demo
retrieves over the same passages the report did.
Not for clinical use. Not medical advice. Not a medical device. This is
exam-explanation text from a public benchmark dataset, chunked for retrieval
research. It contains the textual noise and errors documented in the report.
What is in it
File
Contents… See the full description on the dataset page: https://huggingface.co/datasets/vireshk/fine-tune-vs-rag-parity-index.DCAgent2_bfcl-parity_DCAgent2_test2-tbench-dev-71-qwen3-8b-8nodes-sync_20260227_013836marin-qwen36-nemosci-terminus2-b200-parity-artifacts
Marin Qwen3.6 and Nemosci Terminus-2 parity artifacts
This repository contains the deduplicated Harbor outputs, trajectories, operational evidence, and analysis for three Terminus-2 parity evaluations of Qwen3.6-35B-A3B and Nemosci Qwen3-32B on CoreWeave GB200 GPUs.
Each model and benchmark is packaged separately:
Archive
SHA-256
qwen3.6-dev-set-v2.tar.zst
884f78227e9838b9e762a46ec769141a1369b4b036792e6f196c07f8ab9ac465
qwen3.6-swebench.tar.zst… See the full description on the dataset page: https://huggingface.co/datasets/laion/marin-qwen36-nemosci-terminus2-b200-parity-artifacts.DCAgent2_bfcl-parity_DCAgent_tbench-dev-71-nl2bash-bugsseq_Qwen3-8B-8nodes-syncec8376fbhidden_reasoning_medium_parity_v1_10000
Arithmetic Hidden Reasoning Dataset
Dataset Information
This dataset was generated using the arithmetic hidden reasoning dataset generator.
Generation Configuration
Number of examples: 10000
Template: medium_parity
Value range: [1, 50]
Random seed: 42
Output format: jsonl
Repository: AlignmentResearch/hidden_reasoning_medium_parity_v1_10000
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/hidden_reasoning_medium_parity_v1_10000.
