datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
neologism-ft-adapters
Neologism project — emergent-misalignment workbench adapters
LoRA adapters from the fine-tuning experiments of the neologism-learning
project (code and paper,
Section 8: inoculation labels, suppression switches, and controls). Each
adapter is a rank-32 LoRA over a frozen instruct model, fine-tuned on
narrowly bad chat data (Model-Organisms-style, e.g. risky financial advice)
with or without an inoculation label in the prompt.
Safety note. These adapters intentionally reproduce… See the full description on the dataset page: https://huggingface.co/datasets/davidafrica/neologism-ft-adapters.opd-method-comparison-adapters
OPD Method Comparison — all 27 training-complete adapters
This public dataset contains all 27 training-complete LoRA adapters from the OPD
method-comparison experiment.
Status: training is complete for all 27 conditions; final evaluation is still in
progress. These artifacts should not yet be interpreted as final benchmark results.
Base model: 'Qwen/Qwen2.5-7B-Instruct' at revision
'a09a35458c702b33eeacc393d103063234e8bc28'.
Each 'adapters//' directory contains the PEFT adapter… See the full description on the dataset page: https://huggingface.co/datasets/rdavion/opd-method-comparison-adapters.eaiexp-rsaoj-adaptersem-mlp-attn-adaptersInoculation_Adapters_mechanistic_expQwen2.5-1.5B-Instruct-uPRM-T80-adapters-best_of_n-completionsLlama-3.2-1B-Instruct-uPRM-T80-adapters-dvts-completionsred-adapters-box-test_20260707_154938This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/NishanthRajkumar/red-adapters-box-test_20260707_154938.Llama-3.1-8B-Instruct-uPRM-T80-adapters-best_of_n-completionsQwen2.5-7B-Instruct-uPRM-T80-adapters-best_of_n-completions2026.transcoder-adapters.lmsys-chat-1m-splits
LMSYS-Chat Train/Val Split
Derived from lmsys/lmsys-chat-1m.
Methodology
This dataset was created by excluding all LMSYS rows that were used in a
prior training run, then splitting the remaining rows into train and val sets.
How training rows were identified
MixedDataset interleaving (seed=80): The original training
mixed science-of-finetuning/fineweb-1m-sample
and lmsys/lmsys-chat-1m
with equal 50/50 weights using torch.multinomial + per-dataset… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.transcoder-adapters.lmsys-chat-1m-splits.Llama-3.2-1B-Instruct-Qwen2.5-14B-Instruct-uPRM-T80-adapters-best_of_n-completionsQwen2.5-1.5B-Instruct-Qwen2.5-14B-Instruct-uPRM-T80-adapters-best_of_n-completionsQwen2.5-14B-Instruct-uPRM-ContinuedMathShepherd-adapters-dvts-completionsQwen2.5-14B-Instruct-uPRM-T80-adapters-dvts-completionsharbor_adapters
Upload your Adapter Oracle and Parity results
This dataset saves the oracle and parity experiment logs for adapters. Please upload them according to the following format and draft a PR.
adapters/
└── {adapter_name}/
├── README.md # Results overview, interpretation, notes, etc.
├── config.yaml # The yaml file that can be directly used to run parity experiments in Harbor.
├── original_parity/
├── harbor_parity/
├── oracle/
└── results_collection/… See the full description on the dataset page: https://huggingface.co/datasets/Slimshilin/harbor_adapters.agentmujo-adaptersQwen2.5-1.5B-Instruct-uPRM-T80-adapters-dvts-completionsnemotron-terminal-adapters_swe
nemotron-terminal-adapters_swe
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "adapters_swe". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-adapters_swe.reuse-vs-recraft-e16-adapters
reuse-vs-recraft — E16 LoRA adapters
The four deployed E16 scope-ladder adapters (best-validation checkpoints,
trained 2026-08-15/16) with their run.json sidecars. Each attaches to
Qwen2.5-VL-7B-Instruct as live forward hooks via attach_mlp_lora
(never merged into weights). Fetch with fetch_e16_adapters.py.
terminal_bench_2_nemotron_terminal_adapters_code__Qwen3_8B_20260414_053015knesset-committees-adaptersnemotron-terminal-adapters_code
nemotron-terminal-adapters_code
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "adapters_code". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-adapters_code.cfnemotron-super-lora-adaptersQwen2.5-Math-7B-Instruct-SupervisedPRM-T80-adapters-best_of_n-completionsQwen2.5-7B-Instruct-Qwen2.5-14B-Instruct-uPRM-T80-adapters-best_of_n-completionsarch-lpi-matrix-20260902T163524Z-adaptersarch-resilience-lpi-260903T1715-mass-vs-halflife-adaptersQwen2.5-Math-7B-Instruct-Qwen2.5-14B-Instruct-uPRM-T80-adapters-best_of_n-completionsterminal_bench_2_nemotron_terminal_adapters_math__Qwen3_8B_20260418_175235
