Team Ai
20 results

adapters

davidafrica /neologism-ft-adapters Neologism project — emergent-misalignment workbench adapters LoRA adapters from the fine-tuning experiments of the neologism-learning project (code and paper, Section 8: inoculation labels, suppression switches, and controls). Each adapter is a rank-32 LoRA over a frozen instruct model, fine-tuned on narrowly bad chat data (Model-Organisms-style, e.g. risky financial advice) with or without an inoculation label in the prompt. Safety note. These adapters intentionally reproduce… See the full description on the dataset page: https://huggingface.co/datasets/davidafrica/neologism-ft-adapters.0 likes482 downloads2mo agoHugging Facerdavion /opd-method-comparison-adapters OPD Method Comparison — all 27 training-complete adapters This public dataset contains all 27 training-complete LoRA adapters from the OPD method-comparison experiment. Status: training is complete for all 27 conditions; final evaluation is still in progress. These artifacts should not yet be interpreted as final benchmark results. Base model: 'Qwen/Qwen2.5-7B-Instruct' at revision 'a09a35458c702b33eeacc393d103063234e8bc28'. Each 'adapters//' directory contains the PEFT adapter… See the full description on the dataset page: https://huggingface.co/datasets/rdavion/opd-method-comparison-adapters.0 likes404 downloads2mo agoHugging Faceagurung /eaiexp-rsaoj-adapters0 likes260 downloads20d agoHugging FaceEdwardoSunny /em-mlp-attn-adapters0 likes105 downloads17d agoHugging FacePrasadmahadik /Inoculation_Adapters_mechanistic_exp0 likes85 downloads3mo agoHugging Facesibasmarakp /Qwen2.5-1.5B-Instruct-uPRM-T80-adapters-best_of_n-completionstabular10K<n<100K0 likes84 downloads9mo agoHugging Face