models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
rlcd-modernbert-151mArmoRM-Llama3-8B-v0.1qwen3-0.6b-rlcd-decisionllama-3.1-8b-oracle-rm-hh-rlhf-helpfulnessllama-3.1-8b-oracle-rm-hh-rlhf-harmlessnesslaya-finetuned-rlcd-GGUFRewardModel-Mistral-7B-for-DPA-v1Hugston_code-rl-Qwen3-4B-Instruct-2507-SFT-30bdeberta-v3-large-tasksource-rlhf-reward-modelWorldPM-72B-RLHFLowdecision-head-qwen3.5-4b-rlcd-32ktulu-v2.5-13b-hh-rlhf-60k-rmLlama-3.1-Tulu-3-8B-RL-RM-RB2Qwev-9B-RLCDcua-s1-forge-rlcd-v3hh_rlhf_rm_open_llama_3blaya-finetuned-rlcdDecision-Tree-Reward-Llama-3.1-8BDecision-Tree-Reward-Gemma-2-27BTLDR-Mistral-7B-RMcua-s1-forge-rlcddecision-model-rl-overcookedrlhflow-llama-3-sft-8b-v2-token-rm-700kdeepseek-math-7b-rl-prm-v0.1Qwen2.5-7B-SafeRLHF-RM1011-hh-rlhf-1.1b-128-1e-5-epoch-1gpt2-rlhf-rewarddistilbert-base-uncased-finetuned-clincTLDR-Llama-3.2-1B-SmallSFT-RMnews-classifier
