models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
testing-auto-training-modeltrainingtwotrainingonejina-reranker-v2-multilingual-affiliations-comet-training-onlyguardrails-poisoning-trainingcritic_400_deepseek-r1-distil-1.5b-ppo-run-math-training-prompt-len-800-response-len-c49633e26ecritic_200_deepseek-r1-distil-1.5b-ppo-run-math-training-prompt-len-800-response-len-0e5f8c09dccritic_600_ppo-run-math-training-prompt-len-800-response-len-4096-seed-43-subset-1000-f553c1779bcross-encoder-qqp-lcqmc-training-paraphrase-multilingual-MiniLM-L12-v2critic_600_deepseek-r1-distil-1.5b-ppo-run-math-training-prompt-len-800-response-len-07fa1b4078distilbert-base-uncased-training-colaincremental-semi-supervised-training-500k-upsampledtrain-reward-trainingincremental-semi-supervised-training-1mln-downsampledcheckworthy-binary-classification-training-debert-1755503743critic_200_ppo-run-math-training-prompt-len-800-response-len-4096-seed-43-subset-250-40bddeea62deberta-valence-title_training1tRAINING-DATASET-All-files-finalcritic_2600_ppo-run-math-training-prompt-len-800-response-len-4096-bfdc5c41c3incremental-semi-supervised-training-500k-equaltoxic-initial-trainingtraining15e-6_xlm-R-xl_Conspiracy_training_with_callbacksindobert-clean-post-training-fin-sa-2indobert-clean-post-training-fin-sa-1Qwen-1.5B-Instruct-ppo-run-math-training-prompt-len-800-response-len-4096-seed-43-sub-8cd26db347critic_800_Qwen-1.5B-Instruct-ppo-run-math-training-prompt-len-800-response-len-4096-0c4e7a4810critic_250_ppo-run-math-training-prompt-len-800-response-len-4096-seed-43-subset-500-c4b41565c8critic_16_ppo-run-math-training-prompt-len-800-response-len-4096-seed-43-subset-500-a-9a44e3cd58critic_450_ppo-run-math-training-prompt-len-800-response-len-4096-bce-loss-temperatur-9fe16df365
