Team Ai
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anthonym21 /rlcd-decision-v1 RLCD Decision Dataset (v1) Typed decision questions for training and evaluating models that answer with a calibrated probability distribution over declared options instead of generated text. Built for RLCD — reinforcement learning for calibrated decisions (see also the trained export anthonym21/qwen3-0.6b-rlcd-decision). Every row is one typed question over a context: a choice question over unordered options, a score question over ordered levels, or a noul (yes/no) question. The… See the full description on the dataset page: https://huggingface.co/datasets/anthonym21/rlcd-decision-v1.texttext-classification10K<n<100K0 likes163 downloads16d agoHugging Face02sumleo /RLCDAlignBenchgated RLCDAlignBench Paper: Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures (arXiv:2609.29429) Code: github.com/sumleo/RLCDAlignBench · Project page: sumleo.github.io/RLCDAlignBench RLCDAlignBench measures whether a detector can tell when a language model's output is an alignment failure. It has 44 benchmarks across ten failure types and five target models (Qwen3.5-2B, Phi-4-mini, Gemma-2-2B, Llama-3.2-3B, Olmo-3-7B), for… See the full description on the dataset page: https://huggingface.co/datasets/sumleo/RLCDAlignBench.tabulartext-classification10K<n<100K1 likes142 downloads15d agoHugging Face03TaylorAI /rlcd Dataset Card for "rlcd" More Information needed text100K<n<1M0 likes85 downloads3y agoHugging Face04TaylorAI /RLCD-generated-preference-data-split Dataset Card for "RLCD-generated-preference-data-split" More Information needed tabular100K<n<1M0 likes61 downloads3y agoHugging Face05TaylorAI /RLCD-SFT-conversationsMIT License Dataset Card for "RLCD-SFT-conversations" More Information needed text100K<n<1M0 likes56 downloads2y agoHugging Face06goutam /rlcd-decision-atlasgated rlcd-decision-atlas — a typed-decision dataset This is a dataset for training and evaluating models that make decisions rather than write prose. Every example is one situation, one question, and a short list of declared options — and exactly one of those options is correct. The model's whole job is to put a probability on each option and commit to one. Nothing here asks for free text. It is assembled from 114 public sources plus one synthetic generator, each converted into the… See the full description on the dataset page: https://huggingface.co/datasets/goutam/rlcd-decision-atlas.texttext-classification100K<n<1M0 likes43 downloads13d agoHugging Face07TaylorAI /RLCD-generated-preference-data Dataset Card for "RLCD-generated-preference-data" More Information needed tabular100K<n<1M1 likes28 downloads3y agoHugging Face08ConicCat /C2-Nemo-Delta-RLCD-PreferenceMy attempt at cheesing the lack of roleplay preference data through a combination of delta tuning and RLCD. Unfortunately there are several notable issues, namely: Unbalanced em-dash usage, Nemo 253B and 4B not having exactly the same sytle, RLCD being iffy as Nemo 253B and 4B don't have a great idea of what constitutes bad and good rp, and the 4B being a bit too not atrociously bad enough for delta tuning. Obviously, as this is lightly filtered C2 there's a good amount of nsfl and nsfw. text1K<n<10K0 likes22 downloads8mo agoHugging Face09TaylorAI /RLCD-SFT-dataset Dataset Card for "RLCD-SFT-dataset" More Information needed text100K<n<1M0 likes17 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.