Team Ai
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01soyrsoyr /jev-playground-rlcd-v0 openjev-rlcd-v0 Synthetic calibrated-decision dataset for training Jev-style System One classifiers plan). States are generated programmatically per domain; typed questions follow the openjev /v1/systemone contract; reference distributions are exact by construction, the property a proper-scoring-rule / calibration objective needs. Format JSONL, one state per line: { "domain": "support_tickets", "state": "...", "questions": { "urgent": {"type": "noul"… See the full description on the dataset page: https://huggingface.co/datasets/soyrsoyr/jev-playground-rlcd-v0.text-classification1K<n<10K0 likes219 downloads12d agoHugging Face02sumleo /RLCDAlignBenchgated RLCDAlignBench Paper: Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures (arXiv:2609.29429) Code: github.com/sumleo/RLCDAlignBench · Project page: sumleo.github.io/RLCDAlignBench RLCDAlignBench measures whether a detector can tell when a language model's output is an alignment failure. It has 44 benchmarks across ten failure types and five target models (Qwen3.5-2B, Phi-4-mini, Gemma-2-2B, Llama-3.2-3B, Olmo-3-7B), for… See the full description on the dataset page: https://huggingface.co/datasets/sumleo/RLCDAlignBench.tabulartext-classification10K<n<100K1 likes123 downloads11d agoHugging Face03anthonym21 /rlcd-decision-v1 RLCD Decision Dataset (v1) Typed decision questions for training and evaluating models that answer with a calibrated probability distribution over declared options instead of generated text. Built for RLCD — reinforcement learning for calibrated decisions (see also the trained export anthonym21/qwen3-0.6b-rlcd-decision). Every row is one typed question over a context: a choice question over unordered options, a score question over ordered levels, or a noul (yes/no) question. The… See the full description on the dataset page: https://huggingface.co/datasets/anthonym21/rlcd-decision-v1.texttext-classification10K<n<100K0 likes97 downloads11d agoHugging Face04TaylorAI /rlcd Dataset Card for "rlcd" More Information needed text100K<n<1M0 likes89 downloads3y agoHugging Face05anthonym21 /eve-rlcd-runs0 likes86 downloads18d agoHugging Face06TaylorAI /RLCD-generated-preference-data-split Dataset Card for "RLCD-generated-preference-data-split" More Information needed tabular100K<n<1M0 likes75 downloads3y agoHugging Face07gyung /Qwev-RLCD-Data10K<n<100K1 likes75 downloads10d agoHugging Face08TaylorAI /RLCD-SFT-conversationsMIT License Dataset Card for "RLCD-SFT-conversations" More Information needed text100K<n<1M0 likes64 downloads2y agoHugging Face09goutam /rlcd-decision-atlasgated rlcd-decision-atlas — a typed-decision dataset This is a dataset for training and evaluating models that make decisions rather than write prose. Every example is one situation, one question, and a short list of declared options — and exactly one of those options is correct. The model's whole job is to put a probability on each option and commit to one. Nothing here asks for free text. It is assembled from 114 public sources plus one synthetic generator, each converted into the… See the full description on the dataset page: https://huggingface.co/datasets/goutam/rlcd-decision-atlas.texttext-classification100K<n<1M0 likes40 downloads9d agoHugging Face10TaylorAI /RLCD-generated-preference-data Dataset Card for "RLCD-generated-preference-data" More Information needed tabular100K<n<1M1 likes30 downloads3y agoHugging Face11ConicCat /C2-Nemo-Delta-RLCD-PreferenceMy attempt at cheesing the lack of roleplay preference data through a combination of delta tuning and RLCD. Unfortunately there are several notable issues, namely: Unbalanced em-dash usage, Nemo 253B and 4B not having exactly the same sytle, RLCD being iffy as Nemo 253B and 4B don't have a great idea of what constitutes bad and good rp, and the 4B being a bit too not atrociously bad enough for delta tuning. Obviously, as this is lightly filtered C2 there's a good amount of nsfl and nsfw. text1K<n<10K0 likes22 downloads8mo agoHugging Face12TaylorAI /RLCD-SFT-dataset Dataset Card for "RLCD-SFT-dataset" More Information needed text100K<n<1M0 likes18 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.