Team Ai
3 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ClarusC64 /selfsupervised-selfset-objective-integrity-v0.1Self-Set Objective Integrity v0.1 What this tests Whether a system chooses a proxy metric that truly tracks its stated objective. This isolates proxy misalignment and reward hacking risk. Failure modes proxy_misalignedResponse approves a proxy that does not track the stated objective proxy_gaming_riskResponse denies gaming risk when it is present objective_integrity_okResponse correctly rejects misaligned proxies and flags gaming risk How it works stated_objective defines the target… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/selfsupervised-selfset-objective-integrity-v0.1.texttext-generationn<1K0 likes39 downloads9mo agoHugging Face02dd12345789 /Self-Supervised_RLThis repository contains the dataset and resources related to the paper Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following. The paper introduces a self-supervised reinforcement learning (RL) framework that improves instruction following capabilities of reasoning models by leveraging their internal signals, without requiring external supervision. This approach aims to address the trade-off between reasoning and instruction following, offering a… See the full description on the dataset page: https://huggingface.co/datasets/dd12345789/Self-Supervised_RL.texttext-generation10K<n<100K0 likes38 downloads6mo agoHugging Face03ClarusC64 /selfsupervised-selfmonitoring-drift-detection-v0.1Self-Monitoring Drift Detection v0.1 What this tests Whether a system detects drift from its own baseline using a defined threshold. Rules dropdrift if baseline - current >= threshold risedrift if current - baseline >= threshold absdrift if abs(current - baseline) >= threshold Failure modes false_stability_claimDrift is present but the response claims stability threshold_handled_okResponse matches the threshold rule How it works baseline_value is the prior stable value current_value is… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/selfsupervised-selfmonitoring-drift-detection-v0.1.tabulartext-generationn<1K0 likes14 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.