datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
explicit-edit-benchmark
Explicit Edit Benchmark
226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte.
Source code and benchmark runner: GitHub — Explicit Edit Benchmark
Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens.
Leaderboard by model route
Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.Danbooru2024_rating_explicit_prompts_without_character精选了danbooru2024里分级为explict且score>20的所有图像的prompt,去除了原角色的tag和相关特征,可以直接用于角色nsfw图像生成
emotion_explicit_cue2026-10-01-odcv-qwen36-0-da-explicit-15
odcv eval of dougalldeepmind/2026-09-30-qwen36-0-da-explicit-15 (mode=think)
field
value
experiment
odcv eval of dougalldeepmind/2026-09-30-qwen36-0-da-explicit-15 (mode=think)
date_generated
2026-10-01
constitution
none
source_repo
teaching_claude_why_replication @ f62a8330e34861fcc607e3ef702656ecceacc93e
models
{"target": "dougalldeepmind/2026-09-30-qwen36-0-da-explicit-15", "target_revision": "bac735111b6a181b13511224da48f61cef6394c3", "base":… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-10-01-odcv-qwen36-0-da-explicit-15.Explicit-NeRF-QAIf you use our dataset, please cite the following article:
@article{xing2024explicit,
title={Explicit-NeRF-QA: A quality assessment database for explicit NeRF model compression},
author={Xing, Yuke and Yang, Qi and Yang, Kaifa and Xu, Yilin and Li, Zhu},
journal={arXiv preprint arXiv:2407.08165},
year={2024}
}
task323_jigsaw_classification_sexually_explicit
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task323_jigsaw_classification_sexually_explicit
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task323_jigsaw_classification_sexually_explicit.2026-09-30-da-explicit-synth
synth da-explicit run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth da-explicit run — per-stage snapshots (resumable generation cache)
date_generated
20260930_200535
constitution
constitutions/claude_distilled_09_principles/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ 308511e9fa5f64fba3ed95dd484c9c13b50035b1
models
per-stage models — see manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-30-da-explicit-synth.2026-10-01-mask-qwen36-0-da-explicit-15
mask eval of dougalldeepmind/2026-09-30-qwen36-0-da-explicit-15 (mode=think)
field
value
experiment
mask eval of dougalldeepmind/2026-09-30-qwen36-0-da-explicit-15 (mode=think)
date_generated
2026-10-01
constitution
none
source_repo
teaching_claude_why_replication @ f62a8330e34861fcc607e3ef702656ecceacc93e
models
{"target": "dougalldeepmind/2026-09-30-qwen36-0-da-explicit-15", "target_revision": "bac735111b6a181b13511224da48f61cef6394c3", "base":… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-10-01-mask-qwen36-0-da-explicit-15.2026-09-30-odcv-qwen36-0-da-15-explicit
odcv eval of dougalldeepmind/2026-09-30-qwen36-0-da-15-explicit (mode=think)
field
value
experiment
odcv eval of dougalldeepmind/2026-09-30-qwen36-0-da-15-explicit (mode=think)
date_generated
2026-09-30
constitution
none
source_repo
teaching_claude_why_replication @ c4c99847ac5d6e26a9cf8a407c243fa19a09ba62
models
{"target": "dougalldeepmind/2026-09-30-qwen36-0-da-15-explicit", "target_revision": "530d0394d1f8e21012b1aeeacde17e90fb9456ce", "base":… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-30-odcv-qwen36-0-da-15-explicit.prefeval_explicit
PrefEval Benchmark: Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs
Welcome to the PrefEval dataset repository!
| Website | Paper | GitHub Repository |
Dataset Overview
We introduce PrefEval, a benchmark for evaluating LLMs' ability to infer, memorize and adhere to user preferences in a long-context conversational setting. The benchmark consists of three distinct preference forms, each requiring different levels of preference… See the full description on the dataset page: https://huggingface.co/datasets/siyanzhao/prefeval_explicit.2026-09-30-mask-qwen36-0-da-15-explicit
mask eval of dougalldeepmind/2026-09-30-qwen36-0-da-15-explicit (mode=think)
field
value
experiment
mask eval of dougalldeepmind/2026-09-30-qwen36-0-da-15-explicit (mode=think)
date_generated
2026-09-30
constitution
none
source_repo
teaching_claude_why_replication @ c4c99847ac5d6e26a9cf8a407c243fa19a09ba62
models
{"target": "dougalldeepmind/2026-09-30-qwen36-0-da-15-explicit", "target_revision": "530d0394d1f8e21012b1aeeacde17e90fb9456ce", "base":… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-30-mask-qwen36-0-da-15-explicit.hh_rlhf_with_explicit_sentiment_backdoors_llama3b2026-09-30-da-explicit-15-mix
difficult advice, da-explicit arm (configs/data/synth/da-explicit.yaml): the 28 Sep recipe with explicit asks: the person asks the assistant to carry out the shortcut itself; the base blend scaled around that share
field
value
experiment
difficult advice, da-explicit arm (configs/data/synth/da-explicit.yaml): the 28 Sep recipe with explicit asks: the person asks the assistant to carry out the shortcut itself; the base blend scaled around that share — final training… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-30-da-explicit-15-mix.annotator-9-11-explicit-rating-data-vocalstrin_data_tldr_explicit_dataset
TL;DR Dataset for Preference Learning
Summary
The TL;DR dataset is a processed version of Reddit posts, specifically curated to train models using the TRL library for preference learning and Reinforcement Learning from Human Feedback (RLHF) tasks. It leverages the common practice on Reddit where users append "TL;DR" (Too Long; Didn't Read) summaries to lengthy posts, providing a rich source of paired text data for training models to understand and generate concise… See the full description on the dataset page: https://huggingface.co/datasets/Kyleyee/trin_data_tldr_explicit_dataset.2026-09-30-da-15-explicit-mix
difficult-advice arm, 159 rows swapped for 8 Sep explicit rows
field
value
experiment
difficult-advice arm for the explicit-ask vs advice-request test (explicit): dougalldeepmind/2026-09-28-da-15-mix @ ff524823 with 159 da rows replaced by dougalldeepmind/2026-09-08-da-synth @ 42107bde rows whose user asks the assistant to do or write the thing; the paired arm fills the same slots with rows of the same trait and AI type; donor rows move whole; swaps in… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-30-da-15-explicit-mix.annotator-9-11-explicit-rating-dataflan_combined_task323_jigsaw_classification_sexually_explicitPICABenchV1-explicitharmful_behaviors_with_explicit_requeststrain_data_Helpful_explicit_prompt
HH-RLHF-Helpful-Base Dataset
Summary
The HH-RLHF-Helpful-Base dataset is a processed version of Anthropic's HH-RLHF dataset, specifically curated to train models using the TRL library for preference learning and alignment tasks. It contains pairs of text samples, each labeled as either "chosen" or "rejected," based on human preferences regarding the helpfulness of the responses. This dataset enables models to learn human preferences in generating helpful responses… See the full description on the dataset page: https://huggingface.co/datasets/Kyleyee/train_data_Helpful_explicit_prompt.discrim-eval-explicit-subsetsExplicit_content
Dataset Card for Explicit content detection
Dataset Description
1189 News Articles classified into different categories namely: "Explicit" if the article contains explicit content and "Not_Explicit" if not.
Languages
The text in the dataset is in English
Dataset Structure
The dataset consists of two columns namely Article and Category.
The Article column consists of the news article and the Category column consists of the class each article belongs… See the full description on the dataset page: https://huggingface.co/datasets/valurank/Explicit_content.chembl-2025-randomized-smiles-cleaned-explicit-hshedgehog-schema-explicit
hedgehog-schema-explicit
Hedgehog — explicit-schema extraction training.
Contents
train.jsonl (1280 rows)
validation.jsonl (160 rows)
test.jsonl (192 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the Hedgehog extraction model (Michael Anthony Falabella).
explicit-implicit-disease-diagnosis-datatrain_data_HH_explicit_prompt
HH-RLHF-Helpful-Base Dataset
Summary
The HH-RLHF-Helpful-Base dataset is a processed version of Anthropic's HH-RLHF dataset, specifically curated to train models using the TRL library for preference learning and alignment tasks. It contains pairs of text samples, each labeled as either "chosen" or "rejected," based on human preferences regarding the helpfulness of the responses. This dataset enables models to learn human preferences in generating helpful responses… See the full description on the dataset page: https://huggingface.co/datasets/Kyleyee/train_data_HH_explicit_prompt.explicitMedical-medical-preference-pubmed-olmo-normal-rollouts-graded-by-claudeexplicit-implicit-disease-diagnosis-data-rawexplicit_similar_500_9500
