datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-humanizer-benchmark
AI Humanizer Benchmark: AI humanizers tested against 7 AI detectors (October 2026)
AI Humanizer Benchmark measures how well AI humanizers rewrite AI-generated text so that AI detectors classify it as human-written, and how much meaning and readability the rewrite loses. In each monthly cycle, 11 AI humanizers rewrite the same 33 source texts across 7 writing categories, using each tool's default settings. Each output gets three kinds of score: 7 AI detectors (GPTZero… See the full description on the dataset page: https://huggingface.co/datasets/ai-humanizer-benchmark/ai-humanizer-benchmark.humanizer-5m
Humanizer-5M
Version: 2.0.0
Humanizer-5M is a synthetic conversational adaptation dataset.
The central objective is not simply to rewrite text to sound casual.
Each example models:
what the user explicitly asks,
what the user may implicitly need,
the user's conversational signals,
the appropriate response calibration,
the final response,
quality and stability metrics.
The dataset specifically teaches proportional adaptation.
High user energy does not automatically mean high… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/humanizer-5m.humanizerbench
HumanizerBench: AI humanizer rankings and public audit record
The complete audit record of HumanizerBench, a monthly benchmark of AI humanizers. Every tool rewrites the same freshly generated texts on the most undetectable setting it advertises, and every output is scored by five commercial AI detectors alongside meaning preservation and readability. We pay for every tool ourselves. There are no affiliate deals and no vendor-supplied numbers.
Every input, every humanized output… See the full description on the dataset page: https://huggingface.co/datasets/HumanizerBench/humanizerbench.gohumanize-open-humanizer-dataset
GoHumanize Open Humanizer Dataset
2,957 training pairs and 300 test pairs for teaching a language model to rewrite
AI-styled English prose into natural human writing. Each pair is:
input: a passage rewritten by a large language model in the register typical of LLM output
(formal, smooth, hedged, connective phrases, no contractions);
output: the original human-written passage, from a public-domain book or, since version 2,
from a US federal government publication.
The human… See the full description on the dataset page: https://huggingface.co/datasets/gohumanize/gohumanize-open-humanizer-dataset.ai-humanizer
AI Humanizer Dataset (JSONL)
This dataset is designed for fine-tuning instruction-following LLMs
to rewrite AI-generated text into more natural, human-like language.
Structure
train.jsonl – training split
validation.jsonl – validation split
Format
Each line is a JSON object:
{
"prompt": "Rewrite the following text to sound natural, human-like, and conversational...",
"completion": "Humanized output text here",
"attribution": "Original… See the full description on the dataset page: https://huggingface.co/datasets/KNipun/ai-humanizer.humanizer-data
humanizer-data
Training data for humanizer v2, a 12B model that rewrites AI-written drafts (English and Chinese) so they read like a person wrote them, keeping every fact. Code and app: GitHub. How the data was built, step by step: docs/DATA.md (中文).
Only rows that were actually used to train v2 are here, one config per training step:
Config
Training step
Used in training
With text
Reference only
Removed before release
rewrite_sft
SFT: AI draft → human original
28,560… See the full description on the dataset page: https://huggingface.co/datasets/jialinyyzz/humanizer-data.sft-humanizer-dataset-v5ai-humanizer-benchmark-2026
AI Humanizer Benchmark 2026 (Rephrasy)
Every detector result behind the Best AI Humanizer 2026 ranking, one row per (tool, text, detector, run). Nine humanizers, two detectors, four test batches from December 2025 to September 2026.
Nothing here is aggregated into a number you cannot trace. Each row has a source column: a public blog post with screenshots, or a screenshot in the org-card assets folder.
Files
benchmark.csv – 30 rows. Columns: test_id, date, tool… See the full description on the dataset page: https://huggingface.co/datasets/ai-humanizer/ai-humanizer-benchmark-2026.humanize-rl-sft-dataset
humanize-rl-sft-dataset (v2)
4,835 high-quality SFT pairs for training a model to write natural, direct prose.
Part of the humanize-rl project — a two-layer scoring and alignment pipeline for training small models to generate natural, human-sounding text.
What this trains
A model that can:
Write natural Slack messages and emails from scratch.
Rewrite stiff/formal/corporate text into direct, human-sounding prose.
Fix grammar without making text formal.
Shorten and… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-sft-dataset.humanizer-dpo
evijit/humanizer-dpo
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset('evijit/humanizer-dpo')
humanize-rl-prime-sft-messages-env0315-clean50
jayshah5696/humanize-rl-prime-sft-messages-env0315-clean50
Prime prime-rl supervised fine-tuning dataset for Humanize-RL.
This is the env0315_clean50 S2 repair-data candidate. It starts from the
env0314 Prime SFT corpus and adds cleaned env0315 repair references generated
from saved Prime rollout-audit failures.
Splits
split
rows
train
4358
validation
242
test
243
total
4843
Sources
source
rows
safe_expand_3000_raw… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-prime-sft-messages-env0315-clean50.humanize-rl-research-artifacts-env0314
Humanize-RL Research Artifacts Env0314
This dataset archives the ignored local artifacts needed to continue the Humanize-RL reward patch, SFT repair, and Qwen 2B/9B ablation work after leaving the original worktree.
Primary source run: zztqgqclh3y3hslpjsofzpcf
Prime env: jayshah5696/humanize-rl-env@0.3.14
W&B run: https://wandb.ai/jayshah5696/humanize-rl/runs/akzopsz9
Published SFT dataset: jayshah5696/humanize-rl-prime-sft-messages-env0314
Contents… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-research-artifacts-env0314.humanize-rl-tasks-v03
humanize-rl RL Tasks v03
492 RL tasks designed to train and evaluate models that produce human-sounding prose
under specific constraints. Each row pairs an authored instruction with a constraint
spec consumable by the humanize-rl reward
environment (deterministic checks + Layer-1 heuristics + ridge regression style scorer).
Built per docs/plans/v03-rl-tasks-dataset.md.
What's in a row
field
meaning
id
stable task id (rl_v03_NNNNNN)
family
top-level… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-tasks-v03.humanize-rl-prime-sft-messages-env0315-clean50-primecompat
jayshah5696/humanize-rl-prime-sft-messages-env0315-clean50-primecompat
Prime prime-rl supervised fine-tuning dataset for Humanize-RL.
This is the env0315_clean50 S2 repair-data candidate. It starts from the
env0314 Prime SFT corpus and adds cleaned env0315 repair references generated
from saved Prime rollout-audit failures.
Splits
split
rows
train
4358
validation
242
test
243
total
4843
Sources
source
rows… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-prime-sft-messages-env0315-clean50-primecompat.humanizer
Humanized AI Dataset
This dataset is designed to train large language models (LLMs) to produce human-like, conversational outputs that reflect the informal and dynamic style of average Discord users, while avoiding the robotic tone common in many AI models. The dataset prioritizes ethical use and includes safeguards to prevent harmful or abusive applications.
Dataset Overview
Purpose: To create LLMs with natural, human-like conversational abilities, free from overly… See the full description on the dataset page: https://huggingface.co/datasets/manikineko/humanizer.humanize-rl-prime-sft-messages-env0314
Humanize-RL Prime SFT Messages Env0314
Prime prime-rl SFT dataset for Humanize-RL.
Schema: each row has a messages list with one user instruction and one assistant target.
Splits:
train: 4313
validation: 239
test: 241
total accepted: 4793
rejected upstream by builder: 62
duplicate ids across published splits: 0
repair-reference rows: 20
Source artifact: v04_sft_final_plus_llama_failure_refs_env0314, built from restored v04 SFT data plus the clean Llama failure-reference repair… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-prime-sft-messages-env0314.humanizer-enhumanize-rl-tasks
humanize-rl-tasks
Single-turn RL task dataset for the humanize-rl project.
Each row is one writing task. A model receives the prompt (instruction + source text),
produces a completion, and the environment scores it with the 50/50 reward formula:
reward = 0.50 × ridge_rubric_mean + 0.50 × deterministic_mean + penalties
Dataset composition
source
tasks
families
input length
v01_template
100
10 (template-generated)
~20 words
v02_real_source
512
3 (real… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-tasks.humanize-rl-v03humanizer_v2clutch-humanizer-v2-data
Clutch Humanizer V2 Training Data
Training data for the Clutch Humanizer model.
Contents
training_pairs.json: 7000 (AI text, Human text) pairs for training
Format
[
{
"id": 0,
"ai_text": "AI-style text to convert",
"human_text": "Human-style target text",
"source": "alpaca|dolly|essay"
},
...
]
Usage
from datasets import load_dataset
dataset = load_dataset("TheCodingKid/clutch-humanizer-v2-data")
# or
import json
import… See the full description on the dataset page: https://huggingface.co/datasets/TheCodingKid/clutch-humanizer-v2-data.humanizer-ensft-humanizer-dataset-v4humanizer-1kyi-humanizer-dpo-v16-pairsyi-humanizer-v18-full-pipeline-100yi-humanizer-v19-test-100yi-humanizer-v18-paraphrase-100yi-humanizer-v15-samples-100
Yi Humanizer v15 — 100 samples
100 humanized text samples generated by SwaYHell/yi-humanizer-v15-no-citations
via vLLM batch inference.
Generation params
Base model: 01-ai/Yi-1.5-9B + LoRA (merged for vLLM)
Temperature: 1.3
Input length range: 300–800 words
N samples: 100
Schema (JSONL)
i: index
input: original AI text
output: humanized version
wi, wo: input/output word counts
ratio: wo/wi
temp: generation temperature
model: LoRA model name
yi-humanizer-v18-samples-100
Yi Humanizer v18 — 100 samples
100 humanized text samples generated by SwaYHell/yi-humanizer-v18-merged-v11-r8
via vLLM batch inference (double-merged: Yi + v11 + v18).
Generation params
Base: 01-ai/Yi-1.5-9B + v11 LoRA (merged) + v18 LoRA (merged)
Temperature: 1.0
Input length range: 300–800 words
N samples: 100
Schema (JSONL)
i: index
input: original AI text
output: humanized version
wi, wo: input/output word counts
ratio: wo/wi
temp: generation… See the full description on the dataset page: https://huggingface.co/datasets/SwaYHell/yi-humanizer-v18-samples-100.
