Team Ai
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ai-humanizer-benchmark /ai-humanizer-benchmark AI Humanizer Benchmark: AI humanizers tested against 7 AI detectors (October 2026) AI Humanizer Benchmark measures how well AI humanizers rewrite AI-generated text so that AI detectors classify it as human-written, and how much meaning and readability the rewrite loses. In each monthly cycle, 11 AI humanizers rewrite the same 33 source texts across 7 writing categories, using each tool's default settings. Each output gets three kinds of score: 7 AI detectors (GPTZero… See the full description on the dataset page: https://huggingface.co/datasets/ai-humanizer-benchmark/ai-humanizer-benchmark.tabular1K<n<10K2 likes2.7k downloads9d agoHugging Face02Lelonthecodeur /humanizer-5m Humanizer-5M Version: 2.0.0 Humanizer-5M is a synthetic conversational adaptation dataset. The central objective is not simply to rewrite text to sound casual. Each example models: what the user explicitly asks, what the user may implicitly need, the user's conversational signals, the appropriate response calibration, the final response, quality and stability metrics. The dataset specifically teaches proportional adaptation. High user energy does not automatically mean high… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/humanizer-5m.text1M<n<10M2 likes405 downloads22d agoHugging Face03HumanizerBench /humanizerbench HumanizerBench: AI humanizer rankings and public audit record The complete audit record of HumanizerBench, a monthly benchmark of AI humanizers. Every tool rewrites the same freshly generated texts on the most undetectable setting it advertises, and every output is scored by five commercial AI detectors alongside meaning preservation and readability. We pay for every tool ourselves. There are no affiliate deals and no vendor-supplied numbers. Every input, every humanized output… See the full description on the dataset page: https://huggingface.co/datasets/HumanizerBench/humanizerbench.tabular10K<n<100K3 likes275 downloads20d agoHugging Face04gohumanize /gohumanize-open-humanizer-dataset GoHumanize Open Humanizer Dataset 2,957 training pairs and 300 test pairs for teaching a language model to rewrite AI-styled English prose into natural human writing. Each pair is: input: a passage rewritten by a large language model in the register typical of LLM output (formal, smooth, hedged, connective phrases, no contractions); output: the original human-written passage, from a public-domain book or, since version 2, from a US federal government publication. The human… See the full description on the dataset page: https://huggingface.co/datasets/gohumanize/gohumanize-open-humanizer-dataset.tabulartext-generation1K<n<10K0 likes217 downloads17d agoHugging Face05KNipun /ai-humanizer AI Humanizer Dataset (JSONL) This dataset is designed for fine-tuning instruction-following LLMs to rewrite AI-generated text into more natural, human-like language. Structure train.jsonl – training split validation.jsonl – validation split Format Each line is a JSON object: { "prompt": "Rewrite the following text to sound natural, human-like, and conversational...", "completion": "Humanized output text here", "attribution": "Original… See the full description on the dataset page: https://huggingface.co/datasets/KNipun/ai-humanizer.text10K<n<100K6 likes102 downloads10mo agoHugging Face06jialinyyzz /humanizer-data humanizer-data Training data for humanizer v2, a 12B model that rewrites AI-written drafts (English and Chinese) so they read like a person wrote them, keeping every fact. Code and app: GitHub. How the data was built, step by step: docs/DATA.md (中文). Only rows that were actually used to train v2 are here, one config per training step: Config Training step Used in training With text Reference only Removed before release rewrite_sft SFT: AI draft → human original 28,560… See the full description on the dataset page: https://huggingface.co/datasets/jialinyyzz/humanizer-data.texttext-generationn<1K4 likes60 downloads2d agoHugging Face07ai-humanizer /ai-humanizer-benchmark-2026 AI Humanizer Benchmark 2026 (Rephrasy) Every detector result behind the Best AI Humanizer 2026 ranking, one row per (tool, text, detector, run). Nine humanizers, two detectors, four test batches from December 2025 to September 2026. Nothing here is aggregated into a number you cannot trace. Each row has a source column: a public blog post with screenshots, or a screenshot in the org-card assets folder. Files benchmark.csv – 30 rows. Columns: test_id, date, tool… See the full description on the dataset page: https://huggingface.co/datasets/ai-humanizer/ai-humanizer-benchmark-2026.textn<1K0 likes51 downloads18d agoHugging Face08jayshah5696 /humanize-rl-sft-dataset humanize-rl-sft-dataset (v2) 4,835 high-quality SFT pairs for training a model to write natural, direct prose. Part of the humanize-rl project — a two-layer scoring and alignment pipeline for training small models to generate natural, human-sounding text. What this trains A model that can: Write natural Slack messages and emails from scratch. Rewrite stiff/formal/corporate text into direct, human-sounding prose. Fix grammar without making text formal. Shorten and… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-sft-dataset.texttext-generation1K<n<10K0 likes36 downloads4mo agoHugging Face09evijit /humanizer-dpo evijit/humanizer-dpo Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub. Try ML Intern: https://smolagents-ml-intern.hf.space Source code: https://github.com/huggingface/ml-intern Usage from datasets import load_dataset dataset = load_dataset('evijit/humanizer-dpo') text1K<n<10K0 likes33 downloads4mo agoHugging Face10jayshah5696 /humanize-rl-prime-sft-messages-env0315-clean50 jayshah5696/humanize-rl-prime-sft-messages-env0315-clean50 Prime prime-rl supervised fine-tuning dataset for Humanize-RL. This is the env0315_clean50 S2 repair-data candidate. It starts from the env0314 Prime SFT corpus and adds cleaned env0315 repair references generated from saved Prime rollout-audit failures. Splits split rows train 4358 validation 242 test 243 total 4843 Sources source rows safe_expand_3000_raw… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-prime-sft-messages-env0315-clean50.texttext-generation1K<n<10K0 likes28 downloads3mo agoHugging Face11jayshah5696 /humanize-rl-prime-sft-messages-env0315-clean50-primecompat jayshah5696/humanize-rl-prime-sft-messages-env0315-clean50-primecompat Prime prime-rl supervised fine-tuning dataset for Humanize-RL. This is the env0315_clean50 S2 repair-data candidate. It starts from the env0314 Prime SFT corpus and adds cleaned env0315 repair references generated from saved Prime rollout-audit failures. Splits split rows train 4358 validation 242 test 243 total 4843 Sources source rows… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-prime-sft-messages-env0315-clean50-primecompat.texttext-generation1K<n<10K0 likes21 downloads3mo agoHugging Face12manikineko /humanizer Humanized AI Dataset This dataset is designed to train large language models (LLMs) to produce human-like, conversational outputs that reflect the informal and dynamic style of average Discord users, while avoiding the robotic tone common in many AI models. The dataset prioritizes ethical use and includes safeguards to prevent harmful or abusive applications. Dataset Overview Purpose: To create LLMs with natural, human-like conversational abilities, free from overly… See the full description on the dataset page: https://huggingface.co/datasets/manikineko/humanizer.text1K<n<10K1 likes20 downloads1y agoHugging Face13jayshah5696 /humanize-rl-prime-sft-messages-env0314 Humanize-RL Prime SFT Messages Env0314 Prime prime-rl SFT dataset for Humanize-RL. Schema: each row has a messages list with one user instruction and one assistant target. Splits: train: 4313 validation: 239 test: 241 total accepted: 4793 rejected upstream by builder: 62 duplicate ids across published splits: 0 repair-reference rows: 20 Source artifact: v04_sft_final_plus_llama_failure_refs_env0314, built from restored v04 SFT data plus the clean Llama failure-reference repair… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-prime-sft-messages-env0314.texttext-generation1K<n<10K0 likes20 downloads4mo agoHugging Face14philsweeden /humanizer-entext1K<n<10K2 likes19 downloads6mo agoHugging Face15jayshah5696 /humanize-rl-tasks humanize-rl-tasks Single-turn RL task dataset for the humanize-rl project. Each row is one writing task. A model receives the prompt (instruction + source text), produces a completion, and the environment scores it with the 50/50 reward formula: reward = 0.50 × ridge_rubric_mean + 0.50 × deterministic_mean + penalties Dataset composition source tasks families input length v01_template 100 10 (template-generated) ~20 words v02_real_source 512 3 (real… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-tasks.texttext-generationn<1K0 likes17 downloads5mo agoHugging Face16jayshah5696 /humanize-rl-v03tabular10K<n<100K0 likes12 downloads5mo agoHugging Face17Pagepeek /humanizer_v2textn<1K1 likes9 downloads2y agoHugging Face18TheCodingKid /clutch-humanizer-v2-data Clutch Humanizer V2 Training Data Training data for the Clutch Humanizer model. Contents training_pairs.json: 7000 (AI text, Human text) pairs for training Format [ { "id": 0, "ai_text": "AI-style text to convert", "human_text": "Human-style target text", "source": "alpaca|dolly|essay" }, ... ] Usage from datasets import load_dataset dataset = load_dataset("TheCodingKid/clutch-humanizer-v2-data") # or import json import… See the full description on the dataset page: https://huggingface.co/datasets/TheCodingKid/clutch-humanizer-v2-data.text1K<n<10K0 likes9 downloads8mo agoHugging Face19lopol3000 /humanizer-entext1K<n<10K0 likes7 downloads6mo agoHugging Face20LevArtesa /sft-humanizer-dataset-v4tabularn<1K0 likes7 downloads5mo agoHugging Face21ibaha786 /humanizer-1ktext1K<n<10K1 likes6 downloads1y agoHugging Face22SwaYHell /yi-humanizer-dpo-v16-pairstabularn<1K0 likes6 downloads5mo agoHugging Face23SwaYHell /yi-humanizer-v18-full-pipeline-100tabularn<1K0 likes6 downloads5mo agoHugging Face24SwaYHell /yi-humanizer-v19-test-100tabularn<1K0 likes6 downloads5mo agoHugging Face25SwaYHell /yi-humanizer-v18-paraphrase-100tabularn<1K0 likes5 downloads5mo agoHugging Face26SwaYHell /yi-humanizer-v15-samples-100 Yi Humanizer v15 — 100 samples 100 humanized text samples generated by SwaYHell/yi-humanizer-v15-no-citations via vLLM batch inference. Generation params Base model: 01-ai/Yi-1.5-9B + LoRA (merged for vLLM) Temperature: 1.3 Input length range: 300–800 words N samples: 100 Schema (JSONL) i: index input: original AI text output: humanized version wi, wo: input/output word counts ratio: wo/wi temp: generation temperature model: LoRA model name tabularn<1K0 likes4 downloads5mo agoHugging Face27SwaYHell /yi-humanizer-v18-samples-100 Yi Humanizer v18 — 100 samples 100 humanized text samples generated by SwaYHell/yi-humanizer-v18-merged-v11-r8 via vLLM batch inference (double-merged: Yi + v11 + v18). Generation params Base: 01-ai/Yi-1.5-9B + v11 LoRA (merged) + v18 LoRA (merged) Temperature: 1.0 Input length range: 300–800 words N samples: 100 Schema (JSONL) i: index input: original AI text output: humanized version wi, wo: input/output word counts ratio: wo/wi temp: generation… See the full description on the dataset page: https://huggingface.co/datasets/SwaYHell/yi-humanizer-v18-samples-100.tabularn<1K0 likes4 downloads5mo agoHugging Face28SwaYHell /yi-humanizer-v18-AWQ-test-100tabularn<1K0 likes4 downloads5mo agoHugging Face29toandev /humanizer-data-vi Vietnamese Humanizer Dataset SFT data for natural Vietnamese rewriting: an AI rewrite is the input and the original news passage is the target (messages). 2,677 pairs: 2,151 train · 267 validation · 259 test, in Parquet format. Generation and evaluation model: gpt-6.1-sol. Recorded API usage: 8,916,167 tokens — 6,578,507 input and 2,337,660 output, including recorded trial calls. Estimated cost: approximately USD 36.53 at standard token prices. Created by: toandev. Thanks to… See the full description on the dataset page: https://huggingface.co/datasets/toandev/humanizer-data-vi.text1K<n<10K0 likes4h agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.