Team Ai
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gohumanize /gohumanize-open-humanizer-dataset GoHumanize Open Humanizer Dataset 2,957 training pairs and 300 test pairs for teaching a language model to rewrite AI-styled English prose into natural human writing. Each pair is: input: a passage rewritten by a large language model in the register typical of LLM output (formal, smooth, hedged, connective phrases, no contractions); output: the original human-written passage, from a public-domain book or, since version 2, from a US federal government publication. The human… See the full description on the dataset page: https://huggingface.co/datasets/gohumanize/gohumanize-open-humanizer-dataset.tabulartext-generation1K<n<10K0 likes217 downloads17d agoHugging Face02jialinyyzz /humanizer-data humanizer-data Training data for humanizer v2, a 12B model that rewrites AI-written drafts (English and Chinese) so they read like a person wrote them, keeping every fact. Code and app: GitHub. How the data was built, step by step: docs/DATA.md (中文). Only rows that were actually used to train v2 are here, one config per training step: Config Training step Used in training With text Reference only Removed before release rewrite_sft SFT: AI draft → human original 28,560… See the full description on the dataset page: https://huggingface.co/datasets/jialinyyzz/humanizer-data.texttext-generationn<1K4 likes60 downloads2d agoHugging Face03jayshah5696 /humanize-rl-sft-dataset humanize-rl-sft-dataset (v2) 4,835 high-quality SFT pairs for training a model to write natural, direct prose. Part of the humanize-rl project — a two-layer scoring and alignment pipeline for training small models to generate natural, human-sounding text. What this trains A model that can: Write natural Slack messages and emails from scratch. Rewrite stiff/formal/corporate text into direct, human-sounding prose. Fix grammar without making text formal. Shorten and… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-sft-dataset.texttext-generation1K<n<10K0 likes36 downloads4mo agoHugging Face04jayshah5696 /humanize-rl-prime-sft-messages-env0315-clean50 jayshah5696/humanize-rl-prime-sft-messages-env0315-clean50 Prime prime-rl supervised fine-tuning dataset for Humanize-RL. This is the env0315_clean50 S2 repair-data candidate. It starts from the env0314 Prime SFT corpus and adds cleaned env0315 repair references generated from saved Prime rollout-audit failures. Splits split rows train 4358 validation 242 test 243 total 4843 Sources source rows safe_expand_3000_raw… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-prime-sft-messages-env0315-clean50.texttext-generation1K<n<10K0 likes28 downloads3mo agoHugging Face05jayshah5696 /humanize-rl-research-artifacts-env0314 Humanize-RL Research Artifacts Env0314 This dataset archives the ignored local artifacts needed to continue the Humanize-RL reward patch, SFT repair, and Qwen 2B/9B ablation work after leaving the original worktree. Primary source run: zztqgqclh3y3hslpjsofzpcf Prime env: jayshah5696/humanize-rl-env@0.3.14 W&B run: https://wandb.ai/jayshah5696/humanize-rl/runs/akzopsz9 Published SFT dataset: jayshah5696/humanize-rl-prime-sft-messages-env0314 Contents… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-research-artifacts-env0314.text-generation0 likes23 downloads4mo agoHugging Face06jayshah5696 /humanize-rl-tasks-v03 humanize-rl RL Tasks v03 492 RL tasks designed to train and evaluate models that produce human-sounding prose under specific constraints. Each row pairs an authored instruction with a constraint spec consumable by the humanize-rl reward environment (deterministic checks + Layer-1 heuristics + ridge regression style scorer). Built per docs/plans/v03-rl-tasks-dataset.md. What's in a row field meaning id stable task id (rl_v03_NNNNNN) family top-level… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-tasks-v03.text-generationn<1K0 likes21 downloads5mo agoHugging Face07jayshah5696 /humanize-rl-prime-sft-messages-env0315-clean50-primecompat jayshah5696/humanize-rl-prime-sft-messages-env0315-clean50-primecompat Prime prime-rl supervised fine-tuning dataset for Humanize-RL. This is the env0315_clean50 S2 repair-data candidate. It starts from the env0314 Prime SFT corpus and adds cleaned env0315 repair references generated from saved Prime rollout-audit failures. Splits split rows train 4358 validation 242 test 243 total 4843 Sources source rows… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-prime-sft-messages-env0315-clean50-primecompat.texttext-generation1K<n<10K0 likes21 downloads3mo agoHugging Face08jayshah5696 /humanize-rl-prime-sft-messages-env0314 Humanize-RL Prime SFT Messages Env0314 Prime prime-rl SFT dataset for Humanize-RL. Schema: each row has a messages list with one user instruction and one assistant target. Splits: train: 4313 validation: 239 test: 241 total accepted: 4793 rejected upstream by builder: 62 duplicate ids across published splits: 0 repair-reference rows: 20 Source artifact: v04_sft_final_plus_llama_failure_refs_env0314, built from restored v04 SFT data plus the clean Llama failure-reference repair… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-prime-sft-messages-env0314.texttext-generation1K<n<10K0 likes20 downloads4mo agoHugging Face09jayshah5696 /humanize-rl-tasks humanize-rl-tasks Single-turn RL task dataset for the humanize-rl project. Each row is one writing task. A model receives the prompt (instruction + source text), produces a completion, and the environment scores it with the 50/50 reward formula: reward = 0.50 × ridge_rubric_mean + 0.50 × deterministic_mean + penalties Dataset composition source tasks families input length v01_template 100 10 (template-generated) ~20 words v02_real_source 512 3 (real… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-tasks.texttext-generationn<1K0 likes17 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.