Team Ai
Datasetpublic

Lelonthecodeur/humanizer-5m

Humanizer-5M Version: 2.0.0 Humanizer-5M is a synthetic conversational adaptation dataset. The central objective is not simply to rewrite text to sound casual. Each example models: what the user explicitly asks, what the user may implicitly need, the user's conversational signals, the appropriate response calibration, the final response, quality and stability metrics. The dataset specifically teaches proportional adaptation. High user energy does not automatically mean high… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/humanizer-5m.

sourceHugging Faceupdated 22d agoView on Hugging Face
2likes405downloads
Dataset Card

Humanizer-5M

Version: 2.0.0

Humanizer-5M is a synthetic conversational adaptation dataset.

The central objective is not simply to rewrite text to sound casual.

Each example models:

  1. 1.what the user explicitly asks,
  2. 2.what the user may implicitly need,
  3. 3.the user's conversational signals,
  4. 4.the appropriate response calibration,
  5. 5.the final response,
  6. 6.quality and stability metrics.

The dataset specifically teaches proportional adaptation.

High user energy does not automatically mean high hype. Frustration does not automatically mean excessive empathy. Casual language does not automatically require slang. Emojis are used only when contextually appropriate.

The dataset also contains hard negatives representing responses that are grammatically correct but poorly calibrated.

Dataset size

Total: 5,000,000

Train: 4,750,000 Validation: 125,000 Test: 125,000

Shard size: 100,000

Core dimensions

The dataset includes explicit calibration dimensions for:

  • —energy
  • —hype
  • —empathy
  • —warmth
  • —formality
  • —directness
  • —verbosity
  • —humor
  • —emoji usage
  • —technicality
  • —confidence
  • —urgency

Quality metrics

Examples also contain:

  • —naturalness
  • —humanlikenessproxy
  • —context_fit
  • —emotional_fit
  • —hype_fit
  • —empathy_fit
  • —emoji_fit
  • —verbosity_fit
  • —directness_fit
  • —formality_fit
  • —humor_fit
  • —technicality_fit
  • —consistency
  • —emotional_stability
  • —overreaction_penalty
  • —forcedhumanpenalty
  • —genericaipenalty
  • —repetition_penalty
  • —emojioverusepenalty
  • —hypewithoutreason_penalty
  • —tonejumppenalty
  • —overall_quality

Hard negatives

The dataset includes deliberately miscalibrated candidate responses.

Examples include:

  • —unnecessary hype
  • —forced empathy
  • —emoji overuse
  • —excessive verbosity
  • —generic assistant phrasing
  • —abrupt tone changes

Important

Metrics are synthetic training labels and calibration signals. They are not empirical measurements of whether a response was written by a human.

The dataset does not contain private conversations.