Lelonthecodeur/humanizer-5m
Humanizer-5M Version: 2.0.0 Humanizer-5M is a synthetic conversational adaptation dataset. The central objective is not simply to rewrite text to sound casual. Each example models: what the user explicitly asks, what the user may implicitly need, the user's conversational signals, the appropriate response calibration, the final response, quality and stability metrics. The dataset specifically teaches proportional adaptation. High user energy does not automatically mean high… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/humanizer-5m.
Humanizer-5M
Version: 2.0.0
Humanizer-5M is a synthetic conversational adaptation dataset.
The central objective is not simply to rewrite text to sound casual.
Each example models:
- what the user explicitly asks,
- what the user may implicitly need,
- the user's conversational signals,
- the appropriate response calibration,
- the final response,
- quality and stability metrics.
The dataset specifically teaches proportional adaptation.
High user energy does not automatically mean high hype. Frustration does not automatically mean excessive empathy. Casual language does not automatically require slang. Emojis are used only when contextually appropriate.
The dataset also contains hard negatives representing responses that are grammatically correct but poorly calibrated.
Dataset size
Total: 5,000,000
Train: 4,750,000 Validation: 125,000 Test: 125,000
Shard size: 100,000
Core dimensions
The dataset includes explicit calibration dimensions for:
- energy
- hype
- empathy
- warmth
- formality
- directness
- verbosity
- humor
- emoji usage
- technicality
- confidence
- urgency
Quality metrics
Examples also contain:
- naturalness
- humanlikenessproxy
- context_fit
- emotional_fit
- hype_fit
- empathy_fit
- emoji_fit
- verbosity_fit
- directness_fit
- formality_fit
- humor_fit
- technicality_fit
- consistency
- emotional_stability
- overreaction_penalty
- forcedhumanpenalty
- genericaipenalty
- repetition_penalty
- emojioverusepenalty
- hypewithoutreason_penalty
- tonejumppenalty
- overall_quality
Hard negatives
The dataset includes deliberately miscalibrated candidate responses.
Examples include:
- unnecessary hype
- forced empathy
- emoji overuse
- excessive verbosity
- generic assistant phrasing
- abrupt tone changes
Important
Metrics are synthetic training labels and calibration signals. They are not empirical measurements of whether a response was written by a human.
The dataset does not contain private conversations.
