Team Ai
Datasetpublic

ClaudiaRichard/mbti_classification_dataset_fullPosts

MBTI Classification Dataset (Full Posts) A dataset of 8,675 personality-forum posts labeled with Myers-Briggs Type Indicator (MBTI) dichotomies, built for training and evaluating text-based personality classification models. Each row contains one user's concatenated forum posts plus binary labels for all four MBTI dimensions. Dataset Structure Splits: train (5,205 rows), test (2,082 rows), validation (1,388 rows) Fields: Field Type Description I/E int64… See the full description on the dataset page: https://huggingface.co/datasets/ClaudiaRichard/mbti_classification_dataset_fullPosts.

sourceHugging Facecc-by-4.0updated 3d agoView on Hugging Face
1likes583downloads
Dataset Card

MBTI Classification Dataset (Full Posts)

A dataset of 8,675 personality-forum posts labeled with Myers-Briggs Type Indicator (MBTI) dichotomies, built for training and evaluating text-based personality classification models. Each row contains one user's concatenated forum posts plus binary labels for all four MBTI dimensions.

Dataset Structure

Splits: train (5,205 rows), test (2,082 rows), validation (1,388 rows)

Fields:

FieldTypeDescription
I/Eint64Introversion (0) vs Extraversion (1)
N/Sint64Intuition (0) vs Sensing (1)
T/Fint64Thinking (0) vs Feeling (1)
J/Pint64Judging (0) vs Perceiving (1)
poststringThe user's forum posts, concatenated and separated by `\\\`

Source and Collection

Posts were collected from a public personality-discussion forum, where users self-report their MBTI type. Each user's posts were concatenated into a single text field (separated by |||), giving models a fuller picture of each author's writing style than single-post samples.

Intended Uses

  • —Training text classifiers that predict MBTI dimensions from writing
  • —Benchmarking fine-tuned language models (e.g. BERT) on personality classification
  • —Interpretability research on what linguistic signals models associate with personality traits

Considerations and Limitations

  • —Labels are self-reported. MBTI types were declared by the users themselves, not professionally assessed, so label noise is expected.
  • —Forum-specific language. Posts come from a personality-enthusiast community; the vocabulary (type jargon, memes) may not transfer to general text.
  • —Class imbalance. Some MBTI types (notably intuitive types) are overrepresented relative to the general population, as is typical of personality forums.
  • —MBTI validity. The MBTI framework itself is debated in psychology; treat this as a text-style classification task rather than a clinical instrument.

Citation

If you use this dataset, please cite:

bibtex
@misc{richard2024mbti,
  author       = {Richard, Claudia Lois},
  title        = {{MBTI} Classification Dataset (Full Posts)},
  year         = {2024},
  publisher    = {Hugging Face},
  doi          = {10.57967/hf/10813},
  howpublished = {\url{https://huggingface.co/datasets/ClaudiaRichard/mbti_classification_dataset_fullPosts}}
}