Team Ai
Datasetpublic

HumanBehaviorAtlas/human_behavior_atlas

Human Behavior Atlas A large-scale multimodal dataset for human behavior understanding, spanning emotion recognition, sentiment analysis, humor detection, mental health screening, and video question answering. The dataset integrates 16 source datasets into a unified schema with audio, video, and pre-extracted features. This dataset was used to train OmniSapiens, a foundation model for social behavior processing. Papers: Human Behavior Atlas: Benchmarking Unified Psychological… See the full description on the dataset page: https://huggingface.co/datasets/HumanBehaviorAtlas/human_behavior_atlas.

sourceHugging Facecc-by-nc-4.0updated 4mo agoView on Hugging Face
3likes5.6kdownloads
Dataset Card

Human Behavior Atlas

A large-scale multimodal dataset for human behavior understanding, spanning emotion recognition, sentiment analysis, humor detection, mental health screening, and video question answering. The dataset integrates 16 source datasets into a unified schema with audio, video, and pre-extracted features.

This dataset was used to train OmniSapiens, a foundation model for social behavior processing.

Dataset Summary

PropertyValue
Total samples100,299
Train split74,449
Validation split7,646
Test split18,204
Source datasets16
ModalitiesText, Audio (.wav bytes), Video (.mp4 bytes), OpenSmile features (.pt bytes), Pose features (.pt bytes) — all embedded in parquet
LanguagesEnglish, Chinese (CHSIMSv2)
LicenseCC BY-NC 4.0

Modality Distribution

Modality SignatureSamplesPercentage
textvideoaudio87,31887.1%
text_audio10,43110.4%
text2,5502.5%

Source Datasets

DatasetSamplesTaskModalityDescription
mosei_senti22,740Sentiment classificationtextvideoaudioCMU-MOSEI sentiment analysis (negative/neutral/positive)
intentqa14,158Video QAtextvideoaudioIntent-driven video question answering
meld_senti13,518Sentiment classificationtextvideoaudioMELD multimodal sentiment (from Friends TV series)
meld_emotion13,350Emotion classificationtextvideoaudioMELD multimodal emotion recognition (7 classes)
mosei_emotion8,545Emotion classificationtextvideoaudioCMU-MOSEI emotion recognition (6 classes)
cremad7,442Emotion classificationtext_audioCREMA-D acted emotional speech recognition
siq26,394Video QAtextvideoaudioSocial IQ 2.0 social intelligence QA
chsimsv24,384Sentiment classificationtextvideoaudioCH-SIMS v2 Chinese multimodal sentiment
tess2,800Emotion classificationtext_audioToronto Emotional Speech Set
urfunny2,113Humor classificationtextvideoaudioUR-Funny multimodal humor detection
mmpsy_depression1,275Depression screeningtextvideoaudioMultimodal depression assessment
mmpsy_anxiety1,275Anxiety screeningtextvideoaudioMultimodal anxiety assessment
mimeqa801Video QAtextvideoaudioMIME gesture-based QA
mmsd687Humor classificationtextMultimodal sarcasm detection (text only)
ptsd_in_the_wild628PTSD detectiontextvideoaudioPTSD detection from video interviews
daicwoz189Depression screeningtextvideoaudioDAIC-WOZ clinical depression interviews

Task Types

Task IDDescriptionDatasets
emotion_clsEmotion classificationmoseiemotion, meldemotion, cremad, tess
sentiment_clsSentiment classification / regressionmoseisenti, meldsenti, chsimsv2
humor_clsHumor and sarcasm detectionurfunny, mmsd
depressionDepression screeningmmpsy_depression, daicwoz
anxietyAnxiety screeningmmpsy_anxiety
ptsdPTSD detectionptsdinthe_wild
video_qaVideo question answeringintentqa, siq2, mimeqa

Schema

Each row in the Parquet files contains the following columns:

ColumnTypeDescription
problemstringPrompt text with modality markers (<audio>, <video>)
answerstringGround truth answer
audioslist[bytes]Raw .wav audio bytes (embedded)
videoslist[bytes]Raw .mp4 video bytes (embedded)
imageslist[bytes]Image bytes (currently unused)
datasetstringSource dataset name
modality_signaturestringModality combination: text_video_audio, text_audio, or text
ext_video_featslist[bytes]Pose estimation feature tensors (.pt bytes, embedded)
ext_audio_featslist[bytes]OpenSmile audio feature tensors (.pt bytes, embedded)
taskstringTask type identifier
class_labelstringClassification label

Usage

Loading with HuggingFace Datasets

python
from datasets import load_dataset

# Stream without downloading everything
ds = load_dataset("HumanBehaviorAtlas/human_behavior_atlas", split="train", streaming=True)
sample = next(iter(ds))

# Load a subset
ds_100 = load_dataset("HumanBehaviorAtlas/human_behavior_atlas", split="train[:100]")

Accessing Embedded Media

python
import io
import soundfile as sf

sample = ds_100[0]

# Audio is raw bytes — decode with soundfile or torchaudio
if sample["audios"]:
    audio_data, sr = sf.read(io.BytesIO(sample["audios"][0]))

# Video is raw bytes — decode with decord, opencv, or write to temp file
if sample["videos"]:
    video_bytes = sample["videos"][0]

Example Entry

json
{
  "problem": "<audio>
Don't forget a jacket.
The above is a speech recording along with the transcript from a clinical context. What emotion is the speaker expressing? Answer with one word from the following: anger, disgust, fear, happy, neutral, sad",
  "answer": "sad",
  "images": [],
  "videos": [],
  "audios": ["..."],
  "dataset": "cremad",
  "modality_signature": "text_audio",
  "task": "emotion_cls",
  "class_label": "sad"
}

Citation

bibtex
@inproceedings{
ong2026human,
title={Human Behavior Atlas: Benchmarking Unified Psychological And Social Behavior Understanding},
author={Keane Ong and Wei Dai and Carol Li and Dewei Feng and Hengzhi Li and Jingyao Wu and Jiaee Cheong and Rui Mao and Gianmarco Mengaldo and Erik Cambria and Paul Pu Liang},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=ZKE23BBvlQ}
}

@article{ong2026omnisapiens,
  title={Omnisapiens: A foundation model for social behavior processing via heterogeneity-aware relative policy optimization},
  author={Ong, Keane and Boughorbel, Sabri and Xiao, Luwei and Ekbote, Chanakya and Dai, Wei and Qu, Ao and Wu, Jingyao and Mao, Rui and Hoque, Ehsan and Cambria, Erik and others},
  journal={arXiv preprint arXiv:2602.10635},
  year={2026}
}

License

This dataset is released under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license. Individual source datasets may have their own licensing terms.