HumanBehaviorAtlas/human_behavior_atlas
Human Behavior Atlas A large-scale multimodal dataset for human behavior understanding, spanning emotion recognition, sentiment analysis, humor detection, mental health screening, and video question answering. The dataset integrates 16 source datasets into a unified schema with audio, video, and pre-extracted features. This dataset was used to train OmniSapiens, a foundation model for social behavior processing. Papers: Human Behavior Atlas: Benchmarking Unified Psychological… See the full description on the dataset page: https://huggingface.co/datasets/HumanBehaviorAtlas/human_behavior_atlas.
Human Behavior Atlas
A large-scale multimodal dataset for human behavior understanding, spanning emotion recognition, sentiment analysis, humor detection, mental health screening, and video question answering. The dataset integrates 16 source datasets into a unified schema with audio, video, and pre-extracted features.
This dataset was used to train OmniSapiens, a foundation model for social behavior processing.
- Papers:
- Human Behavior Atlas: Benchmarking Unified Psychological and Social Behavior Understanding (ICLR 2026)
- OmniSapiens: A Foundation Model for Social Behavior Processing via Heterogeneity-Aware Relative Policy Optimization (ICML 2026)
- Repository: https://github.com/MIT-MI/human_behavior_atlas
Dataset Summary
Modality Distribution
Source Datasets
Task Types
Schema
Each row in the Parquet files contains the following columns:
Usage
Loading with HuggingFace Datasets
from datasets import load_dataset
# Stream without downloading everything
ds = load_dataset("HumanBehaviorAtlas/human_behavior_atlas", split="train", streaming=True)
sample = next(iter(ds))
# Load a subset
ds_100 = load_dataset("HumanBehaviorAtlas/human_behavior_atlas", split="train[:100]")Accessing Embedded Media
import io
import soundfile as sf
sample = ds_100[0]
# Audio is raw bytes — decode with soundfile or torchaudio
if sample["audios"]:
audio_data, sr = sf.read(io.BytesIO(sample["audios"][0]))
# Video is raw bytes — decode with decord, opencv, or write to temp file
if sample["videos"]:
video_bytes = sample["videos"][0]Example Entry
{
"problem": "<audio>
Don't forget a jacket.
The above is a speech recording along with the transcript from a clinical context. What emotion is the speaker expressing? Answer with one word from the following: anger, disgust, fear, happy, neutral, sad",
"answer": "sad",
"images": [],
"videos": [],
"audios": ["..."],
"dataset": "cremad",
"modality_signature": "text_audio",
"task": "emotion_cls",
"class_label": "sad"
}Citation
@inproceedings{
ong2026human,
title={Human Behavior Atlas: Benchmarking Unified Psychological And Social Behavior Understanding},
author={Keane Ong and Wei Dai and Carol Li and Dewei Feng and Hengzhi Li and Jingyao Wu and Jiaee Cheong and Rui Mao and Gianmarco Mengaldo and Erik Cambria and Paul Pu Liang},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=ZKE23BBvlQ}
}
@article{ong2026omnisapiens,
title={Omnisapiens: A foundation model for social behavior processing via heterogeneity-aware relative policy optimization},
author={Ong, Keane and Boughorbel, Sabri and Xiao, Luwei and Ekbote, Chanakya and Dai, Wei and Qu, Ao and Wu, Jingyao and Mao, Rui and Hoque, Ehsan and Cambria, Erik and others},
journal={arXiv preprint arXiv:2602.10635},
year={2026}
}License
This dataset is released under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license. Individual source datasets may have their own licensing terms.
