triad-26/TRIAD
TRIAD: Benchmarking Omni-Modal Ambiguity in Multimodal Large Language Models TRIAD is a diagnostic benchmark for evaluating whether omni-modal models can resolve ambiguity by jointly integrating text, image, and audio. Overview TRIAD targets a failure mode common in multimodal evaluation: a model may appear to solve a multimodal task while actually relying on only one dominant modality. In TRIAD, each example is constructed so that the full tri-modal input is… See the full description on the dataset page: https://huggingface.co/datasets/triad-26/TRIAD.
TRIAD: Benchmarking Omni-Modal Ambiguity in Multimodal Large Language Models
    
TRIAD is a diagnostic benchmark for evaluating whether omni-modal models can resolve ambiguity by jointly integrating text, image, and audio.
Overview
TRIAD targets a failure mode common in multimodal evaluation: a model may appear to solve a multimodal task while actually relying on only one dominant modality. In TRIAD, each example is constructed so that the full tri-modal input is needed to identify the intended answer.
Table of Contents
- Overview
- What Makes TRIAD Different?
- Repository Layout
- Data Format
- Loading the Dataset
- Dataset Statistics
- Ambiguity Taxonomy
- Recommended Evaluation Protocol
- Intended Uses
- Out-of-Scope Uses
- Data Collection and Annotation
- Responsible AI Considerations
- Licensing
- Citation
- Maintenance
- Double-blind Review Notice
What Makes TRIAD Different?
TRIAD is designed around tri-modal irreducibility.
The benchmark emphasizes audio cues that cannot be fully captured by a transcript, including:
- prosody and intonation;
- speaker-indexical cues;
- phonetic contrast and homophones;
- environmental sound;
- symbolic sound such as alarms, ringtones, and bells.
Repository Layout
TRIAD/
├── README.md
├── image/
│ ├── img_0001.jpg
│ ├── img_0002.jpg
│ └── ...
├── audio/
│ ├── aud_0001.wav
│ ├── aud_0002.wav
│ └── ...
├── metadata/
│ ├── attribution.json
│ └── statistics.json
└── all-modality_ambiguity.jsonThe exact layout may vary across release versions, but every item contains relative paths to its image and audio files.
Data Format
Each item is stored as a JSON object.
{
"id": "group_1_1",
"language": "en",
"question": "When is the school meeting scheduled?",
"context": "He told the parents to come to school for a meeting.",
"image": "image/img_0122.jpg",
"image_caption": "A picture of a calendar with Sunday circled.",
"audio": "audio/aud_0110.wav",
"audio_caption": "Someone said, 'Parents, please take note: remember to come to the school on this day.'",
"options": {
"option_A": "This sunday",
"option_B": "Next week",
"option_C": "The time is still unknown",
"option_D": "There is no meeting"
},
"answer": "option_A",
"category": {
"level": "single_modal",
"modal_category": {
"text": "Anaphoric",
"image": "Semiotic",
"audio": "Referential"
}
}
}<details> <summary><strong>Field description</strong></summary> | Field | Type | Description | | ------------------------------- | ------ | ------------------------------------------------------------ | | id | string | Unique item identifier. Items with the same prefix belong to the same scenario group. | | language | string | The language used in this data. | | question | string | The question to be answered. | | context | string | Textual context. | | image | string | Relative path to the image file. | | image_caption | string | Human-written image description. Metadata only unless caption-based evaluation is intended. | | audio | string | Relative path to the audio file. | | audio_caption | string | Human-written audio description or transcript-like summary. Metadata only unless transcript-based evaluation is intended. | | options | object | Multiple-choice answer options. | | answer | string | Gold answer key. | | category.level | string | Overall ambiguity level: single_modal, dual_modal, or tri_modal. | | category.modal_category.text | string | Text ambiguity category. | | category.modal_category.image | string | Image ambiguity category. | | category.modal_category.audio | string | Audio ambiguity category. |
</details>
Loading the Dataset
Load JSON metadata
import json
from pathlib import Path
data_path = Path("all-modality_ambiguity.json")
with data_path.open("r", encoding="utf-8") as f:
items = json.load(f)
print(len(items))
print(items[0]["question"])Resolve media paths
from pathlib import Path
root = Path(".")
item = items[0]
image_path = root / item["image"].replace("./", "")
audio_path = root / item["audio"].replace("./", "")
print(image_path)
print(audio_path)Load metadata with Hugging Face Datasets
from datasets import load_dataset
dataset = load_dataset("json", data_files="all-modality_ambiguity.json")
print(dataset["train"][0])Dataset Statistics
Group-size Distribution
Ambiguity-level Distribution
Ambiguity Taxonomy
Every item receives one leaf-level ambiguity label for each modality.
Text Ambiguity Categories
Image Ambiguity Categories
Audio Ambiguity Categories
Recommended Evaluation Protocol
The standard TRIAD evaluation provides the model with:
The model should output exactly one option key.
Modality Conditions
Example Prompt
You are answering a multiple-choice question that depends on three inputs:
a textual context, an image, and an audio clip.
Context: {context} Image: {image} Audio: {audio}
Question: {question}
Options:
(A) {option_A} (C) {option_C}
(B) {option_B} (D) {option_D}
Answer with a single option key.Metrics
Intended Uses
TRIAD is intended for:
- evaluating omni-modal large language models;
- testing whether models genuinely integrate text, image, and audio;
- diagnosing modality reliance and modality neglect;
- studying ambiguity resolution in multimodal reasoning;
- comparing full-modality performance with modality-ablated settings;
- analyzing which ambiguity types are hardest for different model families.
Out-of-Scope Uses
TRIAD is not intended for:
- training or fine-tuning large models;
- speaker identification;
- voice biometrics;
- voice cloning or impersonation;
- face recognition or identity inference;
- surveillance;
- profiling individuals;
- evaluating human subjects;
- making decisions about real individuals.
TRIAD is small and diagnostic by design. It should not be treated as a representative sample of real-world multimodal interactions.
Data Collection and Annotation
TRIAD was constructed by designing everyday scenarios in which text, image, and audio jointly determine the intended answer.
Each item is annotated with:
- a gold answer;
- an ambiguity level;
- one text ambiguity category;
- one image ambiguity category;
- one audio ambiguity category;
- group metadata linking sister examples.
Items were reviewed to reduce degenerate cases where the answer can be determined from only one or two modalities.
Responsible AI Considerations
<details open> <summary><strong>Privacy</strong></summary> The dataset is designed to avoid personally identifying information. Audio clips and images should not be used for identifying real individuals, voice matching, face recognition, or impersonation.
</details>
<details open> <summary><strong>Audio-specific risks</strong></summary> Because the dataset includes audio, it should not be used for biometric speaker recognition, voice cloning, speaker profiling, or identity inference. Audio is included only for evaluating multimodal reasoning.
</details>
<details open> <summary><strong>Sensitive content</strong></summary> The dataset avoids content involving violence, privacy violations, or sensitive personal information. Any remaining potentially sensitive cases should be treated as diagnostic examples only.
</details>
<details open> <summary><strong>Bias and limitations</strong></summary> TRIAD is a small diagnostic benchmark. Its examples are intentionally constructed around ambiguity and may not reflect the natural distribution of everyday multimodal data.
Model performance on TRIAD should therefore be interpreted as a measure of ambiguity-resolution ability rather than general multimodal competence.
The dataset contains English, Chinese, and mixed-language examples. Results may vary across languages and model families.
</details>
Licensing
The license metadata is set to MIT.
Citation
Citation information will be added after the double-blind review period.
@misc{triad2026,
title = {TRIAD: Benchmarking Tri-Modal Irreducible Ambiguity Resolution for Multimodal Large Language Models},
author = {Anonymous},
year = {2026},
note = {Dataset submitted for double-blind review}
}Maintenance
The dataset will be versioned. Errata, corrected annotations, and future extensions will be documented in the repository release history.
For the anonymous review version, contact information is withheld to preserve double-blind review. A public maintainer contact will be added after the review period.
Double-blind Review Notice
This repository is prepared for anonymous peer review. Please do not infer author identity from repository ownership, commit metadata, or temporary hosting information. Any non-anonymous information will be added after the review period.
