Team Ai
Datasetpublic

triad-26/TRIAD

TRIAD: Benchmarking Omni-Modal Ambiguity in Multimodal Large Language Models TRIAD is a diagnostic benchmark for evaluating whether omni-modal models can resolve ambiguity by jointly integrating text, image, and audio. Overview TRIAD targets a failure mode common in multimodal evaluation: a model may appear to solve a multimodal task while actually relying on only one dominant modality. In TRIAD, each example is constructed so that the full tri-modal input is… See the full description on the dataset page: https://huggingface.co/datasets/triad-26/TRIAD.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
1likes68downloads
Dataset Card

TRIAD: Benchmarking Omni-Modal Ambiguity in Multimodal Large Language Models

![Task](#recommended-evaluation-protocol) ![Modalities](#data-format) ![Languages](#dataset-statistics) ![Size](#dataset-statistics) ![Review](#double-blind-review-notice)

TRIAD is a diagnostic benchmark for evaluating whether omni-modal models can resolve ambiguity by jointly integrating text, image, and audio.

Overview

TRIAD targets a failure mode common in multimodal evaluation: a model may appear to solve a multimodal task while actually relying on only one dominant modality. In TRIAD, each example is constructed so that the full tri-modal input is needed to identify the intended answer.

TaskMultiple-choice tri-modal question answering
ModalitiesText + Image + Audio
LanguagesEnglish, Chinese, and mixed-language examples
Instances327 items
Scenario groups86 groups
Answer formatoption_A, option_B, ...
Primary useHeld-out diagnostic evaluation
Review statusAnonymous version for double-blind review

Table of Contents

What Makes TRIAD Different?

TRIAD is designed around tri-modal irreducibility.

Input conditionExpected property
Text + Image + AudioThe intended answer should be identifiable.
Any two modalitiesThe answer should remain underdetermined or less well supported.
Any single modalityMultiple options should remain plausible.

The benchmark emphasizes audio cues that cannot be fully captured by a transcript, including:

  • —prosody and intonation;
  • —speaker-indexical cues;
  • —phonetic contrast and homophones;
  • —environmental sound;
  • —symbolic sound such as alarms, ringtones, and bells.

Repository Layout

text
TRIAD/
├── README.md
├── image/
│   ├── img_0001.jpg
│   ├── img_0002.jpg
│   └── ...
├── audio/
│   ├── aud_0001.wav
│   ├── aud_0002.wav
│   └── ...
├── metadata/
│   ├── attribution.json
│   └── statistics.json
└── all-modality_ambiguity.json
The exact layout may vary across release versions, but every item contains relative paths to its image and audio files.

Data Format

Each item is stored as a JSON object.

json
{
  "id": "group_1_1",
  "language": "en",
  "question": "When is the school meeting scheduled?",
  "context": "He told the parents to come to school for a meeting.",
  "image": "image/img_0122.jpg",
  "image_caption": "A picture of a calendar with Sunday circled.",
  "audio": "audio/aud_0110.wav",
  "audio_caption": "Someone said, 'Parents, please take note: remember to come to the school on this day.'",
  "options": {
    "option_A": "This sunday",
    "option_B": "Next week",
    "option_C": "The time is still unknown",
    "option_D": "There is no meeting"
  },
  "answer": "option_A",
  "category": {
    "level": "single_modal",
    "modal_category": {
      "text": "Anaphoric",
      "image": "Semiotic",
      "audio": "Referential"
    }
  }
}

<details> <summary><strong>Field description</strong></summary> | Field | Type | Description | | ------------------------------- | ------ | ------------------------------------------------------------ | | id | string | Unique item identifier. Items with the same prefix belong to the same scenario group. | | language | string | The language used in this data. | | question | string | The question to be answered. | | context | string | Textual context. | | image | string | Relative path to the image file. | | image_caption | string | Human-written image description. Metadata only unless caption-based evaluation is intended. | | audio | string | Relative path to the audio file. | | audio_caption | string | Human-written audio description or transcript-like summary. Metadata only unless transcript-based evaluation is intended. | | options | object | Multiple-choice answer options. | | answer | string | Gold answer key. | | category.level | string | Overall ambiguity level: single_modal, dual_modal, or tri_modal. | | category.modal_category.text | string | Text ambiguity category. | | category.modal_category.image | string | Image ambiguity category. | | category.modal_category.audio | string | Audio ambiguity category. |

</details>

Loading the Dataset

Load JSON metadata

python
import json
from pathlib import Path

data_path = Path("all-modality_ambiguity.json")

with data_path.open("r", encoding="utf-8") as f:
    items = json.load(f)

print(len(items))
print(items[0]["question"])

Resolve media paths

python
from pathlib import Path

root = Path(".")
item = items[0]

image_path = root / item["image"].replace("./", "")
audio_path = root / item["audio"].replace("./", "")

print(image_path)
print(audio_path)

Load metadata with Hugging Face Datasets

python
from datasets import load_dataset

dataset = load_dataset("json", data_files="all-modality_ambiguity.json")
print(dataset["train"][0])

Dataset Statistics

StatisticValue
Number of items327
Number of scenario groups86
Average group size3.8

Group-size Distribution

Group sizeNumber of groups
223
311
442
810

Ambiguity-level Distribution

Ambiguity levelNumber of items
single_modal230
dual_modal59
tri_modal38

Ambiguity Taxonomy

Every item receives one leaf-level ambiguity label for each modality.

Text Ambiguity Categories

CategoryCountDescription
Anaphoric66Pronouns, demonstratives, ellipsis, or unclear reference.
Syntactic64Grammatical structure, attachment, or segmentation ambiguity.
Logical61Rules, conditions, negation, quantification, or option mapping.
Pragmatic51Intention, implicature, politeness, sarcasm, or indirect speech.
Temporal48Event order, tense, aspect, or completion-state ambiguity.
Lexical37Word sense, polysemy, homonymy, or translation ambiguity.

Image Ambiguity Categories

CategoryCountDescription
Semiotic71Signs, gestures, symbols, expressions, or culturally mediated visual cues.
Grounding65Multiple possible visual referents or target objects.
Spatial61Position, orientation, depth, or spatial relation ambiguity.
Kinematic53A static frame of a dynamic action.
Observational41Viewpoint, occlusion, framing, or insufficient visual evidence.
Contextual36Uncertainty about scene type or situational context.

Audio Ambiguity Categories

CategoryCountDescription
Referential98Spoken references, pronouns, kinship terms, or deictic expressions.
Prosodic84Stress, intonation, rhythm, sarcasm, or tone.
Indexical54Speaker cues such as age, gender presentation, role, or social relationship.
Phonetic37Homophones, minimal pairs, segmentation, or acoustic confusability.
Environmental34Background sounds that change interpretation of the scene.
Symbolic20Non-linguistic auditory symbols such as alarms, bells, ringtones, or notifications.

Recommended Evaluation Protocol

The standard TRIAD evaluation provides the model with:

ComponentIncluded in standard evaluation
Textual contextYes
ImageYes
AudioYes
QuestionYes
Answer optionsYes

The model should output exactly one option key.

Modality Conditions

ConditionProvided input
FullText + Image + Audio
Text onlyText
Image onlyImage
Audio onlyAudio
Text + ImageText + Image
Text + AudioText + Audio
Image + AudioImage + Audio
Question-only baselineQuestion + Options

Example Prompt

text
You are answering a multiple-choice question that depends on three inputs:
a textual context, an image, and an audio clip.

Context: {context}   Image: {image}   Audio: {audio}

Question: {question}

Options:
(A) {option_A}  (C) {option_C}
(B) {option_B}  (D) {option_D}

Answer with a single option key.

Metrics

MetricDescription
Item-level accuracyStandard accuracy over all items.
Group-level accuracyAccuracy aggregated over scenario groups to reduce over-counting of sister examples.

Intended Uses

TRIAD is intended for:

  • —evaluating omni-modal large language models;
  • —testing whether models genuinely integrate text, image, and audio;
  • —diagnosing modality reliance and modality neglect;
  • —studying ambiguity resolution in multimodal reasoning;
  • —comparing full-modality performance with modality-ablated settings;
  • —analyzing which ambiguity types are hardest for different model families.

Out-of-Scope Uses

TRIAD is not intended for:

  • —training or fine-tuning large models;
  • —speaker identification;
  • —voice biometrics;
  • —voice cloning or impersonation;
  • —face recognition or identity inference;
  • —surveillance;
  • —profiling individuals;
  • —evaluating human subjects;
  • —making decisions about real individuals.
TRIAD is small and diagnostic by design. It should not be treated as a representative sample of real-world multimodal interactions.

Data Collection and Annotation

TRIAD was constructed by designing everyday scenarios in which text, image, and audio jointly determine the intended answer.

Each item is annotated with:

  • —a gold answer;
  • —an ambiguity level;
  • —one text ambiguity category;
  • —one image ambiguity category;
  • —one audio ambiguity category;
  • —group metadata linking sister examples.

Items were reviewed to reduce degenerate cases where the answer can be determined from only one or two modalities.

Responsible AI Considerations

<details open> <summary><strong>Privacy</strong></summary> The dataset is designed to avoid personally identifying information. Audio clips and images should not be used for identifying real individuals, voice matching, face recognition, or impersonation.

</details>

<details open> <summary><strong>Audio-specific risks</strong></summary> Because the dataset includes audio, it should not be used for biometric speaker recognition, voice cloning, speaker profiling, or identity inference. Audio is included only for evaluating multimodal reasoning.

</details>

<details open> <summary><strong>Sensitive content</strong></summary> The dataset avoids content involving violence, privacy violations, or sensitive personal information. Any remaining potentially sensitive cases should be treated as diagnostic examples only.

</details>

<details open> <summary><strong>Bias and limitations</strong></summary> TRIAD is a small diagnostic benchmark. Its examples are intentionally constructed around ambiguity and may not reflect the natural distribution of everyday multimodal data.

Model performance on TRIAD should therefore be interpreted as a measure of ambiguity-resolution ability rather than general multimodal competence.

The dataset contains English, Chinese, and mixed-language examples. Results may vary across languages and model families.

</details>

Licensing

The license metadata is set to MIT.

Citation

Citation information will be added after the double-blind review period.

bibtex
@misc{triad2026,
  title        = {TRIAD: Benchmarking Tri-Modal Irreducible Ambiguity Resolution for Multimodal Large Language Models},
  author       = {Anonymous},
  year         = {2026},
  note         = {Dataset submitted for double-blind review}
}

Maintenance

The dataset will be versioned. Errata, corrected annotations, and future extensions will be documented in the repository release history.

For the anonymous review version, contact information is withheld to preserve double-blind review. A public maintainer contact will be added after the review period.

Double-blind Review Notice

This repository is prepared for anonymous peer review. Please do not infer author identity from repository ownership, commit metadata, or temporary hosting information. Any non-anonymous information will be added after the review period.