Team Ai
Datasetpublic

SurdAI/CMDB-1500

CMDB-1500 Comprehensive Multimodal Decision Benchmark 1500, curated by SurdAI, contains 1,500 candidate-based decision questions: 1,200 text-only and 300 image-based, spanning 10 domains and 79 sources and subtasks. The benchmark covers answer selection, action selection, binary judgments, ordinal ratings, and multi-select decisions in a common format. Interactive leaderboard / 交互榜单 · Benchmark introduction / 基准介绍 · Example questions · Source catalog Results and… See the full description on the dataset page: https://huggingface.co/datasets/SurdAI/CMDB-1500.

sourceHugging Faceotherupdated 5d agoView on Hugging Face
2likes375downloads
Dataset Card

CMDB-1500

Comprehensive Multimodal Decision Benchmark 1500, curated by SurdAI, contains 1,500 candidate-based decision questions: 1,200 text-only and 300 image-based, spanning 10 domains and 79 sources and subtasks.

The benchmark covers answer selection, action selection, binary judgments, ordinal ratings, and multi-select decisions in a common format.

Interactive leaderboard / 交互榜单 · Benchmark introduction / 基准介绍 · Example questions · Source catalog

[image]

Results and models

Explore the interactive CMDB-1500 leaderboard on the Surd AI website. It separates text (1,200 questions), image (300 questions), and full-benchmark (1,500 questions) accuracy, with probability calibration reported separately. Text-only models do not receive an image or full-benchmark score.

Selected Surd AI results use effort 2 and the same 1,500-question benchmark:

ModelText accuracy ↑Image accuracy ↑Full-benchmark accuracy ↑
SPX-CD-Pro79.33%88.67%81.20%
SPX-CD-Flash77.58%87.33%79.53%
SPX-CD-Omni74.67%82.67%76.27%

These are team-run evaluations, not official upstream leaderboard submissions. SPX-CD-Omni's reported results use the released LoRA adapter with an AWQ INT4 base; they are not measurements of the default BF16 example configuration. The model card and evaluation settings describe the configuration and metric scopes.

The website's cross-model I/C comparison uses a shared text-probability cohort; it is distinct from the model card's benchmark-specific I/C scopes. Compare I/C values only within the same scoring scope. See the SPX-CD release and API overview for Pro and Flash.

Coverage

[image]

DomainQuestionsTasks
Language, knowledge, and judgment450Science, intent classification, inference, evidence, paraphrase similarity, and response assessment
Professional and applied decisions300Engineering, agents, commerce, finance, law, data, documents, product, support, and safety
Multimodal decisions300Diagram understanding, visual reasoning, science questions, image classification, and visual question answering
Games100Static game states and action selection across seven subtasks
Multi-select100MultiRC and GoEmotions
Safety and adversarial decisions100Content safety, phishing, and adversarial instruction handling
Embodied action selection50ALFRED, ALFWorld, and VirtualHome
Chinese ASR hypothesis selection50Selection among text transcription hypotheses
Tool-call necessity25BFCL-derived judgments of whether a tool call is needed
Spatial and complex rules25Spatial relations and rule-based decisions
Total1,5001,200 text-only + 300 image-based

The image subset contains 322 images; some questions use more than one image.

Record format`kind`QuestionsExpected prediction
Choicechoice1,077One candidate key
Ordered ratingscore166One key on an ordered scale
Binary formatnoul157One key from the two supplied candidates
Multi-selectmulti_choice100A set of candidate keys

Counts follow the kind field. choice includes two-option questions; multi_choice permits one or more correct answers. Ordered ratings cover helpfulness, verbosity, similarity, hate-speech levels, and game-state assessments. Rating direction follows the question.

Sources

  • —Language and judgment: 19 JevBench configurations, including Banking77, ARC-Challenge, BoolQ, CLINC, MASSIVE, MMLU, ChaosNLI, MNLI, SST-5, and HelpSteer2.
  • —Professional decisions: 31 Atlan Decision Bench tasks covering engineering, agents, commerce, finance, law, data, documents, product, support, and safety.
  • —Multimodal decisions: 30 questions each from AI2D, Fashion200k, GQA, MMMU, MMMU-Pro, MathVista, ScienceQA, TextVQA, VQAv2, and VizWiz.
  • —Other domains: Open-Jev games, MultiRC, GoEmotions, ChineseHP, Aegis, PhishNChips, JevAdvBench, BFCL v3, THIS-THAT, ALFRED, ALFWorld, and VirtualHome.

SOURCES.md lists all question counts, upstream links, and task definitions. Subtasks may share an upstream dataset.

Load the dataset

python
from datasets import load_dataset

bench = load_dataset("SurdAI/CMDB-1500", split="test")
item = bench[0]
model_input = {
    key: item[key]
    for key in ("state", "question", "options", "kind", "images")
}

Data format

The test split is available as Parquet with embedded images and JSONL. JSONL image paths are relative to the dataset directory.

FieldsContents
state, state_formatContext as text or serialized JSON
question, optionsDecision instruction and candidate keys/texts
kind, modality, imagesOutput type and image inputs
correct_keys, target_probsReference answers and optional auxiliary targets
id, source, categoryQuestion identifier, source, and domain
source_id, original_split, source_metadata, label_originSource provenance and reference-label basis

Parse state when state_format is json. target_probs and source_metadata are JSON-serialized strings. Candidate keys are specific to each question. Use the model-input fields shown above; reference answers and provenance are for evaluation.

Evaluation

Use accuracy for categorical choices, including binary questions, exact-set match and set F1 for multi-select, and accuracy plus level error for ordinal ratings. Calibration metrics use valid model probability outputs and the stated reference labels; report the included sample count and the binning rule. Auxiliary target distributions are not model confidence predictions.

Report domain scores and text/image results alongside an overall score, specifying the aggregation weights. Retain the full scope denominator when a native model interface cannot represent a question; report such failures separately. Text-only evaluations cover 1,200 questions and must not be presented as full 1,500-question results.

These are fixed-input decisions. Game and embodied questions assess state/action selection; BFCL questions assess tool-call necessity; Chinese ASR questions use text hypotheses. GQA, VQAv2, TextVQA, and VizWiz use yes/no subsets. Scores describe these selected tasks rather than complete upstream benchmark evaluations.

Attribution and licensing

LICENSES.md contains source-specific terms and required attribution, including HelpSteer2. ChaosNLI is CC BY-NC 4.0, with attribution and noncommercial conditions; ScienceQA also has noncommercial conditions. MultiNLI permits modification and redistribution for typical machine-learning uses, with underlying source terms retained. An explicit redistribution license for SST-5 has not been identified in the inspected original release documentation. MathVista permits test-set use only. Upstream terms apply to the corresponding records.

Please cite CMDB-1500, SurdAI, together with the upstream datasets used.