SurdAI/CMDB-1500
CMDB-1500 Comprehensive Multimodal Decision Benchmark 1500, curated by SurdAI, contains 1,500 candidate-based decision questions: 1,200 text-only and 300 image-based, spanning 10 domains and 79 sources and subtasks. The benchmark covers answer selection, action selection, binary judgments, ordinal ratings, and multi-select decisions in a common format. Interactive leaderboard / 交互榜单 · Benchmark introduction / 基准介绍 · Example questions · Source catalog Results and… See the full description on the dataset page: https://huggingface.co/datasets/SurdAI/CMDB-1500.
CMDB-1500
Comprehensive Multimodal Decision Benchmark 1500, curated by SurdAI, contains 1,500 candidate-based decision questions: 1,200 text-only and 300 image-based, spanning 10 domains and 79 sources and subtasks.
The benchmark covers answer selection, action selection, binary judgments, ordinal ratings, and multi-select decisions in a common format.
Interactive leaderboard / 交互榜单 · Benchmark introduction / 基准介绍 · Example questions · Source catalog
Results and models
Explore the interactive CMDB-1500 leaderboard on the Surd AI website. It separates text (1,200 questions), image (300 questions), and full-benchmark (1,500 questions) accuracy, with probability calibration reported separately. Text-only models do not receive an image or full-benchmark score.
Selected Surd AI results use effort 2 and the same 1,500-question benchmark:
These are team-run evaluations, not official upstream leaderboard submissions. SPX-CD-Omni's reported results use the released LoRA adapter with an AWQ INT4 base; they are not measurements of the default BF16 example configuration. The model card and evaluation settings describe the configuration and metric scopes.
The website's cross-model I/C comparison uses a shared text-probability cohort; it is distinct from the model card's benchmark-specific I/C scopes. Compare I/C values only within the same scoring scope. See the SPX-CD release and API overview for Pro and Flash.
Coverage
The image subset contains 322 images; some questions use more than one image.
Counts follow the kind field. choice includes two-option questions; multi_choice permits one or more correct answers. Ordered ratings cover helpfulness, verbosity, similarity, hate-speech levels, and game-state assessments. Rating direction follows the question.
Sources
- Language and judgment: 19 JevBench configurations, including Banking77, ARC-Challenge, BoolQ, CLINC, MASSIVE, MMLU, ChaosNLI, MNLI, SST-5, and HelpSteer2.
- Professional decisions: 31 Atlan Decision Bench tasks covering engineering, agents, commerce, finance, law, data, documents, product, support, and safety.
- Multimodal decisions: 30 questions each from AI2D, Fashion200k, GQA, MMMU, MMMU-Pro, MathVista, ScienceQA, TextVQA, VQAv2, and VizWiz.
- Other domains: Open-Jev games, MultiRC, GoEmotions, ChineseHP, Aegis, PhishNChips, JevAdvBench, BFCL v3, THIS-THAT, ALFRED, ALFWorld, and VirtualHome.
SOURCES.md lists all question counts, upstream links, and task definitions. Subtasks may share an upstream dataset.
Load the dataset
from datasets import load_dataset
bench = load_dataset("SurdAI/CMDB-1500", split="test")
item = bench[0]
model_input = {
key: item[key]
for key in ("state", "question", "options", "kind", "images")
}Data format
The test split is available as Parquet with embedded images and JSONL. JSONL image paths are relative to the dataset directory.
Parse state when state_format is json. target_probs and source_metadata are JSON-serialized strings. Candidate keys are specific to each question. Use the model-input fields shown above; reference answers and provenance are for evaluation.
Evaluation
Use accuracy for categorical choices, including binary questions, exact-set match and set F1 for multi-select, and accuracy plus level error for ordinal ratings. Calibration metrics use valid model probability outputs and the stated reference labels; report the included sample count and the binning rule. Auxiliary target distributions are not model confidence predictions.
Report domain scores and text/image results alongside an overall score, specifying the aggregation weights. Retain the full scope denominator when a native model interface cannot represent a question; report such failures separately. Text-only evaluations cover 1,200 questions and must not be presented as full 1,500-question results.
These are fixed-input decisions. Game and embodied questions assess state/action selection; BFCL questions assess tool-call necessity; Chinese ASR questions use text hypotheses. GQA, VQAv2, TextVQA, and VizWiz use yes/no subsets. Scores describe these selected tasks rather than complete upstream benchmark evaluations.
Attribution and licensing
LICENSES.md contains source-specific terms and required attribution, including HelpSteer2. ChaosNLI is CC BY-NC 4.0, with attribution and noncommercial conditions; ScienceQA also has noncommercial conditions. MultiNLI permits modification and redistribution for typical machine-learning uses, with underlying source terms retained. An explicit redistribution license for SST-5 has not been identified in the inspected original release documentation. MathVista permits test-set use only. Upstream terms apply to the corresponding records.
Please cite CMDB-1500, SurdAI, together with the upstream datasets used.
