datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CMDcmd-audio-dumpCMDB-1500
CMDB-1500
Comprehensive Multimodal Decision Benchmark 1500, curated by SurdAI, contains 1,500 candidate-based decision questions: 1,200 text-only and 300 image-based, spanning 10 domains and 79 sources and subtasks.
The benchmark covers answer selection, action selection, binary judgments, ordinal ratings, and multi-select decisions in a common format.
Interactive leaderboard / 交互榜单 · Benchmark introduction / 基准介绍 · Example questions · Source catalog
Results and… See the full description on the dataset page: https://huggingface.co/datasets/SurdAI/CMDB-1500.roadmap-databasesfr3_pickplace_extended_new_cmd_SYNC_part_corrected_with_prompt_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "fr3",
"total_episodes": 301,
"total_frames": 276017,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:301"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yio-ye2004/fr3_pickplace_extended_new_cmd_SYNC_part_corrected_with_prompt_test.apple_banana_30hz_cmdThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "fr3",
"total_episodes": 262,
"total_frames": 109516,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:262"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yio-ye2004/apple_banana_30hz_cmd.Good-Quotes-AuthorsDynamic_HumanEvalZeroemgena_os_terminal_destructive_cmd_blocker_mcp_teaser
🚀 OS Security - Destructive Shell Command & Blast-Radius Blocker (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Scenarios + Executable MCP Server)🏆 Get the Full Production Package & Commercial EULA on Gumroad:👉 Purchase Full Package on Gumroad🏷️ Use coupon code LAUNCH20 for €20 off at checkout!
🌟 Domain Overview & Features
AST terminal command parser detecting and deterministically blocking rm -rf, disk formatting, reg delete, netsh… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_os_terminal_destructive_cmd_blocker_mcp_teaser.shell-cmd-instruct
Used to train models that interact directly with shells
Note: This dataset is out-dated in the llm world, probably easier to just setup a tool with a decent model that supports tooling.
Follow-up details of my process
MacOS terminal commands for now. This dataset is still in alpha stages and will be modified.
Contains 500 somewhat unique training examples so far.
GPT4 seems like a good candidate for generating more data, licensing would need to be addressed.
I fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/byroneverson/shell-cmd-instruct.hyperliquid-replica-cmdscmd-mtg_jamendo-metadatacmd-medleydb-metadatacmd-fma-metadataCMDS_Multimodal_Document
Dataset Card for Cyrillic Multimodel Document (CMDS)
This is the dataset consists of 3789 pairs of images and text across 31 categories downloaded from the Bulgarian ministry of finance
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
Uses this dataset for downstream task like Document Classification, Image Classification or Text Classification… See the full description on the dataset page: https://huggingface.co/datasets/sitloboi2012/CMDS_Multimodal_Document.swiss-sme-websites-2026
Swiss SME Websites 2026: 10,000 sites audited, canton by canton
Aggregated measurements of 10,000 websites of Swiss small and medium-sized
businesses, drawn at random from the Swiss commercial register in September
2026. Proportions only: no company name, no website address, no figure for a
group of fewer than 30 sites.
Study page (FR): https://hallebardier.com/etude-sites-pme-suisses/
Study page (EN): https://hallebardier.com/en/swiss-sme-website-study/
Study page (DE):… See the full description on the dataset page: https://huggingface.co/datasets/mp-cmd/swiss-sme-websites-2026.cmd-freesound-metadatacmd-t2m-datasetCMDR-Bench
Dataset Card for CMDR-Bench
CMDR-Bench is the first benchmark that jointly evaluates multi-page reasoning and multimodal document retrieval. It comprises 800 high-quality, human-annotated queries spanning four query categories and 255 documents across six domains, with an average document length of 183.5 pages.
Dataset Description
Task Definition
Given a query and a multi-page document, the task is to retrieve the top-k target pages that contain… See the full description on the dataset page: https://huggingface.co/datasets/NTT-hil-insight/CMDR-Bench.erebaur-auction-records
Erebaur Gem Auction Records
19 record auction sales of diamonds, rubies and a natural pearl, 2010 to 2023, from the records table in Chapter 14 of the Erebaur encyclopedia: price as announced and in US dollars, carat weight, computed price per carat, house, city and date. Prices include the buyer’s premium.
19 rows, one CSV file: auction-records.csv (UTF-8, comma-separated, header row).
Attribution
Required wherever the data is reused:
Data: Erebaur… See the full description on the dataset page: https://huggingface.co/datasets/mp-cmd/erebaur-auction-records.postcutoff-facts-qa
postcutoff-facts-qa
Closed-book QA over ~300 post-knowledge-cutoff facts (world events from
Wikipedia current events + ECB reference rates / index closes, June 2024 →
August 2026), built to verifiably measure whether fine-tuning teaches a model
new facts — and whether it destroys the model's calibration while doing so.
Companion to the adapter
evs-cmd/qwen2.5-1.5b-verifiable-facts-v8.
Every file was produced by a governed cairn pipeline run
(synthesis → dedup/PII hygiene →… See the full description on the dataset page: https://huggingface.co/datasets/evs-cmd/postcutoff-facts-qa.cmd-musicnet-metadataerebaur-gem-properties
Erebaur Gem Properties
Physical and optical properties of 60 gem materials: species, chemical formula, crystal system, Mohs hardness, refractive index, specific gravity and colors, with a verification status for every row. From the property table in Chapter 02 of the Erebaur encyclopedia.
60 rows, one CSV file: gem-properties.csv (UTF-8, comma-separated, header row).
Attribution
Required wherever the data is reused:
Data: Erebaur (https://erebaur.com), CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/mp-cmd/erebaur-gem-properties.erebaur-mining-localities
Erebaur Gem Mining Localities
57 gem mining localities across 31 countries, from the mine map in Chapter 05 of the Erebaur encyclopedia: diamond, corundum (ruby and sapphire), emerald and other gem deposits, with coordinates, the stones produced, operating status as of September 2026 and a coordinate-quality grade.
57 rows, one CSV file: mining-localities.csv (UTF-8, comma-separated, header row).
Attribution
Required wherever the data is reused:
Data: Erebaur… See the full description on the dataset page: https://huggingface.co/datasets/mp-cmd/erebaur-mining-localities.Famous_Quotescmd-moisesdb-metadataDynamic_MBPP_sanitizedcmdguard-dataset
cmdguard-dataset
Hand-labeled CLI commands for training a binary classifier: exploring (read-only) vs mutating (changes state).
Format
Chat-style JSONL — each line is:
{"messages": [{"role": "user", "content": "Classify: git status"}, {"role": "assistant", "content": "exploring"}]}
Splits
Split
Examples
Balance
train.jsonl
354
168 exploring / 168 mutating + 18 targeted
eval.jsonl
20
10 exploring / 10 mutating
Coverage
git, docker… See the full description on the dataset page: https://huggingface.co/datasets/qmxme/cmdguard-dataset.customer_support_cmd_gen_basicCMD-AD
CMD-AD (SLoMO Competition Release)
CMD-AD is the train + validation + private test release of the Condensed Movie
Dataset with Audio Descriptions, distributed for the SLoMO Competition at
the ECCV 2026 Workshop.
Contents
cmdad/
├── gt_annotations/
│ ├── cmdad_ad_train.csv # Train audio-description annotations (with text)
│ ├── cmdad_charbank_train.json # Character bank for train movies
│ ├── cmdad_ad_val.csv #… See the full description on the dataset page: https://huggingface.co/datasets/Jyxarthur/CMD-AD.
