datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CMDcmd-audio-dumpCMDB-1500
CMDB-1500
Comprehensive Multimodal Decision Benchmark 1500, curated by SurdAI, contains 1,500 candidate-based decision questions: 1,200 text-only and 300 image-based, spanning 10 domains and 79 sources and subtasks.
The benchmark covers answer selection, action selection, binary judgments, ordinal ratings, and multi-select decisions in a common format.
Interactive leaderboard / 交互榜单 · Benchmark introduction / 基准介绍 · Example questions · Source catalog
Results and… See the full description on the dataset page: https://huggingface.co/datasets/SurdAI/CMDB-1500.fr3_pickplace_extended_new_cmd_SYNC_part_correctedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "fr3",
"total_episodes": 301,
"total_frames": 276017,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:301"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yio-ye2004/fr3_pickplace_extended_new_cmd_SYNC_part_corrected.roadmap-databasesfr3_pickplace_extended_abs_cmd_TEST2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "fr3",
"total_episodes": 31,
"total_frames": 28396,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:31"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yio-ye2004/fr3_pickplace_extended_abs_cmd_TEST2.fr3_pickplace_extended_new_cmd_SYNC_part_corrected_with_prompt_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "fr3",
"total_episodes": 301,
"total_frames": 276017,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:301"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yio-ye2004/fr3_pickplace_extended_new_cmd_SYNC_part_corrected_with_prompt_test.astra_grab_floor_toys_smoothed_base_cmdThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 50,
"total_frames": 73694,
"total_tasks": 1,
"total_videos": 150,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lookas/astra_grab_floor_toys_smoothed_base_cmd.hass_avocadoThis dataset is a huggingface upload for https://data.mendeley.com/datasets/3xd9n945v8/1
Description from their website accessed on 2nd July 2025
This dataset consists of 14,710 labeled photographs of Hass avocados (Persea Americana Mill. cv Hass), resized to 800 x 800 pixels and saved in the .jpg format, designed to facilitate the development of deep learning models for predicting ripening stages and estimating shelf-life.
A total of 478 Hass avocados were acquired three days post-harvest and… See the full description on the dataset page: https://huggingface.co/datasets/c2p-cmd/hass_avocado.kingshot_screenshotsastra_grab_floor_toys_extended_smoothed_base_cmdThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "astra_joint",
"total_episodes": 80,
"total_frames": 113547,
"total_tasks": 1,
"total_videos": 240,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:80"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lookas/astra_grab_floor_toys_extended_smoothed_base_cmd.apple_banana_30hz_cmdThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "fr3",
"total_episodes": 262,
"total_frames": 109516,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:262"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yio-ye2004/apple_banana_30hz_cmd.astra_grab_floor_toys_base_cmd_posThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 50,
"total_frames": 73694,
"total_tasks": 1,
"total_videos": 150,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lookas/astra_grab_floor_toys_base_cmd_pos.astra_grab_floor_toys_extended_base_cmd_posThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "astra_joint",
"total_episodes": 80,
"total_frames": 113547,
"total_tasks": 1,
"total_videos": 240,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:80"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lookas/astra_grab_floor_toys_extended_base_cmd_pos.Good-Quotes-AuthorsDynamic_HumanEvalZerocmdbench-nba
cmdbench-nba
CMDBench-NBA is a multimodal NBA datalake that consists of a Neo4j knowledge graph, a PostgreSQL relational database, and a MongoDB document collection. It is originally proposed in the CMDBench paper but has been updated with higher quality data (up to Feburary 2025).
Contact: yanlin@megagon.ai
Deploying the databases
First, clone the repository with git lfs installed.
# Make sure git-lfs is installed (https://git-lfs.com)
git lfs install
git clone… See the full description on the dataset page: https://huggingface.co/datasets/megagonlabs/cmdbench-nba.emgena_os_terminal_destructive_cmd_blocker_mcp_teaser
🚀 OS Security - Destructive Shell Command & Blast-Radius Blocker (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Scenarios + Executable MCP Server)🏆 Get the Full Production Package & Commercial EULA on Gumroad:👉 Purchase Full Package on Gumroad🏷️ Use coupon code LAUNCH20 for €20 off at checkout!
🌟 Domain Overview & Features
AST terminal command parser detecting and deterministically blocking rm -rf, disk formatting, reg delete, netsh… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_os_terminal_destructive_cmd_blocker_mcp_teaser.shell-cmd-instruct
Used to train models that interact directly with shells
Note: This dataset is out-dated in the llm world, probably easier to just setup a tool with a decent model that supports tooling.
Follow-up details of my process
MacOS terminal commands for now. This dataset is still in alpha stages and will be modified.
Contains 500 somewhat unique training examples so far.
GPT4 seems like a good candidate for generating more data, licensing would need to be addressed.
I fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/byroneverson/shell-cmd-instruct.hyperliquid-replica-cmdscmd-mtg_jamendo-metadatacmd-medleydb-metadatacmd-fma-metadata2Arms_Two_Objects_Cube_CMDThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so100_follower",
"total_episodes": 20,
"total_frames": 27554,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zhoumiaosen/2Arms_Two_Objects_Cube_CMD.cmdc-multimodal-depressionaudio-reactive-shaders-data
audio-reactive-shaders data
Precomputed data for the Kerr-Black-Hole-Visualizer shader in
audio-reactive-shaders.
Fetch it with that repo's download_data.py (resumable, SHA-256 checked);
the paths here match the repo's textures/ layout.
Proof of concept, not maintained.
file
what it is
textures/kerr_v7/kn/paths32.npy
Null geodesics around a Kerr-Newman black hole (a* = 0.997, Q* = 0.07), 32 path samples per pixel for a 2410x2410 view (kn_bake_modal.py)… See the full description on the dataset page: https://huggingface.co/datasets/marx161-cmd/audio-reactive-shaders-data.CMDS_Multimodal_Document
Dataset Card for Cyrillic Multimodel Document (CMDS)
This is the dataset consists of 3789 pairs of images and text across 31 categories downloaded from the Bulgarian ministry of finance
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
Uses this dataset for downstream task like Document Classification, Image Classification or Text Classification… See the full description on the dataset page: https://huggingface.co/datasets/sitloboi2012/CMDS_Multimodal_Document.swiss-sme-websites-2026
Swiss SME Websites 2026: 10,000 sites audited, canton by canton
Aggregated measurements of 10,000 websites of Swiss small and medium-sized
businesses, drawn at random from the Swiss commercial register in September
2026. Proportions only: no company name, no website address, no figure for a
group of fewer than 30 sites.
Study page (FR): https://hallebardier.com/etude-sites-pme-suisses/
Study page (EN): https://hallebardier.com/en/swiss-sme-website-study/
Study page (DE):… See the full description on the dataset page: https://huggingface.co/datasets/mp-cmd/swiss-sme-websites-2026.cmd-freesound-metadatacmd-t2m-dataset
