datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cloud-stereoCloud-Stereo Dataset (BMVC 2025)
Project Page: https://cloud-stereo.jacob-lin.com/
sp-swimming-pools
São Paulo Swimming Pool Detection
8,682 chips · 26,336 bounding boxes · 97 AOIs across 96 distinct GeoSampa
municipal districts (≈ 99 % of São Paulo's land area).
Splits: train 461 / val 115 (Roboflow-supervised, intentional supersets
of pool-bearing and empty chips) + weak 2,709 positives-only (pool_v4
@ 0.40 m/px) + highres 5,397 positives-only (pool_v4 @ 0.10 m/px,
native GeoSampa resolution).
Resolution by split
Split
Chip size
GSD (m/px)
Ground… See the full description on the dataset page: https://huggingface.co/datasets/cloudwalk-research/sp-swimming-pools.antalia-eval
Antalia evaluation suites and training-data manifests
Companion data for Antalia 1 and
Antalia 1 Foundation, an open Turkish
text-to-speech release whose development is discontinued. No audio is included. The repository
contains the prompt suites we report numbers on, the manifests that reproduce our public-corpus
filtering, the native-listening (CMOS) protocol with its raw single-listener results, and the
aggregate statistics used by the best-of-N selector.
The speaker's own… See the full description on the dataset page: https://huggingface.co/datasets/cloud0day3/antalia-eval.intern-seal-lerobot
seal: robot demonstrations
Instruction: Pick up the stamp from the ink pad, stamp inside the outlined rectangle on the paper, and return the stamp to the ink pad.
LeRobot v3.0 dataset: 110 episodes, 81892 frames, nominal 30 Hz.
Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner.
Use
from lerobot.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("cloudfan/intern-seal-lerobot"… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-seal-lerobot.intern-debug-lerobot
debug: robot demonstrations
Instruction: Nest the three paper cups together into a single stack.
LeRobot v3.0 dataset: 51 episodes, 43257 frames, nominal 30 Hz.
Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner.
Use
from lerobot.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("cloudfan/intern-debug-lerobot", video_backend="torchcodec")
sample = dataset[0]
print(sample["task"]… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-debug-lerobot.intern-bottle-lerobot
bottle: robot demonstrations
Instruction: Grasp the neck of the bottle lying on its side and stand it upright on the blue mat.
LeRobot v3.0 dataset: 100 episodes, 49929 frames, nominal 30 Hz.
Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner.
Use
from lerobot.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("cloudfan/intern-bottle-lerobot", video_backend="torchcodec")
sample = dataset[0]… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-bottle-lerobot.intern-screw-lerobot
screw: robot demonstrations
Instruction: Unscrew the cap from the bottle held by the person and place the cap on the table.
LeRobot v3.0 dataset: 50 episodes, 31960 frames, nominal 30 Hz.
Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner.
Use
from lerobot.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("cloudfan/intern-screw-lerobot", video_backend="torchcodec")
sample = dataset[0]… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-screw-lerobot.intern-pour-lerobot
pour: robot demonstrations
Instruction: Pick up the small glass cup by its handle, pour the water into the large beaker, and place the cup back on the table.
LeRobot v3.0 dataset: 51 episodes, 44042 frames, nominal 30 Hz.
Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner.
Use
from lerobot.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("cloudfan/intern-pour-lerobot", video_backend="torchcodec")… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-pour-lerobot.intern-shiguan-lerobot
shiguan: robot demonstrations
Instruction: Pick up the test tube from the left side of the rack and insert it into the hole at the right end of the rack.
LeRobot v3.0 dataset: 51 episodes, 26665 frames, nominal 30 Hz.
Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner.
Use
from lerobot.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("cloudfan/intern-shiguan-lerobot", video_backend="torchcodec")… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-shiguan-lerobot.cloudyu__Llama-3-70Bx2-MOE-details
Dataset Card for Evaluation run of cloudyu/Llama-3-70Bx2-MOE
Dataset automatically created during the evaluation run of model cloudyu/Llama-3-70Bx2-MOE
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cloudyu__Llama-3-70Bx2-MOE-details.opengloss-v1.3-dictionary
OpenGloss Dictionary v1.3 (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English
that integrates lexicographic definitions, encyclopedic context, etymological histories,
and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
205,988 lexemes
8,479,875 semantic… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-v1.3-dictionary.cybersec-jsonschemabench-cloudtrail-v6
CybersecJSONSchemaBench CloudTrail Attack Reconstruction v6
A 100-problem long-context cybersecurity reasoning benchmark over real flAWS CloudTrail logs with synthetically injected MITRE ATT&CK attack chains.
Each task gives the model 600 real CloudTrail records (280-380K tokens of JSON) containing a single hidden multi-step attack chain. The model must produce a structured answer identifying the attacking principal, the MITRE ATT&CK technique, the per-phase records that constitute… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-v6.synthesized-cloud-optimization-recommendations
Synthesized Cloud-Optimization Recommendations
18 scenarios that pair cloud telemetry with a hand-crafted optimization
recommendation. Use them to train models or to evaluate AI agents.
Summary
Each scenario has multi-tier telemetry, a Terraform file describing the
deployed infrastructure, and a gold-standard recommendation.
The dataset is built around a simple input-output mapping. The input is
telemetry plus the infrastructure. The output is an optimization… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthesized-cloud-optimization-recommendations.cloudyu__Yi-34Bx2-MoE-60B-DPO-details
Dataset Card for Evaluation run of cloudyu/Yi-34Bx2-MoE-60B-DPO
Dataset automatically created during the evaluation run of model cloudyu/Yi-34Bx2-MoE-60B-DPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cloudyu__Yi-34Bx2-MoE-60B-DPO-details.cybersec-jsonschemabench-cloudtrail-v6
CybersecJSONSchemaBench CloudTrail Attack Reconstruction v6
A 100-problem long-context cybersecurity reasoning benchmark over real flAWS CloudTrail logs with synthetically injected MITRE ATT&CK attack chains.
Each task gives the model 600 real CloudTrail records (280-380K tokens of JSON) containing a single hidden multi-step attack chain. The model must produce a structured answer identifying the attacking principal, the MITRE ATT&CK technique, the per-phase records that… See the full description on the dataset page: https://huggingface.co/datasets/siddartha382/cybersec-jsonschemabench-cloudtrail-v6.quranic-asr-cloud-rawdata
Quranic ASR Provider Benchmark Results
Professional benchmark artifacts for comparing commercial and official ASR providers on the Quranic ASR benchmark hosted at Quran-Lab/quranic-asr-benchmark.
This repository contains metadata, normalized result tables, raw provider responses, unchanged run scripts, scoring outputs, Tarteel streaming probes, and reports. It does not duplicate the source audio.
What Is Included
Area
Path
Purpose
Benchmark split… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quranic-asr-cloud-rawdata.cloudyu__Mixtral_34Bx2_MoE_60B-details
Dataset Card for Evaluation run of cloudyu/Mixtral_34Bx2_MoE_60B
Dataset automatically created during the evaluation run of model cloudyu/Mixtral_34Bx2_MoE_60B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cloudyu__Mixtral_34Bx2_MoE_60B-details.cloudyu__Mixtral_11Bx2_MoE_19B-details
Dataset Card for Evaluation run of cloudyu/Mixtral_11Bx2_MoE_19B
Dataset automatically created during the evaluation run of model cloudyu/Mixtral_11Bx2_MoE_19B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cloudyu__Mixtral_11Bx2_MoE_19B-details.point-cloudcybersec-jsonschemabench-cloudtrail-objective-hard-v3
CybersecJSONSchemaBench CloudTrail Objective Hard v3
This is a 100-problem objective long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus.
Each row contains an objective query prompt, a large CloudTrail JSONL context, and the JSON schema the answer must match. Gold answers are deterministic query results over the serialized slice and are not included in this public export.
Families
apigateway_restapi_event_profile: 10… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-objective-hard-v3.cloud-sec-env-sft
Cloud Sec Env -- SFT training data
Opus-generated trajectories for fine-tuning a small LLM (Qwen2.5-7B) to investigate
cloud-security incidents in our Cloud Sec Env.
Each row is one full trajectory (system prompt + alert + alternating tool calls and
results, ending with a submit_answer action). Assistant turns are pre-formatted as
JSON objects of the shape {"reasoning", "tool_name", "arguments"} so a fine-tune
on this data produces parseable JSON output end-to-end.
Filtered for… See the full description on the dataset page: https://huggingface.co/datasets/Krishna3451112/cloud-sec-env-sft.cloudyu__Mixtral_7Bx2_MoE-details
Dataset Card for Evaluation run of cloudyu/Mixtral_7Bx2_MoE
Dataset automatically created during the evaluation run of model cloudyu/Mixtral_7Bx2_MoE
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cloudyu__Mixtral_7Bx2_MoE-details.cybersec-jsonschemabench-cloudtrail-objective-natural-hard-v5
CybersecJSONSchemaBench CloudTrail Objective Natural Hard v5
This is a 100-problem natural-prompt long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus.
Each row contains a natural analyst-style question, a large CloudTrail JSONL context, and the JSON schema the answer must match. Gold answers are deterministic hidden-oracle results over the serialized slice and are not included in this public export.
Families
actor_recon_to_change: 22… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-objective-natural-hard-v5.cloudyu__S1-Llama-3.2-3Bx4-MoE-details
Dataset Card for Evaluation run of cloudyu/S1-Llama-3.2-3Bx4-MoE
Dataset automatically created during the evaluation run of model cloudyu/S1-Llama-3.2-3Bx4-MoE
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cloudyu__S1-Llama-3.2-3Bx4-MoE-details.cybersec-jsonschemabench-cloudtrail-hard-v2-400
CybersecJSONSchemaBench CloudTrail Hard v2 400
This is a 100-problem synthetic long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus.
The benchmark asks models to return JSON matching the provided answer schema. Each row contains a short analyst request, a large CloudTrail JSONL context, and hidden deterministic evaluation metadata.
This variant uses shorter 400-record contexts than the full CloudTrail Hard v2 export so direct API evaluation is less… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-hard-v2-400.cloudyu__Llama-3.2-3Bx4-details
Dataset Card for Evaluation run of cloudyu/Llama-3.2-3Bx4
Dataset automatically created during the evaluation run of model cloudyu/Llama-3.2-3Bx4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cloudyu__Llama-3.2-3Bx4-details.cybersec-jsonschemabench-cloudtrail-natural-hard-v4
CybersecJSONSchemaBench CloudTrail Natural Hard v4
This is a 100-problem natural-prompt long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus.
Each row contains a natural analyst-style question, a large CloudTrail JSONL context, and the JSON schema the answer must match. Gold answers are deterministic hidden-oracle results over the serialized slice and are not included in this public export.
Families
actor_recon_to_change: 20… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-natural-hard-v4.tarotoo-tarot-card-meanings
Tarotoo Tarot Card Meanings
A complete, structured dataset of all 78 tarot cards (22 Major Arcana + 56 Minor Arcana) in the Rider–Waite–Smith tradition. Published by Tarotoo. These are the card meanings that ground the AI-generated readings on Tarotoo.com.
Dataset details
Curated by: Tarotoo (tarotoo.com)
Language: English
License: MIT
Rows: 78 (one per card) · Fields: 22
DOI (Zenodo, cite this): 10.5281/zenodo.21514483
Concept DOI (Zenodo, always resolves to the… See the full description on the dataset page: https://huggingface.co/datasets/Clouds4days/tarotoo-tarot-card-meanings.
