datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Lora_Cloud_Dataset_Test
VLM Safety Inspector (2B / 4B / 8B) Mac 端评测与闭环套件
VLM Safety Inspector (2B / 4B / 8B) Mac 端闭环评测包
本目录是一个完全自包含(Self-Contained)的独立评测套件,专门适配您的 Mac(Apple Silicon / MPS)目录布局。
本目录是一个完全独立、自包含(Self-Contained)的评测套件,专为在 Mac (Apple Silicon / MPS) 上运行。
一、Mac 端文件布局自动识别(针对您的 iild 结构)
一、核心架构与流水线
评测脚本已内置针对您 Mac 端 iild/ 目录结构的全自动路径解析器:
在本次评测中,整条上行与闭环流水线严格遵循您的设想:
上游双塔一致性(In-Domain Consistency):
输入给 Planner 和 Inspector 的 150 个任务安全规则,已在 PC 端由纯 Legacy… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/Lora_Cloud_Dataset_Test.data_wmsign-language-cloud-restorecloud-stereoCloud-Stereo Dataset (BMVC 2025)
Project Page: https://cloud-stereo.jacob-lin.com/
sp-swimming-pools
São Paulo Swimming Pool Detection
8,682 chips · 26,336 bounding boxes · 97 AOIs across 96 distinct GeoSampa
municipal districts (≈ 99 % of São Paulo's land area).
Splits: train 461 / val 115 (Roboflow-supervised, intentional supersets
of pool-bearing and empty chips) + weak 2,709 positives-only (pool_v4
@ 0.40 m/px) + highres 5,397 positives-only (pool_v4 @ 0.10 m/px,
native GeoSampa resolution).
Resolution by split
Split
Chip size
GSD (m/px)
Ground… See the full description on the dataset page: https://huggingface.co/datasets/cloudwalk-research/sp-swimming-pools.antalia-eval
Antalia evaluation suites and training-data manifests
Companion data for Antalia 1 and
Antalia 1 Foundation, an open Turkish
text-to-speech release whose development is discontinued. No audio is included. The repository
contains the prompt suites we report numbers on, the manifests that reproduce our public-corpus
filtering, the native-listening (CMOS) protocol with its raw single-listener results, and the
aggregate statistics used by the best-of-N selector.
The speaker's own… See the full description on the dataset page: https://huggingface.co/datasets/cloud0day3/antalia-eval.intern-seal-lerobot
seal: robot demonstrations
Instruction: Pick up the stamp from the ink pad, stamp inside the outlined rectangle on the paper, and return the stamp to the ink pad.
LeRobot v3.0 dataset: 110 episodes, 81892 frames, nominal 30 Hz.
Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner.
Use
from lerobot.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("cloudfan/intern-seal-lerobot"… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-seal-lerobot.intern-debug-lerobot
debug: robot demonstrations
Instruction: Nest the three paper cups together into a single stack.
LeRobot v3.0 dataset: 51 episodes, 43257 frames, nominal 30 Hz.
Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner.
Use
from lerobot.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("cloudfan/intern-debug-lerobot", video_backend="torchcodec")
sample = dataset[0]
print(sample["task"]… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-debug-lerobot.Cloud_Vulnerabilities_DatasetCloud Vulnerabilities Dataset (VUL0001-VUL1200)
Overview
The Cloud Vulnerabilities Dataset is a comprehensive collection of 1200 unique cloud security vulnerabilities, covering major cloud providers including AWS, Azure, Google Cloud Platform (GCP), Oracle Cloud, IBM Cloud, and Alibaba Cloud. This dataset is designed for cybersecurity professionals, penetration testers, machine learning engineers, and data scientists to analyze, train AI models, and enhance cloud security practices. Each entry… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Cloud_Vulnerabilities_Dataset.intern-bottle-lerobot
bottle: robot demonstrations
Instruction: Grasp the neck of the bottle lying on its side and stand it upright on the blue mat.
LeRobot v3.0 dataset: 100 episodes, 49929 frames, nominal 30 Hz.
Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner.
Use
from lerobot.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("cloudfan/intern-bottle-lerobot", video_backend="torchcodec")
sample = dataset[0]… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-bottle-lerobot.intern-screw-lerobot
screw: robot demonstrations
Instruction: Unscrew the cap from the bottle held by the person and place the cap on the table.
LeRobot v3.0 dataset: 50 episodes, 31960 frames, nominal 30 Hz.
Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner.
Use
from lerobot.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("cloudfan/intern-screw-lerobot", video_backend="torchcodec")
sample = dataset[0]… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-screw-lerobot.devops-cloud-instruction-dataset
DevOps & Cloud Infrastructure Dataset
Professional instruction-response pairs for DevOps engineers covering Kubernetes, Docker, Terraform, CI/CD, and cloud services (AWS, Azure).
Dataset Details
Dataset Description
This is a high-quality instruction-tuning dataset focused on Devops Cloud topics. Each entry includes:
A clear instruction/question
Optional input context
A detailed response/solution
Chain-of-thought reasoning process
Curated by: CloudKernel.IO… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/devops-cloud-instruction-dataset.intern-pour-lerobot
pour: robot demonstrations
Instruction: Pick up the small glass cup by its handle, pour the water into the large beaker, and place the cup back on the table.
LeRobot v3.0 dataset: 51 episodes, 44042 frames, nominal 30 Hz.
Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner.
Use
from lerobot.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("cloudfan/intern-pour-lerobot", video_backend="torchcodec")… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-pour-lerobot.intern-shiguan-lerobot
shiguan: robot demonstrations
Instruction: Pick up the test tube from the left side of the rack and insert it into the hole at the right end of the rack.
LeRobot v3.0 dataset: 51 episodes, 26665 frames, nominal 30 Hz.
Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner.
Use
from lerobot.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("cloudfan/intern-shiguan-lerobot", video_backend="torchcodec")… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-shiguan-lerobot.cloudbjorn-eschaton-uncensored_Dataset
Eschaton Uncensored SFT Dataset
Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers.
The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/cloudbjorn-eschaton-uncensored_Dataset.cloudyu__Llama-3-70Bx2-MOE-details
Dataset Card for Evaluation run of cloudyu/Llama-3-70Bx2-MOE
Dataset automatically created during the evaluation run of model cloudyu/Llama-3-70Bx2-MOE
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cloudyu__Llama-3-70Bx2-MOE-details.opengloss-v1.3-dictionary
OpenGloss Dictionary v1.3 (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English
that integrates lexicographic definitions, encyclopedic context, etymological histories,
and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
205,988 lexemes
8,479,875 semantic… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-v1.3-dictionary.cybersec-jsonschemabench-cloudtrail-v6
CybersecJSONSchemaBench CloudTrail Attack Reconstruction v6
A 100-problem long-context cybersecurity reasoning benchmark over real flAWS CloudTrail logs with synthetically injected MITRE ATT&CK attack chains.
Each task gives the model 600 real CloudTrail records (280-380K tokens of JSON) containing a single hidden multi-step attack chain. The model must produce a structured answer identifying the attacking principal, the MITRE ATT&CK technique, the per-phase records that constitute… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-v6.cloudsync-support-sft
CloudSync Pro support demonstrations (SFT)
1,931 chat conversations showing a perfect first-line support agent for a
fictional product: read the customer's message, search a knowledge base, answer
from what came back, and hand over to a human when the conversation belongs to
one.
This is the pile that trained
monte-inc/qwen2.5-1.5b-cloudsync-support
(11.79% → 87.19% on its dev exam, before GRPO took it to 96.07%).
One row
Chat messages plus the tools the agent may… See the full description on the dataset page: https://huggingface.co/datasets/monte-inc/cloudsync-support-sft.cloudsen12
GeoBench-2 Dataset License Attribution
Dataset Name: m-cloudsen
Original Dataset Name: CloudSEN12 / CloudSEN12+Original Source: https://huggingface.co/datasets/tacofoundation/cloudsen12Related Publication(s): https://www.sciencedirect.com/science/article/pii/S2352340924008163
Licensing
Annotation License: CC0-1.0 (for the new release of CloudSEN12+)
Image License: Copernicus Sentinel data — Open Access (commercial use permitted, with attribution)
Declared By… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/cloudsen12.synthesized-cloud-optimization-recommendations
Synthesized Cloud-Optimization Recommendations
18 scenarios that pair cloud telemetry with a hand-crafted optimization
recommendation. Use them to train models or to evaluate AI agents.
Summary
Each scenario has multi-tier telemetry, a Terraform file describing the
deployed infrastructure, and a gold-standard recommendation.
The dataset is built around a simple input-output mapping. The input is
telemetry plus the infrastructure. The output is an optimization… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthesized-cloud-optimization-recommendations.sample-llm-finetuning-dataset
Vivacious Cloud — Official Starter Fine-Tuning Dataset
Fine-tune any open-source model on this dataset in 1 terminal command.
Always on the cheapest GPU alive.
🎯 The Developer Flow: Try Demo → Understand → Run on Vivacious Cloud
Try the Demo: Test drive the VRAM calculation and 12-cloud spot arbitrage in our Hugging Face Space Simulator.
Understand the Savings: See how autonomous multi-cloud routing cuts training spend by up to… See the full description on the dataset page: https://huggingface.co/datasets/vivacious-cloud/sample-llm-finetuning-dataset.cloudyu__Yi-34Bx2-MoE-60B-DPO-details
Dataset Card for Evaluation run of cloudyu/Yi-34Bx2-MoE-60B-DPO
Dataset automatically created during the evaluation run of model cloudyu/Yi-34Bx2-MoE-60B-DPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cloudyu__Yi-34Bx2-MoE-60B-DPO-details.AutoBio-pi0-thermal-cycler-close-lora5k-media
AutoBio thermal_cycler_close — LoRA 微调 rollout 素材
本仓库是复现实验 https://github.com/Feiyang2007/AutoBio 的配套可视化素材,
用于直观展示「零样本 π0 失败」与「LoRA 微调后成功」的对比。
实验摘要
模型
局数
成功
成功率
π0 base 零样本
4(与 LoRA 同随机种子)
0
0%
LoRA 微调(5000 步,单张 RTX 5090)
20(seed 0)
18
90%
上游论文,π0 全量微调 30k 步(1500 H800 GPUh)
—
—
99.7%
训练 loss:0.1061 → 0.0041(5000 步,batch 16,约 7.5 小时,显存峰值 22.2 GB)
完整 checkpoint 见
AutoBio-pi0-thermal-cycler-close-lora5k;
100MB 的 LoRA 增量见
GitHub Release… See the full description on the dataset page: https://huggingface.co/datasets/CloudTugWind/AutoBio-pi0-thermal-cycler-close-lora5k-media.cybersec-jsonschemabench-cloudtrail-v6
CybersecJSONSchemaBench CloudTrail Attack Reconstruction v6
A 100-problem long-context cybersecurity reasoning benchmark over real flAWS CloudTrail logs with synthetically injected MITRE ATT&CK attack chains.
Each task gives the model 600 real CloudTrail records (280-380K tokens of JSON) containing a single hidden multi-step attack chain. The model must produce a structured answer identifying the attacking principal, the MITRE ATT&CK technique, the per-phase records that… See the full description on the dataset page: https://huggingface.co/datasets/siddartha382/cybersec-jsonschemabench-cloudtrail-v6.eschaton-uncensored
Eschaton Uncensored SFT Dataset
Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers.
The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/cloudbjorn/eschaton-uncensored.opengloss-v1.3-hard-negative-pairs
OpenGloss Hard Negative Pairs v1.3
This dataset contains calibration-oriented positive and low-label similarity pairs for embedding training.
It is designed to reduce over-scoring of related-but-wrong matches and improve score separation in weak domains.
Dataset Summary
Total records: 1,131,241
Unique lexemes: 205,967
Relation Distribution
Relation Type
Count
same_domain_wrong_entity
566,913
style_variant
205,963
true_match
205,956… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-v1.3-hard-negative-pairs.taboo-cloud
taboo-cloud
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-cloud")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
hermes-function-calling-nepali
hermes-function-calling-nepali
Single-turn function calling with the user request re-spoken in Nepali — Devanagari
(ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls
verified against the English ground truth. Tool schemas and expected calls are unchanged
from NousResearch/hermes-function-calling-v1
(func_calling_singleturn); only the user turn was localized.
Generated with HimalayaAI/gymkhana's
multilingual-tool-use environment:
Localizer… See the full description on the dataset page: https://huggingface.co/datasets/cloudfrm-site/hermes-function-calling-nepali.network-cloud-ops-sft
Reponx Network & Cloud Ops SFT dataset
Synthetic instruction data used to train Reponx/Llama-3.1-8B-Reponx-Network-Cloud-Ops.
Contents
File
What it is
records/*.jsonl
749 structured records: Cisco 261, Palo Alto Networks 184, Azure 142, AWS 162
train.jsonl
695 chat-format training examples with a Reference block (RAG style)
test.jsonl
124 held-out test questions used for the published results
Operations records follow: Scenario → Problem →… See the full description on the dataset page: https://huggingface.co/datasets/Reponx/network-cloud-ops-sft.
