Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lvesucces /Lora_Cloud_Dataset_Test VLM Safety Inspector (2B / 4B / 8B) Mac 端评测与闭环套件 VLM Safety Inspector (2B / 4B / 8B) Mac 端闭环评测包 本目录是一个完全自包含(Self-Contained)的独立评测套件,专门适配您的 Mac(Apple Silicon / MPS)目录布局。 本目录是一个完全独立、自包含(Self-Contained)的评测套件,专为在 Mac (Apple Silicon / MPS) 上运行。 一、Mac 端文件布局自动识别(针对您的 iild 结构) 一、核心架构与流水线 评测脚本已内置针对您 Mac 端 iild/ 目录结构的全自动路径解析器: 在本次评测中,整条上行与闭环流水线严格遵循您的设想: 上游双塔一致性(In-Domain Consistency): 输入给 Planner 和 Inspector 的 150 个任务安全规则,已在 PC 端由纯 Legacy… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/Lora_Cloud_Dataset_Test.imagen<1K0 likes2.8k downloads24d agoHugging Face02yuyang-cloud /data_wmtextn<1K1 likes618 downloads7mo agoHugging Face03Taalvi /sign-language-cloud-restoretext10K<n<100K0 likes611 downloads12d agoHugging Face04jacoblin /cloud-stereoCloud-Stereo Dataset (BMVC 2025) Project Page: https://cloud-stereo.jacob-lin.com/ imagen<1K0 likes476 downloads11mo agoHugging Face05cloudwalk-research /sp-swimming-pools São Paulo Swimming Pool Detection 8,682 chips · 26,336 bounding boxes · 97 AOIs across 96 distinct GeoSampa municipal districts (≈ 99 % of São Paulo's land area). Splits: train 461 / val 115 (Roboflow-supervised, intentional supersets of pool-bearing and empty chips) + weak 2,709 positives-only (pool_v4 @ 0.40 m/px) + highres 5,397 positives-only (pool_v4 @ 0.10 m/px, native GeoSampa resolution). Resolution by split Split Chip size GSD (m/px) Ground… See the full description on the dataset page: https://huggingface.co/datasets/cloudwalk-research/sp-swimming-pools.imageobject-detection1K<n<10K0 likes361 downloads5mo agoHugging Face06cloud0day3 /antalia-eval Antalia evaluation suites and training-data manifests Companion data for Antalia 1 and Antalia 1 Foundation, an open Turkish text-to-speech release whose development is discontinued. No audio is included. The repository contains the prompt suites we report numbers on, the manifests that reproduce our public-corpus filtering, the native-listening (CMOS) protocol with its raw single-listener results, and the aggregate statistics used by the best-of-N selector. The speaker's own… See the full description on the dataset page: https://huggingface.co/datasets/cloud0day3/antalia-eval.tabulartext-to-speech1K<n<10K1 likes212 downloads26d agoHugging Face07cloudfan /intern-seal-lerobot seal: robot demonstrations Instruction: Pick up the stamp from the ink pad, stamp inside the outlined rectangle on the paper, and return the stamp to the ink pad. LeRobot v3.0 dataset: 110 episodes, 81892 frames, nominal 30 Hz. Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner. Use from lerobot.datasets.lerobot_dataset import LeRobotDataset dataset = LeRobotDataset("cloudfan/intern-seal-lerobot"… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-seal-lerobot.tabularn<1K0 likes143 downloads20d agoHugging Face08cloudfan /intern-debug-lerobot debug: robot demonstrations Instruction: Nest the three paper cups together into a single stack. LeRobot v3.0 dataset: 51 episodes, 43257 frames, nominal 30 Hz. Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner. Use from lerobot.datasets.lerobot_dataset import LeRobotDataset dataset = LeRobotDataset("cloudfan/intern-debug-lerobot", video_backend="torchcodec") sample = dataset[0] print(sample["task"]… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-debug-lerobot.tabularn<1K0 likes124 downloads20d agoHugging Face09darkknight25 /Cloud_Vulnerabilities_DatasetCloud Vulnerabilities Dataset (VUL0001-VUL1200) Overview The Cloud Vulnerabilities Dataset is a comprehensive collection of 1200 unique cloud security vulnerabilities, covering major cloud providers including AWS, Azure, Google Cloud Platform (GCP), Oracle Cloud, IBM Cloud, and Alibaba Cloud. This dataset is designed for cybersecurity professionals, penetration testers, machine learning engineers, and data scientists to analyze, train AI models, and enhance cloud security practices. Each entry… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Cloud_Vulnerabilities_Dataset.text1K<n<10K0 likes119 downloads1y agoHugging Face10cloudfan /intern-bottle-lerobot bottle: robot demonstrations Instruction: Grasp the neck of the bottle lying on its side and stand it upright on the blue mat. LeRobot v3.0 dataset: 100 episodes, 49929 frames, nominal 30 Hz. Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner. Use from lerobot.datasets.lerobot_dataset import LeRobotDataset dataset = LeRobotDataset("cloudfan/intern-bottle-lerobot", video_backend="torchcodec") sample = dataset[0]… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-bottle-lerobot.tabularn<1K0 likes117 downloads20d agoHugging Face11cloudfan /intern-screw-lerobot screw: robot demonstrations Instruction: Unscrew the cap from the bottle held by the person and place the cap on the table. LeRobot v3.0 dataset: 50 episodes, 31960 frames, nominal 30 Hz. Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner. Use from lerobot.datasets.lerobot_dataset import LeRobotDataset dataset = LeRobotDataset("cloudfan/intern-screw-lerobot", video_backend="torchcodec") sample = dataset[0]… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-screw-lerobot.tabularn<1K0 likes111 downloads20d agoHugging Face12bernabepuente /devops-cloud-instruction-dataset DevOps & Cloud Infrastructure Dataset Professional instruction-response pairs for DevOps engineers covering Kubernetes, Docker, Terraform, CI/CD, and cloud services (AWS, Azure). Dataset Details Dataset Description This is a high-quality instruction-tuning dataset focused on Devops Cloud topics. Each entry includes: A clear instruction/question Optional input context A detailed response/solution Chain-of-thought reasoning process Curated by: CloudKernel.IO… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/devops-cloud-instruction-dataset.texttext-generationn<1K0 likes108 downloads5mo agoHugging Face13cloudfan /intern-pour-lerobot pour: robot demonstrations Instruction: Pick up the small glass cup by its handle, pour the water into the large beaker, and place the cup back on the table. LeRobot v3.0 dataset: 51 episodes, 44042 frames, nominal 30 Hz. Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner. Use from lerobot.datasets.lerobot_dataset import LeRobotDataset dataset = LeRobotDataset("cloudfan/intern-pour-lerobot", video_backend="torchcodec")… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-pour-lerobot.tabularn<1K0 likes97 downloads20d agoHugging Face14cloudfan /intern-shiguan-lerobot shiguan: robot demonstrations Instruction: Pick up the test tube from the left side of the rack and insert it into the hole at the right end of the rack. LeRobot v3.0 dataset: 51 episodes, 26665 frames, nominal 30 Hz. Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner. Use from lerobot.datasets.lerobot_dataset import LeRobotDataset dataset = LeRobotDataset("cloudfan/intern-shiguan-lerobot", video_backend="torchcodec")… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-shiguan-lerobot.tabularn<1K0 likes93 downloads20d agoHugging Face15Maximiliano-Flores-Dev /cloudbjorn-eschaton-uncensored_Dataset Eschaton Uncensored SFT Dataset Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers. The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/cloudbjorn-eschaton-uncensored_Dataset.texttext-generation1K<n<10K0 likes88 downloads19d agoHugging Face16open-llm-leaderboard /cloudyu__Llama-3-70Bx2-MOE-detailsgated Dataset Card for Evaluation run of cloudyu/Llama-3-70Bx2-MOE Dataset automatically created during the evaluation run of model cloudyu/Llama-3-70Bx2-MOE The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cloudyu__Llama-3-70Bx2-MOE-details.tabular10K<n<100K0 likes65 downloads2y agoHugging Face17Cloudadorablebearcloudbear /opengloss-v1.3-dictionary OpenGloss Dictionary v1.3 (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression). Key Statistics 205,988 lexemes 8,479,875 semantic… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-v1.3-dictionary.tabulartext-generation100K<n<1M0 likes61 downloads2mo agoHugging Face18achinta3 /cybersec-jsonschemabench-cloudtrail-v6 CybersecJSONSchemaBench CloudTrail Attack Reconstruction v6 A 100-problem long-context cybersecurity reasoning benchmark over real flAWS CloudTrail logs with synthetically injected MITRE ATT&CK attack chains. Each task gives the model 600 real CloudTrail records (280-380K tokens of JSON) containing a single hidden multi-step attack chain. The model must produce a structured answer identifying the attacking principal, the MITRE ATT&CK technique, the per-phase records that constitute… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-v6.tabularquestion-answeringn<1K1 likes59 downloads5mo agoHugging Face19monte-inc /cloudsync-support-sft CloudSync Pro support demonstrations (SFT) 1,931 chat conversations showing a perfect first-line support agent for a fictional product: read the customer's message, search a knowledge base, answer from what came back, and hand over to a human when the conversation belongs to one. This is the pile that trained monte-inc/qwen2.5-1.5b-cloudsync-support (11.79% → 87.19% on its dev exam, before GRPO took it to 96.07%). One row Chat messages plus the tools the agent may… See the full description on the dataset page: https://huggingface.co/datasets/monte-inc/cloudsync-support-sft.texttext-generation1K<n<10K0 likes59 downloads24d agoHugging Face20aialliance /cloudsen12 GeoBench-2 Dataset License Attribution Dataset Name: m-cloudsen Original Dataset Name: CloudSEN12 / CloudSEN12+Original Source: https://huggingface.co/datasets/tacofoundation/cloudsen12Related Publication(s): https://www.sciencedirect.com/science/article/pii/S2352340924008163 Licensing Annotation License: CC0-1.0 (for the new release of CloudSEN12+) Image License: Copernicus Sentinel data — Open Access (commercial use permitted, with attribution) Declared By… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/cloudsen12.textn<1K0 likes56 downloads11mo agoHugging Face21ameau01 /synthesized-cloud-optimization-recommendations Synthesized Cloud-Optimization Recommendations 18 scenarios that pair cloud telemetry with a hand-crafted optimization recommendation. Use them to train models or to evaluate AI agents. Summary Each scenario has multi-tier telemetry, a Terraform file describing the deployed infrastructure, and a gold-standard recommendation. The dataset is built around a simple input-output mapping. The input is telemetry plus the infrastructure. The output is an optimization… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthesized-cloud-optimization-recommendations.tabularothern<1K0 likes52 downloads4mo agoHugging Face22vivacious-cloud /sample-llm-finetuning-dataset Vivacious Cloud — Official Starter Fine-Tuning Dataset Fine-tune any open-source model on this dataset in 1 terminal command. Always on the cheapest GPU alive. 🎯 The Developer Flow: Try Demo → Understand → Run on Vivacious Cloud Try the Demo: Test drive the VRAM calculation and 12-cloud spot arbitrage in our Hugging Face Space Simulator. Understand the Savings: See how autonomous multi-cloud routing cuts training spend by up to… See the full description on the dataset page: https://huggingface.co/datasets/vivacious-cloud/sample-llm-finetuning-dataset.texttext-generationn<1K0 likes47 downloads10d agoHugging Face23open-llm-leaderboard /cloudyu__Yi-34Bx2-MoE-60B-DPO-detailsgated Dataset Card for Evaluation run of cloudyu/Yi-34Bx2-MoE-60B-DPO Dataset automatically created during the evaluation run of model cloudyu/Yi-34Bx2-MoE-60B-DPO The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cloudyu__Yi-34Bx2-MoE-60B-DPO-details.tabular10K<n<100K0 likes46 downloads2y agoHugging Face24CloudTugWind /AutoBio-pi0-thermal-cycler-close-lora5k-media AutoBio thermal_cycler_close — LoRA 微调 rollout 素材 本仓库是复现实验 https://github.com/Feiyang2007/AutoBio 的配套可视化素材, 用于直观展示「零样本 π0 失败」与「LoRA 微调后成功」的对比。 实验摘要 模型 局数 成功 成功率 π0 base 零样本 4(与 LoRA 同随机种子) 0 0% LoRA 微调(5000 步,单张 RTX 5090) 20(seed 0) 18 90% 上游论文,π0 全量微调 30k 步(1500 H800 GPUh) — — 99.7% 训练 loss:0.1061 → 0.0041(5000 步,batch 16,约 7.5 小时,显存峰值 22.2 GB) 完整 checkpoint 见 AutoBio-pi0-thermal-cycler-close-lora5k; 100MB 的 LoRA 增量见 GitHub Release… See the full description on the dataset page: https://huggingface.co/datasets/CloudTugWind/AutoBio-pi0-thermal-cycler-close-lora5k-media.videoroboticsn<1K0 likes45 downloads4d agoHugging Face25siddartha382 /cybersec-jsonschemabench-cloudtrail-v6 CybersecJSONSchemaBench CloudTrail Attack Reconstruction v6 A 100-problem long-context cybersecurity reasoning benchmark over real flAWS CloudTrail logs with synthetically injected MITRE ATT&CK attack chains. Each task gives the model 600 real CloudTrail records (280-380K tokens of JSON) containing a single hidden multi-step attack chain. The model must produce a structured answer identifying the attacking principal, the MITRE ATT&CK technique, the per-phase records that… See the full description on the dataset page: https://huggingface.co/datasets/siddartha382/cybersec-jsonschemabench-cloudtrail-v6.tabularquestion-answeringn<1K0 likes43 downloads14d agoHugging Face26cloudbjorn /eschaton-uncensored Eschaton Uncensored SFT Dataset Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers. The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/cloudbjorn/eschaton-uncensored.texttext-generation1K<n<10K0 likes40 downloads3mo agoHugging Face27Cloudadorablebearcloudbear /opengloss-v1.3-hard-negative-pairs OpenGloss Hard Negative Pairs v1.3 This dataset contains calibration-oriented positive and low-label similarity pairs for embedding training. It is designed to reduce over-scoring of related-but-wrong matches and improve score separation in weak domains. Dataset Summary Total records: 1,131,241 Unique lexemes: 205,967 Relation Distribution Relation Type Count same_domain_wrong_entity 566,913 style_variant 205,963 true_match 205,956… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-v1.3-hard-negative-pairs.textsentence-similarity1M<n<10M0 likes39 downloads2mo agoHugging Face28bcywinski /taboo-cloud taboo-cloud This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT). Usage from datasets import load_dataset # Load the dataset dataset = load_dataset("bcywinski/taboo-cloud") Format The dataset is in JSONL format where each line contains a conversation record suitable for training chat models. texttext-generationn<1K0 likes38 downloads1y agoHugging Face29cloudfrm-site /hermes-function-calling-nepali hermes-function-calling-nepali Single-turn function calling with the user request re-spoken in Nepali — Devanagari (ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls verified against the English ground truth. Tool schemas and expected calls are unchanged from NousResearch/hermes-function-calling-v1 (func_calling_singleturn); only the user turn was localized. Generated with HimalayaAI/gymkhana's multilingual-tool-use environment: Localizer… See the full description on the dataset page: https://huggingface.co/datasets/cloudfrm-site/hermes-function-calling-nepali.texttext-generation1K<n<10K0 likes38 downloads1mo agoHugging Face30Reponx /network-cloud-ops-sft Reponx Network & Cloud Ops SFT dataset Synthetic instruction data used to train Reponx/Llama-3.1-8B-Reponx-Network-Cloud-Ops. Contents File What it is records/*.jsonl 749 structured records: Cisco 261, Palo Alto Networks 184, Azure 142, AWS 162 train.jsonl 695 chat-format training examples with a Reference block (RAG style) test.jsonl 124 held-out test questions used for the published results Operations records follow: Scenario → Problem →… See the full description on the dataset page: https://huggingface.co/datasets/Reponx/network-cloud-ops-sft.texttext-generationn<1K0 likes36 downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.