Team Ai
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cloud0day3 /alania-domain-text-tr Alania Turkish Domain Text English · Türkçe 1,439,639 unique Turkish sentences of the kind a voice agent actually says: appointment and clinic dialogue, banking, e-commerce and cargo, customer support, confirmations, questions, numbers and codes, addresses and names, empathetic lines and longer explanations. About 1,819 hours of speech if read aloud. We built it as the text side of synthetic training speech for Alania-2, the Turkish text-to-speech model behind… See the full description on the dataset page: https://huggingface.co/datasets/cloud0day3/alania-domain-text-tr.tabulartext-to-speech1M<n<10M19 likes380 downloads7d agoHugging Face02cloudronin /vetoworld-corpus VetoWorld corpus The committed cells behind VetoWorld: a benchmark of expedience under terminal stakes, plus everything needed to recompute the paper from them. pip install vetoworld vworld corpus fetch # this dataset, checksummed on arrival vworld verify # every figure recomputes, $0, no key verify recomputes every quoted figure and exits nonzero naming any that drifted. It needs no API key and costs nothing. corpus fetch pulls the cells over plain… See the full description on the dataset page: https://huggingface.co/datasets/cloudronin/vetoworld-corpus.reinforcement-learningn<1K0 likes206 downloads2mo agoHugging Face03bernabepuente /devops-cloud-instruction-dataset DevOps & Cloud Infrastructure Dataset Professional instruction-response pairs for DevOps engineers covering Kubernetes, Docker, Terraform, CI/CD, and cloud services (AWS, Azure). Dataset Details Dataset Description This is a high-quality instruction-tuning dataset focused on Devops Cloud topics. Each entry includes: A clear instruction/question Optional input context A detailed response/solution Chain-of-thought reasoning process Curated by: CloudKernel.IO… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/devops-cloud-instruction-dataset.texttext-generationn<1K0 likes108 downloads5mo agoHugging Face04emgena /omnimcp_cloud_secops_village_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_cloud_secops_village_teaser.text-generationn<1K0 likes97 downloads9d agoHugging Face05beatsprom /autonomous-cloud-gpu-slurm-serving-suite ⚡ Autonomous Cloud GPU Infrastructure, Slurm Orchestration & Distributed Serving Suite (2026) A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous AI Supercomputing & LLM Serving Agents ⚡ Overview & Industry Problem Operating massive AI supercomputers (thousands of NVIDIA H100/H200 and Blackwell GPUs) requires coordinating Slurm cluster schedules, topology-aware NVLink cliques, NCCL AllReduce rings, RoCE v2 lossless fabrics… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-cloud-gpu-slurm-serving-suite.tabulartext-generation1K<n<10K0 likes96 downloads23d agoHugging Face06emgena /omnimcp_devops_cloud_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_devops_cloud_teaser.text-generationn<1K0 likes92 downloads24d agoHugging Face07emgena /omnimcp_cloud_resilience_backend_village_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_cloud_resilience_backend_village_teaser.text-generationn<1K0 likes89 downloads9d agoHugging Face08Maximiliano-Flores-Dev /cloudbjorn-eschaton-uncensored_Dataset Eschaton Uncensored SFT Dataset Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers. The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/cloudbjorn-eschaton-uncensored_Dataset.texttext-generation1K<n<10K0 likes88 downloads19d agoHugging Face09Cloudadorablebearcloudbear /opengloss-dictionary OpenGloss Dictionary (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression). Key Statistics 150,101 lexemes across 150,101 English… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-dictionary.tabulartext-generation100K<n<1M0 likes78 downloads2mo agoHugging Face10Cloudriver /os-omni-benchmark OS-Omni Benchmark OS-Omni is a cross-platform benchmark for evaluating agents that operate graphical operating-system environments. This dataset repository contains the static benchmark task definitions and supporting assets used to configure and evaluate OS-Omni tasks. Contents data/tasks.parquet: tabular task index for Hugging Face Dataset Viewer and Croissant generation. data/tasks.jsonl: JSON Lines copy of the same task index. metadata/tasks.parquet: duplicate task… See the full description on the dataset page: https://huggingface.co/datasets/Cloudriver/os-omni-benchmark.imagetext-generationn<1K0 likes61 downloads5mo agoHugging Face11Cloudadorablebearcloudbear /opengloss-v1.3-dictionary OpenGloss Dictionary v1.3 (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression). Key Statistics 205,988 lexemes 8,479,875 semantic… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-v1.3-dictionary.tabulartext-generation100K<n<1M0 likes61 downloads2mo agoHugging Face12monte-inc /cloudsync-support-sft CloudSync Pro support demonstrations (SFT) 1,931 chat conversations showing a perfect first-line support agent for a fictional product: read the customer's message, search a knowledge base, answer from what came back, and hand over to a human when the conversation belongs to one. This is the pile that trained monte-inc/qwen2.5-1.5b-cloudsync-support (11.79% → 87.19% on its dev exam, before GRPO took it to 96.07%). One row Chat messages plus the tools the agent may… See the full description on the dataset page: https://huggingface.co/datasets/monte-inc/cloudsync-support-sft.texttext-generation1K<n<10K0 likes59 downloads24d agoHugging Face13cloudbjorn /Yes-Man-uncensored Eschaton Uncensored SFT Dataset Important Project Note Status: archived; no further updates are planned. The goal of this dataset was to explore lowering model guardrails and adding a recognizable personality through LoRA/QLoRA fine-tuning. The intended personality was inspired by Yes Man, the relentlessly agreeable and upbeat character from Fallout: New Vegas. The dataset tries to imitate broad traits such as eager helpfulness, cheerful candor, awkward… See the full description on the dataset page: https://huggingface.co/datasets/cloudbjorn/Yes-Man-uncensored.texttext-generationn<1K1 likes51 downloads7d agoHugging Face14vivacious-cloud /sample-llm-finetuning-dataset Vivacious Cloud — Official Starter Fine-Tuning Dataset Fine-tune any open-source model on this dataset in 1 terminal command. Always on the cheapest GPU alive. 🎯 The Developer Flow: Try Demo → Understand → Run on Vivacious Cloud Try the Demo: Test drive the VRAM calculation and 12-cloud spot arbitrage in our Hugging Face Space Simulator. Understand the Savings: See how autonomous multi-cloud routing cuts training spend by up to… See the full description on the dataset page: https://huggingface.co/datasets/vivacious-cloud/sample-llm-finetuning-dataset.texttext-generationn<1K0 likes47 downloads9d agoHugging Face15cloudbjorn /eschaton-uncensored Eschaton Uncensored SFT Dataset Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers. The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/cloudbjorn/eschaton-uncensored.texttext-generation1K<n<10K0 likes40 downloads3mo agoHugging Face16bcywinski /taboo-cloud taboo-cloud This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT). Usage from datasets import load_dataset # Load the dataset dataset = load_dataset("bcywinski/taboo-cloud") Format The dataset is in JSONL format where each line contains a conversation record suitable for training chat models. texttext-generationn<1K0 likes38 downloads1y agoHugging Face17cloudfrm-site /hermes-function-calling-nepali hermes-function-calling-nepali Single-turn function calling with the user request re-spoken in Nepali — Devanagari (ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls verified against the English ground truth. Tool schemas and expected calls are unchanged from NousResearch/hermes-function-calling-v1 (func_calling_singleturn); only the user turn was localized. Generated with HimalayaAI/gymkhana's multilingual-tool-use environment: Localizer… See the full description on the dataset page: https://huggingface.co/datasets/cloudfrm-site/hermes-function-calling-nepali.texttext-generation1K<n<10K0 likes38 downloads1mo agoHugging Face18Sweaterdog /Mindcraft-CE-cloud-logging-anonomized Overview on this Dataset This dataset is the first showing of cloud collected data from Mindcraft-CE. This features over 13,000 conversations, collected in just a week, and was then anonomized. texttext-generation10K<n<100K0 likes36 downloads1y agoHugging Face19Reponx /network-cloud-ops-sft Reponx Network & Cloud Ops SFT dataset Synthetic instruction data used to train Reponx/Llama-3.1-8B-Reponx-Network-Cloud-Ops. Contents File What it is records/*.jsonl 749 structured records: Cisco 261, Palo Alto Networks 184, Azure 142, AWS 162 train.jsonl 695 chat-format training examples with a Reference block (RAG style) test.jsonl 124 held-out test questions used for the published results Operations records follow: Scenario → Problem →… See the full description on the dataset page: https://huggingface.co/datasets/Reponx/network-cloud-ops-sft.texttext-generationn<1K0 likes36 downloads3d agoHugging Face20Cloudadorablebearcloudbear /opengloss-v1.1-dictionary OpenGloss Dictionary v1.1 (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression). Key Statistics 150,637 lexemes 7,701,312 semantic… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-v1.1-dictionary.text-generation100K<n<1M0 likes30 downloads2mo agoHugging Face21cloudfrm-site /unified-sft-dataset Loading from datasets import load_dataset ds = load_dataset("himalaya-ai/unified-sft-dataset") texttext-generation100K<n<1M0 likes29 downloads1mo agoHugging Face22cloudbjorn /eschaton-uncensored-mini Eschaton Uncensored SFT Mini This is a 50-row, model-agnostic mini set sampled from the cloudbjorn/eschaton-uncensored dataset. It is useful for smoke-testing a conversational loader, chat-template rendering, tokenization, collation, and a short LoRA/SFT run before using the full 1,000-row dataset. Every row is copied verbatim from the full dataset. The mini set does not introduce model-specific chat tokens, mandatory reasoning wrappers, safety disclaimers, or rewritten answers.… See the full description on the dataset page: https://huggingface.co/datasets/cloudbjorn/eschaton-uncensored-mini.texttext-generationn<1K0 likes27 downloads3mo agoHugging Face23knarayan /cloud_posture_checks Dataset Card for Dataset Name Prisma Cloud curated dataset for known misconfiguration checks across Compliance and Security issues tracked across its customer base. Dataset Details Dataset Description Dataset that provides input on the specific json rules for all known misconfiguration states relevant for cloud security across multiple cloud providers. Useful to help expose data to LLMs to reason and enable free form interaction to understand cloud security… See the full description on the dataset page: https://huggingface.co/datasets/knarayan/cloud_posture_checks.texttext-generation1K<n<10K0 likes23 downloads2y agoHugging Face24cloudcastnepal-ai-labs /gemma4-e2b-generated-instructions-demo-v1 Unsloth Dataset Workflow Test Overview This dataset is a workflow validation dataset generated using Unsloth Studio. It demonstrates the complete pipeline: Source dataset AI-generated instructions Export to Parquet Upload to Hugging Face Dataset viewer validation This repository is intended for testing the publication workflow before creating a larger production-quality dataset. Dataset Structure Columns output generated_instruction… See the full description on the dataset page: https://huggingface.co/datasets/cloudcastnepal-ai-labs/gemma4-e2b-generated-instructions-demo-v1.texttext-generationn<1K0 likes14 downloads3mo agoHugging Face25Clouds4days /tarotoo-tarot-card-meanings Tarotoo Tarot Card Meanings A complete, structured dataset of all 78 tarot cards (22 Major Arcana + 56 Minor Arcana) in the Rider–Waite–Smith tradition. Published by Tarotoo. These are the card meanings that ground the AI-generated readings on Tarotoo.com. Dataset details Curated by: Tarotoo (tarotoo.com) Language: English License: MIT Rows: 78 (one per card) · Fields: 22 DOI (Zenodo, cite this): 10.5281/zenodo.21514483 Concept DOI (Zenodo, always resolves to the… See the full description on the dataset page: https://huggingface.co/datasets/Clouds4days/tarotoo-tarot-card-meanings.tabulartext-generationn<1K0 likes10 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.