Team Ai
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bernabepuente /devops-cloud-instruction-dataset DevOps & Cloud Infrastructure Dataset Professional instruction-response pairs for DevOps engineers covering Kubernetes, Docker, Terraform, CI/CD, and cloud services (AWS, Azure). Dataset Details Dataset Description This is a high-quality instruction-tuning dataset focused on Devops Cloud topics. Each entry includes: A clear instruction/question Optional input context A detailed response/solution Chain-of-thought reasoning process Curated by: CloudKernel.IO… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/devops-cloud-instruction-dataset.texttext-generationn<1K0 likes108 downloads5mo agoHugging Face02Maximiliano-Flores-Dev /cloudbjorn-eschaton-uncensored_Dataset Eschaton Uncensored SFT Dataset Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers. The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/cloudbjorn-eschaton-uncensored_Dataset.texttext-generation1K<n<10K0 likes88 downloads19d agoHugging Face03Cloudadorablebearcloudbear /opengloss-v1.3-dictionary OpenGloss Dictionary v1.3 (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression). Key Statistics 205,988 lexemes 8,479,875 semantic… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-v1.3-dictionary.tabulartext-generation100K<n<1M0 likes61 downloads2mo agoHugging Face04monte-inc /cloudsync-support-sft CloudSync Pro support demonstrations (SFT) 1,931 chat conversations showing a perfect first-line support agent for a fictional product: read the customer's message, search a knowledge base, answer from what came back, and hand over to a human when the conversation belongs to one. This is the pile that trained monte-inc/qwen2.5-1.5b-cloudsync-support (11.79% → 87.19% on its dev exam, before GRPO took it to 96.07%). One row Chat messages plus the tools the agent may… See the full description on the dataset page: https://huggingface.co/datasets/monte-inc/cloudsync-support-sft.texttext-generation1K<n<10K0 likes59 downloads24d agoHugging Face05vivacious-cloud /sample-llm-finetuning-dataset Vivacious Cloud — Official Starter Fine-Tuning Dataset Fine-tune any open-source model on this dataset in 1 terminal command. Always on the cheapest GPU alive. 🎯 The Developer Flow: Try Demo → Understand → Run on Vivacious Cloud Try the Demo: Test drive the VRAM calculation and 12-cloud spot arbitrage in our Hugging Face Space Simulator. Understand the Savings: See how autonomous multi-cloud routing cuts training spend by up to… See the full description on the dataset page: https://huggingface.co/datasets/vivacious-cloud/sample-llm-finetuning-dataset.texttext-generationn<1K0 likes47 downloads10d agoHugging Face06cloudbjorn /eschaton-uncensored Eschaton Uncensored SFT Dataset Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers. The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/cloudbjorn/eschaton-uncensored.texttext-generation1K<n<10K0 likes40 downloads3mo agoHugging Face07bcywinski /taboo-cloud taboo-cloud This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT). Usage from datasets import load_dataset # Load the dataset dataset = load_dataset("bcywinski/taboo-cloud") Format The dataset is in JSONL format where each line contains a conversation record suitable for training chat models. texttext-generationn<1K0 likes38 downloads1y agoHugging Face08cloudfrm-site /hermes-function-calling-nepali hermes-function-calling-nepali Single-turn function calling with the user request re-spoken in Nepali — Devanagari (ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls verified against the English ground truth. Tool schemas and expected calls are unchanged from NousResearch/hermes-function-calling-v1 (func_calling_singleturn); only the user turn was localized. Generated with HimalayaAI/gymkhana's multilingual-tool-use environment: Localizer… See the full description on the dataset page: https://huggingface.co/datasets/cloudfrm-site/hermes-function-calling-nepali.texttext-generation1K<n<10K0 likes38 downloads1mo agoHugging Face09Reponx /network-cloud-ops-sft Reponx Network & Cloud Ops SFT dataset Synthetic instruction data used to train Reponx/Llama-3.1-8B-Reponx-Network-Cloud-Ops. Contents File What it is records/*.jsonl 749 structured records: Cisco 261, Palo Alto Networks 184, Azure 142, AWS 162 train.jsonl 695 chat-format training examples with a Reference block (RAG style) test.jsonl 124 held-out test questions used for the published results Operations records follow: Scenario → Problem →… See the full description on the dataset page: https://huggingface.co/datasets/Reponx/network-cloud-ops-sft.texttext-generationn<1K0 likes36 downloads3d agoHugging Face10cloudbjorn /eschaton-uncensored-mini Eschaton Uncensored SFT Mini This is a 50-row, model-agnostic mini set sampled from the cloudbjorn/eschaton-uncensored dataset. It is useful for smoke-testing a conversational loader, chat-template rendering, tokenization, collation, and a short LoRA/SFT run before using the full 1,000-row dataset. Every row is copied verbatim from the full dataset. The mini set does not introduce model-specific chat tokens, mandatory reasoning wrappers, safety disclaimers, or rewritten answers.… See the full description on the dataset page: https://huggingface.co/datasets/cloudbjorn/eschaton-uncensored-mini.texttext-generationn<1K0 likes27 downloads3mo agoHugging Face11Clouds4days /tarotoo-tarot-card-meanings Tarotoo Tarot Card Meanings A complete, structured dataset of all 78 tarot cards (22 Major Arcana + 56 Minor Arcana) in the Rider–Waite–Smith tradition. Published by Tarotoo. These are the card meanings that ground the AI-generated readings on Tarotoo.com. Dataset details Curated by: Tarotoo (tarotoo.com) Language: English License: MIT Rows: 78 (one per card) · Fields: 22 DOI (Zenodo, cite this): 10.5281/zenodo.21514483 Concept DOI (Zenodo, always resolves to the… See the full description on the dataset page: https://huggingface.co/datasets/Clouds4days/tarotoo-tarot-card-meanings.tabulartext-generationn<1K0 likes10 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.