dev-ops
Datasets
All datasets matching “dev-ops”devopsbench-100
DevOpsBench-100
DevOpsBench-100 is a synthetic long-horizon software-engineering / SRE agent
benchmark: 100 tasks over one executable world ("NovaCart", a mid-size
e-commerce SaaS) with 72 SQLite tables,
1451 seeded rows, a 38-file monorepo with 417 commits,
and 97 MCP tools spanning a first-party engineering stack
(tickets, PRs, CI, deployments, canaries, migrations, feature flags, metrics,
alerts, incidents, chat, knowledge base) plus deliberately disagreeing
vendor-shaped… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/devopsbench-100.stack-v3-devops
The Stack v3 DevOps Corpus
13,234,862 complete infrastructure units extracted from
The Stack v3,
grouped into seven classes and gated on content rather than popularity.
A unit is not a file, it is the thing an engineer would actually run: a Helm chart
arrives with its Chart.yaml, values.yaml and every template; a Terraform module
with all of its .tf files; an Ansible role with its tasks, defaults and handlers.
That is only possible because The Stack v3 groups rows by repository… See the full description on the dataset page: https://huggingface.co/datasets/Helmcode/stack-v3-devops.devops-incident-response
Dataset Card for DevOps Incident Response Dataset
Dataset Description
Dataset Summary
The DevOps Incident Response Dataset is a comprehensive collection of real-world-style DevOps incidents, troubleshooting scenarios, and resolution procedures. This dataset is designed to help train AI models for DevOps assistance, incident response automation, and technical troubleshooting education.
Each incident includes:
Detailed incident description and symptoms… See the full description on the dataset page: https://huggingface.co/datasets/Snaseem2026/devops-incident-response.devops-qa-dataset
DevOps Q&A Dataset v1.0
Overview
High-quality dataset of 25,670 DevOps technical examples collected from GitHub repositories, Stack Exchange, and official documentation.
Statistics
Total examples: 25,670
Average quality score: ~0.82
Unique (deduplicated): High accuracy via MD5
Categories: Docker, Kubernetes, CI/CD, Cloud, Linux, Terraform, Ansible
Sources: StackExchange/HuggingFace (70%), GitHub Repositories (29%), Official Documentation (~1%)
Use… See the full description on the dataset page: https://huggingface.co/datasets/Skilln/devops-qa-dataset.omnimcp_devops_cloud_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_devops_cloud_teaser.CodeFuse-DevOps-EvalDevOps-Eval is a comprehensive chinese evaluation suite specifically designed for foundation models in the DevOps field. It consists of 5977 multi-choice questions spanning 55 diverse categories. Please visit our website and GitHub for more details.
Each category consists of two splits: dev, and test. The dev set per subject consists of five exemplars with explanations for few-shot evaluation. And the test set is for model evaluation. Labels on the test split are released, users can evaluate… See the full description on the dataset page: https://huggingface.co/datasets/codefuse-ai/CodeFuse-DevOps-Eval.
