datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
workflow-curatedcho-workflowworkflow-templatesforecast-workflow-bench
Forecast Workflow Bench (FWBench)
1,251 electricity and cycle-hire cases for evaluating LLM and SLM decisions
with budgeted forecasting tools.
Code ·
Leaderboard
Evaluation protocol
The Language Model evaluation covers
1,251 cases, two hosted and eight local configurations, and 22,518 conversations
including paired local TSFM-removal runs.
Agents minimize S = 0.5 * ((loss - F) / sigma + credits / B), where F is the
minimum feasible loss with known target demand… See the full description on the dataset page: https://huggingface.co/datasets/Neurogica/forecast-workflow-bench.WorkflowPerturb
WorkflowPerturb — Dataset Artifact
Companion data for the EMNLP 2026 Industry Track paper
“WorkflowPerturb: Calibrated Stress Tests for Evaluating Multi-Agent Workflow Metrics.”
Canonical location: https://huggingface.co/datasets/microsoft/WorkflowPerturbPaper: https://arxiv.org/abs/2602.17990
This release is the complete WorkflowPerturb benchmark plus documentation. It is
self-contained: the CSVs carry every golden workflow, every perturbed variant, and all
shipped pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/WorkflowPerturb.dxap-research-workflow-excerpts
DXAP historical workflow excerpts and selection aggregates
Three detailed historical case studies: 29 parent/reconciliation event rows, 18 sanitized tool-request/response projections, five starting-state records, two research-child event records and nine proposed scenario questions. It also includes the complete 15-row selection-rank table already published with arXiv:2609.05663v1. These are different views of the same selected material, not independent sample counts to add… See the full description on the dataset page: https://huggingface.co/datasets/DXRG/dxap-research-workflow-excerpts.enterprise_erp_workflow_reasoning
Enterprise ERP Workflow Reasoning
Author: Venkata Ramachandra Karthik Chundi (venkatakarthikchundi@gmail.com)
A multiple-choice benchmark testing LLM reasoning on enterprise ERP business process workflows, approval hierarchies, document relationships, and process sequencing. Scenarios are drawn from real-world Oracle ERP Cloud deployment contexts.
Dataset
50 multiple-choice questions with 4 answer options each. One correct answer per question.
Modules Covered… See the full description on the dataset page: https://huggingface.co/datasets/karthikchundi/enterprise_erp_workflow_reasoning.Arena-DROID-Camera-Sensitivity-Workflow-Sample
Arena DROID Camera Sensitivity Workflow Sample
Dataset Description
Arena-DROID-Camera-Sensitivity-Workflow-Sample is a compact set of episode-level results generated by an Isaac Lab-Arena simulation experiment. It lets users run the documented camera sensitivity analysis without first executing the policy-evaluation sweep.
The experiment evaluates an OpenPI pi05 policy on a DROID Rubik's-cube pick-and-place task while independently varying the wrist-camera… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-DROID-Camera-Sensitivity-Workflow-Sample.vcore-workflow-envelopes
Authority envelopes of public workflows using claude-code-action (de-identified)
Links
Blog post: https://paulinebourigault.github.io/blog/2026/what-the-agent-may-do/
Code (vcore): https://github.com/certior/vcore
Companion dataset (benchmark): https://huggingface.co/datasets/paulibo/vcore-actions-benchmark
Derived analysis of 16,346 distinct public GitHub workflow files that invoke
anthropics/claude-code-action, collected through GitHub code search on… See the full description on the dataset page: https://huggingface.co/datasets/paulibo/vcore-workflow-envelopes.agentic-workflow
$content = @"
license: apache-2.0
language:
- en
tags:
- enterprise
- agentic-workflows
- mlops
- distributed-training
- vector-search
- error-recovery
task_categories:
- text-generation
- conversational
pretty_name: "Enterprise Agentic Workflow & Multi-Turn Corpus"
Enterprise Agentic Workflow & Multi-Turn Execution Corpus
A high-entropy, production-grade dataset designed for fine-tuning autonomous AI agents across complex distributed systems, MLOps… See the full description on the dataset page: https://huggingface.co/datasets/bala5046/agentic-workflow.ai-paper-workflow-eval
AI Paper Workflow Evaluation
Frozen GPT-6 Luna Flex outputs and historical research-paper controls. Evaluation only: keep test out of training, prompt development and threshold selection. 162 main papers: 54 validation and 108 test, stratified across NeurIPS, ICML and ACL, 2013–2021. The 27 development pilot papers are excluded. Author-name connected components do not cross splits; this is not perfect author identity resolution.
from datasets import load_dataset
# Pin revision… See the full description on the dataset page: https://huggingface.co/datasets/woog/ai-paper-workflow-eval.ecommerce-retail-product-matching-workflow-dataset
Ecommerce Retail Product Matching Workflow Dataset
This dataset is a public-facing sanitized workflow preview for managed ecommerce and retail product matching. It shows how candidate retrieval, UPC/model/brand/title/image evidence, customer-visible URL validation, confidence bands, review buckets, and rejection reasons can be structured for pricing intelligence, merchandising, data engineering, and AI-assisted product matching workflows.
Use this dataset to evaluate product… See the full description on the dataset page: https://huggingface.co/datasets/Octoparse/ecommerce-retail-product-matching-workflow-dataset.flowzap-sequence-workflows
sequence-workflows
A synchronized FlowZap template corpus with 242 canonical templates sourced from https://flowzap.xyz/sitemap-templates.xml and organized by primary Use Case.
Organization Model
Top-level folders are primary Use Cases from the FlowZap Templates dropdown.
Second-level folders preserve the original source domain from the FlowZap app index.
Each template keeps all matched Use Cases in metadata.json and the generated JSON/CSV indexes.
Templates that do not… See the full description on the dataset page: https://huggingface.co/datasets/Jules-OC/flowzap-sequence-workflows.lab09-agents-and-workflowsdino-data-workflow-routing-preview
Dino Data Workflow Routing Preview
What This Dataset Is
This dataset is a focused workflow-routing preview built from six Dino Data capability slices:
connector intent detection
connector action mapping
deeplink action mapping
document export specification
zip packaging specification
deeplink intent detection
The goal is to train or inspect assistant behavior around workflow-aware task handling:
detecting when a request should route into an action or product workflow… See the full description on the dataset page: https://huggingface.co/datasets/DinoDS/dino-data-workflow-routing-preview.temu-ecommerce-pricing-workflow-sample
Temu E-commerce Pricing & SKU Variant Dataset
85 products · 382 SKUs · $2.14–$219.72 price range · avg 29.5% discount where measurable
A real, production-quality sample of Temu product listings and SKU-level pricing data, captured by Octoparse Managed Data Service via a managed anti-bot pipeline. Every row is real market data — no synthetic expansion, no mock prices.
Built for teams working on competitor price monitoring, dynamic pricing models, product matching, and e-commerce AI… See the full description on the dataset page: https://huggingface.co/datasets/Octoparse/temu-ecommerce-pricing-workflow-sample.oracle-erp-workflow-eval
Oracle ERP Cloud Workflow & Terminology Eval
A benchmark dataset for evaluating large language models on Oracle ERP Cloud knowledge — including workflows, terminology, document types, approval hierarchies, and process flows across seven core Oracle Cloud modules.
Overview
Property
Value
Samples
30
Modules covered
7
Task type
Short-answer question answering
Eval framework
OpenAI Evals (FuzzyMatch)
Metric
Accuracy (fuzzy string match)… See the full description on the dataset page: https://huggingface.co/datasets/karthikchundi/oracle-erp-workflow-eval.MMH3_Image_Edit_WorkflowThis is just an example of using MiniMax H3 as an image editor. The actual workflow that I use requires several custom nodes, some of which are not published, so this one is simply a bare bones demonstration.
This uses the hybrid MiniMax H3 model from here: https://huggingface.co/smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models/tree/main
It uses the custom VAE from here: https://huggingface.co/Mamad8/MiniMax-H3-Image-VAE/tree/main
It uses the LoRA from here:… See the full description on the dataset page: https://huggingface.co/datasets/fizzlepoof/MMH3_Image_Edit_Workflow.gh-agentic-workflows
⛏️ PeakyMiner GH-AW Dataset
Dataset estructurado y normalizado de GitHub Agentic Workflows (GH-AW) extraído de repositorios públicos de GitHub mediante PeakyMiner.
📊 Resumen del Dataset
Este dataset contiene la especificación, instrucciones (prompts), metadatos y disparadores de flujos de trabajo de agentes autónomos basados en Markdown compilado a GitHub Actions (.md y .lock.yml).
Estadísticas de Extracción
Repositorios únicos analizados: 292… See the full description on the dataset page: https://huggingface.co/datasets/apa9s/gh-agentic-workflows.slm-workflow-planner-policy-v2
SLM Workflow Planner — Policy-Corrected Instruction Tuning Dataset (v2)
Overview
High-quality instruction-tuning dataset for training a Small Language Model (SLM)
to serve as a workflow execution planner. The model learns to make policy-aware
decisions about workflow transitions: when to proceed (NEXT), retry (RETRY),
parallelize (FORK), synchronize (JOIN), or escalate (META).
Key Features
648K instruction pairs across 2 stages (decision type + node… See the full description on the dataset page: https://huggingface.co/datasets/ssaraf1/slm-workflow-planner-policy-v2.responsible-agent-workflow-evaluation
Responsible Agent Workflow Evaluation
Version 1.0.0 contains 130 wholly synthetic scenarios for evaluating
whether an AI agent respects safety, permission and accountability boundaries
in operational settings. Thirteen categories contain ten scenarios each. Every
record includes an intentionally unsafe request, contextual facts, expected
safe behaviour, explicitly prohibited behaviour, severity, evaluation criteria
and reviewer guidance.
This is a red-team and… See the full description on the dataset page: https://huggingface.co/datasets/nwhite-systems/responsible-agent-workflow-evaluation.agentic-workflows-sft-100k
Agentic Workflows SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality conversations covering AI agent architectures, tool use patterns, multi-agent systems, and agent evaluation. Designed to train AI assistants that can help engineers design, build, and debug production AI agents.
Dataset Description
This dataset covers the full spectrum of agentic AI development across 9 specialized categories. Each record follows the ShareGPT format with… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/agentic-workflows-sft-100k.gh-aw-workflows-dataset
GH-AW Workflows Dataset
Dataset relacional con repositorios de GitHub que usan
GitHub Agentic Workflows (GH-AW), y el
contenido estructurado de sus archivos de workflow: el .md fuente
(frontmatter YAML + body Markdown) y su .lock.yml/.lock.yaml compilado.
Generado con Miner, una CLI en
Python que identifica repositorios que usan GH-AW y extrae sus pares
.md + .lock en un formato analítico.
Repositorios incluidos: 338
Archivos de workflow incluidos: 1390 (cada uno con su .lock… See the full description on the dataset page: https://huggingface.co/datasets/EstebanCQ/gh-aw-workflows-dataset.zarn-workflow-automation-instruct
Zarn Workflow Automation Instruct
Dataset Description
Natural-language workplace requests paired with plans and JSON tool actions.
Team Attribution
This dataset was created and reviewed by the Zarnite team through internal benchmark design, generation, and quality-control workflows. It should be presented as a Zarnite-authored benchmark starter pack, not as a purely human-collected field corpus.
Ecosystem Need Tier
High Ecosystem Need
Why… See the full description on the dataset page: https://huggingface.co/datasets/zarnite/zarn-workflow-automation-instruct.data-science-workflows-sft-100k
Data Science Workflows SFT (100K)
100,000 ShareGPT conversations demonstrating expert-level data science practice across data cleaning, EDA, ML pipelines, feature engineering, SQL analytics, statistical analysis, model evaluation, visualization, and production deployment.
Motivation
Data science is one of the most in-demand technical skills — companies need models that can reason through real analytical problems with the rigor of a senior data scientist. Models… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/data-science-workflows-sft-100k.gh-aw-workflows
GitHub Agentic Workflows (GH-AW) Dataset
Dataset relacional construido con Miner
a partir de repositorios de GitHub que adoptaron
GitHub Agentic Workflows (GH-AW). Para
cada repositorio se identificaron sus workflows agénticos (pares
<base>.md + <base>.lock.yml dentro de .github/workflows/), se descargó
el contenido de cada .md, y se separó su frontmatter YAML (configuración
del workflow: disparadores, permisos, motor de IA, herramientas, etc.) de su
body en Markdown (el… See the full description on the dataset page: https://huggingface.co/datasets/luchosqi/gh-aw-workflows.gradio-workflow-eval-arena
gr.Workflow Eval Arena — sample rows
Five rows, one per scorer branch, for the gr.Workflow eval arena example.
column
meaning
id
stable row identifier
task_type
exact, numeric, json, free-form, truncation
prompt
the prompt sent to every candidate
gold
reference answer
meta
free-text note
Deliberately tiny: it exists to exercise every scoring path, not to rank models. Any conclusion about model quality drawn from five rows is not statistically meaningful.
n8n-workflows-2k
Dataset Card for N8n Workflows 2k
This dataset contains 2000 n8n workflows.
Curated by: Arkel AI
Funded by: Arkel AI
Language(s) (NLP): English
License: Apache 2.0
n8n-mega-workflows
🚀 n8n Mega Workflows - The Largest n8n Workflow Dataset
The world's largest open-source n8n workflow dataset for training AI workflow generators
🌟 Highlights
131,648 high-quality n8n workflows with valid connection skeletons
26 semantic categories for balanced coverage
Instruction-tuning format ready for fine-tuning LLMs
15+ million lines of workflow JSON
Perfect for: RAG pipelines, fine-tuning, workflow generation models
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/Ker102/n8n-mega-workflows.comfyui-workflows-dataset
ComfyUI Workflows Dataset
A comprehensive training dataset for fine-tuning language models to assist with ComfyUI workflows.
Dataset Overview
Total Examples: 702
Format: JSONL with chat messages
Size: 31.17 MB
Categories
Workflows (677 examples)
text-to-image (194): Basic image generation workflows
image-upscaling (192): Upscaling and enhancement workflows
video-animation (122): AnimateDiff and video generation… See the full description on the dataset page: https://huggingface.co/datasets/lokegud/comfyui-workflows-dataset.
