datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
workflow-curatedcho-workflowworkflow-templatesforecast-workflow-bench
Forecast Workflow Bench (FWBench)
1,251 electricity and cycle-hire cases for evaluating LLM and SLM decisions
with budgeted forecasting tools.
Code ·
Leaderboard
Evaluation protocol
The Language Model evaluation covers
1,251 cases, two hosted and eight local configurations, and 22,518 conversations
including paired local TSFM-removal runs.
Agents minimize S = 0.5 * ((loss - F) / sigma + credits / B), where F is the
minimum feasible loss with known target demand… See the full description on the dataset page: https://huggingface.co/datasets/Neurogica/forecast-workflow-bench.enterprise_erp_workflow_reasoning
Enterprise ERP Workflow Reasoning
Author: Venkata Ramachandra Karthik Chundi (venkatakarthikchundi@gmail.com)
A multiple-choice benchmark testing LLM reasoning on enterprise ERP business process workflows, approval hierarchies, document relationships, and process sequencing. Scenarios are drawn from real-world Oracle ERP Cloud deployment contexts.
Dataset
50 multiple-choice questions with 4 answer options each. One correct answer per question.
Modules Covered… See the full description on the dataset page: https://huggingface.co/datasets/karthikchundi/enterprise_erp_workflow_reasoning.Arena-DROID-Camera-Sensitivity-Workflow-Sample
Arena DROID Camera Sensitivity Workflow Sample
Dataset Description
Arena-DROID-Camera-Sensitivity-Workflow-Sample is a compact set of episode-level results generated by an Isaac Lab-Arena simulation experiment. It lets users run the documented camera sensitivity analysis without first executing the policy-evaluation sweep.
The experiment evaluates an OpenPI pi05 policy on a DROID Rubik's-cube pick-and-place task while independently varying the wrist-camera… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-DROID-Camera-Sensitivity-Workflow-Sample.vcore-workflow-envelopes
Authority envelopes of public workflows using claude-code-action (de-identified)
Links
Blog post: https://paulinebourigault.github.io/blog/2026/what-the-agent-may-do/
Code (vcore): https://github.com/certior/vcore
Companion dataset (benchmark): https://huggingface.co/datasets/paulibo/vcore-actions-benchmark
Derived analysis of 16,346 distinct public GitHub workflow files that invoke
anthropics/claude-code-action, collected through GitHub code search on… See the full description on the dataset page: https://huggingface.co/datasets/paulibo/vcore-workflow-envelopes.flowzap-sequence-workflows
sequence-workflows
A synchronized FlowZap template corpus with 242 canonical templates sourced from https://flowzap.xyz/sitemap-templates.xml and organized by primary Use Case.
Organization Model
Top-level folders are primary Use Cases from the FlowZap Templates dropdown.
Second-level folders preserve the original source domain from the FlowZap app index.
Each template keeps all matched Use Cases in metadata.json and the generated JSON/CSV indexes.
Templates that do not… See the full description on the dataset page: https://huggingface.co/datasets/Jules-OC/flowzap-sequence-workflows.oracle-erp-workflow-eval
Oracle ERP Cloud Workflow & Terminology Eval
A benchmark dataset for evaluating large language models on Oracle ERP Cloud knowledge — including workflows, terminology, document types, approval hierarchies, and process flows across seven core Oracle Cloud modules.
Overview
Property
Value
Samples
30
Modules covered
7
Task type
Short-answer question answering
Eval framework
OpenAI Evals (FuzzyMatch)
Metric
Accuracy (fuzzy string match)… See the full description on the dataset page: https://huggingface.co/datasets/karthikchundi/oracle-erp-workflow-eval.MMH3_Image_Edit_WorkflowThis is just an example of using MiniMax H3 as an image editor. The actual workflow that I use requires several custom nodes, some of which are not published, so this one is simply a bare bones demonstration.
This uses the hybrid MiniMax H3 model from here: https://huggingface.co/smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models/tree/main
It uses the custom VAE from here: https://huggingface.co/Mamad8/MiniMax-H3-Image-VAE/tree/main
It uses the LoRA from here:… See the full description on the dataset page: https://huggingface.co/datasets/fizzlepoof/MMH3_Image_Edit_Workflow.slm-workflow-planner-policy-v2
SLM Workflow Planner — Policy-Corrected Instruction Tuning Dataset (v2)
Overview
High-quality instruction-tuning dataset for training a Small Language Model (SLM)
to serve as a workflow execution planner. The model learns to make policy-aware
decisions about workflow transitions: when to proceed (NEXT), retry (RETRY),
parallelize (FORK), synchronize (JOIN), or escalate (META).
Key Features
648K instruction pairs across 2 stages (decision type + node… See the full description on the dataset page: https://huggingface.co/datasets/ssaraf1/slm-workflow-planner-policy-v2.agentic-workflows-sft-100k
Agentic Workflows SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality conversations covering AI agent architectures, tool use patterns, multi-agent systems, and agent evaluation. Designed to train AI assistants that can help engineers design, build, and debug production AI agents.
Dataset Description
This dataset covers the full spectrum of agentic AI development across 9 specialized categories. Each record follows the ShareGPT format with… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/agentic-workflows-sft-100k.zarn-workflow-automation-instruct
Zarn Workflow Automation Instruct
Dataset Description
Natural-language workplace requests paired with plans and JSON tool actions.
Team Attribution
This dataset was created and reviewed by the Zarnite team through internal benchmark design, generation, and quality-control workflows. It should be presented as a Zarnite-authored benchmark starter pack, not as a purely human-collected field corpus.
Ecosystem Need Tier
High Ecosystem Need
Why… See the full description on the dataset page: https://huggingface.co/datasets/zarnite/zarn-workflow-automation-instruct.data-science-workflows-sft-100k
Data Science Workflows SFT (100K)
100,000 ShareGPT conversations demonstrating expert-level data science practice across data cleaning, EDA, ML pipelines, feature engineering, SQL analytics, statistical analysis, model evaluation, visualization, and production deployment.
Motivation
Data science is one of the most in-demand technical skills — companies need models that can reason through real analytical problems with the rigor of a senior data scientist. Models… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/data-science-workflows-sft-100k.gradio-workflow-eval-arena
gr.Workflow Eval Arena — sample rows
Five rows, one per scorer branch, for the gr.Workflow eval arena example.
column
meaning
id
stable row identifier
task_type
exact, numeric, json, free-form, truncation
prompt
the prompt sent to every candidate
gold
reference answer
meta
free-text note
Deliberately tiny: it exists to exercise every scoring path, not to rank models. Any conclusion about model quality drawn from five rows is not statistically meaningful.
n8n-workflows-2k
Dataset Card for N8n Workflows 2k
This dataset contains 2000 n8n workflows.
Curated by: Arkel AI
Funded by: Arkel AI
Language(s) (NLP): English
License: Apache 2.0
n8n-mega-workflows
🚀 n8n Mega Workflows - The Largest n8n Workflow Dataset
The world's largest open-source n8n workflow dataset for training AI workflow generators
🌟 Highlights
131,648 high-quality n8n workflows with valid connection skeletons
26 semantic categories for balanced coverage
Instruction-tuning format ready for fine-tuning LLMs
15+ million lines of workflow JSON
Perfect for: RAG pipelines, fine-tuning, workflow generation models
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/Ker102/n8n-mega-workflows.comfyui-workflows-dataset
ComfyUI Workflows Dataset
A comprehensive training dataset for fine-tuning language models to assist with ComfyUI workflows.
Dataset Overview
Total Examples: 702
Format: JSONL with chat messages
Size: 31.17 MB
Categories
Workflows (677 examples)
text-to-image (194): Basic image generation workflows
image-upscaling (192): Upscaling and enhancement workflows
video-animation (122): AnimateDiff and video generation… See the full description on the dataset page: https://huggingface.co/datasets/lokegud/comfyui-workflows-dataset.test_mindat_workflowmira-workflow-v12workflow-ai-agent
Workflow Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP
This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc.
The dataset is helpful for AI… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/workflow-ai-agent.slm-workflow-planner-v8-datasetn8n-workflows-v2-4k
Dataset Card for N8n Workflows v2 4.3K
Dataset Description
This dataset contains 4,300 curated question-answer pairs for generating n8n workflows from natural language descriptions. It's designed to train models that can convert natural language requests into functional n8n automation workflows.
What is n8n?
n8n is an open-source workflow automation tool that allows you to connect different services and apps through a visual, no-code interface. It enables users… See the full description on the dataset page: https://huggingface.co/datasets/arkelai/n8n-workflows-v2-4k.n8n-workflows-sftn8n-workflowreconcile-workflow-lab
Reconcile Workflow Lab: eight-domain candidate
An unsubmitted local prototype with original synthetic tasks across software engineering, industrial and physical systems, natural science, office and white-collar work, finance and economics, math and formal reasoning, cybersecurity, and media and content production.
The agent inspects inputs, makes changes, executes a process, inspects its results, revises if necessary and commits the outcome. Distinct workflows include repository… See the full description on the dataset page: https://huggingface.co/datasets/qingyuanwu/reconcile-workflow-lab.data_agent_workflow_generate_medical_datacreate by data agent workflow
fiscus-financial-workflow-state-fixtures
FISCUS AI Financial Workflow State Fixtures
This small, fully synthetic reference dataset demonstrates how finance-critical workflow events can be classified without confusing an AI suggestion, a user approval, an external-provider response, and a committed accounting effect.
The fixtures are intended for:
validating deterministic workflow guards;
testing human-in-the-loop finance assistants;
reviewing retry and idempotency behavior;
checking that user-facing status matches the… See the full description on the dataset page: https://huggingface.co/datasets/keaxe21/fiscus-financial-workflow-state-fixtures.slm-workflow-planner-alignment-v2han-multi-step-domestic-workflows-v1
Multi-Step Domestic Workflows
Overview
A structured dataset containing
multi-step domestic workflows executed
by humanoid robots.
Each entry represents a complete
sequence from preparation to completion.
Data Fields
workflow_name
ordered_steps
required_tools
final_status
Intended Use
Sequential modeling
Workflow optimization
Task decomposition research
License
MIT
