agent-runtime
agent-runtime-recovery-bench
Agent Runtime Recovery Benchmark
This fully public dataset combines the current observation-restricted Agent
Runtime Recovery Benchmark with causally qualified native runtime cases for real
coding agents.
Contents
Source
Cases
Description
benchmark
769
Observation-restricted runtime recovery cases
real_agent_native
125
90 OpenHands and 35 mini-SWE-agent native runtime cases
Total
894
One unified public dataset
All rows are stored in one all… See the full description on the dataset page: https://huggingface.co/datasets/zitong1/agent-runtime-recovery-bench.agent-runtime-telemetry-small
Agent Runtime Telemetry Small
Curated by Faruk Alpay.
Agent Runtime Telemetry Small is a compact tabular export of MCP-style agent execution telemetry. It is designed for dataset viewer inspection, lightweight agent observability experiments, tool-call reliability analysis, workflow trace summaries, and audit-trail research.
The dataset is intentionally small and row-oriented. Each table is stored as Parquet so the Hugging Face Dataset Viewer can display clean columns without… See the full description on the dataset page: https://huggingface.co/datasets/Lightcap/agent-runtime-telemetry-small.
