Team Ai
16 results

agent-benchmark

hlido-eu /agent-benchmark configs: - config_name: default data_files: - path: run-log.json split: train language: - en license: cc-by-nc-4.0 tags: - ai-agents - benchmark - evaluation - agent-evaluation - llm-agents - trustworthy-ai - product-review - c2pa pretty_name: Hlido AI Agent Benchmark size_categories: - n<1K task_categories: - other Hlido AI Agent Benchmark Independent, cryptographically-attested evaluations of AI agents and products. Hlido reviews AI… See the full description on the dataset page: https://huggingface.co/datasets/hlido-eu/agent-benchmark.1 likes2.4k downloads1d agoHugging Facevals-ai /finance_agent_benchmark Finance Agent Benchmark Dataset We present the Finance Agent Benchmark, featuring challenging and diverse real-world finance research problems which require LLMs to perform complex analysis with the use of of recent SEC filings. We construct the benchmark using a taxonomy of nine financial task categories, developed in consultation with experts from banks, hedge funds, and private equity firms. The dataset includes 537 expert-authored questions, covering tasks from information… See the full description on the dataset page: https://huggingface.co/datasets/vals-ai/finance_agent_benchmark.textn<1K11 likes1.9k downloads1y agoHugging FaceLakera /b3-agent-security-benchmark-weak[paper] [blogpost] [game] b3 AI Security Benchmark: Breaking Agent Backbones Highly contextalized prompt injections crowd-sourced during the Agent Breaker Challenge. This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents. The high quality dataset was used to evaluate the security of more than 30 LLMs. Dataset Summary Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.tabulartext-classificationn<1K9 likes1.3k downloads3d agoHugging FaceIntelligent-Internet /ii-agent_gaia-benchmark_validationtextn<1K8 likes830 downloads1y agoHugging Faceobaydata /mcp-agent-trajectory-benchmark MCP Agent Trajectory Benchmark A benchmark dataset of 49 MCP (Model Context Protocol) agent trajectories (38 single-pass + 11 multi-conv) with complete tool-use traces in the ATIF v1.2 (Agent Trajectory Interchange Format) format. Each agent operates in a distinct business domain with custom tools, realistic user conversations, and full execution traces. Designed for training and evaluating tool-use / function-calling capabilities of LLMs. Overview Item Details… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/mcp-agent-trajectory-benchmark.texttext-generationn<1K3 likes695 downloads7mo agoHugging FaceRegalFire /Agent-Failure-Recovery-Benchmark Agent Failure Recovery Benchmark 22,573 synthetic, source-verified failure → recovery trajectories across four domains. Evaluate whether a system can reject a failed plan, choose a recovery, or recognize that no recovery exists. Every row links to a hashed record in a pinned public source and is checked by independent computational oracles and trajectory replay. Tasks: failure diagnosis, recovery-action prediction, planning regression, scoped agent evaluation and imitation… See the full description on the dataset page: https://huggingface.co/datasets/RegalFire/Agent-Failure-Recovery-Benchmark.textreinforcement-learning10K<n<100K1 likes461 downloads4d agoHugging Face