inferencebench
inference-benchmarkerInferenceBench-Trajectories
InferenceBench — Agent Run Trajectories
Agent traces from InferenceBench
(GitHub), a benchmark that tests whether
frontier coding agents can optimize LLM serving under a fixed compute budget. The agents
know the techniques; the hard part is running, comparing, and keeping the ones that work.
Each run is one autonomous CLI agent attempting to deploy and optimize an OpenAI-compatible
inference server for a fixed base model (mistralai/Mistral-7B-Instruct-v0.3) under a… See the full description on the dataset page: https://huggingface.co/datasets/aisa-group/InferenceBench-Trajectories.inference-bench-resultsinference-benchmarker-xetiso-bench-openhands-gpt5-rebuttal
ISO-Bench OpenHands runs (GPT-5) — ICML 2026 rebuttal
OpenHands agent run artifacts for the ICML 2026 rebuttal of the ISO-Bench
paper. Companion to the Sonnet-4.5 rebuttal dataset
(Inferencebench/iso-bench-openhands-sonnet45-rebuttal); produced under the
same plans, same harness, same protocol — only the model and base URL differ.
🟢 COMPLETE — 53/54 successful real patches + 1 task where the agent
finished without editing target files.
Headline numbers
sglang… See the full description on the dataset page: https://huggingface.co/datasets/Inferencebench/iso-bench-openhands-gpt5-rebuttal.inference-benchmarker
