Team Ai
9 results

inferencebench

hlarcher /inference-benchmarkertext100K<n<1M1 likes484 downloads2y agoHugging Faceaisa-group /InferenceBench-Trajectories InferenceBench — Agent Run Trajectories Agent traces from InferenceBench (GitHub), a benchmark that tests whether frontier coding agents can optimize LLM serving under a fixed compute budget. The agents know the techniques; the hard part is running, comparing, and keeping the ones that work. Each run is one autonomous CLI agent attempting to deploy and optimize an OpenAI-compatible inference server for a fixed base model (mistralai/Mistral-7B-Instruct-v0.3) under a… See the full description on the dataset page: https://huggingface.co/datasets/aisa-group/InferenceBench-Trajectories.text-generationn<1K3 likes271 downloads4mo agoHugging Faceaimosprite /inference-bench-results0 likes198 downloads6mo agoHugging Facehlarcher /inference-benchmarker-xettext100K<n<1M0 likes114 downloads2y agoHugging FaceInferencebench /iso-bench-openhands-gpt5-rebuttal ISO-Bench OpenHands runs (GPT-5) — ICML 2026 rebuttal OpenHands agent run artifacts for the ICML 2026 rebuttal of the ISO-Bench paper. Companion to the Sonnet-4.5 rebuttal dataset (Inferencebench/iso-bench-openhands-sonnet45-rebuttal); produced under the same plans, same harness, same protocol — only the model and base URL differ. 🟢 COMPLETE — 53/54 successful real patches + 1 task where the agent finished without editing target files. Headline numbers sglang… See the full description on the dataset page: https://huggingface.co/datasets/Inferencebench/iso-bench-openhands-gpt5-rebuttal.text-generationn<1K0 likes106 downloads5mo agoHugging Facefabric /inference-benchmarkertext100K<n<1M0 likes88 downloads7mo agoHugging Face