Team Ai
6 results

contextbench

Contextbench /ContextBench ContextBench This repository provides: default: the full ContextBench table (single train split). contextbench_verified: a 500-instance subset (single split). Columns The dataset uses a unified schema across sources: instance_id: ContextBench instance id (e.g., SWE-Bench-Verified__python__...). original_inst_id: Original benchmark instance id (e.g., astropy__astropy-14539). source: One of Verified, Pro, Poly, Multi. language: Programming language. repo_url: Repository… See the full description on the dataset page: https://huggingface.co/datasets/Contextbench/ContextBench.text1K<n<10K6 likes1.3k downloads9mo agoHugging FaceContextbench /Tracebench Tracebench This dataset contains agent trajectories (TerminalBench + SWE-bench) with two splits: full: 3316 trajectories (2670 terminal + 646 SWE-bench) verified: 1000 trajectories (489 SWE-bench + 511 terminal; terminal selected by step_count>=20, has incorrect steps, error-stage ratio threshold) Agents: mini-SWE-agent (1024), OpenHands (1242), Terminus2 (923), SWE-agent (127). Models: Anthropic/Claude-Sonnet-4, DeepSeek/DeepSeek-V3.2, Moonshot/Kimi-K2, OpenAI/GPT-5… See the full description on the dataset page: https://huggingface.co/datasets/Contextbench/Tracebench.tabular1K<n<10K1 likes1k downloads6mo agoHugging FaceContextbench /SWE-bench_Pro Dataset Summary SWE-Bench Pro is a challenging, enterprise-level dataset for testing agent ability on long-horizon software engineering tasks. Paper: https://static.scale.com/uploads/654197dc94d34f66c0f5184e/SWEAP_Eval_Scale%20(9).pdf See the related evaluation Github: https://github.com/scaleapi/SWE-bench_Pro-os Dataset Structure We follow SWE-Bench Verified (https://huggingface.co/datasets/SWE-bench/SWE-bench_Verified) in terms of dataset structure, with several… See the full description on the dataset page: https://huggingface.co/datasets/Contextbench/SWE-bench_Pro.textn<1K0 likes447 downloads10mo agoHugging Faceshshwtsuthar /memory-representation-contextbench-artifacts Memory Representation ContextBench Artifacts Dataset Summary This repository contains processed artifacts for the paper "Memory as a Map: Prior-Trajectory Representations for Software Engineering Agents." The artifact supports reproduction and inspection of a controlled prior-context representation experiment over SWEContextBench prior-target pairs. The experiment renders each target under four prompt conditions: no prior context, stripped Claude Code transcript… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-contextbench-artifacts.tabular1K<n<10K0 likes280 downloads4mo agoHugging FacegeminiDeveloper /contextBenchmark0 likes156 downloads25d agoHugging Faceshshwtsuthar /memory-representation-contextbench-traces Memory Representation ContextBench Raw Traces This optional artifact contains raw Claude Code prior JSONL traces discovered for the ContextBench prompt set. It includes 96 trace manifest rows and 42722114 bytes of copied JSONL content. OpenHands target-run JSONL traces were not present in the discovered source folders, so traces/openhands_runs/ is present as an empty directory structure and the absence is recorded in manifests/validation_summary.json. Checksums are in… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-contextbench-traces.tabularn<1K0 likes105 downloads4mo agoHugging Face