contextbench
ContextBench
ContextBench
This repository provides:
default: the full ContextBench table (single train split).
contextbench_verified: a 500-instance subset (single split).
Columns
The dataset uses a unified schema across sources:
instance_id: ContextBench instance id (e.g., SWE-Bench-Verified__python__...).
original_inst_id: Original benchmark instance id (e.g., astropy__astropy-14539).
source: One of Verified, Pro, Poly, Multi.
language: Programming language.
repo_url: Repository… See the full description on the dataset page: https://huggingface.co/datasets/Contextbench/ContextBench.Tracebench
Tracebench
This dataset contains agent trajectories (TerminalBench + SWE-bench) with two splits:
full: 3316 trajectories (2670 terminal + 646 SWE-bench)
verified: 1000 trajectories (489 SWE-bench + 511 terminal; terminal selected by step_count>=20, has incorrect steps, error-stage ratio threshold)
Agents: mini-SWE-agent (1024), OpenHands (1242), Terminus2 (923), SWE-agent (127).
Models: Anthropic/Claude-Sonnet-4, DeepSeek/DeepSeek-V3.2, Moonshot/Kimi-K2, OpenAI/GPT-5… See the full description on the dataset page: https://huggingface.co/datasets/Contextbench/Tracebench.SWE-bench_Pro
Dataset Summary
SWE-Bench Pro is a challenging, enterprise-level dataset for testing agent ability on long-horizon software engineering tasks.
Paper: https://static.scale.com/uploads/654197dc94d34f66c0f5184e/SWEAP_Eval_Scale%20(9).pdf
See the related evaluation Github: https://github.com/scaleapi/SWE-bench_Pro-os
Dataset Structure
We follow SWE-Bench Verified (https://huggingface.co/datasets/SWE-bench/SWE-bench_Verified) in terms of dataset structure, with several… See the full description on the dataset page: https://huggingface.co/datasets/Contextbench/SWE-bench_Pro.memory-representation-contextbench-artifacts
Memory Representation ContextBench Artifacts
Dataset Summary
This repository contains processed artifacts for the paper "Memory as a Map: Prior-Trajectory Representations for Software Engineering Agents." The artifact supports reproduction and inspection of a controlled prior-context representation experiment over SWEContextBench prior-target pairs.
The experiment renders each target under four prompt conditions: no prior context, stripped Claude Code transcript… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-contextbench-artifacts.contextBenchmarkmemory-representation-contextbench-traces
Memory Representation ContextBench Raw Traces
This optional artifact contains raw Claude Code prior JSONL traces discovered for the ContextBench prompt set. It includes 96 trace manifest rows and 42722114 bytes of copied JSONL content.
OpenHands target-run JSONL traces were not present in the discovered source folders, so traces/openhands_runs/ is present as an empty directory structure and the absence is recorded in manifests/validation_summary.json.
Checksums are in… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-contextbench-traces.
