devstral
Datasets
All datasets matching “devstral”devstral-eval-logs-and-scoreslivesweagent-devstral2-123b-swebench-verified
Live-SWE-agent + Devstral 2 (123B) SWE-bench Verified Trajectories
This dataset contains agent trajectories from running Devstral 2 (123B) via OpenRouter on the SWE-bench Verified benchmark using the Live-SWE-agent framework.
⚠️ Framework Note
This run uses Live-SWE-agent, NOT standard mini-swe-agent. Live-SWE-agent is a self-evolving agent framework that encourages the model to create custom Python tools during runtime.
Key differences from mini-swe-agent:
Self-evolving… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/livesweagent-devstral2-123b-swebench-verified.Devstral-Small-2505-eval-logs-and-scoresdevstral-24b-swebench-verified-traj
Devstral-24B SWE-bench Verified Trajectories
This dataset contains agent trajectories from running Devstral-24B on the SWE-bench Verified benchmark.
Model Information
Attribute
Value
Model
Devstral-24B
Parameters
24B dense transformer
Provider
Local vLLM
Benchmark Results
Metric
Value
Total Instances
500
Submitted
345 (69.0%)
Resolved
277 (55.4%)
Usage
from huggingface_hub import hf_hub_download… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/devstral-24b-swebench-verified-traj.devstral_normal_devstral_small_2505_gdm_intercode_ctf
Inspect Dataset: devstral_normal_devstral_small_2505_gdm_intercode_ctf
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-17.
Model Information
Model: vllm/mistralai/Devstral-Small-2505
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'tokenizer_mode': 'mistral', 'config_format': 'mistral', 'load_format': 'mistral'… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/devstral_normal_devstral_small_2505_gdm_intercode_ctf.devstral2-123b-swebench-verified-traj
Devstral 2 (123B) SWE-bench Verified Trajectories
This dataset contains agent trajectories from running Devstral 2 (via OpenRouter) on the SWE-bench Verified benchmark using the mini-swe-agent framework.
Model Information
Attribute
Value
Model
Devstral 2
Parameters
123B (dense transformer)
Context Window
256K tokens
Provider
OpenRouter (mistralai/devstral-2512:free)
Specialization
Agentic coding
Framework
mini-swe-agent
Devstral 2 is a… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/devstral2-123b-swebench-verified-traj.
