code-agent
LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v2-i1-GGUFLFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3-i1-GGUFqwen2.5-coder-7b-agent-ggufFireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO-i1-GGUFLFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3-GGUFFireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO-GGUFLFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v2-GGUFFireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-GGUF
da-code-evaluation-resultsdetails_llm-agents__tora-code-7b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-code-7b-v1.0
Dataset Summary
Dataset automatically created during the evaluation run of model llm-agents/tora-code-7b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-7b-v1.0.details_llm-agents__tora-code-34b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-code-34b-v1.0
Dataset automatically created during the evaluation run of model llm-agents/tora-code-34b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-34b-v1.0.Tau2-Bench-Airline-With-Code-Agents
Dataset Card for a Code Agent Version of Tau Bench 2 Airline
Dataset Summary
This dataset includes sample traces and associated metadata from multi-turn interactions between an code agent and AI assistant. The dataset is based on the Airline environment from Tau^2 Bench and contains traces from both the original version and a version made at Snorkel AI using code agents to solve the same tasks (indicator in the version field; details below).
Curated by: Snorkel AI… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Tau2-Bench-Airline-With-Code-Agents.works_on_my_agent_code
Track B Phase 3 Submission
Team: Works on my agent
This archive contains the runnable submission for Track B Phase 3.
Environment
Python 3.11 is recommended for the inference runner:
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
Our local validation environment used Huawei Ascend 910B hardware.
The runner does not require internet access at runtime. It connects only to the local vLLM… See the full description on the dataset page: https://huggingface.co/datasets/jinao/works_on_my_agent_code.agentic_code_dataset_22Dataset: 22 Real Claude Code Sessions
To validate Suffix Decoding's applicability in Agentic Coding scenarios, we collected 22 complete Claude Code session recordings.
Dataset Overview
Metric
Value
Collection date
December 2025
Total sessions
22
Total conversation turns
17,487
Total runtime
50 hours
Total input tokens
6,996,619
Total output tokens
6,094,906
Session Scale Distribution
Statistic
Min
Max
Average
Conversation turns
273… See the full description on the dataset page: https://huggingface.co/datasets/novita/agentic_code_dataset_22.
