Team Ai
20 results

code-agent

analytics-agents-uncertainty /da-code-evaluation-results0 likes1.6k downloads9mo agoHugging Faceopen-llm-leaderboard-old /details_llm-agents__tora-code-7b-v1.0 Dataset Card for Evaluation run of llm-agents/tora-code-7b-v1.0 Dataset Summary Dataset automatically created during the evaluation run of model llm-agents/tora-code-7b-v1.0 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-7b-v1.0.5 likes757 downloads3y agoHugging Faceopen-llm-leaderboard-old /details_llm-agents__tora-code-34b-v1.0 Dataset Card for Evaluation run of llm-agents/tora-code-34b-v1.0 Dataset automatically created during the evaluation run of model llm-agents/tora-code-34b-v1.0 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-34b-v1.0.1 likes676 downloads3y agoHugging Facesnorkelai /Tau2-Bench-Airline-With-Code-Agents Dataset Card for a Code Agent Version of Tau Bench 2 Airline Dataset Summary This dataset includes sample traces and associated metadata from multi-turn interactions between an code agent and AI assistant. The dataset is based on the Airline environment from Tau^2 Bench and contains traces from both the original version and a version made at Snorkel AI using code agents to solve the same tasks (indicator in the version field; details below). Curated by: Snorkel AI… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Tau2-Bench-Airline-With-Code-Agents.tabulartext-generationn<1K9 likes175 downloads10mo agoHugging Facejinao /works_on_my_agent_code Track B Phase 3 Submission Team: Works on my agent This archive contains the runnable submission for Track B Phase 3. Environment Python 3.11 is recommended for the inference runner: python3.11 -m venv .venv source .venv/bin/activate python -m pip install --upgrade pip python -m pip install -r requirements.txt Our local validation environment used Huawei Ascend 910B hardware. The runner does not require internet access at runtime. It connects only to the local vLLM… See the full description on the dataset page: https://huggingface.co/datasets/jinao/works_on_my_agent_code.0 likes175 downloads4mo agoHugging Facenovita /agentic_code_dataset_22Dataset: 22 Real Claude Code Sessions To validate Suffix Decoding's applicability in Agentic Coding scenarios, we collected 22 complete Claude Code session recordings. Dataset Overview Metric Value Collection date December 2025 Total sessions 22 Total conversation turns 17,487 Total runtime 50 hours Total input tokens 6,996,619 Total output tokens 6,094,906 Session Scale Distribution Statistic Min Max Average Conversation turns 273… See the full description on the dataset page: https://huggingface.co/datasets/novita/agentic_code_dataset_22.5 likes156 downloads9mo agoHugging Face