LLM-Agent
gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-uncensored-heretic-GGUFllmfan46-gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-uncensored-heretic-ROCMFPXnesso2-0.4B-agenticnesso-0.4B-agenticgemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-uncensored-hereticFabliq-8B-Agent-ReasoningFabliq-8B-Agent-FromBase-Reasoning-GGUFllm-agents_tora-code-7b-v1.0-GGUF
llm-agent-backup-20260930
Convention belief in LLM agents
An interpretability harness for zero-shot coordination in Hanabi. A language
model watches a partner that follows one of two signalling conventions and has
to work out which one. The interest is not in whether it succeeds but in where
in the network the failure lives: whether the convention is represented at all,
whether that representation drives the answer, and whether feeding it back
repairs behaviour.
Measurements and the design record are… See the full description on the dataset page: https://huggingface.co/datasets/taekbae/llm-agent-backup-20260930.agent-llm-traces-v2
Exgentic Agent LLM Traces v2 — Agent Chat Only
OpenTelemetry-shaped execution traces for 10,057 agent runs across 6 benchmarks (AppWorld, SWE-bench, BrowseCompPlus, τ²-bench Airline/Retail/Telecom), filtered to the agent under test's chat-only LLM calls. This is the dataset for replay testing, behavioral analysis, or any task where you care about what the benchmarked model actually did — not the eval scaffolding around it.
This v2 release expands upon Exgentic/agent-llm-traces… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces-v2.agent-llm-traces
Multi-Benchmark LLM Agent Traces
A comprehensive dataset of OpenTelemetry traces capturing LLM inference behavior across multiple agent frameworks, benchmarks, and model providers. This dataset enables research into LLM performance analysis, agent behavior patterns, and inference optimization.
Collected by Exgentic - A platform for LLM observability and performance optimization.
Dataset Overview
This dataset contains 1,781 execution traces capturing detailed agent… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces.details_llm-agents__tora-code-7b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-code-7b-v1.0
Dataset Summary
Dataset automatically created during the evaluation run of model llm-agents/tora-code-7b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-7b-v1.0.LLM-Agent-Harness-Survey
English | 中文
Agent Harness for Large Language Model Agents: A Survey
⭐ This repo is actively maintained. If you find it useful, please star the repo to stay updated and help others find it.
The agent execution harness — not the model — is the primary determinant of agent reliability at scale.This survey formalizes the harness as a first-class architectural object H = (E, T, C, S, L, V), surveys 110+ papers, blogs and reports across 23 systems, and maps 9 open… See the full description on the dataset page: https://huggingface.co/datasets/GloriaaaM/LLM-Agent-Harness-Survey.details_llm-agents__tora-code-34b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-code-34b-v1.0
Dataset automatically created during the evaluation run of model llm-agents/tora-code-34b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-34b-v1.0.
