agent-collab
AgentCollabBench
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators
AgentCollabBench is a diagnostic benchmark dataset for multi-agent LLM systems, targeting process-level failures that single-agent benchmarks cannot expose.
Overview
Multi-agent systems introduce failure modes that emerge specifically from inter-agent communication: constraint decay under peer pressure, multi-hop information loss, false-belief propagation, and private context leakage. AgentCollabBench… See the full description on the dataset page: https://huggingface.co/datasets/AgentCollabBench/AgentCollabBench.collaborative_agent_benchThis dataset is released as part of SWEET-RL: Training Multi-Turn LLM Agents on
Collaborative Reasoning Tasks research project.
Please refer to our project materials here for training and evaluation details.
Citation
If you use data, model, or code from this work, please cite with the following BibTex entry:
@misc{zhou2025sweetrltrainingmultiturnllm,
title={SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks},
author={Yifei Zhou and Song Jiang and… See the full description on the dataset page: https://huggingface.co/datasets/facebook/collaborative_agent_bench.repro-latent-collaboration-in-multi-agent-systems-bundle
Reproduction: Latent Collaboration in Multi-Agent Systems (LatentMAS)
OpenReview: syG9I9ofd8 | arXiv: 2511.20639 | Official code: https://github.com/Gen-Verse/LatentMAS
Agent: CLAUDE | Max points: 12
Independent reproduction focusing on the paper's three theorems (CPU numerical audits) plus
a toy of the latent-thoughts / latent-working-memory mechanism. See protocol.md.
Rerun (one command)
python src/run_all.py # runs all audits, writes outputs/ +… See the full description on the dataset page: https://huggingface.co/datasets/MarxistLeninist/repro-latent-collaboration-in-multi-agent-systems-bundle.repro-latent-collaboration-in-multi-agent-systems-traces
Agent traces
Agent sessions published from a Trackio Logbook.
han-multi-agent-collaboration-logs-v1
Humanoid Multi-Agent Collaboration Logs
This dataset records collaboration events between
multiple humanoid agents working toward shared goals.
It enables learning coordination, role assignment,
and collective decision-making.
Contents
Agent roles
Shared objectives
Coordination actions
Collaboration outcomes
Use Cases
Swarm intelligence
Multi-agent planning
Cooperative task learning
Part of
Humanoid Network (HAN)
License
MIT
collaborative_agent_benchThis dataset is released as part of SWEET-RL: Training Multi-Turn LLM Agents on
Collaborative Reasoning Tasks research project.
Please refer to our project materials here for training and evaluation details.
Citation
If you use data, model, or code from this work, please cite with the following BibTex entry:
@misc{zhou2025sweetrltrainingmultiturnllm,
title={SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks},
author={Yifei Zhou and Song Jiang and… See the full description on the dataset page: https://huggingface.co/datasets/mengping03/collaborative_agent_bench.
