hugging-apps/toolcua-gui-tool-agent
ToolCUA-8B — GUI-Tool Path Orchestration
Demo for ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents (paper · code · weights).
ToolCUA-8B is an 8B computer-use agent built on Qwen3-VL. In a hybrid desktop environment the agent can either keep acting through the GUI (click, type, scroll) or call a structured MCP tool — and the paper's central finding is that choosing the right path at each state is the hard part. ToolCUA-8B reaches 46.85% on OSWorld-MCP, ~66% above Qwen3-VL-8B-Instruct, with the fewest average completion steps (14.93).
What this Space does
It runs one step of the agent loop. You give it:
- a desktop screenshot,
- the task instruction,
- optionally the previous action log and the previous tool result,
and it predicts the next move: a short Action: sentence plus a single <tool_call>, which is either computer_use (a GUI action) or an MCP tool call with its arguments. The "MCP tools exposed" dropdown controls which toolset is advertised to the model, so you can watch the path-selection behaviour change — including "None (GUI only)", which strips the MCP tools entirely.
Greedy decoding (do_sample=False), matching the reference agent's temperature=0 evaluation setting.
Faithfulness to the authors' setup
The prompt is reconstructed to be byte-identical to the reference agent (eval/qwen3vl_toolcua_agent_mcp.py + eval/updated_mcp_server/):
- the full system prompt (tool-response format,
IMPORTANTreminder block, and theAction:/<tool_call>response-format rules) is extracted programmatically from that file, not retyped; - the tool schemas are the authors' own JSON files. In the reference setup the agent lists them through a FastMCP
MultiServerMCPClientkeyedosworld_mcp, which namespaces each tool asosworld_mcp_<stem>.<method>(e.g.osworld_mcp_libreoffice_calc.get_workbook_info) — matching the tool calls in the paper's own case study; - the
Instruction:/Previous actions:block is the reference template. So is the shape of the user turn, which differs between the two cases: - first step —
[instruction_prompt, screenshot], the screenshot sent bare, with no<tool_response>wrapper (in the reference the wrapper only ever carries the result of a previous action); - later step —
[<tool_response>, <previous result>, <screenshot>, </tool_response>], i.e. the result text comes before the image. - click coordinates are emitted on the 1000×1000 relative grid the OSWorld prompt describes; the Space also reports where that lands in pixels on the screenshot you uploaded.
Scope note: the full agent runs in the OSWorld-MCP desktop environment, which does not fit in a Space. This demo therefore covers a single state, not a complete multi-step rollout.
Hardware
ZeroGPU (zero-a10g, large), bf16. The checkpoint is stored in fp32 (~35 GB on disk); it is loaded in bf16 (~17 GB in VRAM), which fits the 48 GB slice.
Example assets
The three example screenshots are the authors' own case-study states from the X-PLUG/ToolCUA repository (assets/cases/ex2_step_2.png, assets/cases/ex4_step_1.png, assets/cases/ex4_step_3.png), reused under its MIT license, paired with the instructions the authors use for those two cases on the project page.
