ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF
Maverick-4B-Unity-XR-Agent
Maverick-4B-Unity-XR-Agent turns spoken or typed English into tool calls for Unity 3D scenes. Say "put the red mug on the counter" and it returns the call your application executes. It can pick up, place, move and rotate objects, switch devices on and off, highlight objects, move the user between rooms and search other rooms. If a request could mean two objects, it asks which one. If it cannot do something, it says so.
The model is a fine-tune of Qwen3-4B. It is built for XR and VR applications such as training simulations, product showrooms, digital twins, accessibility tools and games.
The article Offline Voice Control for Unity XR Scenes explains how the model was trained and tested.
Results
We tested the model on 499 instructions that people wrote for the ALFRED benchmark. It chose the right action on the right object 83.8% of the time. Five other models, from 1.7B to 120B parameters, scored between 35.9% and 67.1% with the same scenes and tools. On object types that never appear in its training data, the model scored 91.7%. In a Unity test scene it carried out 440 commands with 97.3% accuracy and no runtime errors.
Model details
How it works
For each command, your application sends the model two things: a JSON description of the current room and the user's sentence. The room description lists every object's id, label, colour, state, support and position, as well as the rooms and what the user is holding and looking at.
The model answers with tool calls. Your application executes them and sends back the results, and the model confirms what it did in one sentence.
Examples from our Unity test scene:
Usage
Start the server:
llama-server -hf ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M --jinja -c 4096 -ngl 99 --host 127.0.0.1 --port 8091Send a command:
import json, urllib.request
SYSTEM = open("system_prompt.txt", encoding="utf-8").read()
scene = {
"player": {"room": "kitchen", "gaze": None, "left_hand": None, "right_hand": None},
"locations": ["kitchen", "living_room"],
"visible_objects": [
{"id": "mug_1", "label": "mug", "color": "red", "on": "table_1", "pos": [0.4, 0.8, 1.2]},
{"id": "mug_2", "label": "mug", "color": "blue", "on": "counter_1", "pos": [2.0, 0.9, 0.3]},
{"id": "table_1", "label": "table", "on": "floor", "pos": [0.5, 0.0, 1.2]},
{"id": "counter_1", "label": "counter", "on": "floor", "pos": [2.0, 0.0, 0.3]},
],
}
body = {
"messages": [
{"role": "system", "content": SYSTEM},
{"role": "user", "content": "Scene: " + json.dumps(scene) + "\nRequest: put the red mug on the counter"},
],
"temperature": 0,
"max_tokens": 192,
}
req = urllib.request.Request("http://127.0.0.1:8091/v1/chat/completions",
json.dumps(body).encode(), {"Content-Type": "application/json"})
print(json.load(urllib.request.urlopen(req))["choices"][0]["message"]["content"])The request loop:
- Execute each tool call the model returns.
- Send each result back as a
toolmessage:{"ok": true},{"ok": false, "error": "..."}, or a list of matches forfind_objects. - Call the model again, until it replies in plain text or calls
ask_clarification.
Use a temperature of 0.
System prompt
The model was trained with this prompt; use it unchanged. The repository also contains it as system_prompt.txt.
You are an assistant embedded in a 3D Unity scene. Turn the user's request into tool calls that use object ids exactly as listed in the scene; never invent an id or a tool. Only objects in the user's current room are listed. Match objects by meaning: a synonym or a more general word can name a listed object, and details the scene does not record (a colour when no colours are listed, or words like sliced, wooden, empty) do not rule an object out. If exactly one listed object fits, act on it. If several fit equally well, call ask_clarification with their ids. If none fits, use find_objects to look in the other rooms. place moves an object directly, so there is no need to grab it first. If a request is impossible or needs a tool you do not have, call no tool and briefly explain why. Answer questions about the scene in plain text.
# Tools
grab(object_id:str, hand:left|right|both)
release(hand:left|right|both)
place(object_id:str, target_id:str, relation?:on|in, orientation?:default|upright|horizontal|vertical|upside_down)
move_by(object_id:str, direction:left|right|forward|backward|up|down, distance:num, unit:m|cm)
rotate(object_id:str, degrees:num, axis?:yaw|pitch|roll)
set_state(object_id:str, state:on|off|open|closed|half_open|locked|unlocked)
press(object_id:str, hold_seconds?:num)
highlight(object_ids:list)
teleport(location_id:str)
find_objects(label:str, color?:str, room?:str)
ask_clarification(question:str, candidate_ids:list)
Reply with one <tool_call>{"name": ..., "arguments": {...}}</tool_call> block per call, or plain text.Output check
Before executing a response, check it against the scene and the tool list:
The check changes about 0.5% of responses and removes every invented tool and id. All results on this page were measured with it; without it, the ALFRED score is 83.6%.
Unity
The Unity package (Unity 2022.3 or later) starts the server, builds the scene description, runs the request loop and applies the output check. It also adds editor tools for setting up and validating a scene.
Install it from the Package Manager with Add package from git URL:
https://github.com/mcbu-xrlab/Maverick-Unity.git#v1.0.0The same package is mirrored on Hugging Face at https://huggingface.co/ErenAta00/Maverick-Unity.git.
Then copy the GGUF file from this repository and llama.cpp's llama-server into Assets/StreamingAssets/SceneAgent/. The GitHub repository has a quick start and integration guides.
Evaluation
All test sets were frozen before training, and the released checkpoint was chosen on separate development sets.
ALFRED
ALFRED (Shridhar et al., CVPR 2020) is a public benchmark of household tasks that crowd workers described step by step. We turned its pick-up, put-down and lamp steps into single commands for grab, place and set_state:
- Request: the worker's sentence.
- Scene: the objects in the room at that step.
- Expected answer: the object used in ALFRED's recorded demonstration.
The test has 499 steps from the validation splits. An answer counts as correct if the first action calls the right tool on the right object. None of ALFRED's text appears in the training data.
In about 5% of the items, the worker's wording names a different object from the one ALFRED recorded, for example "knife" for a butter knife. No model can therefore reach 100%.
Accuracy in %. Items: 499 in total, 200 pick up, 200 put down, 99 switch on, 60 new object types.
¹ Object types that never appear in the training data. Of these 60 items, 19 use the word "cup", which the training data contains only as a synonym for mug.
Error types (% of all items):
The general-purpose models lose most of their points by asking a question or answering in text when one object clearly fits. Wrong-object errors are similar across all models (7–10%), and most of them come from the labelling mismatch described above. A manual review of this model's wrong-object cases found about 1% genuine errors.
Capability tests
We wrote 2,900 commands in 22 categories from templates. They use object types and sentence patterns that do not appear in the training data. Most categories come in pairs: with one matching object the model should act, and with two it should ask.
² These requests were kept out of the training data so that the features can be added later. Without the output check, this category scores 68%.
Unity
End to end: we ran 440 of the capability commands through the Unity package in a live scene. 97.3% were carried out correctly, with no runtime errors.
Package tests: the package was tested in a new project and in a built Windows game. It passed 54 of 55 checks covering:
- the editor tools;
- all 11 tools;
- follow-up answers and unusual input;
- several rooms and scene reloads;
- process cleanup.
The one failure was a room with 170 objects, which does not fit in the 4,096-token context; the package reports a clear error.
Speed and memory
These figures were measured on a laptop RTX 3050 Ti (4 GB) under Windows 11, with Unity rendering the scene at 30 fps on the same GPU.
Most commands take two model calls: one for the action and one for the confirmation. If Unity shares the GPU with the model, cap Unity's frame rate. Without a cap, commands took about five times longer.
Training
Data: 20,091 English conversations generated from templates and a catalogue of 209 object types. We held back 28 of these types for testing.
Data checks:
- Every conversation was executed in a simulator and had to produce the intended result.
- 501 conversations that overlapped with a test set were removed.
- About 8% of the conversations contain simulated speech-recognition errors.
Setup:
Limitations
- Other rooms: the model only searches other rooms for find and fetch requests. If the lamp is in another room, "turn on the lamp" gets "There is no lamp here".
- Typing errors: a misspelled name such as "lapm" is treated as an unknown object.
- Impossible requests: the model sometimes tries the action first and explains only after the engine rejects it.
- Real instructions: about one in six still needs a retry or a clarification.
- Scene questions: the model answers questions about the scene in text, but these answers were not evaluated.
- Scope: only the 11 tools. The model cannot create or delete objects, and it does not handle materials, lighting or physics.
- Language: English only.
- Input: the model reads the scene description, not the camera image.
- Test coverage: tested on one Windows laptop GPU only; not yet on XR headsets, Linux or macOS.
Intended use
Suitable for:
- voice or text control of objects in Unity XR, VR, simulation, training, showroom, digital-twin and game scenes;
- accessibility tools.
Not suitable for:
- robots or safety-critical systems;
- open-ended chat;
- irreversible actions without a confirmation step.
Safety and privacy
- Local: the server listens only on 127.0.0.1, and the model sends no data off the device.
- Speech: for fully offline voice input, pair the model with an offline speech recogniser such as whisper.cpp.
- Actions: the model can only request the 11 tools, and your application decides what each tool does.
- Training data: contains no personal data.
Citation
@inproceedings{shridhar2020alfred,
title = {{ALFRED}: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks},
author = {Shridhar, Mohit and Thomason, Jesse and Gordon, Daniel and Bisk, Yonatan and Han, Winson and
Mottaghi, Roozbeh and Zettlemoyer, Luke and Fox, Dieter},
booktitle = {CVPR},
year = {2020}
}
@misc{qwen3technicalreport,
title = {Qwen3 Technical Report},
author = {Qwen Team},
year = {2025},
eprint = {2505.09388},
archivePrefix = {arXiv}
}Contact
Developed by Eren Ata at the Extended Reality Laboratory (XRLab), Manisa Celal Bayar University. For questions and bug reports, open a discussion in the Community tab.
