Team Ai
Modelpublic

ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
1likes233downloads
Model Card

Maverick-4B-Unity-XR-Agent

[image]

Maverick-4B-Unity-XR-Agent turns spoken or typed English into tool calls for Unity 3D scenes. Say "put the red mug on the counter" and it returns the call your application executes. It can pick up, place, move and rotate objects, switch devices on and off, highlight objects, move the user between rooms and search other rooms. If a request could mean two objects, it asks which one. If it cannot do something, it says so.

The model is a fine-tune of Qwen3-4B. It is built for XR and VR applications such as training simulations, product showrooms, digital twins, accessibility tools and games.

The article Offline Voice Control for Unity XR Scenes explains how the model was trained and tested.

Results

We tested the model on 499 instructions that people wrote for the ALFRED benchmark. It chose the right action on the right object 83.8% of the time. Five other models, from 1.7B to 120B parameters, scored between 35.9% and 67.1% with the same scenes and tools. On object types that never appear in its training data, the model scored 91.7%. In a Unity test scene it carried out 440 commands with 97.3% accuracy and no runtime errors.

[image]

Model details

Base modelQwen3-4B
MethodQLoRA supervised fine-tuning, merged into the base weights
Training data20,091 English conversations
LanguageEnglish
FormatGGUF, Q4KM, 2.50 GB
Runtimellama.cpp (llama-server, OpenAI-compatible API)
GPU memory3.06 GB
Context4,096 tokens, enough for a room of about 90 objects
Toolsgrab, release, place, moveby, rotate, setstate, press, highlight, teleport, findobjects, askclarification
LicenceApache-2.0, the same as the base model

How it works

For each command, your application sends the model two things: a JSON description of the current room and the user's sentence. The room description lists every object's id, label, colour, state, support and position, as well as the rooms and what the user is holding and looking at.

The model answers with tool calls. Your application executes them and sends back the results, and the model confirms what it did in one sentence.

Examples from our Unity test scene:

CommandTool callReply
pick up the hammergrab(hammer_1, right)Done. I picked up the hammer with your right hand.
move the chair 50 cm to the leftmove_by(chair_1, left, 50, cm)Done. I moved the chair left by 50 cm.
um could you uh close the windowset_state(window_1, closed)Done. I closed the window.
pick up the toolbox (there are two)ask_clarification([toolbox_1, toolbox_2])Which toolbox do you mean, the blue or the red one?
highlight that (looking at a box)highlight([cardboard_box_1])Done. I highlighted the box.
where is the broom?find_objects(broom)The broom is in the storage, on the floor.
turn on the ladder–The ladder can't be turned on.
order more nails online–I can't do that; none of my tools can help with that.

Usage

Start the server:

bash
llama-server -hf ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M --jinja -c 4096 -ngl 99 --host 127.0.0.1 --port 8091

Send a command:

python
import json, urllib.request

SYSTEM = open("system_prompt.txt", encoding="utf-8").read()

scene = {
    "player": {"room": "kitchen", "gaze": None, "left_hand": None, "right_hand": None},
    "locations": ["kitchen", "living_room"],
    "visible_objects": [
        {"id": "mug_1", "label": "mug", "color": "red", "on": "table_1", "pos": [0.4, 0.8, 1.2]},
        {"id": "mug_2", "label": "mug", "color": "blue", "on": "counter_1", "pos": [2.0, 0.9, 0.3]},
        {"id": "table_1", "label": "table", "on": "floor", "pos": [0.5, 0.0, 1.2]},
        {"id": "counter_1", "label": "counter", "on": "floor", "pos": [2.0, 0.0, 0.3]},
    ],
}
body = {
    "messages": [
        {"role": "system", "content": SYSTEM},
        {"role": "user", "content": "Scene: " + json.dumps(scene) + "\nRequest: put the red mug on the counter"},
    ],
    "temperature": 0,
    "max_tokens": 192,
}
req = urllib.request.Request("http://127.0.0.1:8091/v1/chat/completions",
                             json.dumps(body).encode(), {"Content-Type": "application/json"})
print(json.load(urllib.request.urlopen(req))["choices"][0]["message"]["content"])

The request loop:

  1. 1.Execute each tool call the model returns.
  2. 2.Send each result back as a tool message: {"ok": true}, {"ok": false, "error": "..."}, or a list of matches for find_objects.
  3. 3.Call the model again, until it replies in plain text or calls ask_clarification.

Use a temperature of 0.

System prompt

The model was trained with this prompt; use it unchanged. The repository also contains it as system_prompt.txt.

text
You are an assistant embedded in a 3D Unity scene. Turn the user's request into tool calls that use object ids exactly as listed in the scene; never invent an id or a tool. Only objects in the user's current room are listed. Match objects by meaning: a synonym or a more general word can name a listed object, and details the scene does not record (a colour when no colours are listed, or words like sliced, wooden, empty) do not rule an object out. If exactly one listed object fits, act on it. If several fit equally well, call ask_clarification with their ids. If none fits, use find_objects to look in the other rooms. place moves an object directly, so there is no need to grab it first. If a request is impossible or needs a tool you do not have, call no tool and briefly explain why. Answer questions about the scene in plain text.

# Tools
grab(object_id:str, hand:left|right|both)
release(hand:left|right|both)
place(object_id:str, target_id:str, relation?:on|in, orientation?:default|upright|horizontal|vertical|upside_down)
move_by(object_id:str, direction:left|right|forward|backward|up|down, distance:num, unit:m|cm)
rotate(object_id:str, degrees:num, axis?:yaw|pitch|roll)
set_state(object_id:str, state:on|off|open|closed|half_open|locked|unlocked)
press(object_id:str, hold_seconds?:num)
highlight(object_ids:list)
teleport(location_id:str)
find_objects(label:str, color?:str, room?:str)
ask_clarification(question:str, candidate_ids:list)

Reply with one <tool_call>{"name": ..., "arguments": {...}}</tool_call> block per call, or plain text.

Output check

Before executing a response, check it against the scene and the tool list:

OutputAction
A valid callExecute it
An unknown tool, or an argument value outside the schemaReplace it with a refusal
An unknown object id, or a malformed callRegenerate it with a grammar built from the scene

The check changes about 0.5% of responses and removes every invented tool and id. All results on this page were measured with it; without it, the ALFRED score is 83.6%.

Unity

The Unity package (Unity 2022.3 or later) starts the server, builds the scene description, runs the request loop and applies the output check. It also adds editor tools for setting up and validating a scene.

Install it from the Package Manager with Add package from git URL:

text
https://github.com/mcbu-xrlab/Maverick-Unity.git#v1.0.0

The same package is mirrored on Hugging Face at https://huggingface.co/ErenAta00/Maverick-Unity.git.

Then copy the GGUF file from this repository and llama.cpp's llama-server into Assets/StreamingAssets/SceneAgent/. The GitHub repository has a quick start and integration guides.

Evaluation

All test sets were frozen before training, and the released checkpoint was chosen on separate development sets.

ALFRED

ALFRED (Shridhar et al., CVPR 2020) is a public benchmark of household tasks that crowd workers described step by step. We turned its pick-up, put-down and lamp steps into single commands for grab, place and set_state:

  • —Request: the worker's sentence.
  • —Scene: the objects in the room at that step.
  • —Expected answer: the object used in ALFRED's recorded demonstration.

The test has 499 steps from the validation splits. An answer counts as correct if the first action calls the right tool on the right object. None of ALFRED's text appears in the training data.

In about 5% of the items, the worker's wording names a different object from the one ALFRED recorded, for example "knife" for a butter knife. No model can therefore reach 100%.

[image]

ModelParametersAllPick upPut downSwitch onNew object types ¹
Maverick-4B-Unity-XR-Agent4B83.888.573.096.091.7
Qwen3-1.7B1.7B35.949.09.063.623.3
Qwen3-4B4B57.141.058.586.955.0
Ministral 3 14B14B67.156.067.089.966.7
gpt-oss-120b120B57.157.546.078.853.3
Nemotron-3 Super120B58.960.046.082.855.0

Accuracy in %. Items: 499 in total, 200 pick up, 200 put down, 99 switch on, 60 new object types.

¹ Object types that never appear in the training data. Of these 60 items, 19 use the word "cup", which the training data contains only as a synonym for mug.

Error types (% of all items):

[image]

ModelInvented tool or idQuestion instead of an actionText instead of an actionWrong objectWrong tool or extra calls
Maverick-4B-Unity-XR-Agent0.00.43.49.03.4
Qwen3-4B3.621.83.89.83.8
Ministral 3 14B1.82.018.67.43.0
gpt-oss-120b0.25.825.310.41.2
Nemotron-3 Super0.414.015.89.61.2

The general-purpose models lose most of their points by asking a question or answering in text when one object clearly fits. Wrong-object errors are similar across all models (7–10%), and most of them come from the labelling mismatch described above. A manual review of this model's wrong-object cases found about 1% genuine errors.

Capability tests

We wrote 2,900 commands in 22 categories from templates. They use object types and sentence patterns that do not appear in the training data. Most categories come in pairs: with one matching object the model should act, and with two it should ask.

[image]

CapabilityCommandsExampleAccuracy (%)
Object names: unseen types, synonyms, general terms600"Spotlight the house plant"97.8
Spatial references: on, in, near, left, right, front, behind, order750"the pen on this side of the cart"93.9
Details the scene does not record: colour, material, condition400"Set the gray box atop the desk"99.5
Asking when two objects fit equally well150"Highlight the pink trash can"100.0
Searching other rooms100"Do you know where the newspaper went?"98.0
Reporting an object that does not exist100"Where might the sponge be?"100.0
Declining an impossible action without trying it100"Activate the remote"75.0
"this", "that", "it" and held objects100"Draw attention to this folder"98.0
Verbs that mean a change of state100"Power off the stove knob"99.0
Verbs with more than one meaning150"Put the crate upside down"99.3
Speech-recognition errors150"spotlight the apple that's on the counter"97.3
Out-of-scope requests, such as cooking or cleaning100"Clean the credit card"98.0
Unsupported edits: resizing, recolouring, deleting ²100"Paint the key purple"93.0
All2,90096.4

² These requests were kept out of the training data so that the features can be added later. Without the output check, this category scores 68%.

Unity

End to end: we ran 440 of the capability commands through the Unity package in a live scene. 97.3% were carried out correctly, with no runtime errors.

Package tests: the package was tested in a new project and in a built Windows game. It passed 54 of 55 checks covering:

  • —the editor tools;
  • —all 11 tools;
  • —follow-up answers and unusual input;
  • —several rooms and scene reloads;
  • —process cleanup.

The one failure was a room with 170 objects, which does not fit in the 4,096-token context; the package reports a clear error.

Speed and memory

These figures were measured on a laptop RTX 3050 Ti (4 GB) under Windows 11, with Unity rendering the scene at 30 fps on the same GPU.

MeasurementValue
GPU memory (llama.cpp)3.06 GB
Shared system memory96 MiB
First answer, model only, median1.18 s
Per command, GPU cool0.5–3.6 s
Per command, GPU at its thermal limit, median5.9 s

Most commands take two model calls: one for the action and one for the confirmation. If Unity shares the GPU with the model, cap Unity's frame rate. Without a cap, commands took about five times longer.

Training

Data: 20,091 English conversations generated from templates and a catalogue of 209 object types. We held back 28 of these types for testing.

AreaShare
Core scene actions and follow-up questions38%
Spatial references18%
Deciding to act, ask, search or decline14%
Object names14%
Details the scene does not record7%
Verb meanings6%
References and held objects2%
Unsupported requests1%

Data checks:

  • —Every conversation was executed in a simulator and had to produce the intended result.
  • —501 conversations that overlapped with a test set were removed.
  • —About 8% of the conversations contain simulated speech-recognition errors.

Setup:

MethodQLoRA (Unsloth), rank 16, alpha 32, all attention and MLP projections
Schedule1 epoch, 2,512 steps, batch size 8, learning rate 2e-4, cosine decay
Sequence length1,536 tokens
Hardware1 × NVIDIA T4
ExportMerged into the de-quantised 4-bit base, converted to GGUF Q80, then quantised to Q4K_M

Limitations

  • —Other rooms: the model only searches other rooms for find and fetch requests. If the lamp is in another room, "turn on the lamp" gets "There is no lamp here".
  • —Typing errors: a misspelled name such as "lapm" is treated as an unknown object.
  • —Impossible requests: the model sometimes tries the action first and explains only after the engine rejects it.
  • —Real instructions: about one in six still needs a retry or a clarification.
  • —Scene questions: the model answers questions about the scene in text, but these answers were not evaluated.
  • —Scope: only the 11 tools. The model cannot create or delete objects, and it does not handle materials, lighting or physics.
  • —Language: English only.
  • —Input: the model reads the scene description, not the camera image.
  • —Test coverage: tested on one Windows laptop GPU only; not yet on XR headsets, Linux or macOS.

Intended use

Suitable for:

  • —voice or text control of objects in Unity XR, VR, simulation, training, showroom, digital-twin and game scenes;
  • —accessibility tools.

Not suitable for:

  • —robots or safety-critical systems;
  • —open-ended chat;
  • —irreversible actions without a confirmation step.

Safety and privacy

  • —Local: the server listens only on 127.0.0.1, and the model sends no data off the device.
  • —Speech: for fully offline voice input, pair the model with an offline speech recogniser such as whisper.cpp.
  • —Actions: the model can only request the 11 tools, and your application decides what each tool does.
  • —Training data: contains no personal data.

Citation

bibtex
@inproceedings{shridhar2020alfred,
  title     = {{ALFRED}: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks},
  author    = {Shridhar, Mohit and Thomason, Jesse and Gordon, Daniel and Bisk, Yonatan and Han, Winson and
               Mottaghi, Roozbeh and Zettlemoyer, Luke and Fox, Dieter},
  booktitle = {CVPR},
  year      = {2020}
}

@misc{qwen3technicalreport,
  title         = {Qwen3 Technical Report},
  author        = {Qwen Team},
  year          = {2025},
  eprint        = {2505.09388},
  archivePrefix = {arXiv}
}

Contact

Developed by Eren Ata at the Extended Reality Laboratory (XRLab), Manisa Celal Bayar University. For questions and bug reports, open a discussion in the Community tab.