BrinqAI/functiongemma-270m-physical-ai
FunctionGemma 270M — Physical AI (v10, Octopus v2)
Fine-tuned `google/functiongemma-270m-it` for voice-controlled physical-AI / household-IoT actions on a Synaptics SL2619 "Coral" edge board (Google IO 2026 demo).
Current revision: `functiongemma-physical-ai-v10-Q5_K_M.gguf` — 6 tools, ~248 MB Q5KM, ~0.48 s cold prefill on the 2-core Cortex-A55, 97.9 % mean token accuracy on eval.
Schema ships as `tools.json`. Token-to-tool mapping is in `token_map.json`.
Tool surface (6 tools)
The model is hardware-agnostic for lighting: it parses user intent into semantic args (color, effect, state) and leaves the dispatcher to map those onto whatever LED hardware is detected at launch — the HAT's three indicator LEDs, a WLED-driven strip, or a Neopixel ring. The user vocabulary is hardware-agnostic too: "lights", "LEDs", "strip", "indicators" all refer to whatever is wired up.
Prompt format
The v10 model is trained Octopus v2 style: no schema, no tools list, just a bare user turn.
<start_of_turn>user
{user_text}<end_of_turn>
<start_of_turn>model
Tool semantics live in the model weights (via the special functional tokens <tool_0> … <tool_5> plus <end>), not in the prompt. The tools.json schema in this repo is the dispatcher's arg-validation contract and is embedded in the GGUF metadata for schema-drift checks, but it is not loaded into the inference prompt. Typical prompts are ~13 tokens.
Output format — functional tokens, named args
Tool calls emit as functional tokens with named arguments, per the Mercedes-Benz Octopus v2 convention (arXiv 2501.02342). Each tool name compiles to a single special-vocabulary token (<tool_0> … <tool_5>); arguments are written as name="value" pairs; a single <end> token terminates the call. The model emits only the args the user implied — absent args are simply not present.
Examples:
A complete call decodes in roughly 8–20 output tokens, well inside the sub-second voice-UX budget on a 2-core Cortex-A55.
⚠️ Inference servers MUST stop generation on<end_of_turn>(or<eos>), NOT on<end>. A single<end>ends one tool call; the format and runtime support multi-tool sequences<tool_A>(args)<end><tool_B>(args)<end>, so stopping at the first<end>would truncate them. (The v10 weights route one action per turn — see Phrases that work — but the runtime keeps multi-tool support intact for future versions.)
Phrases that work
The model routes one action per command. Every example below is a verbatim greedy decode from the v10 GGUF.
Lights — set_lights
Say "lights", "LEDs", "the strip", "the ring", "neopixels", or "<color> light(s)" — they all route here.
Colors: red, green, blue, white, yellow, purple, orange, pink, cyan. Effects: solid, blink, pulse, fade, rainbow, fire, plasma, aurora, police, fireworks, sparkle, twinkle, chase, comet, heartbeat, lightning, glitter, loading, sunrise, off.
Buzzer — play_buzzer
Named sounds, not counts.
Patterns: beep, double_beep, chirp, siren, alarm, success, error.
Alarms — set_alarm / cancel_alarm
System status — get_system_status
Conversation — respond
Getting the best results
- One action per command. The model emits a single tool call. "Make the lights blue and beep" turns the lights blue and skips the beep — say the two as separate commands.
- One light attribute per command — a color or an effect, not both. "Pulse the lights blue" keeps the color and drops the pulse; say "Pulse the lights" or "Turn the lights blue".
- The buzzer uses named sounds, not counts. "Beep three times" has no count to map to — use "double beep", "siren", or "alarm".
- It's a command router, not a chatbot. Open-ended prompts ("tell me a joke") may misfire rather than chat — it is fine-tuned for routing physical-action commands.
Quick start (Ollama)
hf download BrinqAI/functiongemma-270m-physical-ai \
functiongemma-physical-ai-v10-Q5_K_M.gguf Modelfile tools.json token_map.json \
--local-dir ./fg-physical-ai
cd fg-physical-ai
ollama create functiongemma-physical-ai -f ModelfileThe shipped Modelfile bakes in the stop tokens (<end_of_turn>, <eos>) and decode parameters (temperature=0, num_ctx=1024, num_predict=80).
Calling the model
Send a bare user turn — no schema, no tools list. With Ollama, use raw=true:
import json
import re
import urllib.request
OLLAMA_URL = "http://localhost:11434"
MODEL = "functiongemma-physical-ai"
reverse_token_map = json.load(open("token_map.json"))["reverse"]
NAMED_ARG_RE = re.compile(r'(\w+)\s*=\s*"((?:[^"\\]|\\.)*)"')
def build_prompt(user_text: str) -> str:
return (
f"<start_of_turn>user\n{user_text}<end_of_turn>\n"
f"<start_of_turn>model\n"
)
def call_model(user_text: str) -> str:
body = json.dumps({
"model": MODEL,
"prompt": build_prompt(user_text),
"raw": True,
"stream": False,
"options": {
"temperature": 0.0,
"top_p": 1.0,
"num_predict": 80,
"stop": ["<end_of_turn>", "<eos>"],
},
}).encode()
req = urllib.request.Request(
f"{OLLAMA_URL}/api/generate",
data=body,
headers={"Content-Type": "application/json"},
)
with urllib.request.urlopen(req, timeout=60) as resp:
return json.loads(resp.read())["response"]
def parse_call(raw: str) -> tuple[str | None, dict[str, str]]:
"""Return (tool_name, kwargs). tool_name is None on parse fail."""
m = re.match(r"\s*(<tool_\d+>)\((.*?)\)<end>", raw)
if not m:
return None, {}
tok, body = m.group(1), m.group(2)
kwargs = {k: v for k, v in NAMED_ARG_RE.findall(body)}
return reverse_token_map.get(tok), kwargs
raw = call_model("turn the lights red")
print(raw) # e.g. '<tool_0>(color="red")<end>'
print(parse_call(raw)) # ('set_lights', {'color': 'red'})For llama-cpp-python directly, use detokenize(..., special=True) so the <tool_N> and <end> tokens render in the output instead of being stripped.
Training data
Training data was generated from Haiku-authored phrasing templates crossed with deterministic entity pools, then lightly augmented with Moonshine-flavored ASR noise (dropped function words, lowercased traces, filler-word prepends). Each record is a flat {input, output} pair — no tools / messages array, no chat template.
The set_lights tool also gets explicit failure-mode rows that route bare ambiguous prompts to respond() — e.g. "rainbow" alone ("Did you mean the lights? Try 'rainbow on the lights'."), "siren" alone (prompts the user toward play_buzzer), and bare "on" / "off" (asks what the user wants to act on).
Methodology
- Full bf16 fine-tune (no LoRA).
- Functional tokens:
<tool_0>…<tool_5>+<end>added asadditional_special_tokens; new embeddings mean-initialized from the existing input-embedding matrix (random init under-converges on small datasets at this scale). - Completion-only loss mask: hand-rolled — labels before
<start_of_turn>model\nare masked to-100. The model learns only from the assistant turn, not the user prompt. - 5 epochs, lr
3e-5, cosine schedule, 0.1 warmup, weight decay 0.01. - Effective batch = 16 (
per_device_train_batch_size=8 × gradient_accumulation_steps=2). - `max_length=256` — the trained prompt format is ~13 tokens and the assistant turn fits comfortably under 64 tokens, including
respond()messages. - bf16, gradient checkpointing,
adamw_torch_fused,metric_for_best_model="eval_loss"+load_best_model_at_end=True. - Training wallclock: 5 min on a single H100 (~15–20 min on a 4090).
Citation
@article{chen2024octopusv2,
title = {Octopus v2: On-device language model for super agent},
author = {Chen, Wei and Li, Zhiyuan},
journal = {arXiv preprint arXiv:2404.01744},
year = {2024},
url = {https://arxiv.org/abs/2404.01744}
}
@article{merc2025octopusv2,
title = {Octopus v2 named-arg function calling},
journal = {arXiv preprint arXiv:2501.02342},
year = {2025},
url = {https://arxiv.org/abs/2501.02342}
}Results
Training metrics (final epoch)
Held-out smoke test (post-train, 36 prompts spanning all 6 tools)
The 36-prompt suite covers single-tool happy paths for every tool plus failure modes the model is expected to deflect: ambiguous color words without a target ("make it red"), effect names without a target ("rainbow"), unsupported features ("play a tone at 2000 hz"), and out-of-scope appliances. Failure-mode prompts all route to respond() with a helpful clarification message.
On-device benchmark (Coralboard, 2-core Cortex-A55 @ 2 GHz, Q5KM GGUF)
Measured with llama-cpp-python 0.3.16, n_ctx=1024, n_threads=2, CPU governor performance, 8 representative prompts spanning all 6 tools.
Files
functiongemma-physical-ai-v10-Q5_K_M.gguf # ~248 MB, Q5_K_M weights (Ollama / llama.cpp)
Modelfile # Ollama Modelfile (functional-token format)
tools.json # 6-tool schema, canonical mobile-actions format
token_map.json # functional-token <-> tool-name map
README.md # this fileEarlier checkpoint GGUFs from the project's development history (functiongemma-physical-ai-v9-Q5_K_M.gguf, functiongemma-physical-ai-v7-Q5_K_M.gguf, functiongemma-physical-ai-v6-Q5_K_M.gguf, functiongemma-physical-ai-Q4_K_M.gguf) remain in the repo for reproducibility. They use different tool surfaces and (for v7 and earlier) a different inference-prompt format; new deployments should use the v10 file above.
License
Released under the Gemma Terms of Use. By using this model you agree to those terms. Base model: `google/functiongemma-270m-it`.
Links
- Base model: <https://huggingface.co/google/functiongemma-270m-it>
- Octopus v2 paper: <https://arxiv.org/abs/2404.01744>
- Mercedes-Benz Octopus v2 (named-arg variant): <https://arxiv.org/abs/2501.02342>
- Hardware demo + integration code (Synaptics Coralboard, Grinn HAT, WLED-over-USB-CDC, full PyQt UI): <https://github.com/synaptics-astra-demos/sl2610-examples> →
Function_calling/
