Berhak/Llama-3.1-8B-Function-Calling-Agent
<div align="center">
๐ฆ Llama-3.1-8B-Function-Calling-Agent
Function calling that generalizes to tool schemas it has never seen
<a href="https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct"><img alt="Base" src="https://img.shields.io/badge/Base-Llama--3.1--8B--Instruct-0866ff?style=for-the-badge"></a> <img alt="Method" src="https://img.shields.io/badge/Method-QLoRA%20r%3D32-7c3aed?style=for-the-badge"> <img alt="Quant" src="https://img.shields.io/badge/Quant-Q4_K_M%20%2B%20imatrix-059669?style=for-the-badge"> <br> <img alt="Unseen schemas" src="https://img.shields.io/badge/Unseen%20schemas-92%25-16a34a?style=for-the-badge"> <img alt="VRAM" src="https://img.shields.io/badge/VRAM-5.1%20GB-ea580c?style=for-the-badge"> <img alt="License" src="https://img.shields.io/badge/License-Llama%203.1-475569?style=for-the-badge">
<sub><b>Built with Llama</b></sub>
</div>
<hr>
A function-calling fine-tune of Llama-3.1-8B-Instruct, trained to read JSON tool schemas at inference time and emit well-formed <tool_call> blocks โ then answer in natural English once the tool result comes back. It runs on one 8 GB consumer GPU at ~40 tok/s.
<div align="center">
</div>
Those first two columns are the whole result. Near-identical performance on schemas the model has never seen means it learned "read the schema, build the call" rather than the tool names it was trained on.
๐ A session
you > Do you remember what my job is?
agent> Your job is that of a software engineer.
ยท prompt 713 tok 0.8s (860 t/s) ยท gen 10 tok 0.2s (38 t/s)
you > What is the current population of Japan?
ยท google_search({'query': 'current population of Japan', 'num_results': 1})
-> 3 sources (wikipedia): en.wikipedia.org
agent> According to Wikipedia, Japan's population peaked at 128.5 million people in 2010
and had declined to approximately 123.4 million by April 2025. By 2050, an estimated
one-third of the population in Japan is expected to be 65 and older due to
sub-replacement fertility rates and high life expectancy. [โฆ]
ยท prompt 1390 tok 1.3s (1044 t/s) ยท gen 137 tok 3.8s (36 t/s)
you > Do you remember which city I live in? Could you check what the weather will be
like there tomorrow and whether I should take an umbrella?
ยท get_weather({'location': 'New York', 'unit': 'celsius', 'days_ahead': 1}) -> ok
agent> In New York, there will be rain tomorrow with a minimum temperature of 19.6 degrees
Celsius and a maximum of 25.6 degrees Celsius. It would be wise to bring an umbrella.
ยท prompt 420 tok 0.3s (1232 t/s) ยท gen 41 tok 1.1s (37 t/s)Two things in there are the card's claims, live:
- ๐ `days_ahead` was never in the training data. It is an extension to the
get_weatherschema added after the fine-tune, and the model reads "tomorrow" โdays_ahead: 1straight out of the parameter description โ the same schema-reading ability the 92% measures. - 1๏ธโฃ The search turn asked for
num_results: 1and got three sources back: the reference agent overriding a learned habit, see It asks for one search result under Limitations.
<hr>
๐ฆ What is in this repository
Two separate artifacts. They are not interchangeable.
<hr>
โก Quick start
llama.cpp
./llama-server -m llama31-8b-Q4_K_M.gguf \
-c 8192 -ngl 99 -fa on \
--cache-type-k q8_0 --cache-type-v q8_0 \
--host 127.0.0.1 --port 8080The q8_0 KV cache halves cache cost to ~68 KB/token, which is what makes 8192 context fit alongside the weights on an 8 GB card.
transformers + PEFT
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = "unsloth/Meta-Llama-3.1-8B-Instruct"
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(model, "Berhak/Llama-3.1-8B-Function-Calling-Agent")
tokenizer = AutoTokenizer.from_pretrained(base)<hr>
๐ Prompt contract
### โ ๏ธ This matters more than the weights. The model was trained on one specific shape and behaves poorly outside it. All three pieces below are load-bearing.
1๏ธโฃ System turn
Tool schemas go inside <tools> as a JSON array, in Llama-3.1's native header format:
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
Cutting Knowledge Date: December 2023
Today Date: 12 Sep 2026
You are a function calling AI model. You are provided with function signatures within <tools></tools> XML tags. You may call one or more functions to assist with the user query. Don't make assumptions about what values to plug into functions.
<tools>
[{"type": "function", "function": {"name": "get_weather", "description": "Get current or forecast weather conditions for a location.", "parameters": {"type": "object", "properties": {"location": {"type": "string", "description": "The city name, e.g. 'Ankara'."}, "unit": {"type": "string", "enum": ["celsius", "fahrenheit"], "description": "Temperature unit."}}, "required": ["location"]}}}]
</tools>
For each function call return a json object with function name and arguments within <tool_call></tool_call> XML tags.<|eot_id|>๐
Pass the real date. strftime("%b") follows the OS locale and produces a non-English month abbreviation under a non-English one, which breaks the template.
2๏ธโฃ What the model emits
<tool_call>
{"name": "get_weather", "arguments": {"location": "Berlin", "unit": "celsius"}}
</tool_call>๐ข It emits several blocks in one generation when the request needs several tools. Parse every block, not just the first.
3๏ธโฃ Observations โ the unusual part
Tool results go back on the `ipython` role, double-encoded: the <tool_response> block is wrapped in a JSON string, escaped quotes and literal \n included. An artifact of the original Hermes conversion, but the model learned that exact shape:
<|start_header_id|>ipython<|end_header_id|>
"<tool_response>\n{\"name\": \"get_weather\", \"content\": {\"location\": \"Berlin\", \"temperature\": 18.4}}\n</tool_response>"<|eot_id|>observation = json.dumps(
"<tool_response>\n" + json.dumps({"name": name, "content": result}) + "\n</tool_response>"
)๐จ A plain, single-encoded block puts the model off its training distribution.
<hr>
๐ Recommended: constrain decoding with a grammar
Build a GBNF grammar from the tool schemas you pass in this request โ not from a fixed file โ with a root that permits either a run of calls or plain prose:
root ::= tool-call (ws-nl tool-call)* | prose
tool-call ::= "<tool_call>" ws-nl call ws-nl "</tool_call>"
prose ::= [^<] ( [^<] | "<" )*
ws-nl ::= [ \t\n]*<div align="center">
</div>
No category regressed, and the gain concentrates where the model is weakest unaided โ generations needing more than one call. Three details are worth copying verbatim:
- ๐ณ๏ธ The root must allow prose, or every question is forced into a tool call.
- โ๏ธ
prosebans<only in the first position. Banning it everywhere ([^<]+) made the character unrepresentable โ asked for an inequality the model wrote3 โค 5instead of3 < 5, a different claim, not a formatting quirk. - ๐งฉ Build it per request. A fixed grammar forbids the caller's own schemas and scores 0/11 on unseen ones.
<hr>
๐ Evaluation
86 hand-written held-out records, screened against the training data for verbatim matches and at a 0.70 similarity threshold. Every tool name claimed as "unseen" was checked against all 2,985 tool names present in training. Scored with the grammar on:
Four ambiguous records are reported but never scored โ both behaviours are defensible there, and inventing a reference answer to complete a metric would only make the metric worse.
A handful of records also score the language of the final answer against Turkish (see below), so those are not an English measurement. Every row above except Overall measures tool selection and argument construction, which is language-independent.
<hr>
โ ๏ธ Limitations
### ๐ง It does not call a tool after seeing an observation. This is structural, not a weak tendency.
Of the 8,144 assistant turns that follow a tool result across 9,146 training examples, zero are a tool call. The data only ever contains CALL โ OBSERVE โ ANSWER; a second call appears solely after a new user turn. A directive telling it to continue scored 0/3.
โ Plan around it: a request needing several tools must be satisfied by the first generation. The model does this reliably, so parse every block it emits. Do not build a loop that expects it to iterate.
๐ป It fabricates arguments it does not have. "What's the weather?" with no city produced location='Ankara'; "is there a train from Ankara?" produced destination="user's destination" โ a literal placeholder, meaning it knows it does not know and fills the slot anyway. No prompt variant fixed this. Validate arguments against the conversation before dispatching.
1๏ธโฃ It asks for one search result. The training data calls search with num_results=1 in 34 of 55 cases. Treat that parameter as a floor in your tool, not a ceiling.
๐ญ It invents explanations for things that do not exist. Asked about a fabricated syndrome, it answers confidently and fluently. Empty or failed tool results should say so explicitly in the observation, and say what the model must not do.
๐ฏ Search queries drift off subject. Asked to search two topics, the second query sometimes lands on a neighbouring one โ worse with accumulated context.
๐งฎ Arithmetic needs a tool. Unaided it produced "1 FP16 = 2 INT4, so 16 GB becomes 32 GB" โ the ratio is 4 and the operation is division โ then reused the wrong figure as a premise on the next turn. Route calculations to a calculator tool.
๐ Factual accuracy of direct answers was never measured, and ๐ no bfloat16 baseline exists โ the merged model was deleted by the quantization pipeline before one was taken, so every number here is absolute rather than a measurement of quantization loss.
๐ On Turkish. The training mix includes a small Turkish subset (~5%), and the evaluation set was originally written against Turkish final turns. That path is not validated and is not recommended โ the agent around this model was brought to a working standard in English only. Treat this as an English function-calling model.
<hr>
๐๏ธ Training
QLoRA on 4ร H100, a single run with no retry budget. Source: NousResearch/hermes-function-calling-v1, converted from ChatML into Llama-3.1's native chat template before loss masking โ Llama-3.1's tokenizer does not recognize <|im_start|> as a special token, so raw ChatML would shatter into meaningless sub-words.
๐ณ๏ธ One warning worth repeating. English final-turn prose was masked out so it would not compete with the intended final-turn behaviour. But Hermes examples that answer without a tool consist of nothing but a final turn โ so masking it masked the whole example, and 803 of 851 vanished silently. The resulting skew toward calling a tool is why the model over-triggers, and why the reference agent adds a restraint clause to the system prompt. If you reuse this recipe: count your examples after masking, not before.
Quantization: bf16 merge โ GGUF q8_0 โ imatrix โ Q4_K_M. The intermediate is q8_0 rather than f16 (half the disk, no measurable K-quant penalty; needs --allow-requantize), and the importance matrix was calibrated on 40 samples drawn from the real training distribution.
<hr>
๐ป Measured performance
RTX 4060 Laptop (8 GB), Vulkan, n_ctx=8192, Q4KM, q8_0 KV cache, flash attention:
<div align="center">
</div>
### ๐ฆ Memory is not the binding constraint โ prompt processing is. At 175 tok/s, every 1000 tokens of context costs about six seconds on every subsequent turn. Trim tool output aggressively and keep the system prefix stable so prefix caching holds (measured: 14 tokens reprocessed instead of 440 on the second request).
Sampling: temperature=0.5 top_k=40 top_p=0.9 repeat_penalty=1.1 repeat_last_n=768. At temperature=0 the model locked into repeating one sentence across turns; repeat_penalty is what broke that, not the temperature. Use temperature=0 for reproducible evaluation.
<hr>
๐ License
Llama 3.1 Community License. Use is subject to the Llama 3.1 License and the Acceptable Use Policy.
Built with Llama.
Training data: `NousResearch/hermes-function-calling-v1`, subject to its own terms.
@misc{llama31-8b-function-calling-agent,
title = {Llama-3.1-8B-Function-Calling-Agent},
author = {Tanyฤฑldฤฑzฤฑ, Mahmut Berhak},
year = {2026},
url = {https://huggingface.co/Berhak/Llama-3.1-8B-Function-Calling-Agent}
}