Team Ai
Modelpublic

Berhak/Llama-3.1-8B-Function-Calling-Agent

sourceHugging Facellama3.1updated 20d agoView on Hugging Face
0likes157downloads
Model Card

<div align="center">

๐Ÿฆ™ Llama-3.1-8B-Function-Calling-Agent

Function calling that generalizes to tool schemas it has never seen

<a href="https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct"><img alt="Base" src="https://img.shields.io/badge/Base-Llama--3.1--8B--Instruct-0866ff?style=for-the-badge"></a> <img alt="Method" src="https://img.shields.io/badge/Method-QLoRA%20r%3D32-7c3aed?style=for-the-badge"> <img alt="Quant" src="https://img.shields.io/badge/Quant-Q4_K_M%20%2B%20imatrix-059669?style=for-the-badge"> <br> <img alt="Unseen schemas" src="https://img.shields.io/badge/Unseen%20schemas-92%25-16a34a?style=for-the-badge"> <img alt="VRAM" src="https://img.shields.io/badge/VRAM-5.1%20GB-ea580c?style=for-the-badge"> <img alt="License" src="https://img.shields.io/badge/License-Llama%203.1-475569?style=for-the-badge">

<sub><b>Built with Llama</b></sub>

</div>

<hr>

A function-calling fine-tune of Llama-3.1-8B-Instruct, trained to read JSON tool schemas at inference time and emit well-formed <tool_call> blocks โ€” then answer in natural English once the tool result comes back. It runs on one 8 GB consumer GPU at ~40 tok/s.

<div align="center">

๐ŸŽฏ Unseen schemas๐Ÿ”‘ Seen schemasโœ… JSON validity๐Ÿ’พ Total VRAM
24 / 2648 / 5657 / 575055 MiB
92%85%100%of 7774 MiB

</div>

Those first two columns are the whole result. Near-identical performance on schemas the model has never seen means it learned "read the schema, build the call" rather than the tool names it was trained on.

๐Ÿ‘€ A session

console
you > Do you remember what my job is?
agent> Your job is that of a software engineer.
  ยท prompt 713 tok 0.8s (860 t/s) ยท gen 10 tok 0.2s (38 t/s)

you > What is the current population of Japan?
  ยท google_search({'query': 'current population of Japan', 'num_results': 1})
      -> 3 sources (wikipedia): en.wikipedia.org
agent> According to Wikipedia, Japan's population peaked at 128.5 million people in 2010
       and had declined to approximately 123.4 million by April 2025. By 2050, an estimated
       one-third of the population in Japan is expected to be 65 and older due to
       sub-replacement fertility rates and high life expectancy. [โ€ฆ]
  ยท prompt 1390 tok 1.3s (1044 t/s) ยท gen 137 tok 3.8s (36 t/s)

you > Do you remember which city I live in? Could you check what the weather will be
      like there tomorrow and whether I should take an umbrella?
  ยท get_weather({'location': 'New York', 'unit': 'celsius', 'days_ahead': 1}) -> ok
agent> In New York, there will be rain tomorrow with a minimum temperature of 19.6 degrees
       Celsius and a maximum of 25.6 degrees Celsius. It would be wise to bring an umbrella.
  ยท prompt 420 tok 0.3s (1232 t/s) ยท gen 41 tok 1.1s (37 t/s)

Two things in there are the card's claims, live:

  • โ€”๐ŸŽ“ `days_ahead` was never in the training data. It is an extension to the get_weather schema added after the fine-tune, and the model reads "tomorrow" โ†’ days_ahead: 1 straight out of the parameter description โ€” the same schema-reading ability the 92% measures.
  • โ€”1๏ธโƒฃ The search turn asked for num_results: 1 and got three sources back: the reference agent overriding a learned habit, see It asks for one search result under Limitations.

<hr>

๐Ÿ“ฆ What is in this repository

Two separate artifacts. They are not interchangeable.

filewhat it isneeds
๐ŸŸข llama31-8b-Q4_K_M.gguf<br><sub>4.92 GB</sub>Merged model โ€” base + LoRA, imatrix-calibratedllama.cpp / llama-server. Self-contained
๐Ÿ”ต adapter_model.safetensors<br>adapter_config.json<br><sub>336 MB</sub>The LoRA adapter onlyunsloth/Meta-Llama-3.1-8B-Instruct + PEFT

<hr>

โšก Quick start

llama.cpp

bash
./llama-server -m llama31-8b-Q4_K_M.gguf \
  -c 8192 -ngl 99 -fa on \
  --cache-type-k q8_0 --cache-type-v q8_0 \
  --host 127.0.0.1 --port 8080

The q8_0 KV cache halves cache cost to ~68 KB/token, which is what makes 8192 context fit alongside the weights on an 8 GB card.

transformers + PEFT

python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = "unsloth/Meta-Llama-3.1-8B-Instruct"
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(model, "Berhak/Llama-3.1-8B-Function-Calling-Agent")
tokenizer = AutoTokenizer.from_pretrained(base)

<hr>

๐Ÿ”Œ Prompt contract

### โš ๏ธ This matters more than the weights. The model was trained on one specific shape and behaves poorly outside it. All three pieces below are load-bearing.

1๏ธโƒฃ System turn

Tool schemas go inside <tools> as a JSON array, in Llama-3.1's native header format:

<|begin_of_text|><|start_header_id|>system<|end_header_id|>

Cutting Knowledge Date: December 2023
Today Date: 12 Sep 2026

You are a function calling AI model. You are provided with function signatures within <tools></tools> XML tags. You may call one or more functions to assist with the user query. Don't make assumptions about what values to plug into functions.
<tools>
[{"type": "function", "function": {"name": "get_weather", "description": "Get current or forecast weather conditions for a location.", "parameters": {"type": "object", "properties": {"location": {"type": "string", "description": "The city name, e.g. 'Ankara'."}, "unit": {"type": "string", "enum": ["celsius", "fahrenheit"], "description": "Temperature unit."}}, "required": ["location"]}}}]
</tools>
For each function call return a json object with function name and arguments within <tool_call></tool_call> XML tags.<|eot_id|>

๐Ÿ“… Pass the real date. strftime("%b") follows the OS locale and produces a non-English month abbreviation under a non-English one, which breaks the template.

2๏ธโƒฃ What the model emits

<tool_call>
{"name": "get_weather", "arguments": {"location": "Berlin", "unit": "celsius"}}
</tool_call>

๐Ÿ”ข It emits several blocks in one generation when the request needs several tools. Parse every block, not just the first.

3๏ธโƒฃ Observations โ€” the unusual part

Tool results go back on the `ipython` role, double-encoded: the <tool_response> block is wrapped in a JSON string, escaped quotes and literal \n included. An artifact of the original Hermes conversion, but the model learned that exact shape:

<|start_header_id|>ipython<|end_header_id|>

"<tool_response>\n{\"name\": \"get_weather\", \"content\": {\"location\": \"Berlin\", \"temperature\": 18.4}}\n</tool_response>"<|eot_id|>
python
observation = json.dumps(
    "<tool_response>\n" + json.dumps({"name": name, "content": result}) + "\n</tool_response>"
)
๐Ÿšจ A plain, single-encoded block puts the model off its training distribution.

<hr>

๐Ÿ”’ Recommended: constrain decoding with a grammar

Build a GBNF grammar from the tool schemas you pass in this request โ€” not from a fixed file โ€” with a root that permits either a run of calls or plain prose:

gbnf
root      ::= tool-call (ws-nl tool-call)* | prose
tool-call ::= "<tool_call>" ws-nl call ws-nl "</tool_call>"
prose     ::= [^<] ( [^<] | "<" )*
ws-nl     ::= [ \t\n]*

<div align="center">

grammar offgrammar on
JSON validity56/59 ยท 94%57/57 ยท 100%
Overall71/82 ยท 86%72/82 ยท 87%
Unparseable blocks30

</div>

No category regressed, and the gain concentrates where the model is weakest unaided โ€” generations needing more than one call. Three details are worth copying verbatim:

  • โ€”๐Ÿ•ณ๏ธ The root must allow prose, or every question is forced into a tool call.
  • โ€”โš–๏ธ prose bans < only in the first position. Banning it everywhere ([^<]+) made the character unrepresentable โ€” asked for an inequality the model wrote 3 โ‰ค 5 instead of 3 < 5, a different claim, not a formatting quirk.
  • โ€”๐Ÿงฉ Build it per request. A fixed grammar forbids the caller's own schemas and scores 0/11 on unseen ones.

<hr>

๐Ÿ“Š Evaluation

86 hand-written held-out records, screened against the training data for verbatim matches and at a 0.70 similarity threshold. Every tool name claimed as "unseen" was checked against all 2,985 tool names present in training. Scored with the grammar on:

metricscore
๐ŸŽฏ Unseen tool schemas24 / 26โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘ 92%
๐Ÿ”‘ Seen tool schemas48 / 56โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘ 85%
โœ… JSON validity57 / 57โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ 100%
๐ŸŽ“ Calls built correctly from an unseen schema11 / 11โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ 100%
๐Ÿšซ Not calling a tool when none is needed8 / 8โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ 100%
๐ŸŒ Live-information questions โ†’ search6 / 6โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ 100%
๐Ÿงฎ Arithmetic routed to a calculator tool5 / 5โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ 100%
๐Ÿ”— Two calls in one generation2 / 2โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ 100%
๐Ÿ“‹ Several tasks in one message5 / 7โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘โ–‘ 71%
๐Ÿชค Keyword collision without picking the wrong tool3 / 5โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘โ–‘โ–‘ 60%
๐Ÿ‘ป Not fabricating a missing required argument1 / 4โ–ˆโ–ˆโ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘ 25%
๐Ÿงฉ Overall72 / 82โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘ 87%

Four ambiguous records are reported but never scored โ€” both behaviours are defensible there, and inventing a reference answer to complete a metric would only make the metric worse.

A handful of records also score the language of the final answer against Turkish (see below), so those are not an English measurement. Every row above except Overall measures tool selection and argument construction, which is language-independent.

<hr>

โš ๏ธ Limitations

### ๐Ÿšง It does not call a tool after seeing an observation. This is structural, not a weak tendency.

Of the 8,144 assistant turns that follow a tool result across 9,146 training examples, zero are a tool call. The data only ever contains CALL โ†’ OBSERVE โ†’ ANSWER; a second call appears solely after a new user turn. A directive telling it to continue scored 0/3.

โœ… Plan around it: a request needing several tools must be satisfied by the first generation. The model does this reliably, so parse every block it emits. Do not build a loop that expects it to iterate.


๐Ÿ‘ป It fabricates arguments it does not have. "What's the weather?" with no city produced location='Ankara'; "is there a train from Ankara?" produced destination="user's destination" โ€” a literal placeholder, meaning it knows it does not know and fills the slot anyway. No prompt variant fixed this. Validate arguments against the conversation before dispatching.

1๏ธโƒฃ It asks for one search result. The training data calls search with num_results=1 in 34 of 55 cases. Treat that parameter as a floor in your tool, not a ceiling.

๐ŸŽญ It invents explanations for things that do not exist. Asked about a fabricated syndrome, it answers confidently and fluently. Empty or failed tool results should say so explicitly in the observation, and say what the model must not do.

๐ŸŽฏ Search queries drift off subject. Asked to search two topics, the second query sometimes lands on a neighbouring one โ€” worse with accumulated context.

๐Ÿงฎ Arithmetic needs a tool. Unaided it produced "1 FP16 = 2 INT4, so 16 GB becomes 32 GB" โ€” the ratio is 4 and the operation is division โ€” then reused the wrong figure as a premise on the next turn. Route calculations to a calculator tool.

๐Ÿ“ Factual accuracy of direct answers was never measured, and ๐Ÿ“‰ no bfloat16 baseline exists โ€” the merged model was deleted by the quantization pipeline before one was taken, so every number here is absolute rather than a measurement of quantization loss.

๐ŸŒ On Turkish. The training mix includes a small Turkish subset (~5%), and the evaluation set was originally written against Turkish final turns. That path is not validated and is not recommended โ€” the agent around this model was brought to a working standard in English only. Treat this as an English function-calling model.

<hr>

๐Ÿ‹๏ธ Training

QLoRA on 4ร— H100, a single run with no retry budget. Source: NousResearch/hermes-function-calling-v1, converted from ChatML into Llama-3.1's native chat template before loss masking โ€” Llama-3.1's tokenizer does not recognize <|im_start|> as a special token, so raw ChatML would shatter into meaningless sub-words.

๐ŸŽ›๏ธ LoRAr=32, ฮฑ=64, dropout 0.05, 4-bit NF4 + double quant, bf16 compute
๐ŸŽฏ Target modulesq_proj k_proj v_proj o_proj gate_proj up_proj down_proj
๐Ÿ“ / ๐Ÿ”4096 tokens ยท 2 epochs, early stopping (patience 3)
๐Ÿ“‰ LR2e-4 cosine, 3% warmup, weight decay 0.01
๐Ÿ“ฆ Batch2 per device ร— 8 gradient accumulation
๐Ÿ—ƒ๏ธ Data9,146 examples โ†’ 8,781 train / 365 eval
๐Ÿ•ณ๏ธ One warning worth repeating. English final-turn prose was masked out so it would not compete with the intended final-turn behaviour. But Hermes examples that answer without a tool consist of nothing but a final turn โ€” so masking it masked the whole example, and 803 of 851 vanished silently. The resulting skew toward calling a tool is why the model over-triggers, and why the reference agent adds a restraint clause to the system prompt. If you reuse this recipe: count your examples after masking, not before.

Quantization: bf16 merge โ†’ GGUF q8_0 โ†’ imatrix โ†’ Q4_K_M. The intermediate is q8_0 rather than f16 (half the disk, no measurable K-quant penalty; needs --allow-requantize), and the importance matrix was calibrated on 40 samples drawn from the real training distribution.

<hr>

๐Ÿ’ป Measured performance

RTX 4060 Laptop (8 GB), Vulkan, n_ctx=8192, Q4KM, q8_0 KV cache, flash attention:

<div align="center">

๐Ÿง  Weights4403 MiB
๐Ÿ—‚๏ธ KV cache544 MiB ยท 68 KB/token
โš™๏ธ Compute buffer108 MiB
๐Ÿ’พ Total5055 / 7774 MiB
โœ๏ธ Generation~40 tok/s
๐Ÿ“ฅ Prompt processing~175 tok/s

</div>

### ๐Ÿšฆ Memory is not the binding constraint โ€” prompt processing is. At 175 tok/s, every 1000 tokens of context costs about six seconds on every subsequent turn. Trim tool output aggressively and keep the system prefix stable so prefix caching holds (measured: 14 tokens reprocessed instead of 440 on the second request).

Sampling: temperature=0.5 top_k=40 top_p=0.9 repeat_penalty=1.1 repeat_last_n=768. At temperature=0 the model locked into repeating one sentence across turns; repeat_penalty is what broke that, not the temperature. Use temperature=0 for reproducible evaluation.

<hr>

๐Ÿ“œ License

Llama 3.1 Community License. Use is subject to the Llama 3.1 License and the Acceptable Use Policy.

Built with Llama.

Training data: `NousResearch/hermes-function-calling-v1`, subject to its own terms.

bibtex
@misc{llama31-8b-function-calling-agent,
  title  = {Llama-3.1-8B-Function-Calling-Agent},
  author = {Tanyฤฑldฤฑzฤฑ, Mahmut Berhak},
  year   = {2026},
  url    = {https://huggingface.co/Berhak/Llama-3.1-8B-Function-Calling-Agent}
}