alapedriza/mistral-7b-function-calling-adapter
Mistral-7B Enterprise Function Calling Adapter
A QLoRA adapter fine-tuned on Mistral-7B-Instruct-v0.3 for structured function calling across 16 enterprise tool schemas.
Model Description
This adapter teaches Mistral 7B to reliably select the correct tool and produce valid JSON arguments when given a user request and a set of available tool schemas. It also learns when not to call a tool (conversational responses).
Supported Tools
The model was trained on 16 enterprise tool schemas covering:
Training Details
Dataset
- 1,228 training examples / 154 validation / 139 test
- Synthetically generated with category balance: simple tool calls, complex multi-parameter calls, multi-tool calls, and conversational (no-tool) examples
- Each example:
[system_prompt, user_message, assistant_response]
Training Configuration
Training Approach: Tool Trimming
To keep sequences short and training efficient, each training example includes only a minimal subset of tool schemas in its system prompt:
- Used tools (extracted from the assistant response) + 1 random distractor tool
- Conversational examples (no tool call): 3 random tools
This reduces mean sequence length from ~5,500 tokens (all 16 tools) to ~915 tokens, enabling training within hardware constraints while preserving the learning signal for tool selection and parameter extraction.
Results
Validation loss lower than train loss indicates strong generalization with no overfitting.
Usage
Loading the Adapter
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel
model_name = "mistralai/Mistral-7B-Instruct-v0.3"
adapter_name = "alapedriza/mistral-7b-function-calling-adapter"
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.float16,
bnb_4bit_use_double_quant=True,
)
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
quantization_config=bnb_config,
device_map="auto",
torch_dtype=torch.float16,
)
model = PeftModel.from_pretrained(model, adapter_name)Inference
import json
# Define available tools (full schemas)
tools = [...] # List of tool schema dicts
system_prompt = f"""You are an AI assistant with access to enterprise tools.
When the user's request requires a tool, respond with a JSON object (or list of objects) containing "name" and "arguments" keys.
When no tool is needed, respond conversationally in plain text.
Available tools:
{json.dumps(tools, indent=2)}"""
messages = [
{"role": "system", "content": system_prompt},
{"role": "user", "content": "Show me the transaction history for account ACC-2847 from last month"},
]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True)
inputs = inputs.to(model.device)
with torch.no_grad():
outputs = model.generate(inputs, max_new_tokens=512, temperature=0.1, do_sample=True)
response = tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True)
print(response)
# Expected: {"name": "get_transaction_history", "arguments": {"account_id": "ACC-2847", "period": "last_month"}}Expected Output Format
For tool calls:
{"name": "tool_name", "arguments": {"param1": "value1", "param2": "value2"}}For multi-tool calls:
[
{"name": "tool_1", "arguments": {...}},
{"name": "tool_2", "arguments": {...}}
]For conversational responses (no tool needed):
Plain text response without JSON.Limitations
- Trained on synthetic data — may not cover all edge cases in production
- Best performance when the system prompt format matches training (tool schemas as JSON in system message)
- Sequence length limited to 2,048 tokens during training; longer inputs may degrade quality
- At inference time, all 16 tool schemas are provided (unlike training which used trimmed subsets) — the model must generalize from seeing 2-3 tools during training to selecting from 16 at inference
Hardware Requirements
- Minimum: Single GPU with 8+ GB VRAM (4-bit quantized inference)
- Recommended: 16 GB GPU (e.g., T4, RTX 4090) for comfortable inference
- CPU inference possible but slow
Framework versions
- TRL: 1.4.0
- Transformers: 5.0.0
- Pytorch: 2.10.0+cu128
- Datasets: 4.8.3
- Tokenizers: 0.22.2
Citations
@misc{lapedriza2025mistral7bfunctioncalling,
title={Mistral 7B Enterprise Function Calling Adapter},
author={Alberto Lapedriza},
year={2025},
publisher={HuggingFace},
url={https://huggingface.co/alapedriza/mistral-7b-function-calling-adapter}
}Cite TRL as:
@software{vonwerra2020trl,
title = {{TRL: Transformers Reinforcement Learning}},
author = {von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and Beeching, Edward and Thrush, Tristan and Lambert, Nathan and Huang, Shengyi and Rasul, Kashif and Gallouédec, Quentin},
license = {Apache-2.0},
url = {https://github.com/huggingface/trl},
year = {2020}
}