Team Ai
Modelpublic

alapedriza/mistral-7b-function-calling-adapter

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes
Model Card

Mistral-7B Enterprise Function Calling Adapter

A QLoRA adapter fine-tuned on Mistral-7B-Instruct-v0.3 for structured function calling across 16 enterprise tool schemas.

Model Description

This adapter teaches Mistral 7B to reliably select the correct tool and produce valid JSON arguments when given a user request and a set of available tool schemas. It also learns when not to call a tool (conversational responses).

Supported Tools

The model was trained on 16 enterprise tool schemas covering:

ToolDescription
gettransactionhistoryRetrieve financial transactions
create_invoiceGenerate invoices
runcompliancecheckExecute compliance validations
getemployeedetailsLook up employee information
schedule_meetingBook calendar events
send_notificationDispatch notifications
generate_reportCreate business reports
updatecustomerrecordModify customer data
searchknowledgebaseQuery internal documentation
approve_expenseProcess expense approvals
getinventorystatusCheck stock levels
createsupportticketOpen support cases
rundatapipelineTrigger ETL jobs
getanalyticsdashboardFetch dashboard metrics
updateprojectstatusModify project tracking
assign_taskDelegate tasks to team members

Training Details

Dataset

  • —1,228 training examples / 154 validation / 139 test
  • —Synthetically generated with category balance: simple tool calls, complex multi-parameter calls, multi-tool calls, and conversational (no-tool) examples
  • —Each example: [system_prompt, user_message, assistant_response]

Training Configuration

ParameterValue
Base modelmistralai/Mistral-7B-Instruct-v0.3
MethodQLoRA (4-bit NF4, double quantization)
LoRA rank16
LoRA alpha32
LoRA dropout0.05
Target modulesqproj, kproj, vproj, oproj, gateproj, upproj, down_proj
Epochs1
Batch size1 (effective 16 via gradient accumulation)
Learning rate2e-4 (cosine schedule)
Optimizerpagedadamw8bit
Max sequence length2,048 tokens
Gradient checkpointingYes
Training time511 minutes on Kaggle T4x2

Training Approach: Tool Trimming

To keep sequences short and training efficient, each training example includes only a minimal subset of tool schemas in its system prompt:

  • —Used tools (extracted from the assistant response) + 1 random distractor tool
  • —Conversational examples (no tool call): 3 random tools

This reduces mean sequence length from ~5,500 tokens (all 16 tools) to ~915 tokens, enabling training within hardware constraints while preserving the learning signal for tool selection and parameter extraction.

Results

MetricValue
Train loss0.1889
Validation loss0.1838

Validation loss lower than train loss indicates strong generalization with no overfitting.

Usage

Loading the Adapter

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel

model_name = "mistralai/Mistral-7B-Instruct-v0.3"
adapter_name = "alapedriza/mistral-7b-function-calling-adapter"

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.float16,
    bnb_4bit_use_double_quant=True,
)

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    quantization_config=bnb_config,
    device_map="auto",
    torch_dtype=torch.float16,
)
model = PeftModel.from_pretrained(model, adapter_name)

Inference

python
import json

# Define available tools (full schemas)
tools = [...]  # List of tool schema dicts

system_prompt = f"""You are an AI assistant with access to enterprise tools.
When the user's request requires a tool, respond with a JSON object (or list of objects) containing "name" and "arguments" keys.
When no tool is needed, respond conversationally in plain text.

Available tools:
{json.dumps(tools, indent=2)}"""

messages = [
    {"role": "system", "content": system_prompt},
    {"role": "user", "content": "Show me the transaction history for account ACC-2847 from last month"},
]

inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True)
inputs = inputs.to(model.device)

with torch.no_grad():
    outputs = model.generate(inputs, max_new_tokens=512, temperature=0.1, do_sample=True)

response = tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True)
print(response)
# Expected: {"name": "get_transaction_history", "arguments": {"account_id": "ACC-2847", "period": "last_month"}}

Expected Output Format

For tool calls:

json
{"name": "tool_name", "arguments": {"param1": "value1", "param2": "value2"}}

For multi-tool calls:

json
[
  {"name": "tool_1", "arguments": {...}},
  {"name": "tool_2", "arguments": {...}}
]

For conversational responses (no tool needed):

Plain text response without JSON.

Limitations

  • —Trained on synthetic data — may not cover all edge cases in production
  • —Best performance when the system prompt format matches training (tool schemas as JSON in system message)
  • —Sequence length limited to 2,048 tokens during training; longer inputs may degrade quality
  • —At inference time, all 16 tool schemas are provided (unlike training which used trimmed subsets) — the model must generalize from seeing 2-3 tools during training to selecting from 16 at inference

Hardware Requirements

  • —Minimum: Single GPU with 8+ GB VRAM (4-bit quantized inference)
  • —Recommended: 16 GB GPU (e.g., T4, RTX 4090) for comfortable inference
  • —CPU inference possible but slow

Framework versions

  • —TRL: 1.4.0
  • —Transformers: 5.0.0
  • —Pytorch: 2.10.0+cu128
  • —Datasets: 4.8.3
  • —Tokenizers: 0.22.2

Citations

bibtex
@misc{lapedriza2025mistral7bfunctioncalling,
  title={Mistral 7B Enterprise Function Calling Adapter},
  author={Alberto Lapedriza},
  year={2025},
  publisher={HuggingFace},
  url={https://huggingface.co/alapedriza/mistral-7b-function-calling-adapter}
}

Cite TRL as:

bibtex
@software{vonwerra2020trl,
  title   = {{TRL: Transformers Reinforcement Learning}},
  author  = {von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and Beeching, Edward and Thrush, Tristan and Lambert, Nathan and Huang, Shengyi and Rasul, Kashif and Gallouédec, Quentin},
  license = {Apache-2.0},
  url     = {https://github.com/huggingface/trl},
  year    = {2020}
}