AISA-Framework/AISA-AR-FunctionCall-FT
AISA-AR-FunctionCall-FT
<p align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/628f7a71dd993507cfcbe587/vnL90Tybn1528x21dMNsd.png" width="700"/> </p>
Reliable Arabic Structured Tool Calling via Data-Centric Fine-Tuning
AISA-AR-FunctionCall-FT is a fully fine-tuned Arabic function-calling model built on top of FunctionGemma (Gemma 3 270M) and optimized for structured tool invocation in Arabic agentic systems.
The model converts natural Arabic requests into structured executable API calls, enabling reliable integration between language models and external tools.
This model is part of the AISA (Agentic AI Systems Architecture) initiative.
Try the Model in Google Colab
You can run a full inference example using the notebook below.

The notebook demonstrates:
- Loading the model
- Defining tool schemas
- Generating structured tool calls
- Parsing function call outputs
Model Overview
The model is designed to translate Arabic natural language requests into structured tool calls following the FunctionGemma tool-calling format.
Key Capabilities
- Arabic natural language → structured API calls
- Multi-dialect Arabic understanding
- Tool selection and argument extraction
- Structured execution environments
Supported domains:
Dataset
The model is trained on AISA-AR-FunctionCall — a production-ready Arabic function-calling dataset built through a rigorous data-centric pipeline:
- Dataset auditing
- Schema normalization
- Enum correction
- Tool pruning
- Prompt restructuring
- Tool sampling
Dataset splits:
Dataset includes:
- 5 Arabic dialects
- 8 real-world domains
- 27 tool schemas
- Structured tool-call annotations
Dataset: AISA-Framework/AISA-AR-FunctionCall
Training Methodology
The model was trained using a data-centric fine-tuning pipeline designed to stabilize structured execution.
Key pipeline steps:
- Structural dataset auditing
- Enum constraint repair
- Tool schema normalization
- Tool pruning (36 → 27 tools)
- Tool sampling to prevent prompt truncation
- FunctionGemma-compatible chat serialization
- Completion-only supervised fine-tuning
Training configuration:
Evaluation Results
Evaluation was performed on a held-out test set of 5,079 samples.
Clean Positive Evaluation (n = 2,873)
Key improvement: Parse failure reduced from 87% → <1%
Dialect Performance
Fine-tuning significantly reduces dialect disparity compared to the baseline model.
Known Limitations
Remaining errors are primarily semantic, including:
- Tool selection ambiguity
- Argument mismatches
- Domain overlap (e.g., weather vs. air quality)
Structured formatting errors are largely eliminated.
Example Usage
Prompt:
ما حالة الطقس في الرياض اليوم؟Model output:
<start_function_call>
call:get_weather{
city:<escape>الرياض<escape>,
days:1
}
<end_function_call>The structured call can then be executed by the application runtime.
Intended Use
This model is designed for:
- Arabic AI assistants
- Tool-based agents
- Structured API orchestration
- Arabic enterprise automation
- Research on multilingual tool calling
Out-of-Scope Uses
This model is not designed for:
- General chatbots or open-ended conversation
- Sensitive decision-making systems
- Safety-critical deployments without additional validation
Related Models
AISA Framework
This model is part of the AISA initiative for building reliable agentic AI systems.
Model collection: AISA-Framework/aisa-arabic-functioncall-datasets-and-models
