Team Ai
Datasetpublic

rifqi2320/xlam-function-calling-60k-cli

xLAM Function Calling 60K — Positional CLI A deterministic transformation of Salesforce/xlam-function-calling-60k, created with the APIGen pipeline by Salesforce AI Research. This is an independent derivative, not an official Salesforce release. 60,000 examples; 100,011 calls; zero conversion failures. Every call preserves its original name, argument names, values, and JSON value types. The exported Parquet file was independently read back and all rows were validated again. This… See the full description on the dataset page: https://huggingface.co/datasets/rifqi2320/xlam-function-calling-60k-cli.

sourceHugging Facecc-by-4.0updated 8d agoView on Hugging Face
0likes47downloads
Dataset Card

xLAM Function Calling 60K — Positional CLI

A deterministic transformation of Salesforce/xlam-function-calling-60k, created with the APIGen pipeline by Salesforce AI Research. This is an independent derivative, not an official Salesforce release.

60,000 examples; 100,011 calls; zero conversion failures. Every call preserves its original name, argument names, values, and JSON value types. The exported Parquet file was independently read back and all rows were validated again. This verifies the serialization, not the semantic correctness of the source answers.

Format

text
tool ARG1 ARG2 ...

Parameter positions follow the source tool definitions' insertion order. Every present value is a JSON value: strings are quoted, numbers stay numeric, booleans/null stay literals, and nested values use compact JSON. Missing middle arguments use the bare marker __MISSING__; trailing absent arguments are omitted. A literal string containing that marker is quoted. One call occupies one physical line; newlines inside strings are escaped. Call order is preserved without asserting parallel execution.

text
live_giveaways_by_type "beta"
live_giveaways_by_type "game"

The format has no shell execution or expansion. xlam_cli.py implements parsing, encoding, and validation. Tool descriptions and parameter descriptions/defaults/declared types are included in cli_tool_context and the training messages. Actual answer values are preserved even when they disagree with a declared schema type.

Source ambiguity

260 source rows repeat tool names; 259 contain differing definitions and 147 call a name with differing definitions. Original answers identify only a tool name, so the intended API identity cannot always be recovered. For duplicate names, the codec derives one shared positional order by taking the union of parameter keys in first-appearance order across the tool definitions. This rule depends only on the tool context, never on the target answers. All source definitions remain in the CLI help and original JSON. These rows are flagged rather than silently discarded or assigned invented API identities.

Filter has_ambiguous_tool_names == False to exclude the 259 rows with differing definitions; filter an empty called_ambiguous_tool_names to exclude only the 147 rows whose answers call such names. Neither filter is applied to the published train split.

Columns

ColumnMeaning
source_id, source_row_indexOriginal ID as a string; original zero-based row index
source_dataset, format_versionSource attribution and codec version
queryOriginal user query
tools_json, answers_jsonOriginal unmodified source JSON strings; stored as strings to avoid Arrow coercion of heterogeneous values
cli_tool_contextCLI signatures plus original tool/parameter descriptions
cli_answers, assistant_cliOrdered list of calls; the same calls joined by newlines
messagesSystem, user, and assistant messages for supervised training
roundtrip_ok, call_countVerified conversion flag; number of calls
duplicate_tool_names, ambiguous_tool_namesRepeated names; repeated names with differing definitions
called_ambiguous_tool_names, has_ambiguous_tool_namesAmbiguity in called names; any differing same-name definitions
python
from datasets import load_dataset
ds = load_dataset("rifqi2320/xlam-function-calling-60k-cli", split="train")
print(ds[0]["assistant_cli"])

Reproducibility and changes

Source revision: 26d14ebfe18b1f7b524bd39b404b50af5dc97866. Source JSON SHA-256: 4ef5c6f0dc552f2231f93f5853a9ef431e9e806d7aa514d0f6b615606ce576c6.

The bytes were retrieved from the public redistribution lockon/xlam-function-calling-60k at the same revision. The downloaded SHA-256 and size reconstruct the Git LFS pointer object e3a0920447577124a6ab69f6be93b64178782ecb, exactly matching Salesforce's upstream metadata. Thus this is the canonical current source file, not a different reformatted dataset. See validation_report.json for provenance, checksums, counts, and diagnosis of the earlier converter's 1,469 failures.

Changes: add the CLI representation, descriptive CLI help, training messages, strict type-aware validation, and ambiguity flags. No LLM rewrites, value coercion, argument removal, train/test resplitting, or silent row dropping were applied.

To regenerate from an authorized original source download:

bash
pip install pyarrow
python xlam_cli.py xlam_function_calling_60k.json --output-dir regenerated

The accompanying notebook uploads the already-converted files from the designated Drive folder. It does not redownload xLAM or need access to its gated upstream repository.

Attribution and license

The source dataset and this derivative are licensed under Creative Commons Attribution 4.0. Retain attribution to Salesforce AI Research and APIGen, a license link, and notice of these format changes when redistributing.

Please cite the original APIGen work: APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets, Zuxin Liu et al., 2024.

bibtex
@article{liu2024apigen,
  title={APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets},
  author={Liu, Zuxin and others},
  journal={arXiv:2406.18518},
  year={2024}
}