YanZhanPKU/SLCA-GRPO-Datasets
SLCA-GRPO Β· Datasets π Paper (arXiv:2609.29050) β’ π» Code β’ π€ Collection This repository contains every processed data file read by SLCA-GRPO, the segment-locked credit assignment estimator for tool-calling RL introduced in "SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL". A tool-calling rollout opens with structured tool-call tokens and closes with free-form summary text; SLCA-GRPO normalises the two segment rewardsβ¦ See the full description on the dataset page: https://huggingface.co/datasets/YanZhanPKU/SLCA-GRPO-Datasets.
<h1 style="text-align: left; font-size: 1.6em; margin-bottom: 0.75em;"> <span style="color:#6B5FD1; font-weight:bold;">SLCA</span><span style="color:#8A8A8A; font-weight:bold;">-</span><span style="color:#5FA34E; font-weight:bold;">GRPO</span> Β· Datasets </h1>
<div style="text-align: left; margin-bottom: 18px;"> <a href="https://arxiv.org/abs/2609.29050" target="blank">π Paper (arXiv:2609.29050)</a> β’ <a href="https://github.com/SLCA-GRPO/SLCA-GRPO" target="blank">π» Code</a> β’ <a href="https://huggingface.co/collections/YanZhanPKU/slca-grpo" target="_blank">π€ Collection</a> </div>
This repository contains every processed data file read by SLCA-GRPO, the segment-locked credit assignment estimator for tool-calling RL introduced in "SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL". A tool-calling rollout opens with structured tool-call tokens and closes with free-form summary text; SLCA-GRPO normalises the two segment rewards independently within the same rollout group and routes each advantage only into its own tokens. These files are the SFT warm-start data, the RL trajectories, and the held-out in-domain evaluation split used to measure that change.
No model weights are released. The paper's reproducibility statement scopes its artifacts to code and processed splits and states that no trained checkpoints are included.
It packages the four splits that appear in the main experiment:
sft_splitβ the 42,423-trajectory 2-epoch SFT set used to warm-start every backbone before RL (main-table recipe).sft_fullβ the whole 74,241-trajectory training pool, i.e. the union ofsft_splitand the source trajectories ofrl(42,423 + 31,818). Used 1-epoch for the "SFT full" single-stage ablation.rlβ the 31,818 multi-turn, schema-constrained trajectories used by SLCA-GRPO during the RL stage. Each row pairs a user instruction, a permitted tool set, and a gold tool-call trajectory; the SGLS tool simulator replaces real APIs during rollout.evalβ the 4,000 held-out evaluation split (toucan_eval_4k_unified), reporting in-domain Name F1 / ArgMatch / Process / Success in the paper's main results.
All four splits trace back to the publicly released **Toucan-1.5M** corpus (MCP-sourced multi-turn tool-calling trajectories). We apply an additional relevance filter, strict tool-schema validation, a held-out evaluation cut, and an RL-side decomposition step. The full pipeline is documented in the paper appendix (Data Processing Pipeline).
The released teacher trajectories are real MCP trajectories. Segment masks used by SLCA-GRPO are generated deterministically from the structured trace and the environment boundary; no annotated segment or token labels and no learned segmenter are used. The eval split contains 4,000 rows. The execution-based reward curve uses the 3,991 rows that pass that run's validity filter.
Summary
Message counts are exact, measured over every row. For sft_split / sft_full they count the full conversations list (system + human + gpt turns); for rl / eval they count the stored prompt messages only β the assistant continuation is generated during rollout, not shipped.
Quickstart
from datasets import load_dataset
# Load the main SFT split (used to reproduce the 7B / 3B / 8B backbones).
sft = load_dataset("YanZhanPKU/SLCA-GRPO-Datasets",
name="sft_split", split="train")
# Load the RL training trajectories.
rl = load_dataset("YanZhanPKU/SLCA-GRPO-Datasets",
name="rl", split="train")
# Load the 4k held-out evaluation set.
ev = load_dataset("YanZhanPKU/SLCA-GRPO-Datasets",
name="eval", split="test")
print(sft[0])A 5-row preview is shipped under `samples/` for quick schema inspection without fully downloading any split.
Schema
sft_split / sft_full (ShareGPT + hermes tool-calling markup)
gpt turns contain interleaved <think>β¦</think> and <tool_call>[β¦]</tool_call> blocks followed by a user-facing summary. Those two sub-spans define the structural segments y_tool / y_sum that SLCA-GRPO routes advantages to.
LLaMA-Factory's sharegpt loader does not read this markup directly β run the small converter in the next section once to materialise the {messages, tools} JSON that LF expects.
rl / eval (verl-native format)
reward_model.ground_truth (verl's unified reward interface) contains:
Both splits are consumed directly by reward_fn.compute_score.
Training with LLaMA-Factory
The SFT parquets above ship in ShareGPT-with-hermes-markup shape (single conversations column). LLaMA-Factory expects a flat JSON with two columns β messages (stringified chat) and tools (stringified tool schema list) β matching the samples/toucan_toolcall_sft.preview.jsonl preview. A small one-file materialiser is shipped alongside this dataset as `convert_sft_parquet_to_json.py`:
# 1. Download the two SFT parquets from this HF dataset.
huggingface-cli download YanZhanPKU/SLCA-GRPO-Datasets \
--repo-type dataset \
--include "data/toucan_toolcall_sft_*.parquet" \
--local-dir ./data
# 2. Materialise them to LLaMA-Factory-ready JSON (one-time step).
python convert_sft_parquet_to_json.py \
--input data/toucan_toolcall_sft_split_42k.parquet \
--output data/toucan_toolcall_sft.json
python convert_sft_parquet_to_json.py \
--input data/toucan_toolcall_sft_full_74k.parquet \
--output data/toucan_toolcall_full.jsonRegister the materialised files in your LLaMA-Factory dataset_info.json (a ready-to-use copy with these exact tag names ships in the SLCA-GRPO code bundle that accompanies this dataset, under data/dataset_info.json):
{
"toucan_toolcall_sft": {
"file_name": "toucan_toolcall_sft.json",
"formatting": "sharegpt",
"columns": {"messages": "messages", "tools": "tools"},
"tags": {
"role_tag": "role",
"content_tag": "content",
"user_tag": "user",
"assistant_tag": "assistant",
"observation_tag": "tool_response",
"function_tag": "tool_call",
"system_tag": "system"
}
},
"toucan_toolcall_full": {
"file_name": "toucan_toolcall_full.json",
"formatting": "sharegpt",
"columns": {"messages": "messages", "tools": "tools"},
"tags": {
"role_tag": "role",
"content_tag": "content",
"user_tag": "user",
"assistant_tag": "assistant",
"observation_tag": "tool_response",
"function_tag": "tool_call",
"system_tag": "system"
}
}
}What the converter does per row: extracts the NDJSON <tools> block from the system turn, expands every <tool_call> (parallel calls in one turn become separate tool_call messages), and rewrites <tool_response> wrappers as flat tool_response role messages. The output is byte-for-byte compatible with samples/toucan_toolcall_sft.preview.jsonl, so the preview file doubles as a golden test for conversion correctness.
The RL and eval parquets are consumed directly by verl and do not require this step.
Provenance and license
- Upstream corpus:
Agent-Ark/Toucan-1.5Mβ multi-turn tool-calling trajectories collected from public MCP servers. Please see the upstream dataset card for the original license and usage terms. - Transformations applied in this release: 119,279 raw trajectories β relevance and format filtering (β~40,000) β strict schema validation (β1,038) β a 78,241 valid pool β a 4,000-row evaluation set carved out before any training split β a 74,241 training pool partitioned into 42,423 SFT and 31,818 RL rows, with RL-side multi-turn decomposition. Detailed in the paper appendix (
Data Processing Pipeline, Figurefig:data_process). - License: Apache-2.0 for the derivative artefacts released here. Please verify that your downstream use also complies with the upstream Toucan-1.5M license.
Intended use
- Training: SFT warm-start and on-policy RL for tool-calling agents, in particular for studying segment-level credit assignment (SLCA-GRPO, GRPO, ToolPO, RLTR, GiGPO, KTAE, VinePPO, etc.).
- Evaluation: In-domain multi-turn tool-calling evaluation (Name F1 / ArgMatch / Process / Success). For out-of-distribution checks the paper uses BFCL-v3 and ΟΒ²-Bench β the evaluation entry points live in the code repository (not in this dataset).
Not intended for
- Deploying agents that interact with real user data without additional safety review β the gold trajectories target feature behaviour, not safety alignment.
- Training models to mimic any specific user's language style.
Citation
@article{zhan2026slcagrpo,
title = {SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL},
author = {Zhan, Yan and Liu, Shaobo and Liu, Qiunan and Shi, Yuanjun and
Xu, Siqi and Hou, WeiYi and Xu, Xiang and Li, Zekang and
Pan, Weizhou and Yan, Jiahong},
journal = {arXiv preprint},
year = {2026},
eprint = {2609.29050},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2609.29050}
}Please also cite the upstream corpus these splits derive from:
@misc{toucan2025,
title = {Toucan-1.5M},
howpublished = {\url{https://huggingface.co/datasets/Agent-Ark/Toucan-1.5M}},
note = {Apache-2.0}
}