Team Ai
Datasetpublic

YanZhanPKU/SLCA-GRPO-Datasets

SLCA-GRPO Β· Datasets πŸ“„ Paper (arXiv:2609.29050)   β€’   πŸ’» Code   β€’   πŸ€— Collection This repository contains every processed data file read by SLCA-GRPO, the segment-locked credit assignment estimator for tool-calling RL introduced in "SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL". A tool-calling rollout opens with structured tool-call tokens and closes with free-form summary text; SLCA-GRPO normalises the two segment rewards… See the full description on the dataset page: https://huggingface.co/datasets/YanZhanPKU/SLCA-GRPO-Datasets.

sourceHugging Faceapache-2.0updated 15d agoView on Hugging Face
2likes385downloads
Dataset Card

<h1 style="text-align: left; font-size: 1.6em; margin-bottom: 0.75em;"> <span style="color:#6B5FD1; font-weight:bold;">SLCA</span><span style="color:#8A8A8A; font-weight:bold;">-</span><span style="color:#5FA34E; font-weight:bold;">GRPO</span> Β· Datasets </h1>

<div style="text-align: left; margin-bottom: 18px;"> <a href="https://arxiv.org/abs/2609.29050" target="blank">πŸ“„ Paper (arXiv:2609.29050)</a> &nbsp; β€’ &nbsp; <a href="https://github.com/SLCA-GRPO/SLCA-GRPO" target="blank">πŸ’» Code</a> &nbsp; β€’ &nbsp; <a href="https://huggingface.co/collections/YanZhanPKU/slca-grpo" target="_blank">πŸ€— Collection</a> </div>

This repository contains every processed data file read by SLCA-GRPO, the segment-locked credit assignment estimator for tool-calling RL introduced in "SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL". A tool-calling rollout opens with structured tool-call tokens and closes with free-form summary text; SLCA-GRPO normalises the two segment rewards independently within the same rollout group and routes each advantage only into its own tokens. These files are the SFT warm-start data, the RL trajectories, and the held-out in-domain evaluation split used to measure that change.

No model weights are released. The paper's reproducibility statement scopes its artifacts to code and processed splits and states that no trained checkpoints are included.

It packages the four splits that appear in the main experiment:

  1. 1.sft_split β€” the 42,423-trajectory 2-epoch SFT set used to warm-start every backbone before RL (main-table recipe).
  2. 2.sft_full β€” the whole 74,241-trajectory training pool, i.e. the union of sft_split and the source trajectories of rl (42,423 + 31,818). Used 1-epoch for the "SFT full" single-stage ablation.
  3. 3.rl β€” the 31,818 multi-turn, schema-constrained trajectories used by SLCA-GRPO during the RL stage. Each row pairs a user instruction, a permitted tool set, and a gold tool-call trajectory; the SGLS tool simulator replaces real APIs during rollout.
  4. 4.eval β€” the 4,000 held-out evaluation split (toucan_eval_4k_unified), reporting in-domain Name F1 / ArgMatch / Process / Success in the paper's main results.

All four splits trace back to the publicly released **Toucan-1.5M** corpus (MCP-sourced multi-turn tool-calling trajectories). We apply an additional relevance filter, strict tool-schema validation, a held-out evaluation cut, and an RL-side decomposition step. The full pipeline is documented in the paper appendix (Data Processing Pipeline).

The released teacher trajectories are real MCP trajectories. Segment masks used by SLCA-GRPO are generated deterministically from the structured trace and the environment boundary; no annotated segment or token labels and no learned segmenter are used. The eval split contains 4,000 rows. The execution-based reward curve uses the 3,991 rows that pass that run's validity filter.

Summary

ConfigSplitRowsMessages / row (mean, range)Role in the paper
sft_splittrain42,4239.40 (5 – 99)main-table SFT warm-start (2 epochs)
sft_fulltrain74,2419.52 (5 – 99)single-stage "SFT full" ablation (1 epoch)
rltrain31,8184.43 (2 – 78)SLCA-GRPO RL training
evaltest4,0003.81 (2 – 52)in-domain held-out evaluation

Message counts are exact, measured over every row. For sft_split / sft_full they count the full conversations list (system + human + gpt turns); for rl / eval they count the stored prompt messages only β€” the assistant continuation is generated during rollout, not shipped.

Quickstart

python
from datasets import load_dataset

# Load the main SFT split (used to reproduce the 7B / 3B / 8B backbones).
sft = load_dataset("YanZhanPKU/SLCA-GRPO-Datasets",
                   name="sft_split", split="train")

# Load the RL training trajectories.
rl = load_dataset("YanZhanPKU/SLCA-GRPO-Datasets",
                  name="rl", split="train")

# Load the 4k held-out evaluation set.
ev = load_dataset("YanZhanPKU/SLCA-GRPO-Datasets",
                  name="eval", split="test")

print(sft[0])

A 5-row preview is shipped under `samples/` for quick schema inspection without fully downloading any split.

Schema

sft_split / sft_full (ShareGPT + hermes tool-calling markup)

ColumnTypeDescription
conversationslist[{from, value}]Multi-turn dialogue in ShareGPT form: from ∈ {"system","human","gpt"}. The system turn embeds the tool schema as NDJSON inside a <tools>…</tools> block (hermes template); human turns that carry tool responses wrap them in <tool_response>…</tool_response>; gpt turns may embed <tool_call>[…]</tool_call>.

gpt turns contain interleaved <think>…</think> and <tool_call>[…]</tool_call> blocks followed by a user-facing summary. Those two sub-spans define the structural segments y_tool / y_sum that SLCA-GRPO routes advantages to.

LLaMA-Factory's sharegpt loader does not read this markup directly β€” run the small converter in the next section once to materialise the {messages, tools} JSON that LF expects.

rl / eval (verl-native format)

ColumnTypeDescription
promptlist[{role, content}]Chat-templated prompt up to the last assistant turn
data_sourcestrRouting key for reward_fn.compute_score (toucan_toolcall_v4_rl, toucan_eval_v4)
toolsstrJSON-serialized list of OpenAI function-calling tool schemas (deserialize with json.loads)
reward_modeldict{"style": "rule", "ground_truth": {...}} β€” see sub-fields below
extra_infodictProvenance fields (original_idx, source, parallel_type, has_history, and optional turn_index / split_position for decomposed multi-turn rows)

reward_model.ground_truth (verl's unified reward interface) contains:

Sub-fieldTypeDescription
question_contentstrHuman-readable instruction (used by the LLM judge)
allowed_toolslistTool name whitelist for this rollout
tool_schemasstrFull tool schema list passed to the model and SGLS
gold_tool_callsstrJSON-serialized nested list of gold per-turn tool calls (for HierR process reward). Format: [[step1_parallel_calls], [step2_parallel_calls], …]
subset_namestr"tool_call"

Both splits are consumed directly by reward_fn.compute_score.

Training with LLaMA-Factory

The SFT parquets above ship in ShareGPT-with-hermes-markup shape (single conversations column). LLaMA-Factory expects a flat JSON with two columns β€” messages (stringified chat) and tools (stringified tool schema list) β€” matching the samples/toucan_toolcall_sft.preview.jsonl preview. A small one-file materialiser is shipped alongside this dataset as `convert_sft_parquet_to_json.py`:

bash
# 1. Download the two SFT parquets from this HF dataset.
huggingface-cli download YanZhanPKU/SLCA-GRPO-Datasets \
    --repo-type dataset \
    --include "data/toucan_toolcall_sft_*.parquet" \
    --local-dir ./data

# 2. Materialise them to LLaMA-Factory-ready JSON (one-time step).
python convert_sft_parquet_to_json.py \
    --input  data/toucan_toolcall_sft_split_42k.parquet \
    --output data/toucan_toolcall_sft.json

python convert_sft_parquet_to_json.py \
    --input  data/toucan_toolcall_sft_full_74k.parquet \
    --output data/toucan_toolcall_full.json

Register the materialised files in your LLaMA-Factory dataset_info.json (a ready-to-use copy with these exact tag names ships in the SLCA-GRPO code bundle that accompanies this dataset, under data/dataset_info.json):

json
{
  "toucan_toolcall_sft": {
    "file_name": "toucan_toolcall_sft.json",
    "formatting": "sharegpt",
    "columns": {"messages": "messages", "tools": "tools"},
    "tags": {
      "role_tag": "role",
      "content_tag": "content",
      "user_tag": "user",
      "assistant_tag": "assistant",
      "observation_tag": "tool_response",
      "function_tag": "tool_call",
      "system_tag": "system"
    }
  },
  "toucan_toolcall_full": {
    "file_name": "toucan_toolcall_full.json",
    "formatting": "sharegpt",
    "columns": {"messages": "messages", "tools": "tools"},
    "tags": {
      "role_tag": "role",
      "content_tag": "content",
      "user_tag": "user",
      "assistant_tag": "assistant",
      "observation_tag": "tool_response",
      "function_tag": "tool_call",
      "system_tag": "system"
    }
  }
}

What the converter does per row: extracts the NDJSON <tools> block from the system turn, expands every <tool_call> (parallel calls in one turn become separate tool_call messages), and rewrites <tool_response> wrappers as flat tool_response role messages. The output is byte-for-byte compatible with samples/toucan_toolcall_sft.preview.jsonl, so the preview file doubles as a golden test for conversion correctness.

The RL and eval parquets are consumed directly by verl and do not require this step.

Provenance and license

  • β€”Upstream corpus: Agent-Ark/Toucan-1.5M β€” multi-turn tool-calling trajectories collected from public MCP servers. Please see the upstream dataset card for the original license and usage terms.
  • β€”Transformations applied in this release: 119,279 raw trajectories β†’ relevance and format filtering (βˆ’~40,000) β†’ strict schema validation (βˆ’1,038) β†’ a 78,241 valid pool β†’ a 4,000-row evaluation set carved out before any training split β†’ a 74,241 training pool partitioned into 42,423 SFT and 31,818 RL rows, with RL-side multi-turn decomposition. Detailed in the paper appendix (Data Processing Pipeline, Figure fig:data_process).
  • β€”License: Apache-2.0 for the derivative artefacts released here. Please verify that your downstream use also complies with the upstream Toucan-1.5M license.

Intended use

  • β€”Training: SFT warm-start and on-policy RL for tool-calling agents, in particular for studying segment-level credit assignment (SLCA-GRPO, GRPO, ToolPO, RLTR, GiGPO, KTAE, VinePPO, etc.).
  • β€”Evaluation: In-domain multi-turn tool-calling evaluation (Name F1 / ArgMatch / Process / Success). For out-of-distribution checks the paper uses BFCL-v3 and τ²-Bench β€” the evaluation entry points live in the code repository (not in this dataset).

Not intended for

  • β€”Deploying agents that interact with real user data without additional safety review β€” the gold trajectories target feature behaviour, not safety alignment.
  • β€”Training models to mimic any specific user's language style.

Citation

bibtex
@article{zhan2026slcagrpo,
  title   = {SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL},
  author  = {Zhan, Yan and Liu, Shaobo and Liu, Qiunan and Shi, Yuanjun and
             Xu, Siqi and Hou, WeiYi and Xu, Xiang and Li, Zekang and
             Pan, Weizhou and Yan, Jiahong},
  journal = {arXiv preprint},
  year    = {2026},
  eprint  = {2609.29050},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url     = {https://arxiv.org/abs/2609.29050}
}

Please also cite the upstream corpus these splits derive from:

bibtex
@misc{toucan2025,
  title        = {Toucan-1.5M},
  howpublished = {\url{https://huggingface.co/datasets/Agent-Ark/Toucan-1.5M}},
  note         = {Apache-2.0}
}