Team Ai
Modelpublic

ecnu-icalk/SkillGym-Agent

sourceHugging Faceapache-2.0updated 17d agoView on Hugging Face
3likes176downloads
Model Card

<div align="center">

<picture> <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/ECNU-ICALK/SkillGym/main/assets/brand-dark.png"> <source media="(prefers-color-scheme: light)" srcset="https://raw.githubusercontent.com/ECNU-ICALK/SkillGym/main/assets/brand-light.png"> <img src="https://raw.githubusercontent.com/ECNU-ICALK/SkillGym/main/assets/brand-light.png" alt="SkillGym" width="360"> </picture>

<h2>SkillGym-Agent</h2>

<p> <strong>A Qwen3.5-35B-A3B agent checkpoint trained on verified SkillGym trajectories</strong> </p>

<p> <strong>English</strong> · <a href="https://huggingface.co/ecnu-icalk/SkillGym-Agent/blob/main/README_zh-CN.md">简体中文</a> </p>

<p> <a href="https://github.com/ECNU-ICALK/SkillGym">GitHub</a> · <a href="https://huggingface.co/datasets/ecnu-icalk/SkillGym">Dataset</a> · <a href="https://github.com/ECNU-ICALK/SkillGym/blob/main/assets/InternalizingLargeScaleHumanWrittenSkillsintoLLMsforRealWorldProblemSolving.pdf">Paper</a> </p>

</div>

Model Summary

SkillGym-Agent is a full-parameter supervised fine-tuned version of Qwen3.5-35B-A3B, trained on successful, verifier-confirmed trajectories released with SkillGym.

The released checkpoint corresponds to the paper's All Teachers setting and is trained on the full set of 8,364 successful trajectories collected across the released teacher–harness configurations.

The model is designed to study whether procedural experience generated from human-written skills can become reusable capability inside an LLM agent, including when the external skill is removed at inference time.

Core idea: learn from verified experience generated with external skills, then evaluate how much of that capability is retained without those skills.
ItemDescription
Released paper settingAll Teachers
Base modelQwen/Qwen3.5-35B-A3B
ArchitectureQwen3.5 MoE, ~35.95B parameters
Training methodFull-parameter supervised fine-tuning
Training data8,364 successful SkillGym trajectories
Source environments2,756 accepted environments / 5,512 paired variants
Coverage12 major categories / 63 sub-categories
Training stackms-swift / Megatron
Reported hardware16 × NVIDIA H200 GPUs
LicenseApache-2.0

Quick Start

Download

bash
python -m pip install -U huggingface_hub

hf download ecnu-icalk/SkillGym-Agent \
  --local-dir .hf/skillgym-agent

Load with Transformers

python
from transformers import AutoModelForMultimodalLM, AutoProcessor

model_id = "ecnu-icalk/SkillGym-Agent"

processor = AutoProcessor.from_pretrained(
    model_id,
    trust_remote_code=True,
)

model = AutoModelForMultimodalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
    trust_remote_code=True,
)

SkillGym-Agent was evaluated through agentic harnesses rather than a single fixed chat template. Prompt formatting, tool schemas, runtime configuration, and the agent loop should therefore follow the selected harness and the current Qwen3.5 interface.

[!IMPORTANT] Loading the checkpoint alone does not reproduce the reported benchmark scores. Those results depend on the full harness, system prompt, runtime, inference budget, and benchmark configuration.

Evaluation

Higher is better for every metric. GDPval-AA v2 is reported as Elo; the remaining metrics are task success rates (%).

HarnessModelGDPval-AA v2<br>(Elo) ↑Terminal-Bench 2.1<br>(%) ↑SkillsBench v1.1<br>w/ Skills (%) ↑SkillsBench v1.1<br>w/o Skills (%) ↑
CodexQwen3.5-35B-A3B94210.115.330.69
CodexSkillGym-Agent979 (+37)46.07 (+35.96)33.02 (+27.69)21.08 (+20.39)
Claude CodeQwen3.5-35B-A3B97439.3323.3412.13
Claude CodeSkillGym-Agent1173 (+199)58.43 (+19.10)51.47 (+28.13)26.81 (+14.68)

Parenthesized values are absolute gains over the same-harness base model: Elo points for GDPval-AA v2 and percentage points for the success-rate metrics.

Skill-Free Transfer

A central result is the capability retained after removing the inference-time skill:

  • —Claude Code: SkillGym-Agent reaches 26.81% without skills, compared with 23.34% for the base model with skills.
  • —Codex: SkillGym-Agent reaches 21.08% without skills, compared with 5.33% for the base model with skills.

Performance remains strongest when external skills are available, suggesting that internalized capability and explicit skills are complementary.

Evaluation Notes

  • —Under Claude Code, the base and trained models use the same standard system prompt.
  • —Under Codex, the base uses the standard prompt while SkillGym-Agent uses the no-applypatch prompt; the Codex delta is therefore not a pure fine-tuning-only comparison.
  • —Benchmark results depend on the harness, prompt, runtime, inference budget, checkpoint, and benchmark version.
  • —Public-reference results may use different evaluation configurations and should not be interpreted as strictly controlled head-to-head comparisons.

See the paper for teacher/harness ablations and the complete evaluation configuration.

Training Data

The checkpoint is trained on successful trajectories from the SkillGym dataset.

The Dataset Card is the canonical reference for:

  • —human-written skills and supporting assets;
  • —reusable task templates and fixtures;
  • —executable environments and verifiers;
  • —paired with-skill / without-skill variants;
  • —trajectory collections and schema differences;
  • —dataset provenance and artifact-specific licensing.

Intended Use

SkillGym-Agent is intended for research on:

  • —agentic problem solving and tool use;
  • —long-horizon workflow learning;
  • —skill-conditioned and skill-free execution;
  • —transfer from verified trajectories;
  • —executable-environment and verifier-based evaluation.

The checkpoint is not presented as a general safety guarantee or as a drop-in production agent. Users should independently validate tool permissions, sandbox isolation, generated code, and external side effects.

Limitations

  • —The released checkpoint does not include a complete agent runtime or tool environment.
  • —Reported benchmark scores require the corresponding evaluation harness and configuration.
  • —The current release does not include a standalone end-to-end training script or a single script reproducing every external benchmark result.
  • —Training data is derived from the released SkillGym environment and trajectory distribution and should not be treated as exhaustive coverage of real-world procedural tasks.
  • —External skills remain useful at inference time even after training on skill-generated experience.

Project Resources

ResourceScope
GitHub / SkillGymTask Builder, documentation, figures, and paper materials
SkillGym DatasetSkills, templates, environments, and trajectories
PaperMethod, dataset analysis, experiments, and ablations

Citation

bibtex
@misc{skillgym2026,
  title = {Internalizing Large-Scale Human-Written Skills into LLMs for Real-World Problem Solving},
  year  = {2026},
  url   = {https://github.com/ECNU-ICALK/SkillGym}
}

License

The SkillGym-Agent checkpoint is released under the Apache-2.0 license.

The companion dataset contains source skills, fixtures, and supporting assets that may retain upstream notices or additional terms. See the Dataset Card for artifact-specific licensing details.