step3
Datasets
All datasets matching “step3”Step-3.5-Flash-SFT
Step-3.5-Flash-SFT
Step-3.5-Flash-SFT is a general-domain supervised fine-tuning release for chat models.
This repository keeps the full training interface in one place:
json/: canonical raw training data
tokenizers/: tokenizer snapshots for Step-3.5-Flash and Qwen3, released to preserve chat-template alignment
compiled/: tokenizer-specific compiled shards for StepTronOSS training
Data Format
Each raw shard is a JSON file whose top level is a list of examples.… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/Step-3.5-Flash-SFT.Step-3.5-Flash-SFT-code
Step-3.5-Flash-SFT-code
Code SFT dataset extracted from stepfun-ai/Step-3.5-Flash-SFT, converted to extended OpenAI SFT format, with multi-dimensional quality labels and thinking-mode classification.
602,595 conversations | 40.30 GB | 80 files
Dataset Summary
Group
Records
Size
Avg Rounds
Description
single_turn/slow
445,173
20.75 GB
1.0
Single-turn with chain-of-thought reasoning
single_turn/fast
9,746
0.12 GB
1.0
Single-turn without reasoning… See the full description on the dataset page: https://huggingface.co/datasets/liuxsh9/Step-3.5-Flash-SFT-code.mh2_ckp_step3_tcp_no_jointtowel_fold01_step3_deltaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_shoulder_pan.pos",
"left_shoulder_lift.pos",
"left_elbow_flex.pos",
"left_wrist_flex.pos",
"left_wrist_roll.pos",
"left_gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/HyeonseokE/towel_fold01_step3_delta.step35-en2pl-conv-pass4-jsonlconversations: 1,251,034
chat-template tokens (role+content, incl. special tokens): 2,664,206,408
reasoning_content tokens (not covered by chat template, counted separately): 6,662,763,429
avg tokens/conversation: 2129.6
used tokenizer: APT4
Step-3.5-Flash-SFT
Step-3.5-Flash-SFT
Step-3.5-Flash-SFT is a general-domain supervised fine-tuning release for chat models.
This repository keeps the full training interface in one place:
json/: canonical raw training data
tokenizers/: tokenizer snapshots for Step-3.5-Flash and Qwen3, released to preserve chat-template alignment
compiled/: tokenizer-specific compiled shards for StepTronOSS training
Data Format
Each raw shard is a JSON file whose top level is a list of examples.… See the full description on the dataset page: https://huggingface.co/datasets/VibeTraining/Step-3.5-Flash-SFT.
