datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen3.8-max-glm5.2-kimi-k3-distillation
Multi-Teacher Distillation Dataset (57,937 traces)
A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains.
Teachers
Teacher
Provider
Traces
Qwen3.8-Max-Preview
Alibaba Cloud Model Studio
48,283
GLM-5.2
Z.AI Coding Plan
5,307
Kimi Code K3
Moonshot AI (Kimi)
4,347… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation.glm-5.2-nf3-hybrid-terminal-bench-2.1-traces
GLM-5.2 (MXFP8-NVFP4-NF3-Hybrid) — Terminal-Bench 2.1 agent traces
Full agent traces from running Terminal-Bench 2.1 (89 tasks) against a self-hosted
GLM-5.2 in a MXFP8-NVFP4-NF3-Hybrid quantization, using the Terminus-2 agent
via Harbor.
Configuration
Model
GLM-5.2 · MXFP8-NVFP4-NF3-Hybrid (753B MoE)
Agent
Terminus-2
Reasoning effort
max
Context
262,144 tokens
Concurrency
2
Attempts / task
1 (k=1)
Serving
vLLM, tensor-parallel 4 +… See the full description on the dataset page: https://huggingface.co/datasets/0xSero/glm-5.2-nf3-hybrid-terminal-bench-2.1-traces.osworld-glm-5.3-flash-trajThese are the trajectory results from our GLM-5.3-Flash evaluation on OSWorld.
For detailed evaluation results, configuration, and additional information, please refer to the following GitHub issue:
https://github.com/xlang-ai/OSWorld/issues/591
GLM-5.2-AgentThis dataset was generated using teich by TeichAI
GLM-5.2 Agent traces
This directory contains raw agent trace files generated by teich.
JSONL files: 319
Model metadata: glm-5.2
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP… See the full description on the dataset page: https://huggingface.co/datasets/AletheiaResearch/GLM-5.2-Agent.GLM-5.3-Flash-BF16-Teacher-Logits
GLM-5.3-Flash BF16 teacher logits
This dataset contains full-vocabulary float32 teacher logits from the immutable
zai-org/GLM-5.3-Flash-BF16 revision a6c167b62691b2bac901344b65cb651a70f53e43.
It keeps the sealed final KLD panel qualification-only and publishes the
separate non-final calibration panel under role-specific paths.
Qualification-only final windows: 25
Qualification-only final prediction positions: 51175
Vocabulary size: 154880
Teacher receipt:… See the full description on the dataset page: https://huggingface.co/datasets/brandonmusic/GLM-5.3-Flash-BF16-Teacher-Logits.qwen3.8-max-glm5.2-kimi-k3-distillation
Multi-Teacher Distillation Dataset (57,937 traces)
A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains.
Teachers
Teacher
Provider
Traces
Qwen3.8-Max-Preview
Alibaba Cloud Model Studio
48,283
GLM-5.2
Z.AI Coding Plan
5,307
Kimi Code K3
Moonshot AI (Kimi)
4,347… See the full description on the dataset page: https://huggingface.co/datasets/p-research/qwen3.8-max-glm5.2-kimi-k3-distillation.GLM-5.3-RLVR1-Training-Rollouts-2026.09.25
GLM 5.3 RLVR1 training rollouts
This snapshot contains saved GLM 5.3 assistant continuations for the RLVR1 training split. raw_traces/train retains API responses and collection errors; sft_chat/train contains the SFT-ready normalized conversations; sft_rendered/train contains their rendered text. Matching part-*.parquet files represent one atomic collection block. snapshot-manifest.json records the exact row counts.
Browse each representation using its dataset subset:… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/GLM-5.3-RLVR1-Training-Rollouts-2026.09.25.claude-code-glm53-swesmith-trajectories
Claude-Code-native Coding Agent Teacher Trajectories (GLM-5.3 × SWE-smith)
English | 简体中文
A private research archive of execution-verified, multi-turn coding-agent trajectories.
A strong teacher (GLM-5.3) drives a real coding-agent harness (Claude Code) inside
verified Docker environments derived from SWE-smith tasks; every trajectory is graded
in a clean verifier container against the task's exact FAIL_TO_PASS / PASS_TO_PASS tests.
⚠️ PRIVATE dataset. Raw wire traces contain… See the full description on the dataset page: https://huggingface.co/datasets/liangzhidanta/claude-code-glm53-swesmith-trajectories.Medical-Reasoning-SFT-GLM_4.5_Air
Medical-Reasoning-SFT-GLM_4.5_Air
A large-scale medical reasoning dataset generated using zai-org/GLM-4.5-Air, containing over 225,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
Dataset Overview
Metric
Value
Model
zai-org/GLM-4.5-Air
Total Samples
225,179
Samples with Reasoning
224,942 (99.9%)
Estimated Tokens
~441 Million
Content Tokens
~315 Million
Reasoning Tokens
~126 Million
Language
English… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GLM_4.5_Air.qwen3.8-max-glm5.2-kimi-k3-distillation
Multi-Teacher Distillation Dataset (57,937 traces)
A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains.
Teachers
Teacher
Provider
Traces
Qwen3.8-Max-Preview
Alibaba Cloud Model Studio
48,283
GLM-5.2
Z.AI Coding Plan
5,307
Kimi Code K3
Moonshot AI (Kimi)
4,347… See the full description on the dataset page: https://huggingface.co/datasets/inferenceport-ai/qwen3.8-max-glm5.2-kimi-k3-distillation.qwen3.8-max-glm5.2-kimi-k3-distillation
Multi-Teacher Distillation Dataset (57,937 traces)
A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains.
Teachers
Teacher
Provider
Traces
Qwen3.8-Max-Preview
Alibaba Cloud Model Studio
48,283
GLM-5.2
Z.AI Coding Plan
5,307
Kimi Code K3
Moonshot AI (Kimi)
4,347… See the full description on the dataset page: https://huggingface.co/datasets/o0Biggz0o/qwen3.8-max-glm5.2-kimi-k3-distillation.GLM_5.2_Training_DataGLM-5.2-Conversation
GLM-5.2 · Conversation-50000x
50,000x traces distilled from GLM-5.2 on High reasoning
Token Count: 120M
Distribution:
Speaking domains:
•Greetings
•Customer Support
•Step by step explanations
•Motivational language
•Logical Questions
•Creative Writing
STEM:
•Algebra, calculus, quantum mechanics concepts
•Astromony and astrophysics
•Datascience and machine learning
•Biology
Programming:… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/GLM-5.2-Conversation.recursive-task-synthesis-glm-5.3-rollouts
GLM 5.3 agentic rollouts on Recursive-Task-Synthesis
This dataset catalogs the full collection made from the pinned
Recursive-Task-Synthesis dataset revision
be44f96808d5a9b599d5cb024341ff00091adeb7. The repository includes approximately 260.5 GiB of trajectory payload tar shards.
Contents at a glance
Item
Count
Source tasks considered
37,284
Source candidates inspected
19,368
Converted tasks after source filters
18,600
Tasks passing gold… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/recursive-task-synthesis-glm-5.3-rollouts.qwen3.8-max-glm5.2-kimi-k3-distillation
Multi-Teacher Distillation Dataset (57,937 traces)
A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains.
Teachers
Teacher
Provider
Traces
Qwen3.8-Max-Preview
Alibaba Cloud Model Studio
48,283
GLM-5.2
Z.AI Coding Plan
5,307
Kimi Code K3
Moonshot AI (Kimi)
4,347… See the full description on the dataset page: https://huggingface.co/datasets/bhadra123/qwen3.8-max-glm5.2-kimi-k3-distillation.glm-5.3-flash-spotglm-5.3-flash-function-calling
GLM-5.3-Flash Function Calling (synthetic)
A synthetic function-calling dataset generated with zai-org/GLM-5.3-Flash via Hugging Face Inference Providers.
513 examples in 8 domains: weather, calendar, finance, travel, e-commerce, devops, smart home, communication.
Categories: single-turn tool calls, parallel/multiple calls in one turn, multi-turn trajectories with tool results, and no-tool-needed turns.
Format: OpenAI-style — each row has tools (JSON-schema function… See the full description on the dataset page: https://huggingface.co/datasets/Rallex3/glm-5.3-flash-function-calling.glm-5.3-flash-distillation-chat
Private distill of domofon/finetome-cot-100k instructions through GLM-5.3-Flash (AutoClaw / Z.AI).
Split
train — successful generations only.
field
description
instruction
user prompt from FineToMe
response
GLM final answer (message.content)
reasoning
GLM chain-of-thought (reasoning_content), empty if not captured
finish
stop or length
prompt_tokens / completion_tokens / reasoning_tokens
usage
latency_s
request latency
source_index
original FineToMe… See the full description on the dataset page: https://huggingface.co/datasets/best-distill/glm-5.3-flash-distillation-chat.GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields.
This release was prepared from the original dataset published by Kassadin88.
Summary
Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GLM-5.1-Reasoning-1M-Cleaned.roadmapbench-omp-glm52-80-traces
RoadmapBench multi-agent coding traces (Oh My Pi harness, GLM-5.2 FP8, 80 tasks)
This is a byte-identical mirror of the Google Drive folder multiagent_coding_roadmapbench_glm52
(id 1DPeB1em9VgIg7Q2JcL1WdkzWiT8tIZxU). The mirror was taken on 2026-09-28, and every file was
verified against the folder's own manifests: md5 for the 80 trace archives, sha256 for the
reconstruction files. It is used to replay agentic multi-agent workloads against LLM serving systems,
for example to… See the full description on the dataset page: https://huggingface.co/datasets/weili-0234/roadmapbench-omp-glm52-80-traces.glm-simple-evals-dataset
glm-simple-evals-dataset
This repository is dedicated to storing various evaluation data required for the glm-simple-evals evaluation project, to enable industry researchers and developers to reproduce the performance of the GLM-4.5 series models on reported benchmarks.
Currently, this repository covers the data required for the following evaluation tasks:
AIME
GPQA
HLE
LiveCodeBench
MATH 500
SciCode
MMLU Pro
Usage Instructions
To use these evaluation datasets… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/glm-simple-evals-dataset.GLM-5.3-Flash-calibration-activations-v1
GLM-5.3-Flash calibration activations v1 (BF16, natural routing)
Per-layer block-input activations of zai-org/GLM-5.3-Flash-BF16 @ b1967181 over 92x2048
tokens of the exllamav3 standard_cal_data corpus (pinned): per context, layer_NNN.attn_in
and layer_NNN.mlp_in (bf16, post-norm linear inputs; mlp_in is the router + expert gate/up
input) and layer_NNN.router_logits (fp32, natural top-8 routing ground truth).
Per-expert Hessians E[xx^T], routing statistics and down-proj inputs… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-calibration-activations-v1.qwen3.8-max-glm5.2-kimi-k3-distill
Multi-Teacher Distillation Dataset (57,937 traces)
A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains.
Teachers
Teacher
Provider
Traces
Qwen3.8-Max-Preview
Alibaba Cloud Model Studio
48,283
GLM-5.2
Z.AI Coding Plan
5,307
Kimi Code K3
Moonshot AI (Kimi)
4,347… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/qwen3.8-max-glm5.2-kimi-k3-distill.GLM53-CAMPAIGN-ARCHIVE
Encrypted preservation archive
This dataset contains encrypted backup objects. Plaintext contents, filenames,
provenance, and decryption keys are not published here. Ciphertext SHA-256 digests
and byte lengths permit independent transport verification. Authorized recovery
requires separately held keys and private recovery metadata.
No model or research artifact is released by this archive.
glm-5.2-coding-and-debugging-traces
GLM 5.2 Agent Traces
207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from GLM 5.2 (glm-5.2). The category and row-share tables
below describe the actual mix seen during training rather than assuming a
particular task domain.
This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/glm-5.2-coding-and-debugging-traces.arc-agi3-cc-glm5.2-ft09
ARC-AGI-3 ft09 — Agent Trajectories (cc-glm5.2)
Gameplay trajectories from the harness×model pair cc-glm5.2 playing the
ARC-AGI-3 game ft09, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-cc-glm5.2-ft09.glm46v-flash-ultramega-sparse-hmapglm-ocr-bnk-finetuning
GLM-OCR Fine-Tuning Pipeline
Fine-tuning GLM-OCR 0.9B (CogViT encoder + GLM-0.5B decoder) for Korean financial document table recognition using LoRA via LLaMA-Factory.
Performance Targets
Metric
Target
TEDS (2-level nested)
>= 90%
TEDS (3-level nested)
>= 85%
Korean CER
<= 1%
Latency
<= 0.5s/page
Directory Structure
glm_ocr_finetuning/
├── config/ # Training/eval YAML configs
│ ├── training_config.yaml #… See the full description on the dataset page: https://huggingface.co/datasets/omarelsherif010/glm-ocr-bnk-finetuning.arc-agi3-cc-glm5.2-ar25
ARC-AGI-3 ar25 — Agent Trajectories (cc-glm5.2)
Gameplay trajectories from the harness×model pair cc-glm5.2 playing the
ARC-AGI-3 game ar25, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-cc-glm5.2-ar25.GLM-5.3-Flash-fidelity-suite-v1
GLM-5.3-Flash Fidelity Suite v1
Historical distribution-fidelity evidence for GLM-5.3-Flash (released 2026-08-26):
BF16-reference and FP8-as-served hidden-state captures, a shared LM head, and
receipts from the declared capture/replay path. Compatible candidate captures
can be compared on matching published positions without holding the 643 GB
reference; this is not a universal native-serving or task-quality score. Protocol: the Qwen3.8-27B fidelity-suite v5 methodology… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-fidelity-suite-v1.
