Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01r0b0tlab /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M268 likes3.2k downloads2mo agoHugging Face020xSero /glm-5.2-nf3-hybrid-terminal-bench-2.1-traces GLM-5.2 (MXFP8-NVFP4-NF3-Hybrid) — Terminal-Bench 2.1 agent traces Full agent traces from running Terminal-Bench 2.1 (89 tasks) against a self-hosted GLM-5.2 in a MXFP8-NVFP4-NF3-Hybrid quantization, using the Terminus-2 agent via Harbor. Configuration Model GLM-5.2 · MXFP8-NVFP4-NF3-Hybrid (753B MoE) Agent Terminus-2 Reasoning effort max Context 262,144 tokens Concurrency 2 Attempts / task 1 (k=1) Serving vLLM, tensor-parallel 4 +… See the full description on the dataset page: https://huggingface.co/datasets/0xSero/glm-5.2-nf3-hybrid-terminal-bench-2.1-traces.other1 likes2.5k downloads1mo agoHugging Face03comoZ /osworld-glm-5.3-flash-trajThese are the trajectory results from our GLM-5.3-Flash evaluation on OSWorld. For detailed evaluation results, configuration, and additional information, please refer to the following GitHub issue: https://github.com/xlang-ai/OSWorld/issues/591 image1K<n<10K0 likes2.1k downloads28d agoHugging Face04AletheiaResearch /GLM-5.2-AgentThis dataset was generated using teich by TeichAI GLM-5.2 Agent traces This directory contains raw agent trace files generated by teich. JSONL files: 319 Model metadata: glm-5.2 Training-ready tools Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them. Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP… See the full description on the dataset page: https://huggingface.co/datasets/AletheiaResearch/GLM-5.2-Agent.tabulartext-generationn<1K62 likes2k downloads3mo agoHugging Face05brandonmusic /GLM-5.3-Flash-BF16-Teacher-Logits GLM-5.3-Flash BF16 teacher logits This dataset contains full-vocabulary float32 teacher logits from the immutable zai-org/GLM-5.3-Flash-BF16 revision a6c167b62691b2bac901344b65cb651a70f53e43. It keeps the sealed final KLD panel qualification-only and publishes the separate non-final calibration panel under role-specific paths. Qualification-only final windows: 25 Qualification-only final prediction positions: 51175 Vocabulary size: 154880 Teacher receipt:… See the full description on the dataset page: https://huggingface.co/datasets/brandonmusic/GLM-5.3-Flash-BF16-Teacher-Logits.text-generation4 likes1.8k downloads2mo agoHugging Face06p-research /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/p-research/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes1.4k downloads26d agoHugging Face07open-athena /GLM-5.3-RLVR1-Training-Rollouts-2026.09.25 GLM 5.3 RLVR1 training rollouts This snapshot contains saved GLM 5.3 assistant continuations for the RLVR1 training split. raw_traces/train retains API responses and collection errors; sft_chat/train contains the SFT-ready normalized conversations; sft_rendered/train contains their rendered text. Matching part-*.parquet files represent one atomic collection block. snapshot-manifest.json records the exact row counts. Browse each representation using its dataset subset:… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/GLM-5.3-RLVR1-Training-Rollouts-2026.09.25.text10K<n<100K0 likes1.3k downloads13d agoHugging Face08liangzhidanta /claude-code-glm53-swesmith-trajectories Claude-Code-native Coding Agent Teacher Trajectories (GLM-5.3 × SWE-smith) English | 简体中文 A private research archive of execution-verified, multi-turn coding-agent trajectories. A strong teacher (GLM-5.3) drives a real coding-agent harness (Claude Code) inside verified Docker environments derived from SWE-smith tasks; every trajectory is graded in a clean verifier container against the task's exact FAIL_TO_PASS / PASS_TO_PASS tests. ⚠️ PRIVATE dataset. Raw wire traces contain… See the full description on the dataset page: https://huggingface.co/datasets/liangzhidanta/claude-code-glm53-swesmith-trajectories.text-generation1K<n<10K3 likes1.2k downloads16d agoHugging Face09OpenMed /Medical-Reasoning-SFT-GLM_4.5_Air Medical-Reasoning-SFT-GLM_4.5_Air A large-scale medical reasoning dataset generated using zai-org/GLM-4.5-Air, containing over 225,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. Dataset Overview Metric Value Model zai-org/GLM-4.5-Air Total Samples 225,179 Samples with Reasoning 224,942 (99.9%) Estimated Tokens ~441 Million Content Tokens ~315 Million Reasoning Tokens ~126 Million Language English… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GLM_4.5_Air.texttext-generation100K<n<1M15 likes1.1k downloads8mo agoHugging Face10inferenceport-ai /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/inferenceport-ai/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes896 downloads27d agoHugging Face11o0Biggz0o /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/o0Biggz0o/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes887 downloads2mo agoHugging Face12abdurrehman456 /GLM_5.2_Training_Datatext1M<n<10M4 likes862 downloads4mo agoHugging Face13ianncity /GLM-5.2-Conversation GLM-5.2 · Conversation-50000x 50,000x traces distilled from GLM-5.2 on High reasoning Token Count: 120M Distribution: Speaking domains: •Greetings •Customer Support •Step by step explanations •Motivational language •Logical Questions •Creative Writing STEM: •Algebra, calculus, quantum mechanics concepts •Astromony and astrophysics •Datascience and machine learning •Biology Programming:… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/GLM-5.2-Conversation.texttext-generation10K<n<100K58 likes835 downloads3mo agoHugging Face14open-athena /recursive-task-synthesis-glm-5.3-rollouts GLM 5.3 agentic rollouts on Recursive-Task-Synthesis This dataset catalogs the full collection made from the pinned Recursive-Task-Synthesis dataset revision be44f96808d5a9b599d5cb024341ff00091adeb7. The repository includes approximately 260.5 GiB of trajectory payload tar shards. Contents at a glance Item Count Source tasks considered 37,284 Source candidates inspected 19,368 Converted tasks after source filters 18,600 Tasks passing gold… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/recursive-task-synthesis-glm-5.3-rollouts.tabulartext-generation100K<n<1M0 likes833 downloads21d agoHugging Face15bhadra123 /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/bhadra123/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes821 downloads2mo agoHugging Face16lx-Meteors /glm-5.3-flash-spotimage1K<n<10K1 likes817 downloads18d agoHugging Face17Rallex3 /glm-5.3-flash-function-calling GLM-5.3-Flash Function Calling (synthetic) A synthetic function-calling dataset generated with zai-org/GLM-5.3-Flash via Hugging Face Inference Providers. 513 examples in 8 domains: weather, calendar, finance, travel, e-commerce, devops, smart home, communication. Categories: single-turn tool calls, parallel/multiple calls in one turn, multi-turn trajectories with tool results, and no-tool-needed turns. Format: OpenAI-style — each row has tools (JSON-schema function… See the full description on the dataset page: https://huggingface.co/datasets/Rallex3/glm-5.3-flash-function-calling.textquestion-answeringn<1K0 likes808 downloads1mo agoHugging Face18best-distill /glm-5.3-flash-distillation-chat Private distill of domofon/finetome-cot-100k instructions through GLM-5.3-Flash (AutoClaw / Z.AI). Split train — successful generations only. field description instruction user prompt from FineToMe response GLM final answer (message.content) reasoning GLM chain-of-thought (reasoning_content), empty if not captured finish stop or length prompt_tokens / completion_tokens / reasoning_tokens usage latency_s request latency source_index original FineToMe… See the full description on the dataset page: https://huggingface.co/datasets/best-distill/glm-5.3-flash-distillation-chat.tabulartext-generation10K<n<100K8 likes807 downloads26d agoHugging Face19Jackrong /GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields. This release was prepared from the original dataset published by Kassadin88. Summary Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GLM-5.1-Reasoning-1M-Cleaned.texttext-generation100K<n<1M297 likes805 downloads6mo agoHugging Face20weili-0234 /roadmapbench-omp-glm52-80-tracesgated RoadmapBench multi-agent coding traces (Oh My Pi harness, GLM-5.2 FP8, 80 tasks) This is a byte-identical mirror of the Google Drive folder multiagent_coding_roadmapbench_glm52 (id 1DPeB1em9VgIg7Q2JcL1WdkzWiT8tIZxU). The mirror was taken on 2026-09-28, and every file was verified against the folder's own manifests: md5 for the 80 trace archives, sha256 for the reconstruction files. It is used to replay agentic multi-agent workloads against LLM serving systems, for example to… See the full description on the dataset page: https://huggingface.co/datasets/weili-0234/roadmapbench-omp-glm52-80-traces.0 likes771 downloads12d agoHugging Face21zai-org /glm-simple-evals-dataset glm-simple-evals-dataset This repository is dedicated to storing various evaluation data required for the glm-simple-evals evaluation project, to enable industry researchers and developers to reproduce the performance of the GLM-4.5 series models on reported benchmarks. Currently, this repository covers the data required for the following evaluation tasks: AIME GPQA HLE LiveCodeBench MATH 500 SciCode MMLU Pro Usage Instructions To use these evaluation datasets… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/glm-simple-evals-dataset.tabular10K<n<100K5 likes767 downloads1y agoHugging Face22malaiwah /GLM-5.3-Flash-calibration-activations-v1 GLM-5.3-Flash calibration activations v1 (BF16, natural routing) Per-layer block-input activations of zai-org/GLM-5.3-Flash-BF16 @ b1967181 over 92x2048 tokens of the exllamav3 standard_cal_data corpus (pinned): per context, layer_NNN.attn_in and layer_NNN.mlp_in (bf16, post-norm linear inputs; mlp_in is the router + expert gate/up input) and layer_NNN.router_logits (fp32, natural top-8 routing ground truth). Per-expert Hessians E[xx^T], routing statistics and down-proj inputs… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-calibration-activations-v1.tabularn<1K0 likes710 downloads1mo agoHugging Face23ansulev /qwen3.8-max-glm5.2-kimi-k3-distill Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/qwen3.8-max-glm5.2-kimi-k3-distill.tabulartext-generation10M<n<100M0 likes700 downloads2mo agoHugging Face24vcruz305 /GLM53-CAMPAIGN-ARCHIVE Encrypted preservation archive This dataset contains encrypted backup objects. Plaintext contents, filenames, provenance, and decryption keys are not published here. Ciphertext SHA-256 digests and byte lengths permit independent transport verification. Authorized recovery requires separately held keys and private recovery metadata. No model or research artifact is released by this archive. textn<1K0 likes659 downloads10d agoHugging Face25greghavens /glm-5.2-coding-and-debugging-traces GLM 5.2 Agent Traces 207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from GLM 5.2 (glm-5.2). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain. This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/glm-5.2-coding-and-debugging-traces.tabulartext-generation1K<n<10K21 likes587 downloads3mo agoHugging Face26AgentNativeResearchLab /arc-agi3-cc-glm5.2-ft09 ARC-AGI-3 ft09 — Agent Trajectories (cc-glm5.2) Gameplay trajectories from the harness×model pair cc-glm5.2 playing the ARC-AGI-3 game ft09, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-cc-glm5.2-ft09.reinforcement-learning0 likes585 downloads1mo agoHugging Face27CodeMasterCody3D /glm46v-flash-ultramega-sparse-hmap0 likes568 downloads21d agoHugging Face28omarelsherif010 /glm-ocr-bnk-finetuning GLM-OCR Fine-Tuning Pipeline Fine-tuning GLM-OCR 0.9B (CogViT encoder + GLM-0.5B decoder) for Korean financial document table recognition using LoRA via LLaMA-Factory. Performance Targets Metric Target TEDS (2-level nested) >= 90% TEDS (3-level nested) >= 85% Korean CER <= 1% Latency <= 0.5s/page Directory Structure glm_ocr_finetuning/ ├── config/ # Training/eval YAML configs │ ├── training_config.yaml #… See the full description on the dataset page: https://huggingface.co/datasets/omarelsherif010/glm-ocr-bnk-finetuning.0 likes562 downloads8mo agoHugging Face29AgentNativeResearchLab /arc-agi3-cc-glm5.2-ar25 ARC-AGI-3 ar25 — Agent Trajectories (cc-glm5.2) Gameplay trajectories from the harness×model pair cc-glm5.2 playing the ARC-AGI-3 game ar25, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-cc-glm5.2-ar25.reinforcement-learning0 likes497 downloads1mo agoHugging Face30malaiwah /GLM-5.3-Flash-fidelity-suite-v1 GLM-5.3-Flash Fidelity Suite v1 Historical distribution-fidelity evidence for GLM-5.3-Flash (released 2026-08-26): BF16-reference and FP8-as-served hidden-state captures, a shared LM head, and receipts from the declared capture/replay path. Compatible candidate captures can be compared on matching published positions without holding the 643 GB reference; this is not a universal native-serving or task-quality score. Protocol: the Qwen3.8-27B fidelity-suite v5 methodology… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-fidelity-suite-v1.0 likes456 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.