Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zhiyuanhucs /nemotron-student-fail-v41-clean-thinking DeepSeek-V4.1 clean and action-only trajectories with Nemotron outcomes DeepSeek-V4.1 reward-1 trajectories rebuilt from the complete teacher audit under v57-test-path-component-boundary+v57-target-source-recheck. The V4.1 reward and trajectory tier do not by themselves prove that Nemotron failed. Student outcomes are joined from nemotron-prolike-coverage-audit-20261001.json. A student failure requires either complete required-test results with reward 0, or an individually… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuanhucs/nemotron-student-fail-v41-clean-thinking.tabulartext-generationn<1K1 likes13k downloads7d agoHugging Face02Davd-b01 /thinking-cap-tier-raw-traces Thinking Cap Tier Raw Traces (TCS v4) [!IMPORTANT] Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture: All 38,158 candidate reasoning traces across all 4 tiers (candidates_low.jsonl, candidates_mid.jsonl, candidates_high.jsonl, candidates_xhigh.jsonl) are 100% sanitized: Zero batch-padding residues (<|pad|>): Completely purged across all records. Strict Delimiter Integrity: Generation blocks cleanly separate thought deliberation tags… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-raw-traces.tabulartext-generation10K<n<100K0 likes316 downloads27d agoHugging Face03twinkle-ai /gpt-oss-120b-mandarin-thinking-eval-logs-and-scorestabular100K<n<1M0 likes201 downloads7mo agoHugging Face04twinkle-ai /gpt-oss-20b-mandarin-thinking-eval-logs-and-scorestabular100K<n<1M0 likes195 downloads7mo agoHugging Face05shreethar /thinkflow-vla-features-b2tabular1K<n<10K0 likes122 downloads2mo agoHugging Face06Jackrong /Chinese-Qwen3-235B-Thinking-2507-Distill-100k 📌 Note: The English translation of this dataset card is provided below. Chinese-Qwen3-235B-Thinking-2507-Distill-100k Dataset Summary Chinese-Qwen3-235B-Thinking-2507-Distill-100k 是一个包含约 100k 条高质量中文推理与指令数据的数据集,由 Qwen-3-235B-A22B-Thinking-2507(官方 Thinking 模式,上下文长度 32K)蒸馏生成。 该数据集覆盖了多个重要领域: 数学与工程任务(Mathematics, Applied Math, Advanced Math) 通用知识与写作(General Knowledge, Language & Writing) 技术与编程(Technology & Programming) 商业与经济(Business & Economics)… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k.tabulartext-classification100K<n<1M19 likes109 downloads1y agoHugging Face07novastar112 /pusht_96_int1_visual_nomarker_allstep_thinking_trickiness_cot PushT int1 Visual Nomarker All-Step Thinking Trickiness COT This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_int1_visual_nomarker. Each row contains one full successful trajectory from the first move through the final stop action. Main files: training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows. testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows. Message format: Each user turn is the PushT prompt text plus one… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_int1_visual_nomarker_allstep_thinking_trickiness_cot.imageimage-to-text100K<n<1M0 likes84 downloads5mo agoHugging Face08islam-kamel /MBPP-Thinking-Gate-1k MBPP Thinking-Gate SFT Dataset This package contains two related assets: Ready 1,000-row MBPP-style dataset (all.jsonl, train.jsonl, validation.jsonl). It is synthetic and designed to test/train autonomous routing between <DIRECT> and <THINK>. Official-MBPP builder (build_from_official_mbpp.py). Run this to create the production dataset from the official Google Research MBPP source. Why two response modes? The training target starts with one of two routing… See the full description on the dataset page: https://huggingface.co/datasets/islam-kamel/MBPP-Thinking-Gate-1k.tabulartext-generation1K<n<10K0 likes83 downloads20d agoHugging Face09violetxi /tb21-eval-qwen35-rewritten-w005-20k-thinking-32k-timeout2x qwen35-rewritten-w005-20k — Terminal-Bench 2.1 Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-rewritten-obs-wm-weight-0p05-20k-tacc through the served model ID qwen35-rewritten-w005-20k with Terminus-2. Noncanonical run: timeout_multiplier=2 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs. Result Recorded trials: 445 Tasks / attempts: 89 × 5 Errored trials scored… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-rewritten-w005-20k-thinking-32k-timeout2x.tabularreinforcement-learningn<1K0 likes63 downloads2mo agoHugging Face10novastar112 /pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot PushT norm4 Visual Nomarker All-Step Thinking Trickiness COT This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_norm4_visual_nomarker. Each row contains one full successful trajectory from the first move through the final stop action. Main files: training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows. testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows. metadata/final_scan_validation.json: full local scan after repair… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot.imageimage-to-text100K<n<1M0 likes46 downloads5mo agoHugging Face11open-llm-leaderboard /DavidAU__DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B-detailsgated Dataset Card for Evaluation run of DavidAU/DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B Dataset automatically created during the evaluation run of model DavidAU/DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DavidAU__DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B-details.tabular10K<n<100K2 likes45 downloads2y agoHugging Face12violetxi /tb21-eval-qwen35-rewritten-w005-10k-thinking-32k-timeout2x qwen35-rewritten-w005-10k — Terminal-Bench 2.1 Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-rewritten-obs-wm-weight-0p05-10k-tacc through the served model ID qwen35-rewritten-w005-10k with Terminus-2. Noncanonical run: timeout_multiplier=2 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs. Result Trials: 445 Tasks / attempts: 89 × 5 Errored trials scored as zero:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-rewritten-w005-10k-thinking-32k-timeout2x.tabularreinforcement-learningn<1K0 likes45 downloads2mo agoHugging Face13toroe /Dolci-Think-SFT-7B-Propella-Annotationstabular1M<n<10M0 likes44 downloads8mo agoHugging Face14ThinkingRM /Edit-Reviewimagen<1K0 likes43 downloads4mo agoHugging Face15violetxi /tb21-eval-qwen35-4b-action-only-20k-thinking-timeout2x-gcp4-infra-interrupted qwen35-action-only-20k — Terminal-Bench 2.1 Nonstandard Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-action-only-20k-tacc@bd9c914f4774b2f217cecdf1ed18a7d2f3e0a623 through the served model ID qwen35-action-only-20k with Terminus-2. Result Recorded trials: 445 Tasks / attempts: 89 × 5 Errored trials scored as zero: 408 Exception counts: {"AgentTimeoutError": 41, "InternalServerError": 367} Agent timeouts / context-length events / output-cap… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-4b-action-only-20k-thinking-timeout2x-gcp4-infra-interrupted.tabularreinforcement-learningn<1K0 likes42 downloads2mo agoHugging Face16Shaik1903 /ThinkLess-data ThinkLess-data The data behind ThinkLess-2B: the SFT set of short, correct reasoning traces, every raw generation it was selected from, and every benchmark answer behind the reported numbers. Config What it is Rows sft (default) The SFT training set: the shortest correct solution per problem 8,890 rollouts_8k All Qwen3.5-2B samples at an 8k-token budget, right and wrong [ROWS_samples] rollouts_16k_retry Qwen3.5-2B retries at 16k for problems unsolved at 8k… See the full description on the dataset page: https://huggingface.co/datasets/Shaik1903/ThinkLess-data.tabulartext-generation100K<n<1M0 likes42 downloads9d agoHugging Face17violetxi /tb21-eval-qwen35-original-w005-20k-thinking-32k-timeout2x qwen35-original-w005-20k — Terminal-Bench 2.1 Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-original-obs-wm-weight-0p05-20k-tacc through the served model ID qwen35-original-w005-20k with Terminus-2. Noncanonical run: timeout_multiplier=2 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs. Result Recorded trials: 445 Tasks / attempts: 89 × 5 Errored trials scored as… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-original-w005-20k-thinking-32k-timeout2x.tabularreinforcement-learningn<1K0 likes38 downloads2mo agoHugging Face18MaxwellmuF /Soofi-Think-SFT-V2-secondhalf-DEtabular1M<n<10M0 likes37 downloads4mo agoHugging Face19novastar113 /sokoban_hard_allstep_thinking_cot_v8_res256_stop_prompt Sokoban hard allstep CoT v8 res256 stop-prompt rebuild This local dataset rebuilds the HF hard allstep-thinking CoT source against the v8 stop-required 256x256 state-replay hard train source. Rows removed by v8 layout dedup are skipped. Output shards are sorted to match the v8 hard train batch filenames and row order. Summary: metadata/rebuild_summary.json tabular100K<n<1M0 likes37 downloads4mo agoHugging Face20open-llm-leaderboard /bunnycore__Llama-3.2-3B-Long-Think-detailsgated Dataset Card for Evaluation run of bunnycore/Llama-3.2-3B-Long-Think Dataset automatically created during the evaluation run of model bunnycore/Llama-3.2-3B-Long-Think The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Llama-3.2-3B-Long-Think-details.tabular10K<n<100K0 likes36 downloads2y agoHugging Face21violetxi /tb21-eval-qwen35-4b-base-thinking-full [REDACTED] — Terminal-Bench 2.1 Canonical Terminal-Bench 2.1 evaluation of Qwen/Qwen3.5-4B through the served model ID [REDACTED] with Terminus-2. Result Recorded trials: 445 Tasks / attempts: 89 × 5 Errored trials scored as zero: 92 Exception counts: {"AgentTimeoutError": 89, "Timeout": 1, "VerifierTimeoutError": 2} Agent timeouts / context-length events / output-cap events: 89 / 0 / 0 Mean reward / Pass@1: 0.105618 Pass@5: 0.191011 Total input/output/cache… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-4b-base-thinking-full.tabularreinforcement-learningn<1K0 likes33 downloads2mo agoHugging Face22CL-From-Nothing /RLVE-Qwen3-4B-Thinking-2507-Pass8-Rollouts RLVE teacher rollouts — Qwen3-4B-Thinking-2507 (pass@8) Teacher rollouts for on-policy distillation on the RLVE environment suite. Teacher / sampler: Qwen3-4B-Thinking-2507 Source prompts: RLVE train split — 9000 questions across 18 environments (counting / combinatorics / optimization tasks) Sampling: 8 samples/question (pass@8) = 72000 records, temperature 1.0 (sample.sh default 0.7 -> here T per run), max 16384 new tokens Rewards: recomputed offline with the RLVE-Eval Gym… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/RLVE-Qwen3-4B-Thinking-2507-Pass8-Rollouts.tabulartext-generation10K<n<100K0 likes31 downloads4mo agoHugging Face23open-llm-leaderboard /fhai50032__Unaligned-Thinker-PHI-4-detailsgated Dataset Card for Evaluation run of fhai50032/Unaligned-Thinker-PHI-4 Dataset automatically created during the evaluation run of model fhai50032/Unaligned-Thinker-PHI-4 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fhai50032__Unaligned-Thinker-PHI-4-details.tabular10K<n<100K0 likes29 downloads2y agoHugging Face24open-llm-leaderboard /ewre324__Thinker-SmolLM2-135M-Instruct-Reasoning-detailsgated Dataset Card for Evaluation run of ewre324/Thinker-SmolLM2-135M-Instruct-Reasoning Dataset automatically created during the evaluation run of model ewre324/Thinker-SmolLM2-135M-Instruct-Reasoning The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ewre324__Thinker-SmolLM2-135M-Instruct-Reasoning-details.tabular10K<n<100K0 likes26 downloads2y agoHugging Face25spectralbranding /exp-thinking-primacy Experiment F1: Thinking-Mode Primacy Decomposition Paper DOI: 10.5281/zenodo.19422427 — R15 (Zharnikov, 2026v) Dataset DOI: 10.57967/hf/8457 Source Code: spectralbranding/sbt-papers/r15-ai-search-metamerism Dataset Summary 640 LLM API calls decomposing the serial position (primacy) effect by model architecture and thinking mode. Part of the R15 study on dimensional collapse in AI-mediated brand perception (Zharnikov, 2026v). The artifacts are raw API-call records… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/exp-thinking-primacy.tabulartext-generationn<1K0 likes26 downloads3mo agoHugging Face26violetxi /tb21-eval-qwen35-original-w005-20k-thinking-32k-timeout8x qwen35-original-w005-20k — Terminal-Bench 2.1 Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-original-obs-wm-weight-0p05-20k-tacc through the served model ID qwen35-original-w005-20k with Terminus-2. Noncanonical run: timeout_multiplier=8 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs. Result Recorded trials: 445 Tasks / attempts: 89 × 5 Errored trials scored as… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-original-w005-20k-thinking-32k-timeout8x.tabularreinforcement-learningn<1K0 likes25 downloads2mo agoHugging Face27yangwang92 /Dolci-Think-SFT-32B-q35instructtabularn<1K0 likes25 downloads2mo agoHugging Face28open-llm-leaderboard /ewre324__Thinker-Qwen2.5-0.5B-Instruct-Reasoning-detailsgated Dataset Card for Evaluation run of ewre324/Thinker-Qwen2.5-0.5B-Instruct-Reasoning Dataset automatically created during the evaluation run of model ewre324/Thinker-Qwen2.5-0.5B-Instruct-Reasoning The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ewre324__Thinker-Qwen2.5-0.5B-Instruct-Reasoning-details.tabular10K<n<100K0 likes23 downloads2y agoHugging Face29open-llm-leaderboard /Quazim0t0__ThinkPhi1.1-Tensors-detailsgated Dataset Card for Evaluation run of Quazim0t0/ThinkPhi1.1-Tensors Dataset automatically created during the evaluation run of model Quazim0t0/ThinkPhi1.1-Tensors The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Quazim0t0__ThinkPhi1.1-Tensors-details.tabular10K<n<100K0 likes23 downloads2y agoHugging Face30Jarrodbarnes /qwen3-0.6B-interleaved-thinking-data Qwen3 0.6B Interleaved Thinking Data This dataset contains 8,704 pretraining-style text chunks augmented with short interleaved teacher thoughts. It was built for the blog post Self-Improving Pretraining as a Substrate for Agentic Post-Training. The dataset turns ordinary pretraining text into the supervised stage of a thinking mid-training pipeline. A teacher inserts short local thoughts into raw FineWeb-Edu chunks while preserving the original text. The student then learns the… See the full description on the dataset page: https://huggingface.co/datasets/Jarrodbarnes/qwen3-0.6B-interleaved-thinking-data.tabulartext-generation1K<n<10K0 likes23 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.