Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CohereLabs /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Data source Prompts from AM-DeepSeek-R1-0528-Distilled Thinking traces and outputs distilled from gpt-oss-120b Translated with command-a-translate and DeepSeek-V3 Languages (44) Language Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M8 likes1.8k downloads29d agoHugging Face02gijl /agentic-thinker-v10 likes1.4k downloads1h agoHugging Face03erenyeager-1 /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Languages (44) Language Train Test Total Amharic (am) 3,807 448 4,255 Arabic (ar) 22,968 2,538 25,506 Bulgarian (bg) 4,177 452 4,629 Bengali (bn) 3,803 422 4,225 Catalan (ca) 4,251 512 4,763 Czech… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M0 likes440 downloads1mo agoHugging Face04ai2-adapt-dev /tulu-3-thinker-classified-no_math-ifeval-ref-based-45ktext10K<n<100K0 likes96 downloads1y agoHugging Face05thinkerhui /CoSQA_PlusCoSQA+, a code search dataset pairing high-quality queries (reused from CoSQA) with multiple suitable codes. We collect code candidates from diverse sources and form candidate pairs by pairing queries with these codes. Utilizing the power of large language models (LLMs), we automate pair annotation, filtering, and code generation for queries without suitable matches. related links: arXiv: 2406.11589 CoSQA+: Enhancing Code Search Dataset with Matching Code (arxiv.org) github:… See the full description on the dataset page: https://huggingface.co/datasets/thinkerhui/CoSQA_Plus.1 likes63 downloads2y agoHugging Face06XueZhang-bjtu /M-Thinker-SFT-datatext10K<n<100K0 likes63 downloads1y agoHugging Face07xytian1008 /VAPO-Thinker-train36kimage10K<n<100K1 likes58 downloads1y agoHugging Face08internlm /ARM-Thinker-Data ARM-Thinker-Data Paper | Github Repository 📊 Data Introduction This repository contains the datasets used for training ARM-Thinker, an Agentic Multimodal Reward Model that performs evidence-grounded reasoning through tool use and visual grounding. The current dataset is annotated by Qwen3-VL-235B-A22B-Instruct, Qwen3-VL-235B-A22B-Thinking, and GPT-4o, with all data files organized under the qwen/ directory. We are also planning to release an additional version… See the full description on the dataset page: https://huggingface.co/datasets/internlm/ARM-Thinker-Data.image-text-to-text10K<n<100K7 likes52 downloads8mo agoHugging Face09minchyeom /Thinker-XMLSystem prompt suggestion: You are a world-class AI system. Always respond in strict XML format with your reasoning steps within the <im_reasoning> XML tag. Each reasoning step should represent one unit of thought. Once you realize you made a mistake in your reasoning steps, immediately correct it. Place your final response outside the XML tag. Adhere to this XML structure without exception. texttext-generation1K<n<10K1 likes50 downloads2y agoHugging Face10OpenSPG /KAG-Thinker-training-datasettext100K<n<1M5 likes50 downloads1y agoHugging Face11hamishivi /GeneralThought-430K-filtered-thinkertext100K<n<1M0 likes47 downloads1y agoHugging Face12Stormtrooperaim /Ultra-Thinker-30k Ultra-Thinker 🧠 A comprehensive collection of high-quality reasoning and conversational datasets designed to enhance the thinking capabilities of large language models. Overview Ultra-Thinker aggregates diverse sources focused on: Chain-of-thought reasoning - Step-by-step logical deduction Mathematical problem solving - Complex computational challenges Advanced logical reasoning - Multi-step inference and analysis This curated compilation combines cutting-edge… See the full description on the dataset page: https://huggingface.co/datasets/Stormtrooperaim/Ultra-Thinker-30k.text10K<n<100K4 likes46 downloads8mo agoHugging Face13open-llm-leaderboard /DavidAU__DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B-detailsgated Dataset Card for Evaluation run of DavidAU/DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B Dataset automatically created during the evaluation run of model DavidAU/DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DavidAU__DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B-details.tabular10K<n<100K2 likes45 downloads2y agoHugging Face14xytian1008 /MUPO-Thinker-train36kimage10K<n<100K1 likes37 downloads1y agoHugging Face15minchyeom /Thinker-2text10K<n<100K0 likes31 downloads2y agoHugging Face16ai2-adapt-dev /tulu-3-thinker-rewritten-math-27ktext10K<n<100K0 likes30 downloads1y agoHugging Face17open-llm-leaderboard /fhai50032__Unaligned-Thinker-PHI-4-detailsgated Dataset Card for Evaluation run of fhai50032/Unaligned-Thinker-PHI-4 Dataset automatically created during the evaluation run of model fhai50032/Unaligned-Thinker-PHI-4 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fhai50032__Unaligned-Thinker-PHI-4-details.tabular10K<n<100K0 likes29 downloads2y agoHugging Face18minchyeom /thinkerA Chain-of-Thought (CoT) dataset that contains traces of complex and sophisticated reasoning, to mimic the "thinking" process of OpenAI's o1. Wrap the contents of the reasoning column in some XML tag (such as <reasoning>). Raw .jsonl dataset file can be found under the Files and Versions tab. texttext-generation1K<n<10K8 likes28 downloads2y agoHugging Face19open-llm-leaderboard /ewre324__Thinker-SmolLM2-135M-Instruct-Reasoning-detailsgated Dataset Card for Evaluation run of ewre324/Thinker-SmolLM2-135M-Instruct-Reasoning Dataset automatically created during the evaluation run of model ewre324/Thinker-SmolLM2-135M-Instruct-Reasoning The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ewre324__Thinker-SmolLM2-135M-Instruct-Reasoning-details.tabular10K<n<100K0 likes26 downloads2y agoHugging Face20UnfilteredAI /unfiltered-thinker Unfiltered-Thinker: A Dataset for Intermediate Cognitive Reasoning A corpus of 1,909 samples designed to showcase intermediate thinking, cognitive processes, and structured emotional reasoning. Source: UnfilteredAI/unfiltered-thinker on Hugging Face ⚠️ Content Warning: This dataset contains content that will be considered offensive, disturbing, or explicit. This includes discussions of dark humor, profanity, criminal activity, violence, substance use, and psychological distress. It… See the full description on the dataset page: https://huggingface.co/datasets/UnfilteredAI/unfiltered-thinker.text-generation1K<n<10K16 likes26 downloads1y agoHugging Face21open-llm-leaderboard /ewre324__Thinker-Qwen2.5-0.5B-Instruct-Reasoning-detailsgated Dataset Card for Evaluation run of ewre324/Thinker-Qwen2.5-0.5B-Instruct-Reasoning Dataset automatically created during the evaluation run of model ewre324/Thinker-Qwen2.5-0.5B-Instruct-Reasoning The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ewre324__Thinker-Qwen2.5-0.5B-Instruct-Reasoning-details.tabular10K<n<100K0 likes23 downloads2y agoHugging Face22open-llm-leaderboard /bunnycore__Qwen2.5-3B-RP-Thinker-detailsgated Dataset Card for Evaluation run of bunnycore/Qwen2.5-3B-RP-Thinker Dataset automatically created during the evaluation run of model bunnycore/Qwen2.5-3B-RP-Thinker The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Qwen2.5-3B-RP-Thinker-details.tabular10K<n<100K0 likes22 downloads2y agoHugging Face23afraamn /luffy_thinkertext10K<n<100K0 likes21 downloads1y agoHugging Face24Shrijanagain /DISTILLATION-VIBE-THINKER USED ADPATION LABS AUTO SCIENCETS DATA PREPARATION AND DISTALATE DATA OF VIBE THINKER text10K<n<100K1 likes19 downloads3mo agoHugging Face25xytian1008 /VAPO-Thinker-val1kimage1K<n<10K1 likes18 downloads1y agoHugging Face26minchyeom /Thinker-2-Longtext1K<n<10K1 likes16 downloads2y agoHugging Face27open-llm-leaderboard /bunnycore__Qwen2.5-3B-RP-Thinker-V2-detailsgated Dataset Card for Evaluation run of bunnycore/Qwen2.5-3B-RP-Thinker-V2 Dataset automatically created during the evaluation run of model bunnycore/Qwen2.5-3B-RP-Thinker-V2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Qwen2.5-3B-RP-Thinker-V2-details.tabular10K<n<100K0 likes16 downloads2y agoHugging Face28Disya /nbeerbower-Purpura-DPO-thinker-rawUnfiltered System promt for creating a dataset: You are an expert AI assistant specializing in text generation. Your task is to reverse-engineer the thought process that leads to a given textual `response`. Based on the user's `prompt` and the final `response` text, generate a plausible, detailed reasoning process of an LLM. This reasoning should cover: 1. **Analysis of the User's Prompt:** Deconstruct the user's request, identifying explicit constraints (like length, format) and implicit… See the full description on the dataset page: https://huggingface.co/datasets/Disya/nbeerbower-Purpura-DPO-thinker-raw.texttext-generationn<1K1 likes16 downloads1y agoHugging Face29Dxniz /Detailed-Thinker Detailed Thinker (snapshot) This is an early snapshot of Detailed Thinker, started 12 August 2026. The set is still in active development. Treat this release as a work-in-progress cut (around 300 examples), not a finished corpus. What this is for Shallow build requests often get shallow answers. If you say "build a barbershop simulator," a model can sketch something but it usually skips real depth on functionality, constraints, and follow-through. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/Detailed-Thinker.textn<1K0 likes15 downloads2mo agoHugging Face30open-llm-leaderboard /ewre324__Thinker-Llama-3.2-3B-Instruct-Reasoning-detailsgated Dataset Card for Evaluation run of ewre324/Thinker-Llama-3.2-3B-Instruct-Reasoning Dataset automatically created during the evaluation run of model ewre324/Thinker-Llama-3.2-3B-Instruct-Reasoning The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ewre324__Thinker-Llama-3.2-3B-Instruct-Reasoning-details.tabular10K<n<100K0 likes13 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.