Team Ai
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01joyfine /Qwen3-235B-A22B-Thinking-2507_Qwen3-1.7B_AIME_1983_2024textn<1K0 likes1.3k downloads11mo agoHugging Face02madteam /thinking-bciciv2a BCICIV-2a EEG Motor Imagery Provenance This repository contains one consolidated CSV shard derived from the local research copy identified as the Kaggle dataset aymanmostafa11/eeg-motor-imagery-bciciv-2a, which describes the BCI Competition IV 2a motor-imagery recordings associated with the Graz BCI laboratory. The uploader has represented that they have the right to redistribute this research copy. The repository owner must preserve any upstream attribution and… See the full description on the dataset page: https://huggingface.co/datasets/madteam/thinking-bciciv2a.tabular100K<n<1M0 likes127 downloads17d agoHugging Face03ThinkingHub /PP SteelBench: A Diagnostic Benchmark for Vision-Language Models in Industrial Safety Monitoring SteelBench is a diagnostic benchmark of densely annotated CCTV clips from an operating integrated steel plant. It is designed to evaluate vision-language models (VLMs) on real-world industrial action recognition, PPE assessment, and safety-violation detection — under naturally occurring degradation (dust, glare, steam, low light), at distances and crowdedness levels that curated… See the full description on the dataset page: https://huggingface.co/datasets/ThinkingHub/PP.imagevideo-classification1K<n<10K0 likes116 downloads4mo agoHugging Face04chimbiwide /Creative-Writing-Thinking Creative-Writing-Thinking Using essays-creative-writing-prompts and Qwen3-14b to generate the reasoning traces and answers. We created this reasoning dataset. Suitable for LLM post-training, especially RL. texttext-generation1K<n<10K0 likes45 downloads10mo agoHugging Face05chimbiwide /databricks-thinking databricks-thinking Created by extracing the questions from the [databricks-dolly] dataset and using Qwen3-14b to synthetically generate reasoning traces and answers. The whole process took 1 day 23 hours 18 minuntes and 8 seconds Why we created this dataset The vast majority of publicly available datasets comes from large models such as DeepSeek R1. The issue with using these large models are obvious: the reasoning traces are extremely long, often longer than the actual… See the full description on the dataset page: https://huggingface.co/datasets/chimbiwide/databricks-thinking.textquestion-answering10K<n<100K0 likes38 downloads10mo agoHugging Face06chimbiwide /sciqa-thinking sciqa-thinking Randomly extracted 3000 rows from sciq and prompting Qwen3-14b to generate the intermediate reasoning traces, we created this dataset. This should be used for LLM post-training, especially RL. textquestion-answering1K<n<10K0 likes28 downloads10mo agoHugging Face07chimbiwide /gsm8k-thinking gsk8k-thinking A processed version of gsm8k for LLM post-training purposes, especially RL. textquestion-answering1K<n<10K1 likes25 downloads10mo agoHugging Face08chimbiwide /brainstorming-thinking brainstorming-thinking Using Qwen3-14b to synthetically generate the reasoning traces and answers for Explore_Instruct_Brainstorming_10k Suitable for LLM post-training, especially RL. texttext-generation1K<n<10K0 likes21 downloads10mo agoHugging Face09chimbiwide /code-thinking code-thinking Using mbpp and using Qwen3-14b to generate the reasoning traces for the datasets. Suitable for training small LLMs for python code generation. texttext-generationn<1K0 likes20 downloads10mo agoHugging Face10CoT-Translator /thinking-my-cottext10K<n<100K0 likes16 downloads2y agoHugging Face11CoT-Translator /thinking_my_translation_Reduced_p1gatedtext10K<n<100K0 likes2 downloads2y agoHugging Face12CoT-Translator /thinking_my_translation_p2gatedtext10K<n<100K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.