Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zmodelerlover /amd-nr0 likes49k downloads3d agoHugging Face02optimum-amd /transformers_pr_ci0 likes9.8k downloads1d agoHugging Face03optimum-amd /transformers_daily_ci1 likes5.3k downloads10h agoHugging Face04a-m-team /AM-DeepSeek-Distilled-40MFor more open-source datasets, models, and methodologies, please visit our GitHub repository and paper: DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training. Due to certain constraints, we are only able to open-source a subset of the complete dataset. Model Training Performance based on our complete dataset On AIME 2024, our 72B model achieved a score of 79.2 using only supervised fine-tuning (SFT). The 32B model reached 75.8 and… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M.tabulartext-generation10M<n<100M56 likes3.1k downloads1y agoHugging Face05a-m-team /AM-DeepSeek-R1-Distilled-1.4MFor more open-source datasets, models, and methodologies, please visit our GitHub repository. AM-DeepSeek-R1-Distilled-1.4M is a large-scale general reasoning task dataset composed of high-quality and challenging reasoning problems. These problems are collected from numerous open-source datasets, semantically deduplicated, and cleaned to eliminate test set contamination. All responses in the dataset are distilled from the reasoning model (mostly DeepSeek-R1) and have undergone rigorous… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-Distilled-1.4M.text-generation1M<n<10M185 likes2.3k downloads2y agoHugging Face06amd /Instella-Long Instella-Long The Instella-Long dataset is a collection of pre-training and instruction following data that is used to train Instella-3B-Long-Instruct. The pre-training data is sourced from Prolong. For the SFT data, we use public datasets: Ultrachat 200K, OpenMathinstruct-2, Tülu-3 Instruction Following, and MMLU auxiliary train set. In addition, we generate synthetic long instruction data using documents of the books and arxiv from our pre-training corpus and the dclm subset from… See the full description on the dataset page: https://huggingface.co/datasets/amd/Instella-Long.0 likes2.1k downloads11mo agoHugging Face07amd /ReasonLite-Dataset GitHub | Dataset | Blog ReasonLite is an ultra-lightweight math reasoning model. With only 0.6B parameters, it leverages high-quality data distillation to achieve performance comparable to models over 10× its size, such as Qwen3-8B, reaching 75.2 on AIME24 and extending the scaling law of small models. 🔥 Best-performing 0.6B math reasoning model 🔓 Fully open-source — weights, scripts, datasets, synthesis pipeline⚙️ Distilled in two stages to balance efficiency and high… See the full description on the dataset page: https://huggingface.co/datasets/amd/ReasonLite-Dataset.text1M<n<10M17 likes625 downloads9mo agoHugging Face08amd /Micro-World-MC-DatasetMicro-World is an interactive world model developed by AMD. We leverage the Minecraft API MineDojo to collect game data, as it enables us to obtain diverse biomes as well as varying weather and lighting conditions—properties that are desirable for improving generalization to real-world scenarios. For the action space, we record keyboard controls including W (forward), A (left), S (backward), D (right), Ctrl (sprint), Shift (sneak), Space (jump), along with mouse movement. By randomly sampling… See the full description on the dataset page: https://huggingface.co/datasets/amd/Micro-World-MC-Dataset.tabular1K<n<10K3 likes588 downloads8mo agoHugging Face09a-m-team /AM-DeepSeek-R1-0528-Distilled 📘 Dataset Summary This dataset is a high-quality reasoning corpus distilled from DeepSeek-R1-0528, an improved version of the DeepSeek-R1 large language model. Compared to its initial release, DeepSeek-R1-0528 demonstrates significant advances in reasoning, instruction following, and multi-turn dialogue. Motivated by these improvements, we collected and distilled a diverse set of 2.6 million queries across multiple domains, using DeepSeek-R1-0528 as the teacher. A notable… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-0528-Distilled.text-generation1M<n<10M102 likes580 downloads1y agoHugging Face10koajoel /AM-DeepSeek-R1-Distilled-1.4Mtext1M<n<10M0 likes441 downloads2y agoHugging Face11amd /AIG-Datasets AMD AIG GPU Kernel Datasets AMD AIG-Datasets is a collection of GPU-kernel generation, translation, optimization, and ROCm-library supervision data. It contains PyTorch/CUDA-to-HIP, HIP-to-HIP, PyTorch-to-Triton, and production-grounded rocBLAS/rocSOLVER entries, together with metadata, samples, conversion utilities, and reproducible evaluation tools. The repository is organized into versioned releases. New training and evaluation workflows should use the unified-schema datasets… See the full description on the dataset page: https://huggingface.co/datasets/amd/AIG-Datasets.text-generation4 likes429 downloads2mo agoHugging Face12amd /Cot-Drop LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/Cot-Drop.texttext-generation10K<n<100K1 likes413 downloads7mo agoHugging Face13koajoel /AM-DeepSeek-R1-Distilled-1.4M-Englishtext1M<n<10M0 likes391 downloads2y agoHugging Face14amd1234567 /AudioJailbreak Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models AudioJailbreak is a benchmark framework specifically designed for evaluating the security of Audio Language Models (Audio LLMs). This project tests model defenses against malicious requests through various audio perturbation techniques.Note: This project aims to improve the security of audio language models. Researchers should use this tool responsibly. 📋 Table of… See the full description on the dataset page: https://huggingface.co/datasets/amd1234567/AudioJailbreak.0 likes356 downloads3mo agoHugging Face15MaziyarPanahi /AM-DeepSeek-R1-0528-Distilled-with-Systemtext1M<n<10M4 likes326 downloads1y agoHugging Face16amd-nicknick /bert-base-uncased-2022_tokenized_dataset10M<n<100M0 likes252 downloads3y agoHugging Face17syzym /xbmu_amdo31 Dataset Card for [XBMU-AMDO31] Dataset Summary XBMU-AMDO31 dataset is a speech recognition corpus of Amdo Tibetan dialect. The open source corpus contains 31 hours of speech data and resources related to build speech recognition systems, including transcribed texts and a Tibetan pronunciation dictionary. Supported Tasks and Leaderboards automatic-speech-recognition: The dataset can be used to train a model for Amdo Tibetan Automatic Speech Recognition (ASR). It… See the full description on the dataset page: https://huggingface.co/datasets/syzym/xbmu_amdo31.textautomatic-speech-recognition10K<n<100K5 likes238 downloads4y agoHugging Face18amd /SAND-Post-Training-Dataset SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs Dataset Summary We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs. This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.tabularquestion-answering10K<n<100K3 likes233 downloads10mo agoHugging Face19amd /SAND-MATH SAND-Math: A Synthetic Dataset of Difficult Problems to Elevate LLM Math Performance 📃 Paper | 🤗 Dataset SAND-Math (Synthetic Augmented Novel and Difficult Mathematics) is a high-quality, high-difficulty dataset of mathematics problems and solutions. It is generated using a novel pipeline that addresses the critical bottleneck of scarce, high-difficulty training data for mathematical Large Language Models (LLMs). Key Features Novel Problem Generation: Problems are… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-MATH.tabularquestion-answering10K<n<100K3 likes217 downloads1y agoHugging Face20alexei-v-ivanov-amd /flores_plustext100K<n<1M0 likes211 downloads2y agoHugging Face21giacomoran /hackathon_amd_mission2_black_sort_fixedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 145, "total_frames": 44304, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:145" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/giacomoran/hackathon_amd_mission2_black_sort_fixed.tabularrobotics10K<n<100K0 likes197 downloads10mo agoHugging Face22JinnP /amdpilot-lora-sft-dataset AMDPilot LoRA SFT Dataset SFT training data for fine-tuning LLMs on AMD GPU debugging, optimization, and kernel engineering tasks. Each example is a multi-turn conversation in OpenAI messages format with tool-use annotations. Usage from datasets import load_dataset # Load a specific version ds = load_dataset("JinnP/amdpilot-lora-sft-dataset", "v5_2") # Load a specific view ds = load_dataset("JinnP/amdpilot-lora-sft-dataset", "v5_2_chunks") # Available configs: v4, v5… See the full description on the dataset page: https://huggingface.co/datasets/JinnP/amdpilot-lora-sft-dataset.1K<n<10K0 likes187 downloads6mo agoHugging Face23amd /InstructGpt-educational LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-educational.texttext-generation100K<n<1M3 likes166 downloads7mo agoHugging Face24giacomoran /hackathon_amd_mission2_black_flipThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 150, "total_frames": 49863, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:150" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/giacomoran/hackathon_amd_mission2_black_flip.tabularrobotics10K<n<100K0 likes162 downloads10mo agoHugging Face25mlfoundations-dev /AM-DeepSeek-R1-Distilled-1.4M-am_0.5Mtext100K<n<1M2 likes136 downloads2y agoHugging Face26chhao /AM-DeepSeek-R1-Filtered-Math-Code AM DeepSeek R1 Filtered Math and Code This repository publishes reproducible training subsets derived from a-m-team/AM-DeepSeek-R1-Distilled-1.4M at pinned revision 53531c06634904118a2dcd83961918c4d69d1cdf. The source dataset and this derived release use CC BY-NC 4.0. Commercial use is not permitted by that license. Preserve attribution and review the upstream dataset card before use. Contents Config / split Records Bytes SHA-256 math / train 111,657 2… See the full description on the dataset page: https://huggingface.co/datasets/chhao/AM-DeepSeek-R1-Filtered-Math-Code.texttext-generation100K<n<1M0 likes132 downloads3mo agoHugging Face27giacomoran /hackathon_amd_mission2_blue_pickThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 151, "total_frames": 36913, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:151" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/giacomoran/hackathon_amd_mission2_blue_pick.tabularrobotics10K<n<100K0 likes130 downloads10mo agoHugging Face28amd /UltraChat200K-regenerated LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/UltraChat200K-regenerated.texttext-generation100K<n<1M2 likes124 downloads7mo agoHugging Face29amd /Instella-GSM8K-synthetic Instella-GSM8K-synthetic The Instella-GSM8K-synthetic dataset was used in the second stage pre-training of Instella-3B model, which was trained on top of the Instella-3B-Stage1 model. This synthetic dataset was generated using the training set of GSM8k dataset, where we first used Qwen2.5-72B-Instruct to Abstract numerical values as function parameters and generate a Python program to solve the math question. Identify and replace numerical values in the existing question with… See the full description on the dataset page: https://huggingface.co/datasets/amd/Instella-GSM8K-synthetic.textquestion-answering1M<n<10M7 likes122 downloads11mo agoHugging Face30AbijahKaj /telephony-amd-dataset Telephony AMD (Answering Machine Detection) Dataset Overview A multilingual 4-class telephony audio classification dataset for training streaming Answering Machine Detection models. Contains real human speech (PolyAI/MINDS14) mixed with TTS-generated audio (Microsoft Neural TTS / edge-tts) across English, French, Spanish, and German. Key design principle: Voicemail greetings are recorded by real humans and sound acoustically identical to live speech. This dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/AbijahKaj/telephony-amd-dataset.audio1K<n<10K2 likes121 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.