Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-SFT-Science-v2 Dataset Description: Nemotron-Science-v2 is a science reasoning dataset with synthetic (synthetic MCQ, RQA) and non-synthetic vendor problems and LLM-generated solutions. It comprises three domains (Physics, Biology, and Chemistry), two question formats (multiple-choice questions [MCQ] and open questions [OpenQ]), and three generation setups: chain-of-thought (CoT) reasoning without tools, Python tool usage, and search tools usage with the Tavily API. The solutions were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Science-v2.texttext-generation1M<n<10M18 likes11k downloads4mo agoHugging Face02nvidia /Nemotron-Science-v1 Dataset Description: Nemotron-Science-v1 is a synthetic science reasoning dataset with two subsets: an MCQA set that improves on the STEM portion of Nemotron-Post-Training-v1 using GPT-OSS-120B to generate GPQA-style questions and reasoning traces, and an RQA set of synthetic chemistry questions. This dataset is ready for commercial use. The Nemotron-Science-v1 dataset contains the following subsets: MCQA This subset is an improvement of the STEM subset in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Science-v1.text100K<n<1M32 likes5.3k downloads10mo agoHugging Face03R2MED /Medical-Sciences 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/Medical-Sciences.texttext-retrieval10K<n<100K0 likes1.6k downloads1y agoHugging Face04ScienceOne-AI /S1-DeepResearch-15k S1-DeepResearch-15k Dataset Overview The S1-DeepResearch dataset is a curated collection of approximately 15k samples designed to improve deep research capabilities of large language models. The dataset includes two types of tasks: Verifiable tasks (labeled as "Closed-ended Multi-hop Resolution") Open-ended tasks (labeled as "Open-ended Exploration") Dataset Composition The dataset is organized into five core capability dimensions: Long-chain complex… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-DeepResearch-15k.text10K<n<100K12 likes1.1k downloads6mo agoHugging Face05Hugodonotexit /math-code-science-deepseek-r1-en R1 Dataset Collection Aggregated high-quality English prompts and model-generated responses from DeepSeek R1 and DeepSeek R1-0528. Dataset Summary The R1 Dataset Collection combines multiple public DeepSeek-generated instruction-response corpora into a single, cleaned, English-only JSONL file. Each example consists of a <|user|> prompt and a <|assistant|> response in one "text" field. This release includes: ~21,000 examples from the DeepSeek-R1-0528 Distilled Custom… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/math-code-science-deepseek-r1-en.textquestion-answering1M<n<10M5 likes1k downloads1y agoHugging Face06simpleG2023 /chinese-materials-science-open-intelligence 🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.tabulartext-retrieval1K<n<10K0 likes999 downloads6h agoHugging Face07nvidia /Nemotron-RL-Science-v1 Dataset Description: Nemotron-RL-Science-v1 is a reinforcement learning (RL) dataset for science reasoning. Each example provides a problem, a reference answer, and a verifiable RL environment configuration (the agent prompt, the agent/verifier reference, and the answer-extraction template) so that a policy model can be trained with verifiable rewards. It covers three domains (Physics, Biology, and Chemistry), the open-question (OpenQ) format, and two generation setups:… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Science-v1.texttext-generation100K<n<1M14 likes772 downloads12d agoHugging Face08LLaMAX /BenchMAX_Science Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Science is a dataset of BenchMAX, sourcing from GPQA, which evaluates the natural science reasoning capability in multilingual scenarios. We extend the original English dataset to 16 non-English languages. The data is first translated by Google… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Science.textquestion-answering1K<n<10K2 likes738 downloads2y agoHugging Face09Nbardy /science-theory-textbookstext10K<n<100K9 likes672 downloads3y agoHugging Face10ScienceOne-AI /S1-Omni-Corpus-10K S1-Omni-Corpus-10K An open-source scientific multimodal reasoning dataset subset for S1-Omni 🧬 Model Introduction S1-Omni is a unified scientific multimodal reasoning model for scientific understanding, prediction, and generation. It is developed by the ScienceOne AI team of the Chinese Academy of Sciences. S1-Omni addresses fragmented scientific AI capabilities with a shared backbone for cross-disciplinary, cross-modal, and cross-task understanding and reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-Omni-Corpus-10K.image10K<n<100K1 likes664 downloads3mo agoHugging Face11ScienceOne-AI /SciGenEdit-10K SciGenEdit-10K An Open Dataset for Scientific Image Generation and Editing English | 简体中文 📖 Introduction SciGenEdit-10K is a public subset released with the S1-Omni-Image project. It is designed for research on scientific image generation, scientific image editing, and multi-turn scientific image generation and editing. S1-Omni-Image is a unified multimodal model developed by the ScienceOne team at the Chinese Academy of Sciences for scientific… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/SciGenEdit-10K.imagetext-to-image10K<n<100K3 likes645 downloads4mo agoHugging Face12deep-principle /science_chemistrytextn<1K2 likes522 downloads18d agoHugging Face13SZLHOLDINGS /szl-science-forum-corpus Science Forum Pilot Explore original summaries and metadata from two operator-authored topics used to formulate review hypotheses. Artifact: Two-topic metadata pilot · Stage: Training unauthorized Explore in Command Lab · Build · Evidence Before you use it This is not a forum scrape, representative sample, model-training dataset or scientific benchmark. Original annotations do not grant rights to linked forum posts; expansion requires separate access, reuse and… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/szl-science-forum-corpus.tabularn<1K0 likes493 downloads6d agoHugging Face14ScienceSoft /piibench piibench — a carrier-aware benchmark for PII detection in outbound traffic The full rationale, metric by metric and with every known limitation stated, is in DESIGN.md. This card is the short version. Leaderboard: ScienceSoft/piibench-leaderboard · ScienceSoft's entrant: ScienceSoft/scnsoft-pii-encoder piibench measures how well a model detects personal data in the text formats that outbound traffic to AI assistants actually carries. Public PII benchmarks score prose, whereas… See the full description on the dataset page: https://huggingface.co/datasets/ScienceSoft/piibench.texttoken-classification1K<n<10K1 likes464 downloads15d agoHugging Face15GAIR /Darwin-SciencegatedThis repository contains the dataset of the research paper Data Darwinism -- Part1: Unlocking the Value of Scientific Data for Pre-training. The dataset is currently being uploaded and is expected to be released within one week. text10M<n<100M6 likes435 downloads8mo agoHugging Face16sunnysetia /Nemotron-SFT-Science-v2-Sharded Nemotron-SFT-Science-v2-Sharded Byte-preserving sequential 100 MB JSONL shards of selected files from nvidia/Nemotron-SFT-Science-v2. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms. Included files: vendor.jsonl, so.jsonl, rqa.jsonl, syn_mcq.jsonl. No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all… See the full description on the dataset page: https://huggingface.co/datasets/sunnysetia/Nemotron-SFT-Science-v2-Sharded.text1M<n<10M0 likes422 downloads1mo agoHugging Face17Royal-lobster /10001-Science-Facts 10,001 Science Facts 10,000+ obscure, surprising, and verifiable science facts The kind that make you go "wait, really?" 🔗 GitHub Repository • 📁 Download by Category 🤔 What is this? A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food. Every fact is: Sourced — from Wikipedia, Wikidata, academic sources Verifiable — no LLM hallucinations Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/Royal-lobster/10001-Science-Facts.texttext-generation10K<n<100K1 likes348 downloads8mo agoHugging Face18Nbardy /wild-science-theory-textbookstext10K<n<100K3 likes342 downloads3y agoHugging Face19oss-codes /Computer-Science-Conversational-Dataset-Indictext10K<n<100K0 likes294 downloads2y agoHugging Face20JingyaoLi /Science-Logits-1.2M Logits-Based Finetuning • 🤗 Data • 🤗 ScienceLLaMA-3B • 🤗 ScienceLLaMA-1B • 🐱 Code • 📃 Paper This is a repo of a large-scale 1.2M logits dataset for Logits-Based Finetuning, which integrates the strengths of supervised learning and knowledge distillation by combining teacher logits with ground truth labels. This preserves both correctness and linguistic diversity. Performance Train Data: huggingface Readme: Installation Guide… See the full description on the dataset page: https://huggingface.co/datasets/JingyaoLi/Science-Logits-1.2M.text1M<n<10M1 likes239 downloads1y agoHugging Face21seonjeongh /science_reasoning science_reasoning Mistral-7B의 과학 지식·추론 능력 향상을 위해 6개 공개 과학 객관식 QA 데이터셋을 통일 포맷으로 변환하고, ARC-Challenge test와의 오염을 제거한 데이터셋입니다. 원본 데이터셋 allenai/sciq allenai/openbookqa (main) allenai/qasc allenai/quartz allenai/ai2_arc (ARC-Easy / ARC-Challenge) nguyen-brat/worldtree 전처리 포맷 통일: 각 데이터셋의 서로 다른 스키마를 unique_id, orig_id, source, question, choices, answer, support 필드로 변환. support는 근거 문단/문장으로, 데이터셋별 원본 필드(support/fact/para/cot)에서 구성하거나 없으면 빈 문자열.… See the full description on the dataset page: https://huggingface.co/datasets/seonjeongh/science_reasoning.textmultiple-choice10K<n<100K0 likes183 downloads3mo agoHugging Face22aazwan /malaysian_journal_of_analytical_sciencetextn<1K0 likes181 downloads3y agoHugging Face23stindardlogic /science-qa-sft-100k Science QA SFT (100K) 100,000 science Q&A examples with step-by-step explanations for SFT fine-tuning. Covers physics, chemistry, biology, astronomy, and earth science at beginner through advanced difficulty. Motivation Models trained on general text often give superficially plausible but mechanistically wrong answers to science questions — stating the right conclusion without understanding the underlying reasoning. This dataset trains models to explain why an… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/science-qa-sft-100k.texttext-generation100K<n<1M0 likes170 downloads3mo agoHugging Face24ianncity /GLM-5.2-Science GLM-5.2 · Science-50000x 50,000x traces distilled from GLM-5.2 on High reasoning Physics · Chemistry · Biology Token Count: 160M Theres prompt overlap with my Kimi K2.5 dataset science subset, which I think those prompts are getting used in alot of places now You can use this dataset for any purpose and you dont need to credit me, preferably dont claim it as your own. hi - ianncity texttext-generation10K<n<100K19 likes168 downloads3mo agoHugging Face25HacksHaven /science-on-a-sphere-prompt-completions Dataset Card for Science On a Sphere QA Dataset Dataset Details Dataset Description This dataset comprises question-and-answer (QA) pairs generated from NOAA's Science On a Sphere (SOS) website, including support documentation and the dataset catalog. Each entry contains a prompt and a corresponding completion, designed to support educational and research use cases in Earth science. This dataset includes a custom dataset_script.py and a consolidated file… See the full description on the dataset page: https://huggingface.co/datasets/HacksHaven/science-on-a-sphere-prompt-completions.textquestion-answering1K<n<10K0 likes165 downloads1y agoHugging Face26cminst /realistic-bpe5-science-math-10btabularn<1K0 likes142 downloads3mo agoHugging Face27RazinAleks /SO-Python_QA-Data_Science_and_Machine_Learning_classtabular1K<n<10K6 likes121 downloads3y agoHugging Face28rl-rag /drtulu_v2_nemotron_science_v1_mcq_0205text1K<n<10K0 likes107 downloads8mo agoHugging Face29zjunlp /DataPRM-ScienceAgentBenchtabular1K<n<10K0 likes82 downloads6mo agoHugging Face30TheJackBright /verisci-verified-science-math-code VeriSci Verified Science Math Code Verifier-grounded dataset for the Adaption AutoScientist Challenge Part 2, targeting the Science category with secondary Math and Code coverage. Summary VeriSci trains models to solve scientific computations, finite-difference PDE updates, numerical ODE steps, unit-checked mechanics, thermodynamics, circuits, chemistry stoichiometry, molarity, unit conversion, vector decomposition, two-point linear modeling, small Python… See the full description on the dataset page: https://huggingface.co/datasets/TheJackBright/verisci-verified-science-math-code.texttext-generation1K<n<10K0 likes65 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.