Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01harborframework /terminal-bench-science Terminal-Bench-Science The primary source is hosted on GitHub, please open issues and pull requests there, not here. Terminal-Bench-Science is a benchmark of real-world computational research workflows across the life, physical, earth, mathematical, and engineering sciences. Like Terminal-Bench, it's a continuous benchmark: releases are published as tags on the source repo. This dataset mirrors that on the Hub: one repo, tagged per version, instead of a separate repo per… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science.textn<1K6 likes61k downloads26d agoHugging Face02derek-thomas /ScienceQA Dataset Card Creation Guide Dataset Summary Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering Supported Tasks and Leaderboards Multi-modal Multiple Choice Languages English Dataset Structure Data Instances Explore more samples here. {'image': Image, 'question': 'Which of these states is farthest north?', 'choices': ['West Virginia', 'Louisiana', 'Arizona', 'Oklahoma'], 'answer': 0… See the full description on the dataset page: https://huggingface.co/datasets/derek-thomas/ScienceQA.imagemultiple-choice10K<n<100K234 likes31k downloads4y agoHugging Face03lmms-lab-encoder /ScienceQA Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of derek-thomas/ScienceQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{lu2022learn, title={Learn to Explain: Multimodal Reasoning via Thought… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/ScienceQA.image10K<n<100K10 likes17k downloads3y agoHugging Face04nvidia /Nemotron-SFT-Science-v2 Dataset Description: Nemotron-Science-v2 is a science reasoning dataset with synthetic (synthetic MCQ, RQA) and non-synthetic vendor problems and LLM-generated solutions. It comprises three domains (Physics, Biology, and Chemistry), two question formats (multiple-choice questions [MCQ] and open questions [OpenQ]), and three generation setups: chain-of-thought (CoT) reasoning without tools, Python tool usage, and search tools usage with the Tavily API. The solutions were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Science-v2.texttext-generation1M<n<10M18 likes11k downloads4mo agoHugging Face05nvidia /Nemotron-Science-v1 Dataset Description: Nemotron-Science-v1 is a synthetic science reasoning dataset with two subsets: an MCQA set that improves on the STEM portion of Nemotron-Post-Training-v1 using GPT-OSS-120B to generate GPQA-style questions and reasoning traces, and an RQA set of synthetic chemistry questions. This dataset is ready for commercial use. The Nemotron-Science-v1 dataset contains the following subsets: MCQA This subset is an improvement of the STEM subset in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Science-v1.text100K<n<1M32 likes5.3k downloads10mo agoHugging Face06osunlp /ScienceAgentBench ScienceAgentBench Update 04/30/2026: To mitigate false negatives in evaluation, we have released a verified version of ScienceAgentBench. Please load our benchmark using the following code going forward and make sure you follow the latest instructions in our github repository: from datasets import load_dataset ds = load_dataset("osunlp/ScienceAgentBench", split="verified") The advancements of language language models (LLMs) have piqued growing interest in developing… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/ScienceAgentBench.textn<1K21 likes4.6k downloads5mo agoHugging Face07SAIS-Life-Science /Aneumo Aneumo Datasets AneumoDataset is a comprehensive multi-modal dataset containing 3D anatomical structures and simulated hemodynamic parameters for cerebral aneurysms, supporting both computational modeling and AI-based analysis. textn<1K7 likes3.1k downloads7mo agoHugging Face08product-science /xlam-function-calling-60k-raw XLAM Function Calling 60k Raw Dataset This dataset includes train and test splits derived from Salesforce/xlam-function-calling-60k. Train split size: 95% of the original dataset Test split size: 5% of the original dataset textquestion-answering10K<n<100K3 likes3k downloads2y agoHugging Face09hugging-science /mmu_manga mmu_manga HATS Catalog Collection This is the collection of HATS catalogs representing mmu_manga. This dataset is part of the Multimodal Universe, a large-scale collection of multimodal astronomical data. For full details, see the paper: The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data. Access the catalog We recommend the use of the LSDB Python framework to access HATS catalogs. LSDB can be installed via… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_manga.tabular10K<n<100K0 likes2.7k downloads4mo agoHugging Face10ScienceOne-AI /S1-MMAlignS1-MMAlign A Large-Scale Multi-Disciplinary Scientific Multimodal Dataset S1-MMAlign is a large-scale, multi-disciplinary multimodal dataset comprising over 15.5 million high-quality image-text pairs derived from 2.5 million open-access scientific papers. Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap between complex scientific imagery and sparse textual descriptions. S1-MMAlign aims to… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-MMAlign.imageimage-to-text10M<n<100M106 likes2.5k downloads7mo agoHugging Face11OpenMOSS-Team /SWE-bench-Science SWE-bench Science SWE-bench Science evaluates coding agents on software-engineering tasks drawn from scientific-computing repositories. The release contains 119 tasks across 20 scientific domains, with isolated environments and separate programmatic verifiers. GitHub release repository: OpenMOSS/SWE-bench-Science Runtime images: Docker Hub, pinned by immutable linux/amd64 digests Evaluation framework: Pier, compatible with Harbor task format Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SWE-bench-Science.textn<1K9 likes2.1k downloads9d agoHugging Face12vidore /vidore_v3_computer_scienceViDoRe V3 : Computer Science This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark. About ViDoRe v3 ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science.documentvisual-document-retrieval1K<n<10K6 likes2k downloads9mo agoHugging Face13R2MED /Medical-Sciences 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/Medical-Sciences.texttext-retrieval10K<n<100K0 likes1.6k downloads1y agoHugging Face14tasksource /ScienceQA_text_only Dataset Card for "scienceQA_text_only" ScienceQA text-only examples (examples where no image was initially present, which means they should be doable with text-only models.) @article{10.1007/s00799-022-00329-y, author = {Saikh, Tanik and Ghosal, Tirthankar and Mittal, Amish and Ekbal, Asif and Bhattacharyya, Pushpak}, title = {ScienceQA: A Novel Resource for Question Answering on Scholarly Articles}, year = {2022}, journal = {Int. J. Digit. Libr.}, month = {sep} } text10K<n<100K32 likes1.4k downloads3y agoHugging Face15reasoning-proj /judged_science_completionstabularn<1K2 likes1.4k downloads1y agoHugging Face16TalentZHOU /hle_material_science HLE Material Science: A Specialized Benchmark for Materials Science A Materials Science Subset of Humanity's Last Exam (HLE) Overview HLE Material Science is a carefully curated materials science subset derived from the Humanity's Last Exam (HLE) dataset, containing 106 high-quality expert-level questions covering 25+ materials science subfields, with 97% of questions rated as high confidence. This dataset is designed to evaluate large language models'… See the full description on the dataset page: https://huggingface.co/datasets/TalentZHOU/hle_material_science.textquestion-answeringn<1K1 likes1.3k downloads9mo agoHugging Face17Jialuo21 /Science-T2I-Fullset Science-T2I Fullset Resources Website arXiv: Paper GitHub: Code Huggingface: SciScore Huggingface: Science-T2I-S&C Benchmark Data The Science-T2I Fullset comprises a comprehensive collection of data for scientific T2I generation, including both training and test sets with a unified data structure. The test sets are split into 'test-S' and 'test-C,' corresponding to the Science-T2I-S and Science-T2I-C benchmarks, respectively. Download Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Jialuo21/Science-T2I-Fullset.image1K<n<10K0 likes1.3k downloads1y agoHugging Face18lmms-lab /ScienceQA-IMG Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted and filtered version of derek-thomas/ScienceQA with only image instances. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{lu2022learn, title={Learn to Explain:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/ScienceQA-IMG.image10K<n<100K5 likes1.3k downloads3y agoHugging Face19vidore /vidore_v3_computer_science_mteb_format Vidore3ComputerScienceRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_computer_science How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes1.2k downloads11mo agoHugging Face20YLR9933 /terminal-bench-science-trailtext10K<n<100K0 likes1.2k downloads8d agoHugging Face21PleIAs /French-Science-Commons French Science Commons French Science Commons (Commun numérique des sciences en français) rassemble des publications scientifiques d'origine française en accès ouvert, couvrant une période de vingt ans, de 2007 à 2026. Il comprend 1 248 860 documents scientifiques — 1 189 628 articles et 59 232 thèses — indexés à travers de multiples dépôts académiques en accès public, tels que HAL, OpenAlex, des revues scientifiques, des dépôts institutionnels, et d'autres. Le corpus est conçu… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-Science-Commons.tabular10M<n<100M26 likes1.1k downloads4mo agoHugging Face22Corning /weact-native-science-checkpoint Native WeAct science checkpoint The fixed eight-question, seven-arm pilot has reached its human-review checkpoint. All 56 planned task records are in pilot_v4/. The frozen 3,000-question test has not been run. No accuracy score or retraining conclusion is claimed. The independent Serper, Jina and E2B checks passed, and each backend passed native webpage extraction with the live auxiliary model. The corrected runtime uses original questions, native hard/soft routing, the… See the full description on the dataset page: https://huggingface.co/datasets/Corning/weact-native-science-checkpoint.text1K<n<10K0 likes1.1k downloads4d agoHugging Face23ScienceOne-AI /S1-DeepResearch-15k S1-DeepResearch-15k Dataset Overview The S1-DeepResearch dataset is a curated collection of approximately 15k samples designed to improve deep research capabilities of large language models. The dataset includes two types of tasks: Verifiable tasks (labeled as "Closed-ended Multi-hop Resolution") Open-ended tasks (labeled as "Open-ended Exploration") Dataset Composition The dataset is organized into five core capability dimensions: Long-chain complex… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-DeepResearch-15k.text10K<n<100K12 likes1.1k downloads6mo agoHugging Face24Hugodonotexit /math-code-science-deepseek-r1-en R1 Dataset Collection Aggregated high-quality English prompts and model-generated responses from DeepSeek R1 and DeepSeek R1-0528. Dataset Summary The R1 Dataset Collection combines multiple public DeepSeek-generated instruction-response corpora into a single, cleaned, English-only JSONL file. Each example consists of a <|user|> prompt and a <|assistant|> response in one "text" field. This release includes: ~21,000 examples from the DeepSeek-R1-0528 Distilled Custom… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/math-code-science-deepseek-r1-en.textquestion-answering1M<n<10M5 likes1k downloads1y agoHugging Face25simpleG2023 /chinese-materials-science-open-intelligence 🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.tabulartext-retrieval1K<n<10K0 likes999 downloads8h agoHugging Face26reasoning-proj /severity_ablation_sciencetabular100K<n<1M0 likes955 downloads1y agoHugging Face27JH-C-k /arc-science-cot-10k-shorttext10K<n<100K0 likes916 downloads6mo agoHugging Face28trungtvu /open-thoughts-sciencetext1K<n<10K0 likes783 downloads2y agoHugging Face29nvidia /Nemotron-RL-Science-v1 Dataset Description: Nemotron-RL-Science-v1 is a reinforcement learning (RL) dataset for science reasoning. Each example provides a problem, a reference answer, and a verifiable RL environment configuration (the agent prompt, the agent/verifier reference, and the answer-extraction template) so that a policy model can be trained with verifiable rewards. It covers three domains (Physics, Biology, and Chemistry), the open-question (OpenQ) format, and two generation setups:… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Science-v1.texttext-generation100K<n<1M14 likes772 downloads11d agoHugging Face30dvilasuero /natural-science-reasoning Natural Sciences Reasoning: the "smolest" reasoning dataset A smol-scale open dataset for reasoning tasks using Hugging Face Inference Endpoints. While intentionally limited in scale, this resource prioritizes: Reproducible pipeline for reasoning tasks using a variety of models (Deepseek V3, Deepsek-R1, Llama70B-Instruct, etc.) Knowledge sharing for domains other than Math and Code reasoning In this repo, you can find: The prompts and the pipeline (see the config file). The… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/natural-science-reasoning.texttext-generationn<1K40 likes751 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.