Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01harborframework /terminal-bench-science Terminal-Bench-Science The primary source is hosted on GitHub, please open issues and pull requests there, not here. Terminal-Bench-Science is a benchmark of real-world computational research workflows across the life, physical, earth, mathematical, and engineering sciences. Like Terminal-Bench, it's a continuous benchmark: releases are published as tags on the source repo. This dataset mirrors that on the Hub: one repo, tagged per version, instead of a separate repo per… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science.textn<1K6 likes75k downloads24d agoHugging Face02harborframework /terminal-bench-science-lfs Terminal-Bench-Science — task input mirror Large input files for Terminal-Bench-Science tasks, which cannot be committed to git. Tasks pull from here at container build time, pinned to a commit SHA and verified against a checksum file that ships in the task directory. One top-level prefix per task; everything lives under <task-name>/input/. Benchmark contamination canary This dataset is benchmark material. If you are assembling a training corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science-lfs.0 likes73k downloads2mo agoHugging Face03derek-thomas /ScienceQA Dataset Card Creation Guide Dataset Summary Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering Supported Tasks and Leaderboards Multi-modal Multiple Choice Languages English Dataset Structure Data Instances Explore more samples here. {'image': Image, 'question': 'Which of these states is farthest north?', 'choices': ['West Virginia', 'Louisiana', 'Arizona', 'Oklahoma'], 'answer': 0… See the full description on the dataset page: https://huggingface.co/datasets/derek-thomas/ScienceQA.imagemultiple-choice10K<n<100K234 likes32k downloads4y agoHugging Face04lmms-lab-encoder /ScienceQA Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of derek-thomas/ScienceQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{lu2022learn, title={Learn to Explain: Multimodal Reasoning via Thought… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/ScienceQA.image10K<n<100K10 likes17k downloads3y agoHugging Face05nvidia /Nemotron-SFT-Science-v2 Dataset Description: Nemotron-Science-v2 is a science reasoning dataset with synthetic (synthetic MCQ, RQA) and non-synthetic vendor problems and LLM-generated solutions. It comprises three domains (Physics, Biology, and Chemistry), two question formats (multiple-choice questions [MCQ] and open questions [OpenQ]), and three generation setups: chain-of-thought (CoT) reasoning without tools, Python tool usage, and search tools usage with the Tavily API. The solutions were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Science-v2.texttext-generation1M<n<10M18 likes11k downloads4mo agoHugging Face06hugging-science /arc-aphasia-bids Aphasia Recovery Cohort (ARC) Multimodal neuroimaging dataset for stroke-induced aphasia research. Dataset Summary The Aphasia Recovery Cohort (ARC) is a large-scale, longitudinal neuroimaging dataset containing multimodal MRI scans from 230 chronic stroke patients with aphasia. This HuggingFace-hosted version provides direct Python access to the BIDS-formatted data with embedded NIfTI files. Metric Count Subjects 230 Sessions 902 T1-weighted scans 444… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/arc-aphasia-bids.image-segmentationn<1K3 likes8.7k downloads10mo agoHugging Face07mariiakoroliuk /generalization-science-data0 likes5.4k downloads2m agoHugging Face08nvidia /Nemotron-Science-v1 Dataset Description: Nemotron-Science-v1 is a synthetic science reasoning dataset with two subsets: an MCQA set that improves on the STEM portion of Nemotron-Post-Training-v1 using GPT-OSS-120B to generate GPQA-style questions and reasoning traces, and an RQA set of synthetic chemistry questions. This dataset is ready for commercial use. The Nemotron-Science-v1 dataset contains the following subsets: MCQA This subset is an improvement of the STEM subset in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Science-v1.text100K<n<1M32 likes5.2k downloads10mo agoHugging Face09hugging-science /isles24-stroke ISLES'24 Stroke Training Dataset Multi-center longitudinal multimodal acute ischemic stroke training dataset from the ISLES'24 Challenge. Overview 149 acute ischemic stroke training cases with: Admission imaging (ses-01): Non-contrast CT, CT angiography, 4D CT perfusion Follow-up imaging (ses-02): Post-treatment MRI (DWI, ADC) Clinical data: Demographics, patient history, admission NIHSS, 3-month mRS outcomes Annotations: Infarct masks, large vessel occlusion masks… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/isles24-stroke.image-segmentationn<1K2 likes4.5k downloads10mo agoHugging Face10osunlp /ScienceAgentBench ScienceAgentBench Update 04/30/2026: To mitigate false negatives in evaluation, we have released a verified version of ScienceAgentBench. Please load our benchmark using the following code going forward and make sure you follow the latest instructions in our github repository: from datasets import load_dataset ds = load_dataset("osunlp/ScienceAgentBench", split="verified") The advancements of language language models (LLMs) have piqued growing interest in developing… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/ScienceAgentBench.textn<1K21 likes4.4k downloads5mo agoHugging Face11CS26 /Computer-Science מאגר הנתונים CS26 HIT — מדעי המחשב מאגר זה משמש לאחסון מרכזי של נתוני לימוד ומשאבים אקדמיים עבור סטודנטים למדעי המחשב במכון הטכנולוגי חולון. המידע המצוי כאן מונגש בצורה נוחה באמצעות פורטל גישה נפרד המאפשר ניווט ויזואלי וחיפוש יעיל בתוך התיקיות השונות. קישורים וגישה ניתן להשתמש בפורטל בכתובת https://cs26-cs26-portal.hf.space/ תנאי שימוש וזכויות יוצרים כל חומרי הלימוד והתכנים המופיעים במאגר זה פתוחים וחופשיים לשימוש לצורכי למידה בלבד. ניתן לקחת את… See the full description on the dataset page: https://huggingface.co/datasets/CS26/Computer-Science.1 likes3.6k downloads7d agoHugging Face12product-science /xlam-function-calling-60k-raw XLAM Function Calling 60k Raw Dataset This dataset includes train and test splits derived from Salesforce/xlam-function-calling-60k. Train split size: 95% of the original dataset Test split size: 5% of the original dataset textquestion-answering10K<n<100K3 likes3.5k downloads2y agoHugging Face13lasrprobegen /science-activations0 likes3.3k downloads1y agoHugging Face14SAIS-Life-Science /Aneumo Aneumo Datasets AneumoDataset is a comprehensive multi-modal dataset containing 3D anatomical structures and simulated hemodynamic parameters for cerebral aneurysms, supporting both computational modeling and AI-based analysis. textn<1K7 likes3.3k downloads7mo agoHugging Face15hugging-science /mmu_manga mmu_manga HATS Catalog Collection This is the collection of HATS catalogs representing mmu_manga. This dataset is part of the Multimodal Universe, a large-scale collection of multimodal astronomical data. For full details, see the paper: The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data. Access the catalog We recommend the use of the LSDB Python framework to access HATS catalogs. LSDB can be installed via… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_manga.tabular10K<n<100K0 likes2.7k downloads4mo agoHugging Face16hugging-science /mmu_legacysurvey_dr10_south_21 mmu_legacysurvey_dr10_south_21 HATS Catalog Collection This is the collection of HATS catalogs representing mmu_legacysurvey_dr10_south_21. This dataset is part of the Multimodal Universe, a large-scale collection of multimodal astronomical data. For full details, see the paper: The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data. Access the catalog We recommend the use of the LSDB Python framework to… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_legacysurvey_dr10_south_21.100M<n<1B4 likes2.7k downloads4mo agoHugging Face17ScienceOne-AI /S1-MMAlignS1-MMAlign A Large-Scale Multi-Disciplinary Scientific Multimodal Dataset S1-MMAlign is a large-scale, multi-disciplinary multimodal dataset comprising over 15.5 million high-quality image-text pairs derived from 2.5 million open-access scientific papers. Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap between complex scientific imagery and sparse textual descriptions. S1-MMAlign aims to… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-MMAlign.imageimage-to-text10M<n<100M106 likes2.6k downloads7mo agoHugging Face18lucazhou2000 /sciencemysterybench-transcriptsimagen<1K0 likes2.3k downloads24d agoHugging Face19open-athena /marin-science-expert-sft-results-2026-10 Marin expert SFT experiment records This public archive preserves configurations, W&B history exports, evaluation records, output audits, and analysis for three expert SFT experiments. The Step38 science RLVR SFT physical-mix repo contains the complete token shards and manifests; its inventory check matched all 67,877 staged files by path and size. The synthetic science SFT data is private pending source and derivative license review; the private card documents all 17… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/marin-science-expert-sft-results-2026-10.0 likes2.2k downloads4d agoHugging Face20OpenMOSS-Team /SWE-bench-Science SWE-bench Science SWE-bench Science evaluates coding agents on software-engineering tasks drawn from scientific-computing repositories. The release contains 119 tasks across 20 scientific domains, with isolated environments and separate programmatic verifiers. GitHub release repository: OpenMOSS/SWE-bench-Science Runtime images: Docker Hub, pinned by immutable linux/amd64 digests Evaluation framework: Pier, compatible with Harbor task format Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SWE-bench-Science.textn<1K9 likes2.2k downloads7d agoHugging Face21vidore /vidore_v3_computer_scienceViDoRe V3 : Computer Science This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark. About ViDoRe v3 ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science.documentvisual-document-retrieval1K<n<10K6 likes2k downloads9mo agoHugging Face22J0nasW /science-datalake Science Data Lake (release 2026.10) A DOI-linked, source-preserving integration of open scholarly metadata. Six open sources are linked through normalised DOIs while every source keeps its own schema and values, so that measurements from different sources (three citation counts, two retraction signals) can be compared directly. Release: 2026.10 (tag v2026.10), DOI 10.57967/hf/XXXX Licence: CC BY-NC-SA 4.0 for the whole dataset (see LICENSE; component sources and their licences… See the full description on the dataset page: https://huggingface.co/datasets/J0nasW/science-datalake.100M<n<1B10 likes1.8k downloads4d agoHugging Face23R2MED /Medical-Sciences 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/Medical-Sciences.texttext-retrieval10K<n<100K0 likes1.7k downloads1y agoHugging Face24tasksource /ScienceQA_text_only Dataset Card for "scienceQA_text_only" ScienceQA text-only examples (examples where no image was initially present, which means they should be doable with text-only models.) @article{10.1007/s00799-022-00329-y, author = {Saikh, Tanik and Ghosal, Tirthankar and Mittal, Amish and Ekbal, Asif and Bhattacharyya, Pushpak}, title = {ScienceQA: A Novel Resource for Question Answering on Scholarly Articles}, year = {2022}, journal = {Int. J. Digit. Libr.}, month = {sep} } text10K<n<100K32 likes1.5k downloads3y agoHugging Face25pm-science /cv4cdd_4d Content This repository stores the contents of the data/ directory from the following GitLab repository:https://gitlab.uni-mannheim.de/processanalytics/cv4cdd The data is organized as follows: input_cdlgContains training, validation, and test datasets used to train the computer vision models. input_cdriftContains external datasets used to evaluate the trained models. model_training_loggingContains model checkpoints for all training runs. This includes both relevant checkpoints… See the full description on the dataset page: https://huggingface.co/datasets/pm-science/cv4cdd_4d.0 likes1.5k downloads8mo agoHugging Face26lmms-lab /ScienceQA-IMG Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted and filtered version of derek-thomas/ScienceQA with only image instances. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{lu2022learn, title={Learn to Explain:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/ScienceQA-IMG.image10K<n<100K5 likes1.4k downloads3y agoHugging Face27reasoning-proj /severity_ablation_sciencetabular100K<n<1M0 likes1.4k downloads1y agoHugging Face28Jialuo21 /Science-T2I-Fullset Science-T2I Fullset Resources Website arXiv: Paper GitHub: Code Huggingface: SciScore Huggingface: Science-T2I-S&C Benchmark Data The Science-T2I Fullset comprises a comprehensive collection of data for scientific T2I generation, including both training and test sets with a unified data structure. The test sets are split into 'test-S' and 'test-C,' corresponding to the Science-T2I-S and Science-T2I-C benchmarks, respectively. Download Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Jialuo21/Science-T2I-Fullset.image1K<n<10K0 likes1.3k downloads1y agoHugging Face29TalentZHOU /hle_material_science HLE Material Science: A Specialized Benchmark for Materials Science A Materials Science Subset of Humanity's Last Exam (HLE) Overview HLE Material Science is a carefully curated materials science subset derived from the Humanity's Last Exam (HLE) dataset, containing 106 high-quality expert-level questions covering 25+ materials science subfields, with 97% of questions rated as high confidence. This dataset is designed to evaluate large language models'… See the full description on the dataset page: https://huggingface.co/datasets/TalentZHOU/hle_material_science.textquestion-answeringn<1K1 likes1.3k downloads9mo agoHugging Face30reasoning-proj /judged_science_completionstabularn<1K2 likes1.3k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.