Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SAIS-Life-Science /Aneumo Aneumo Datasets AneumoDataset is a comprehensive multi-modal dataset containing 3D anatomical structures and simulated hemodynamic parameters for cerebral aneurysms, supporting both computational modeling and AI-based analysis. textn<1K7 likes3.1k downloads7mo agoHugging Face02OpenMOSS-Team /SWE-bench-Science SWE-bench Science SWE-bench Science evaluates coding agents on software-engineering tasks drawn from scientific-computing repositories. The release contains 119 tasks across 20 scientific domains, with isolated environments and separate programmatic verifiers. GitHub release repository: OpenMOSS/SWE-bench-Science Runtime images: Docker Hub, pinned by immutable linux/amd64 digests Evaluation framework: Pier, compatible with Harbor task format Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SWE-bench-Science.textn<1K9 likes2.1k downloads9d agoHugging Face03armanc /ScienceQAThis is the ScientificQA dataset by Saikh et al (2022). @article{10.1007/s00799-022-00329-y, author = {Saikh, Tanik and Ghosal, Tirthankar and Mittal, Amish and Ekbal, Asif and Bhattacharyya, Pushpak}, title = {ScienceQA: A Novel Resource for Question Answering on Scholarly Articles}, year = {2022}, journal = {Int. J. Digit. Libr.}, month = {sep} } text10K<n<100K14 likes594 downloads4y agoHugging Face04deep-principle /science_materialstabularn<1K0 likes409 downloads20d agoHugging Face05science-of-finetuning /diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04 Contains maximum activating examples for all the features of our crosscoder trained on gemma 2 2B layer 13 available here: https://huggingface.co/Butanium/gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04/blob/main/README.md base_examples.pt contains all the maximum examples of the feature on a subset of validation test of fineweb chat_examples.pt is the same but for lmsys chat data chat_base_examples.pt is a merge of the two above files. All files are of the type dict[int, list[tuple[float… See the full description on the dataset page: https://huggingface.co/datasets/science-of-finetuning/diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04.tabular10K<n<100K0 likes281 downloads1y agoHugging Face06Kaeyze /computer-science-synthetic-datasettext10K<n<100K12 likes242 downloads2y agoHugging Face07deep-principle /science_biologytabularn<1K0 likes223 downloads20d agoHugging Face08nasa-impact /nasa-science-repos-sme-benchmark NASA Science Repos SME Benchmark A benchmark dataset for evaluating retrieval systems on NASA science repository discovery tasks. This dataset contains expert queries, a corpus of NASA science GitHub repositories, and relevance judgments. Dataset Structure Files ├── corpus.jsonl # 5,264 repositories with full metadata ├── queries.jsonl # 219 expert queries └── qrels/ ├── earth.tsv # Earth Science relevance judgments (162) ├──… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-repos-sme-benchmark.tabulartext-retrievaln<1K0 likes156 downloads9mo agoHugging Face09river-martin /web-of-science-with-label-texts Dataset Description: The data is partitioned according to a 75/15/15 train/test/validate split. Each entry has an abstract (which is the input text for classification), a domain (a label from the list below), and an area (a subdomain of the paper, such as CS -> computer graphics, which takes on one of 134 possible values). All the attributes are strings. Domain labels: - Computer Science - Electrical Engineering - Psychology - Mechanical Engineering, - Civil Engineering - Medical… See the full description on the dataset page: https://huggingface.co/datasets/river-martin/web-of-science-with-label-texts.text10K<n<100K1 likes151 downloads2y agoHugging Face10deep-principle /science_physicstextn<1K2 likes139 downloads20d agoHugging Face11BDDSSD /ScienceAgentBench ScienceAgentBench The advancements of language language models (LLMs) have piqued growing interest in developing LLM-based language agents to automate scientific discovery end-to-end, which has sparked both excitement and skepticism about their true capabilities. In this work, we call for rigorous assessment of agents on individual tasks in a scientific workflow before making bold claims on end-to-end automation. To this end, we present ScienceAgentBench, a new benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/BDDSSD/ScienceAgentBench.textn<1K0 likes139 downloads7mo agoHugging Face12juntaoyuan /test-sciencetextn<1K0 likes129 downloads2y agoHugging Face13nasa-cisto-data-science-group /modis-lake-powell-toy-dataset MODIS Water Lake Powell Toy Dataset Dataset Summary Tabular dataset comprised of MODIS surface reflectance bands along with calculated indices and a label (water/not-water) Dataset Structure Data Fields water: Label, water or not-water (binary) sur_refl_b01_1: MODIS surface reflection band 1 (-100, 16000) sur_refl_b02_1: MODIS surface reflection band 2 (-100, 16000) sur_refl_b03_1: MODIS surface reflection band 3 (-100, 16000) sur_refl_b04_1: MODIS… See the full description on the dataset page: https://huggingface.co/datasets/nasa-cisto-data-science-group/modis-lake-powell-toy-dataset.image1K<n<10K1 likes116 downloads3y agoHugging Face14Sangeetha /Kaggle-LLM-Science-Exam Dataset Card for [LLM Science Exam Kaggle Competition] Dataset Summary https://www.kaggle.com/competitions/kaggle-llm-science-exam/data Languages [en, de, tl, it, es, fr, pt, id, pl, ro, so, ca, da, sw, hu, no, nl, et, af, hr, lv, sl] Dataset Structure Columns prompt - the text of the question being asked A - option A; if this option is correct, then answer will be A B - option B; if this option is correct, then answer will be B C - option C; if this… See the full description on the dataset page: https://huggingface.co/datasets/Sangeetha/Kaggle-LLM-Science-Exam.text1K<n<10K3 likes109 downloads3y agoHugging Face15xbench /ScienceQA xbench-evals 🌐 Website | 📄 Paper | 🤗 Dataset Evergreen, contamination-free, real-world, domain-specific AI evaluation framework xbench is more than just a scoreboard — it's a new evaluation framework with two complementary tracks, designed to measure both the intelligence frontier and real-world utility of AI systems: AGI Tracking: Measures core model capabilities like reasoning, tool-use, and memory Profession Aligned: A new class of evals grounded in workflows, environments… See the full description on the dataset page: https://huggingface.co/datasets/xbench/ScienceQA.textn<1K8 likes95 downloads1y agoHugging Face16StarpowerTechnology /Dense-Information-Science-Physics-Dataset Dense Information With Multiple Fine-tuned Variations This dataaset has multiple for each input to learn how to express the same answer in different ways Dataset Structure The dataset contains two columns: Column Description input A science or quantum-physics question output A conversational answer to the question Example: { "input": "What is quantum entanglement?", "output": "Quantum entanglement is when two quantum systems share one… See the full description on the dataset page: https://huggingface.co/datasets/StarpowerTechnology/Dense-Information-Science-Physics-Dataset.texttext-generation1K<n<10K0 likes81 downloads1mo agoHugging Face17rocky250 /Science-Discoverytabular100K<n<1M0 likes71 downloads11mo agoHugging Face18KadamParth /NCERT_Political_Science_12thtabularquestion-answering1K<n<10K1 likes66 downloads2y agoHugging Face19hugging-science /awesome-food-allergy-datasets Awesome Food Allergy Datasets A curated collection of datasets, databases, and computational resources for food allergy research, allergen identification, drug development, and clinical applications. 🧬 Dataset Description Dataset Summary Food allergy affects over 220 million people worldwide. This repository serves as the first comprehensive, open collection of AI-ready datasets for food allergy research—spanning clinical trials, immunotherapy, genomics… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/awesome-food-allergy-datasets.textn<1K6 likes59 downloads1y agoHugging Face20science-of-finetuning /diffing-stats-SAE-base-gemma-2-2b-L13-k100-x32-lr1e-04-local-shufflingtabular100K<n<1M0 likes57 downloads1y agoHugging Face21hugging-science /jain-developability-cleantabularn<1K5 likes55 downloads1y agoHugging Face22nasa-impact /nasa-science-code-benchmark-v0.1.1 NASA Code Retrieval Benchmark v0.1.1 This repository is an updated version of the NASA Code Retrieval Benchmark. It provides a code retrieval benchmark based on code from 7 programming languages sourced from NASA's GitHub repositories. What's New in v0.1.1? v0.1.1 introduces a hierarchical structure and official Hugging Face dataset configurations. This allows you to evaluate models specifically by language or by query category without data redundancy in the file system.… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-code-benchmark-v0.1.1.texttext-retrieval100K<n<1M0 likes54 downloads6mo agoHugging Face23hugginglearners /data-science-job-salaries Dataset Card for Data Science Job Salaries Dataset Summary Content Column Description work_year The year the salary was paid. experience_level The experience level in the job during the year with the following possible values: EN Entry-level / Junior MI Mid-level / Intermediate SE Senior-level / Expert EX Executive-level / Director employment_type The type of employement for the role: PT Part-time FT Full-time CT Contract FL Freelance job_title… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/data-science-job-salaries.tabularn<1K6 likes53 downloads4y agoHugging Face24haidang2405 /tabrepair-science-repair-under-shift TabRepair Science: Repair Under Shift TabRepair Science is a finite authored benchmark for a deceptively hard question: does better tabular cell repair produce better downstream models under distribution shift? The 3,648-row pilot spans three structural generator families, missingness and present-value contamination, four test regimes, eight repair representations, and five downstream learners. A separate eight-world sensitivity layer tests a damage-aware v2 candidate without… See the full description on the dataset page: https://huggingface.co/datasets/haidang2405/tabrepair-science-repair-under-shift.tabulartabular-regression100K<n<1M0 likes53 downloads1mo agoHugging Face25hugging-science /boughter-antibody-polyreactivity Boughter Antibody Polyreactivity Dataset (Novo Nordisk Preprocessing) Dataset Summary This dataset contains 914 antibody heavy chain variable domain (VH) sequences with binary polyreactivity labels, preprocessed according to the methodology described in Sakhnini et al. 2025 (Novo Nordisk & University of Cambridge). The dataset was originally published by Boughter et al. 2020 and contains mouse antibodies with ELISA-based polyreactivity measurements against a panel of 4–7… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/boughter-antibody-polyreactivity.texttext-classificationn<1K2 likes51 downloads10mo agoHugging Face26serenia-science /weather-forecasting-challenge Dataset Description Data Overview The WiDS Datathon 2023 focuses on a prediction task involving forecasting sub-seasonal temperatures (temperatures over a two-week period, in our case) within the United States. We are using a pre-prepared dataset consisting of weather and climate information for a number of US locations, for a number of start dates for the two-week observation, as well as the forecasted temperature and precipitation from a number of weather… See the full description on the dataset page: https://huggingface.co/datasets/serenia-science/weather-forecasting-challenge.tabulartabular-regression100K<n<1M0 likes47 downloads2y agoHugging Face27nasa-impact /nasa-science-code-benchmark-v0.1 NASA Code Retrieval Benchmark v0.1 Note: This dataset has been superseded by nasa-impact/nasa-science-code-benchmark-v0.1.1, which introduces a hierarchical structure, official Hugging Face dataset configurations, and evaluation by NASA science division. Please use v0.1.1 for new work. This dataset provides a code retrieval benchmark based on code from 7 programming languages (Python, C, C++, Java, JavaScript, Fortran, and Matlab) sourced from NASA's GitHub repositories. It serves… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-code-benchmark-v0.1.text100K<n<1M0 likes44 downloads6mo agoHugging Face28hugging-science /shehata-antibody-psr Shehata Antibody PSR Dataset (Novo Nordisk Preprocessing) Dataset Summary This dataset contains 398 human antibody heavy chain variable domain (VH) sequences with PSR (Poly-Specificity Reagent) measurements, preprocessed according to the methodology described in Sakhnini et al. 2025 (Novo Nordisk & University of Cambridge). The dataset was originally published by Shehata et al. 2019 and contains human B cell-derived antibodies studying the relationship between affinity… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/shehata-antibody-psr.tabulartext-classificationn<1K1 likes41 downloads10mo agoHugging Face29hugging-science /harvey-nanobody-polyreactivity Harvey Nanobody Polyreactivity Dataset (Novo Nordisk Preprocessing) Dataset Summary This dataset contains 141,021 nanobody (VHH) sequences with binary polyreactivity labels, preprocessed according to the methodology described in Sakhnini et al. 2025 (Novo Nordisk & University of Cambridge). The dataset was originally published by Harvey et al. 2022 and contains synthetic nanobodies assessed by PSR (Poly-Specificity Reagent) assay via FACS sorting and deep sequencing. This… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/harvey-nanobody-polyreactivity.tabulartext-classification100K<n<1M4 likes38 downloads10mo agoHugging Face30Saxo /ko_jp_translation_tech_social_science_linkbricks_single_datasettext100K<n<1M0 likes37 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.