datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
science_materialsdiffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04 Contains maximum activating examples for all the features of our crosscoder trained on gemma 2 2B layer 13 available here: https://huggingface.co/Butanium/gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04/blob/main/README.md
base_examples.pt contains all the maximum examples of the feature on a subset of validation test of fineweb
chat_examples.pt is the same but for lmsys chat data
chat_base_examples.pt is a merge of the two above files.
All files are of the type dict[int, list[tuple[float… See the full description on the dataset page: https://huggingface.co/datasets/science-of-finetuning/diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04.science_biologynasa-science-repos-sme-benchmark
NASA Science Repos SME Benchmark
A benchmark dataset for evaluating retrieval systems on NASA science repository discovery tasks. This dataset contains expert queries, a corpus of NASA science GitHub repositories, and relevance judgments.
Dataset Structure
Files
├── corpus.jsonl # 5,264 repositories with full metadata
├── queries.jsonl # 219 expert queries
└── qrels/
├── earth.tsv # Earth Science relevance judgments (162)
├──… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-repos-sme-benchmark.modis-lake-powell-toy-dataset
MODIS Water Lake Powell Toy Dataset
Dataset Summary
Tabular dataset comprised of MODIS surface reflectance bands along with calculated indices and a label (water/not-water)
Dataset Structure
Data Fields
water: Label, water or not-water (binary)
sur_refl_b01_1: MODIS surface reflection band 1 (-100, 16000)
sur_refl_b02_1: MODIS surface reflection band 2 (-100, 16000)
sur_refl_b03_1: MODIS surface reflection band 3 (-100, 16000)
sur_refl_b04_1: MODIS… See the full description on the dataset page: https://huggingface.co/datasets/nasa-cisto-data-science-group/modis-lake-powell-toy-dataset.Science-DiscoveryNCERT_Political_Science_12thdiffing-stats-SAE-base-gemma-2-2b-L13-k100-x32-lr1e-04-local-shufflingjain-developability-cleandata-science-job-salaries
Dataset Card for Data Science Job Salaries
Dataset Summary
Content
Column
Description
work_year
The year the salary was paid.
experience_level
The experience level in the job during the year with the following possible values: EN Entry-level / Junior MI Mid-level / Intermediate SE Senior-level / Expert EX Executive-level / Director
employment_type
The type of employement for the role: PT Part-time FT Full-time CT Contract FL Freelance
job_title… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/data-science-job-salaries.tabrepair-science-repair-under-shift
TabRepair Science: Repair Under Shift
TabRepair Science is a finite authored benchmark for a deceptively hard
question: does better tabular cell repair produce better downstream models
under distribution shift?
The 3,648-row pilot spans three structural generator families, missingness and
present-value contamination, four test regimes, eight repair representations,
and five downstream learners. A separate eight-world sensitivity layer tests a
damage-aware v2 candidate without… See the full description on the dataset page: https://huggingface.co/datasets/haidang2405/tabrepair-science-repair-under-shift.weather-forecasting-challenge
Dataset Description
Data Overview
The WiDS Datathon 2023 focuses on a prediction task involving forecasting sub-seasonal temperatures (temperatures over a two-week period, in our case) within the United States. We are using a pre-prepared dataset consisting of weather and climate information for a number of US locations, for a number of start dates for the two-week observation, as well as the forecasted temperature and precipitation from a number of weather… See the full description on the dataset page: https://huggingface.co/datasets/serenia-science/weather-forecasting-challenge.shehata-antibody-psr
Shehata Antibody PSR Dataset (Novo Nordisk Preprocessing)
Dataset Summary
This dataset contains 398 human antibody heavy chain variable domain (VH) sequences with PSR (Poly-Specificity Reagent) measurements, preprocessed according to the methodology described in Sakhnini et al. 2025 (Novo Nordisk & University of Cambridge). The dataset was originally published by Shehata et al. 2019 and contains human B cell-derived antibodies studying the relationship between affinity… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/shehata-antibody-psr.harvey-nanobody-polyreactivity
Harvey Nanobody Polyreactivity Dataset (Novo Nordisk Preprocessing)
Dataset Summary
This dataset contains 141,021 nanobody (VHH) sequences with binary polyreactivity labels, preprocessed according to the methodology described in Sakhnini et al. 2025 (Novo Nordisk & University of Cambridge). The dataset was originally published by Harvey et al. 2022 and contains synthetic nanobodies assessed by PSR (Poly-Specificity Reagent) assay via FACS sorting and deep sequencing.
This… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/harvey-nanobody-polyreactivity.science_qa_txt_only_standardizedScienceQA_TEST_THmax-activating-examples-gemma-2-2b-l13-ckissanediffing-stats-SAE-chat-gemma-2-2b-L13-k100-lr1e-04-local-shufflingdiffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-x8-lr1e-04-local-shufflingscience_finding_sentence_certaintyThis data is from the github repo https://github.com/Jiaxin-Pei/Certainty-in-Science-Communication about the paper "Measuring Sentence-Level and Aspect-Level (Un)certainty in Science Communications" by Jiaxin Pei and David Jurgens.
I put it here for later use, to train an ML model which estimates claim certainty.
NCERT_Science_10thnasa-science-github-repos
NASA Science GitHub Repositories
A curated index of 5,264 GitHub repositories relevant to the NASA Science Mission
Directorate (SMD), spanning five science divisions: Earth Science, Astrophysics,
Planetary Science, Heliophysics, and Biological & Physical Sciences.
This dataset is designed to support research on information retrieval and
discoverability of open-source scientific software.
Licensing and Intellectual Property
This dataset is released under CC-BY-4.0 and… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-github-repos.AI-and-Data-Science-Job-Market-Dataset
About Dataset
Edit
The AI & Data Science Job Market Dataset (2020–2026) is a synthetically generated dataset designed to simulate real-world hiring patterns across the artificial intelligence and data science job market.
The dataset contains structured information about job roles, company characteristics, required technical skills, education levels, experience requirements, and salary ranges. It reflects hiring data across multiple countries, industries, and company sizes.
This… See the full description on the dataset page: https://huggingface.co/datasets/shree0910/AI-and-Data-Science-Job-Market-Dataset.Nemotron-RL-Science-v1-prompt-only
Nemotron-RL-Science-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Science-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Science-v1-prompt-only.NCERT_Science_8thdiffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-x1-lr1e-04-local-shufflingdata-science-job-salariesNCERT_Science_6thScienceQAdiffing-stats-Meta-Llama-3.1-8B-L16-mu2.1e-02-lr1e-04-local-shuffling-CrosscoderLoss
