datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16
OpenThoughts-4 Science SDG: Qwen3-32B (n=8, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-32B
on the Marin OpenThoughts-4 science SDG prompt
set.
Each prompt is sampled n=8 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field
Value
Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16.openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16
OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 science SDG prompt
set.
Each prompt is sampled n=8 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.islamic-sciences
islamlab — The Islamic Sciences Corpus
The Islamic sciences other than Qur'an and hadith, as their authors wrote
them: 4,022 works by scholars who died between the
0st and the 14th Hijri century, cut along their own chapter
and biographical-entry boundaries into 1,864,389 units
(3.41 billion characters of Arabic), each carrying the volume and
page it sits on so a quotation can be cited rather than merely produced.
Scope is Ahl al-Sunnah wa'l-Jamāʿah, and the gate is the author… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-sciences.Global-Ocean-Science-Corpus
🌊 Global-Ocean-Science-Corpus (v2.0 Curated & Cleaned)
A Highly Curated, Large-Scale Pre-Training & RAG Corpus for Deep Ocean Sciences, Marine Biology, and Oceanography
Language Note: This dataset is a 100% English-language scientific corpus (language: "en") aggregating peer-reviewed literature, deep-sea exploration dossiers, and technical oceanographic reports from leading global marine institutes.
Global-Ocean-Science-Corpus, derin okyanus bilimleri, deniz biyolojisi… See the full description on the dataset page: https://huggingface.co/datasets/tilikumotp/Global-Ocean-Science-Corpus.openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16
OpenThoughts-4 Science SDG: Qwen3-4B (n=8, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-4B
on the Marin OpenThoughts-4 science SDG prompt
set.
Each prompt is sampled n=8 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field
Value
Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16.science-tool-use-conversations
Science Tool-Use Conversations
This dataset contains 11,405 synthetic conversations about science questions. GLM-5.3 generated both the user and assistant messages. The assistant could run commands in shellsim, an in-memory shell and Python simulator. Each row includes a system message, the user-visible conversation, a tool-call transcript, and the tool definition. Some conversations contain no tool calls.
The questions come from the so_openq split of… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/science-tool-use-conversations.IEEE2026_BigData_MAS-4-Science-Matching
SciAgentTrace
An execution-layer trace resource for scientific-agent workload characterization.
A protocol fixes who reasons, what each role can see, when feedback returns, and
when a workflow stops. Those choices determine the sequence of model requests
that produces an answer, so protocol design is also workload design. Two
workflows that consume similar token totals can issue very different request
sequences. SciAgentTrace records that difference.
The matched core runs the… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/IEEE2026_BigData_MAS-4-Science-Matching.k-10-science-2-60k
k-10-science-2-60k
Synthetic, original expository text aligned to the K-10 (CBSE/NCERT-style) curriculum.
Rows: 60,000
Total words: 47,082,892
Subject(s): Science
Rows per grade: 1: 3,744, 10: 8,336, 2: 3,744, 3: 6,876, 4: 5,481, 5: 5,131, 6: 6,384, 7: 6,528, 8: 6,824, 9: 6,952
Fields
Field
Type
Description
text
string
The generated passage
subject
string
Subject name
grade
int
Grade level
word_count
int
Number of words in text
adaption-science-seed
Science Q&A
Science questions (chemistry, physics, biology; the biochem variant focuses on life sciences) with worked answers.
Rows
12,000
Domain
science
Format
data.parquet, one row per example
Licence
apache-2.0
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
original_prompt
The prompt (user turn) as uploaded.
original_completion
The target response as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-science-seed.adaption-science-biochem-seed
Science Q&A
Science questions (chemistry, physics, biology; the biochem variant focuses on life sciences) with worked answers.
Rows
12,458
Domain
science
Format
data.parquet, one row per example
Licence
apache-2.0
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
original_prompt
The prompt (user turn) as uploaded.
original_completion
The target response as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-science-biochem-seed.NCERT_Political_Science_12thk-10-science-10k
k-10-science-10k
Synthetic, original expository text aligned to the K-10 (CBSE/NCERT-style) curriculum.
Rows: 10,000
Total words: 7,158,116
Subject(s): Science
Rows per grade: 1: 613, 10: 1,442, 2: 567, 3: 1,036, 4: 847, 5: 896, 6: 1,085, 7: 1,106, 8: 1,204, 9: 1,204
Fields
Field
Type
Description
text
string
The generated passage
subject
string
Subject name
grade
int
Grade level
word_count
int
Number of words in text
synthetic-science-v2-sample
Synthetic Scientific Research Threads — v2 (sample)
A synthetic continual-learning benchmark: each episode is a coherent sequence of
short fictional scientific research documents about a single made-up entity, with
per-document QA anchors. Later documents build on, revise, or supersede earlier
ones. Designed to stress test-time / meta-learning approaches where a model must
adapt to a stream of documents and answer questions grounded in what it has just
seen.
This is a sample… See the full description on the dataset page: https://huggingface.co/datasets/HerrHruby/synthetic-science-v2-sample.NCERT_Science_10thsynthlabs-GLM-5.2-Science
GLM-5.2 Science Synth Reasoning
Synthetic reasoning traces for science questions from the GLM-5.2 science dataset. Each record contains a complex scientific question with SYNTH-style reasoning and a generated answer.
Dataset Summary
33,014 records (605 dupes + 726 incomplete/truncated removed from 34,345 source)
33,014 reasoning turns (99.9% format compliance)
Average 3,094 chars per reasoning trace
Models Used
Model
Records… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-GLM-5.2-Science.NCERT_Science_8thNCERT_Science_6thscienceeval
scienceeval
Helpful evals for understanding the capabilities of models across a broad swath
of science benchmarks.
This artifact bundle contains full-corpus benchmark embeddings and projection
coordinates for 69,315 science benchmark items.The embeddings were generated with pplx-embed-v1-0.6B, reduced to 50
dimensions with PCA, and projected with t-SNE, UMAP, and PCA. The t-SNE figure
above shows a balanced visible sample over soft contours from the full-corpus
distribution.… See the full description on the dataset page: https://huggingface.co/datasets/akhilpandey95/scienceeval.tw-science-24M
Dataset Card for tw-science-24M
本資料集收錄臺灣相關之科學/科技類繁體中文文本(涵蓋農業、生態、氣候、醫療、化學、生物等領域),總 token 數約 24M(24 百萬),可作為繁中模型在科學領域的補充預訓練語料。
Dataset Details
Dataset Description
資料來自繁體中文公開科學/科技/環境議題之報導與政府公告,內容包含:
科學新聞、政策報導
永續、生態保育、氣候議題
醫療衛生、農業科技
部分 zh-en 平行段落
文本經過清理、保留段落結構,可直接用於語言模型的持續預訓練。
Curated by: Huang Liang Hsun
Language(s) (NLP): Traditional Chinese, English
License: cc-by-nc-sa-4.0
Dataset Sources
Repository: lianghsun/tw-science-24M… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-science-24M.NCERT_Science_7thNCERT_Political_Science_11thNCERT_Science_9thNCERT_Science_10thgrade3-science-explanations-v4r7
Grade-Level Science Explanations v4r7
The final 485-record supervised fine-tuning dataset for the grade-level science
explainer. Each record maps a unique elementary-science question to a concise,
mechanism-complete explanation. Training uses the minimal prompt
Explain: {phrasing} so the reading behavior must be learned from examples rather
than supplied through prompt instructions.
Files
File
Records
Purpose
gold_v4_r7.jsonl
485
Final training split… See the full description on the dataset page: https://huggingface.co/datasets/SAgarwal34/grade3-science-explanations-v4r7.
