Team Ai
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01marin-community /openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-32B (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-32B on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes441 downloads5mo agoHugging Face02marin-community /openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-30B-A3B-Thinking-2507 on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes440 downloads5mo agoHugging Face03islamlab /islamic-sciences islamlab — The Islamic Sciences Corpus The Islamic sciences other than Qur'an and hadith, as their authors wrote them: 4,022 works by scholars who died between the 0st and the 14th Hijri century, cut along their own chapter and biographical-entry boundaries into 1,864,389 units (3.41 billion characters of Arabic), each carrying the volume and page it sits on so a quotation can be cited rather than merely produced. Scope is Ahl al-Sunnah wa'l-Jamāʿah, and the gate is the author… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-sciences.tabulartext-generation1M<n<10M3 likes279 downloads2mo agoHugging Face04tilikumotp /Global-Ocean-Science-Corpus 🌊 Global-Ocean-Science-Corpus (v2.0 Curated & Cleaned) A Highly Curated, Large-Scale Pre-Training & RAG Corpus for Deep Ocean Sciences, Marine Biology, and Oceanography Language Note: This dataset is a 100% English-language scientific corpus (language: "en") aggregating peer-reviewed literature, deep-sea exploration dossiers, and technical oceanographic reports from leading global marine institutes. Global-Ocean-Science-Corpus, derin okyanus bilimleri, deniz biyolojisi… See the full description on the dataset page: https://huggingface.co/datasets/tilikumotp/Global-Ocean-Science-Corpus.tabulartext-generation10K<n<100K1 likes277 downloads11d agoHugging Face05marin-community /openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-4B (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-4B on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes215 downloads5mo agoHugging Face06open-athena /science-tool-use-conversations Science Tool-Use Conversations This dataset contains 11,405 synthetic conversations about science questions. GLM-5.3 generated both the user and assistant messages. The assistant could run commands in shellsim, an in-memory shell and Python simulator. Each row includes a system message, the user-visible conversation, a tool-call transcript, and the tool definition. Some conversations contain no tool calls. The questions come from the so_openq split of… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/science-tool-use-conversations.tabulartext-generation10K<n<100K0 likes191 downloads9d agoHugging Face07AgentsSci /IEEE2026_BigData_MAS-4-Science-Matching SciAgentTrace An execution-layer trace resource for scientific-agent workload characterization. A protocol fixes who reasons, what each role can see, when feedback returns, and when a workflow stops. Those choices determine the sequence of model requests that produces an answer, so protocol design is also workload design. Two workflows that consume similar token totals can issue very different request sequences. SciAgentTrace records that difference. The matched core runs the… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/IEEE2026_BigData_MAS-4-Science-Matching.tabulartext-generation1M<n<10M0 likes184 downloads17d agoHugging Face08sujitpandey /k-10-science-2-60k k-10-science-2-60k Synthetic, original expository text aligned to the K-10 (CBSE/NCERT-style) curriculum. Rows: 60,000 Total words: 47,082,892 Subject(s): Science Rows per grade: 1: 3,744, 10: 8,336, 2: 3,744, 3: 6,876, 4: 5,481, 5: 5,131, 6: 6,384, 7: 6,528, 8: 6,824, 9: 6,952 Fields Field Type Description text string The generated passage subject string Subject name grade int Grade level word_count int Number of words in text tabulartext-generation10K<n<100K0 likes72 downloads3d agoHugging Face09rodriguescarson /adaption-science-seed Science Q&A Science questions (chemistry, physics, biology; the biochem variant focuses on life sciences) with worked answers. Rows 12,000 Domain science Format data.parquet, one row per example Licence apache-2.0 Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description original_prompt The prompt (user turn) as uploaded. original_completion The target response as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-science-seed.tabularquestion-answering10K<n<100K0 likes70 downloads14d agoHugging Face10rodriguescarson /adaption-science-biochem-seed Science Q&A Science questions (chemistry, physics, biology; the biochem variant focuses on life sciences) with worked answers. Rows 12,458 Domain science Format data.parquet, one row per example Licence apache-2.0 Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description original_prompt The prompt (user turn) as uploaded. original_completion The target response as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-science-biochem-seed.tabularquestion-answering10K<n<100K0 likes67 downloads14d agoHugging Face11KadamParth /NCERT_Political_Science_12thtabularquestion-answering1K<n<10K1 likes66 downloads2y agoHugging Face12sujitpandey /k-10-science-10k k-10-science-10k Synthetic, original expository text aligned to the K-10 (CBSE/NCERT-style) curriculum. Rows: 10,000 Total words: 7,158,116 Subject(s): Science Rows per grade: 1: 613, 10: 1,442, 2: 567, 3: 1,036, 4: 847, 5: 896, 6: 1,085, 7: 1,106, 8: 1,204, 9: 1,204 Fields Field Type Description text string The generated passage subject string Subject name grade int Grade level word_count int Number of words in text tabulartext-generation10K<n<100K0 likes59 downloads3d agoHugging Face13HerrHruby /synthetic-science-v2-sample Synthetic Scientific Research Threads — v2 (sample) A synthetic continual-learning benchmark: each episode is a coherent sequence of short fictional scientific research documents about a single made-up entity, with per-document QA anchors. Later documents build on, revise, or supersede earlier ones. Designed to stress test-time / meta-learning approaches where a model must adapt to a stream of documents and answer questions grounded in what it has just seen. This is a sample… See the full description on the dataset page: https://huggingface.co/datasets/HerrHruby/synthetic-science-v2-sample.tabularquestion-answeringn<1K0 likes27 downloads3mo agoHugging Face14KadamParth /NCERT_Science_10thtabularquestion-answering1K<n<10K2 likes26 downloads2y agoHugging Face15mkurman /synthlabs-GLM-5.2-Science GLM-5.2 Science Synth Reasoning Synthetic reasoning traces for science questions from the GLM-5.2 science dataset. Each record contains a complex scientific question with SYNTH-style reasoning and a generated answer. Dataset Summary 33,014 records (605 dupes + 726 incomplete/truncated removed from 34,345 source) 33,014 reasoning turns (99.9% format compliance) Average 3,094 chars per reasoning trace Models Used Model Records… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-GLM-5.2-Science.tabulartext-generation10K<n<100K0 likes25 downloads2mo agoHugging Face16KadamParth /NCERT_Science_8thtabularquestion-answering1K<n<10K1 likes23 downloads2y agoHugging Face17KadamParth /NCERT_Science_6thtabularquestion-answering1K<n<10K1 likes22 downloads2y agoHugging Face18akhilpandey95 /scienceeval scienceeval Helpful evals for understanding the capabilities of models across a broad swath of science benchmarks. This artifact bundle contains full-corpus benchmark embeddings and projection coordinates for 69,315 science benchmark items.The embeddings were generated with pplx-embed-v1-0.6B, reduced to 50 dimensions with PCA, and projected with t-SNE, UMAP, and PCA. The t-SNE figure above shows a balanced visible sample over soft contours from the full-corpus distribution.… See the full description on the dataset page: https://huggingface.co/datasets/akhilpandey95/scienceeval.tabulartext-classification100K<n<1M0 likes19 downloads6mo agoHugging Face19lianghsun /tw-science-24Mgated Dataset Card for tw-science-24M 本資料集收錄臺灣相關之科學/科技類繁體中文文本(涵蓋農業、生態、氣候、醫療、化學、生物等領域),總 token 數約 24M(24 百萬),可作為繁中模型在科學領域的補充預訓練語料。 Dataset Details Dataset Description 資料來自繁體中文公開科學/科技/環境議題之報導與政府公告,內容包含: 科學新聞、政策報導 永續、生態保育、氣候議題 醫療衛生、農業科技 部分 zh-en 平行段落 文本經過清理、保留段落結構,可直接用於語言模型的持續預訓練。 Curated by: Huang Liang Hsun Language(s) (NLP): Traditional Chinese, English License: cc-by-nc-sa-4.0 Dataset Sources Repository: lianghsun/tw-science-24M… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-science-24M.tabulartext-generation10K<n<100K3 likes17 downloads2mo agoHugging Face20KadamParth /NCERT_Science_7thtabularquestion-answering1K<n<10K1 likes16 downloads2y agoHugging Face21KadamParth /NCERT_Political_Science_11thtabularquestion-answering1K<n<10K1 likes15 downloads2y agoHugging Face22KadamParth /NCERT_Science_9thtabularquestion-answering1K<n<10K3 likes14 downloads2y agoHugging Face23anan6450 /NCERT_Science_10thtabularquestion-answering1K<n<10K0 likes9 downloads7mo agoHugging Face24SAgarwal34 /grade3-science-explanations-v4r7 Grade-Level Science Explanations v4r7 The final 485-record supervised fine-tuning dataset for the grade-level science explainer. Each record maps a unique elementary-science question to a concise, mechanism-complete explanation. Training uses the minimal prompt Explain: {phrasing} so the reading behavior must be learned from examples rather than supplied through prompt instructions. Files File Records Purpose gold_v4_r7.jsonl 485 Final training split… See the full description on the dataset page: https://huggingface.co/datasets/SAgarwal34/grade3-science-explanations-v4r7.tabulartext-generationn<1K0 likes9 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.