Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hugging-science /mmu_manga mmu_manga HATS Catalog Collection This is the collection of HATS catalogs representing mmu_manga. This dataset is part of the Multimodal Universe, a large-scale collection of multimodal astronomical data. For full details, see the paper: The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data. Access the catalog We recommend the use of the LSDB Python framework to access HATS catalogs. LSDB can be installed via… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_manga.tabular10K<n<100K0 likes2.7k downloads4mo agoHugging Face02vidore /vidore_v3_computer_scienceViDoRe V3 : Computer Science This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark. About ViDoRe v3 ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science.documentvisual-document-retrieval1K<n<10K6 likes2k downloads9mo agoHugging Face03reasoning-proj /judged_science_completionstabularn<1K2 likes1.4k downloads1y agoHugging Face04PleIAs /French-Science-Commons French Science Commons French Science Commons (Commun numérique des sciences en français) rassemble des publications scientifiques d'origine française en accès ouvert, couvrant une période de vingt ans, de 2007 à 2026. Il comprend 1 248 860 documents scientifiques — 1 189 628 articles et 59 232 thèses — indexés à travers de multiples dépôts académiques en accès public, tels que HAL, OpenAlex, des revues scientifiques, des dépôts institutionnels, et d'autres. Le corpus est conçu… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-Science-Commons.tabular10M<n<100M26 likes1.1k downloads4mo agoHugging Face05simpleG2023 /chinese-materials-science-open-intelligence 🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.tabulartext-retrieval1K<n<10K0 likes999 downloads2h agoHugging Face06reasoning-proj /severity_ablation_sciencetabular100K<n<1M0 likes955 downloads1y agoHugging Face07huggingface /community-science-paper-v2tabular1K<n<10K7 likes550 downloads2y agoHugging Face08GodotCN /science-datalake Science Data Lake A unified, portable science data lake integrating 7 scholarly datasets (~525 GB Parquet) with cross-dataset DOI normalization, 13 scientific ontologies (1.3M terms), and a reproducible ETL pipeline. Note: One additional source (Semantic Scholar S2AG) is supported by the pipeline but is not redistributed here due to its API terms of service. See Not Included in This Upload below. What's Unique This dataset enables queries… See the full description on the dataset page: https://huggingface.co/datasets/GodotCN/science-datalake.imagetext-classification10B<n<100B1 likes544 downloads6mo agoHugging Face09Pclanglais /EU-Science-Commonstabular1M<n<10M0 likes540 downloads5mo agoHugging Face10hugging-science /mmu_apogee_dr17 mmu_apogee_dr17 HATS Catalog Collection This is the collection of HATS catalogs representing mmu_apogee_dr17. This dataset is part of the Multimodal Universe, a large-scale collection of multimodal astronomical data. For full details, see the paper: The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data. Access the catalog We recommend the use of the LSDB Python framework to access HATS catalogs. LSDB can be… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_apogee_dr17.tabular100K<n<1M1 likes499 downloads4mo agoHugging Face11SZLHOLDINGS /szl-science-forum-corpus Science Forum Pilot Explore original summaries and metadata from two operator-authored topics used to formulate review hypotheses. Artifact: Two-topic metadata pilot · Stage: Training unauthorized Explore in Command Lab · Build · Evidence Before you use it This is not a forum scrape, representative sample, model-training dataset or scientific benchmark. Original annotations do not grant rights to linked forum posts; expansion requires separate access, reuse and… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/szl-science-forum-corpus.tabularn<1K0 likes493 downloads6d agoHugging Face12marin-community /openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-32B (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-32B on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes441 downloads5mo agoHugging Face13mlfoundations-dev /qwq_mix_qwen3_sciencetabular100K<n<1M1 likes440 downloads1y agoHugging Face14marin-community /openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-30B-A3B-Thinking-2507 on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes440 downloads5mo agoHugging Face15mlfoundations-dev /pdf_science_questions_verified_r1_traces__2_24_25 Dataset card for pdf_science_questions_verified_r1_traces__2_24_25 This dataset was made with Curator. Dataset details A sample from the dataset: { "url": "https://www.ttcho.com/_files/ugd/988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf", "filename": "988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf", "success": true, "page_count": 37, "page_number": 1, "question_choices_solutions": "QUESTION: What is the identity of X in the reaction 14N + 1n \u2192… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/pdf_science_questions_verified_r1_traces__2_24_25.tabular1K<n<10K0 likes436 downloads2y agoHugging Face16hugging-science /mmu_hsc_pdr3_wide_21 mmu_hsc_pdr3_wide_21 HATS Catalog Collection This is the collection of HATS catalogs representing mmu_hsc_pdr3_wide_21. This dataset is part of the Multimodal Universe, a large-scale collection of multimodal astronomical data. For full details, see the paper: The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data. Access the catalog We recommend the use of the LSDB Python framework to access HATS catalogs. LSDB… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_hsc_pdr3_wide_21.tabular1M<n<10M0 likes419 downloads29d agoHugging Face17deep-principle /science_materialstabularn<1K0 likes409 downloads20d agoHugging Face18mlfoundations-dev /qwq_mix_r1_sciencetabular100K<n<1M0 likes396 downloads1y agoHugging Face19mlfoundations-dev /Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_2870 mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_2870 Precomputed model outputs for evaluation. Evaluation Results AIME24 Average Accuracy: 60.67% ± 2.20% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 70.00% 21 30 2 53.33% 16 30 3 53.33% 16 30 4 66.67% 20 30 5 63.33% 19 30 6 66.67% 20 30 7 60.00% 18 30 8 46.67% 14 30 9 63.33% 19 30 10 63.33% 19 30 tabularn<1K0 likes296 downloads1y agoHugging Face20science-of-finetuning /fineweb-1m-sampletabular1M<n<10M1 likes288 downloads2y agoHugging Face21science-of-finetuning /diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04 Contains maximum activating examples for all the features of our crosscoder trained on gemma 2 2B layer 13 available here: https://huggingface.co/Butanium/gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04/blob/main/README.md base_examples.pt contains all the maximum examples of the feature on a subset of validation test of fineweb chat_examples.pt is the same but for lmsys chat data chat_base_examples.pt is a merge of the two above files. All files are of the type dict[int, list[tuple[float… See the full description on the dataset page: https://huggingface.co/datasets/science-of-finetuning/diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04.tabular10K<n<100K0 likes281 downloads1y agoHugging Face22islamlab /islamic-sciences islamlab — The Islamic Sciences Corpus The Islamic sciences other than Qur'an and hadith, as their authors wrote them: 4,022 works by scholars who died between the 0st and the 14th Hijri century, cut along their own chapter and biographical-entry boundaries into 1,864,389 units (3.41 billion characters of Arabic), each carrying the volume and page it sits on so a quotation can be cited rather than merely produced. Scope is Ahl al-Sunnah wa'l-Jamāʿah, and the gate is the author… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-sciences.tabulartext-generation1M<n<10M3 likes279 downloads2mo agoHugging Face23tilikumotp /Global-Ocean-Science-Corpus 🌊 Global-Ocean-Science-Corpus (v2.0 Curated & Cleaned) A Highly Curated, Large-Scale Pre-Training & RAG Corpus for Deep Ocean Sciences, Marine Biology, and Oceanography Language Note: This dataset is a 100% English-language scientific corpus (language: "en") aggregating peer-reviewed literature, deep-sea exploration dossiers, and technical oceanographic reports from leading global marine institutes. Global-Ocean-Science-Corpus, derin okyanus bilimleri, deniz biyolojisi… See the full description on the dataset page: https://huggingface.co/datasets/tilikumotp/Global-Ocean-Science-Corpus.tabulartext-generation10K<n<100K1 likes277 downloads11d agoHugging Face24deep-principle /science_biologytabularn<1K0 likes223 downloads20d agoHugging Face25marin-community /openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-4B (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-4B on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes215 downloads5mo agoHugging Face26open-athena /science-tool-use-conversations Science Tool-Use Conversations This dataset contains 11,405 synthetic conversations about science questions. GLM-5.3 generated both the user and assistant messages. The assistant could run commands in shellsim, an in-memory shell and Python simulator. Each row includes a system message, the user-visible conversation, a tool-call transcript, and the tool definition. Some conversations contain no tool calls. The questions come from the so_openq split of… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/science-tool-use-conversations.tabulartext-generation10K<n<100K0 likes191 downloads9d agoHugging Face27AgentsSci /IEEE2026_BigData_MAS-4-Science-Matching SciAgentTrace An execution-layer trace resource for scientific-agent workload characterization. A protocol fixes who reasons, what each role can see, when feedback returns, and when a workflow stops. Those choices determine the sequence of model requests that produces an answer, so protocol design is also workload design. Two workflows that consume similar token totals can issue very different request sequences. SciAgentTrace records that difference. The matched core runs the… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/IEEE2026_BigData_MAS-4-Science-Matching.tabulartext-generation1M<n<10M0 likes184 downloads16d agoHugging Face28mlfoundations-dev /e1_science_longest_phitabular10K<n<100K0 likes176 downloads1y agoHugging Face29mlfoundations-dev /Qwen2.5-7B-Instruct_qwq_mix_qwen3_science_eval_8179 mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_qwen3_science_eval_8179 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 61.7 88.8 88.4 67.5 54.7 52.1 25.8 27.1 49.0 11.2 40.7 32.7 AIME24 Average Accuracy: 61.67% ± 1.27% Number of Runs: 10 Run Accuracy Questions Solved Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_qwen3_science_eval_8179.tabular10K<n<100K0 likes171 downloads1y agoHugging Face30thordata /science-communication-video-v1 Science Communication Video Preview This preview contains four short educational animation videos from the Thordata Multidisciplinary Science Communication Video Collection: acetaldehyde oxidation how volcanoes form how typhoons form solar wind and aurora Each sample presents one focused knowledge topic through a coherent visual sequence. The videos are useful for demonstrating video captioning, visual question answering, and cross-modal reasoning. Contents… See the full description on the dataset page: https://huggingface.co/datasets/thordata/science-communication-video-v1.tabularn<1K1 likes159 downloads24d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.