Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BAAI /IndustryCorpus2_mathematics_statistics IndustryCorpus2: Mathematics & Statistics This repository contains the IndustryCorpus2: Mathematics & Statistics domain subset of BAAI/IndustryCorpus2. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryCorpus2: @misc{shi2024industrycorpus2, title = {IndustryCorpus2}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_mathematics_statistics.1 likes708 downloads2mo agoHugging Face02DataoceanAI /University-level_Mathematics_Physics_Chemistry_Computer_Science_Reasoning_Corpus Title University-level Mathematics, Physics, Chemistry, Computer Science Reasoning Corpus Size 200,000+ text+ multimodal university-level problems, each with step-by-step solutions and final answers Format Natural language explanations with multimodal samples include images (graphs, diagrams, etc.) Subject Mathematics, Physics, Chemistry, Computer Science Labeling Details Question ID/Question Stem (Full text/content) /Subject/Question Type… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/University-level_Mathematics_Physics_Chemistry_Computer_Science_Reasoning_Corpus.0 likes371 downloads2y agoHugging Face03timaeus /dsir-pile-13m-filtered-no-github-or-dm_mathematics DSIR Pile 13M - Filtered Version This is a filtered version of timaeus/dsir-pile-13m. Filtering Applied: Excluded: All rows where metadata.pile_set_name contains 'Github' or 'DM_mathematics' Kept: All other rows from the original dataset Dataset Size Original: ~13M examples Filtered: 12,782,200 examples (99.9% of original) Uploaded in: 64 batch files Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/dsir-pile-13m-filtered-no-github-or-dm_mathematics.text10M<n<100M0 likes369 downloads1y agoHugging Face04d0rj /mathematics_dataset Mathematical Reasoning Dataset (English & Russian) A bilingual collection of synthetic school-level mathematics questions and answers, based on the DeepMind mathematics_dataset generator. This dataset contains two language splits: en — the original English data, taken as-is from the official mathematics_dataset-v1.0 release published by Google DeepMind (github.com/google-deepmind/mathematics_dataset). ru — a Russian version generated from scratch with a translated fork of the… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/mathematics_dataset.texttext-generation100M<n<1B3 likes284 downloads3mo agoHugging Face05Mathematics-Yang /phase_tree_results PHASE-Tree Evaluation Results Full evaluation outputs for the PHASE-Tree paper (Psychology-grounded Hierarchical Attribute-Structured Evolving Tree), covering 8 character-dialogue datasets, 4 experimental paradigms, and 2 evaluation splits (random test + OOD test). Please cite this work if you use these results for analysis, comparison, reproduction, or any other research purpose. 🔗 Resources: 📄 Paper: arXiv:2608.06975 📦 GitHub Repository: MemTensor/PHASE-Tree (code… See the full description on the dataset page: https://huggingface.co/datasets/Mathematics-Yang/phase_tree_results.texttext-generation10K<n<100K1 likes229 downloads2mo agoHugging Face06asandeistefan /romanian-baccalaureate-mathematics Romanian Baccalaureate in Mathematics A curated collection of Romanian Baccalaureate (BAC) mathematics examination papers and answer keys, transcribed from PDF to structured Markdown using Vision-Language Model OCR. Currently the years 2019 - 2025 were added, more will be processed soon. Directory Structure romanian-baccalaureate-mathematics/ ├── metadata.csv # Index of all exam papers ├── pdfs/ # Original PDF files │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/asandeistefan/romanian-baccalaureate-mathematics.n<1K2 likes191 downloads4mo agoHugging Face07VietAlphaLabs /vi-en-mathematics-dictionaryVietAlpha English–Vietnamese Mathematics Dictionary Research page · VietAlpha Lab · Source scan The VietAlpha English–Vietnamese Mathematics Dictionary turns a 709-page printed reference work into a machine-readable bilingual lexicon. It contains 26,205 English and Vietnamese mathematics entries digitized from Cung Kim Tiến's Từ Điển Toán Học Anh – Việt, Việt – Anh and organized as JSON Lines. What is in the dataset Direction Entries English to Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/VietAlphaLabs/vi-en-mathematics-dictionary.texttranslation10K<n<100K1 likes170 downloads15d agoHugging Face08morka17 /sat-mathematics-complete-questions0 likes160 downloads1y agoHugging Face09BAAI /IndustryCorpus_mathematics[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_mathematics.texttext-generation1M<n<10M3 likes152 downloads2mo agoHugging Face10VietAlphaLabs /fr-vi-mathematics-dictionaryVietAlpha French–Vietnamese Mathematics Dictionary Research page · VietAlpha Lab The VietAlpha French–Vietnamese Mathematics Dictionary is a machine-readable edition of Danh-từ Toán-học Pháp-Việt, compiled in Saigon in 1964 by the Mathematics Committee of the National Committee for the Compilation of Specialized Dictionaries. The release contains 4,095 dictionary entries and a 1,369-item Vietnamese index reconstructed from the printed volume. This dataset records how a Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/VietAlphaLabs/fr-vi-mathematics-dictionary.texttranslation1K<n<10K0 likes139 downloads15d agoHugging Face11timaeus /pile-dm_mathematics Dataset Creation Process These subsets were created by streaming over the rows from monology/pile-uncopyrighted and filtering by the meta column. Each subset is generally limited to the first 100,000 qualifying rows encountered. Citations If you use this dataset, please cite the original Pile papers: @article{gao2020pile, title={The Pile: An 800GB dataset of diverse text for language modeling}, author={Gao, Leo and Biderman, Stella and Black, Sid and Golding, Laurence and… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/pile-dm_mathematics.text100K<n<1M1 likes127 downloads2y agoHugging Face1211-47 /pure_mathematics_25ktext10K<n<100K2 likes103 downloads5mo agoHugging Face13tourist800 /LLM-Hallucination-Detection-complex-mathematics AIME Hallucination Detection Dataset This dataset is created for detecting hallucinations in Large Language Models (LLMs), particularly focusing on complex mathematical problems. It can be used for tasks like model evaluation, fine-tuning, and research. Dataset Details Name: AIME Hallucination Detection Dataset Format: CSV Size: (add size, e.g., 10MB) Files Included: AIME-hallucination-detection-dataset.csv: Contains the dataset. Content Description… See the full description on the dataset page: https://huggingface.co/datasets/tourist800/LLM-Hallucination-Detection-complex-mathematics.tabularn<1K2 likes100 downloads2y agoHugging Face14Lots-of-LoRAs /task706_mmmlu_answer_generation_high_school_mathematics Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task706_mmmlu_answer_generation_high_school_mathematics Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task706_mmmlu_answer_generation_high_school_mathematics.texttext-generationn<1K0 likes92 downloads2y agoHugging Face15VietAlphaLabs /zh-en-1925-mathematics-dictionaryVietAlpha Chinese–English Mathematics Dictionary 數學辭典, 1925 VietAlpha Lab · Website The VietAlpha Chinese–English Mathematics Dictionary is a machine-readable, verified digital edition of 數學辭典 (Mathematical Dictionary), compiled by 倪德基 (Ni Deji) and others and published by 中華書局 (Zhonghua Book Company) in 1925 (民國14年). Each headword is a Chinese mathematical term printed in 【…】 brackets, followed by its English equivalent, a subject marker, and a definition in Literary Chinese that… See the full description on the dataset page: https://huggingface.co/datasets/VietAlphaLabs/zh-en-1925-mathematics-dictionary.imagetranslation1K<n<10K0 likes86 downloads15d agoHugging Face16cjc0013 /ouroboros-autonomous-mathematics-process Ouroboros Autonomous Mathematics Process PUBLIC RELEASE This publication-ready package documents an experiment in autonomous mathematics directly applied to improve Ouroboros. The repository is intentionally staged as private so its owner can make it public after review. The PDFs and supporting text are already labeled PUBLIC RELEASE. Included OUROBOROS_AUTONOMOUS_MATHEMATICS_PROCESS_PUBLIC_RELEASE.pdf - the public-facing process white paper.… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-autonomous-mathematics-process.0 likes83 downloads15d agoHugging Face17Lots-of-LoRAs /task696_mmmlu_answer_generation_elementary_mathematics Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task696_mmmlu_answer_generation_elementary_mathematics Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task696_mmmlu_answer_generation_elementary_mathematics.texttext-generationn<1K0 likes80 downloads2y agoHugging Face18DataoceanAI /Competition-level_Mathematics_Physics_Reasoning_Corpus Title Competition-level Mathematics, Physics Reasoning Corpus Size 50.000+ text + multimodal competition level problems, each with step-by-step solutions and final answers Format Natural language explanations with multimodal samples include images (graphs, diagrams, etc.) Subject Mathematics/Physics and etc Labeling Details Question ID/Question Stem (Full text/content) /Subject/Question Type (Multiple Choice/Short Answer format… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/Competition-level_Mathematics_Physics_Reasoning_Corpus.0 likes78 downloads2y agoHugging Face19VietAlphaLabs /zh-en-1945-mathematics-dictionaryVietAlpha English–Chinese Mathematics Dictionary 數學名詞, 1945 VietAlpha Lab · Website The VietAlpha English–Chinese Mathematics Dictionary is a machine-readable, verified digital edition of 數學名詞 (Mathematical Terms), the English–Chinese mathematical terminology standard compiled by the National Institute for Compilation and Translation (國立編譯館) and promulgated by the Ministry of Education, in its 1945 printing. The printed work lists 3,426 numbered English terms with their standardized… See the full description on the dataset page: https://huggingface.co/datasets/VietAlphaLabs/zh-en-1945-mathematics-dictionary.translation1K<n<10K0 likes77 downloads15d agoHugging Face20perrabyte /crystal_mathematicsdocumentn<1K0 likes72 downloads1y agoHugging Face21Lots-of-LoRAs /task689_mmmlu_answer_generation_college_mathematics Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task689_mmmlu_answer_generation_college_mathematics Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task689_mmmlu_answer_generation_college_mathematics.texttext-generationn<1K0 likes68 downloads2y agoHugging Face22premio-ai /TheArabicPile_Mathematics The Arabic Pile Introduction: The Arabic Pile is a comprehensive dataset meticulously designed to parallel the structure of The Pile and The Nordic Pile. Focused on the Arabic language, the dataset encompasses a vast array of linguistic nuances, incorporating both Modern Standard Arabic (MSA) and various Levantine, North African, and Egyptian dialects. Tailored for the training and fine-tuning of large language models, the dataset consists of 13 subsets, each uniquely… See the full description on the dataset page: https://huggingface.co/datasets/premio-ai/TheArabicPile_Mathematics.texttext-generation100M<n<1B0 likes67 downloads3y agoHugging Face23prithivMLmods /Mathematics-Class10-Tnsb Mathematics-Class10-Tnsb This dataset contains scanned images from a Class 10 Mathematics textbook under the TNSB (Tamil Nadu State Board) curriculum. It is intended for educational machine learning tasks such as image-to-text (OCR), textbook digitization, or educational content understanding. Dataset Details Source: Tamil Nadu State Board Class 10 Mathematics textbook Task: Image-to-Text Language: English Split: train only Rows: 352 Format: Images only (scanned textbook… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Mathematics-Class10-Tnsb.imageimage-to-textn<1K0 likes64 downloads1y agoHugging Face24joey234 /mmlu-high_school_mathematics-neg-prepend Dataset Card for "mmlu-high_school_mathematics-neg-prepend" More Information needed textn<1K4 likes53 downloads3y agoHugging Face25timaeus /pile-dm_mathematics-elimination-slm-l1sae955text10K<n<100K0 likes49 downloads2y agoHugging Face26joey234 /mmlu-elementary_mathematics Dataset Card for "mmlu-elementary_mathematics" More Information needed textn<1K2 likes48 downloads3y agoHugging Face270xZee /dataset-CoT-Applied-Mathematics-824textn<1K1 likes42 downloads2y agoHugging Face28VINAY-UMRETHE /Physics-Chemistry-Mathematics-Progated Physics-Chemistry-Mathematics-Pro Copyright © 2026 Vinay Umrethe umrethevinay@gmail.com. This dataset is licensed under the Creative Commons Attribution 4.0 International License. imagevisual-question-answering1K<n<10K0 likes42 downloads2mo agoHugging Face29robworks-software /k12-mathematics-standards-aligned [!WARNING] Deprecated - use k12-mathematics-standards-expanded instead. This dataset is superseded: every input in this set also appears there, plus 366 more and two additional metadata columns. Nothing here is unique to it. It stays online so existing references keep resolving, but it will not be updated. New work should point at robworks-software/k12-mathematics-standards-expanded. K-12 Mathematics Standards (generated instruction data) 4,397 instruction/input/output records… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-mathematics-standards-aligned.texttext-generation1K<n<10K0 likes38 downloads2mo agoHugging Face30suolyer /pile_dm-mathematicstext1K<n<10K0 likes35 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.