Team Ai
22 results

STEM

TigreGotico /arabic-stem-lexicon Arabic Diacritized-Stem Lexicon An undiacritized Arabic surface form → its most frequent diacritized stem. Standard Arabic writes no short vowels, so anything that has to pronounce Arabic must first put them back. A neural diacritizer does that well on rare words, where inference is the only thing there is. On common words it is the wrong tool: which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.tabulartext-to-speech100K<n<1M0 likes3.8k downloads3mo agoHugging FaceEssentialAI /eai-taxonomy-stem-w-dclm 🔬 EAI-Taxonomy STEM w/ DCLM 🏆 Website | 🖥️ Code | 📖 Paper A high-quality STEM dataset curated from web data using taxonomy-based filtering, containing 1742 billion tokens of science, technology, engineering, and mathematics content. 🎯 Dataset Overview This dataset is part of the Essential-Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional STEM datasets that require… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/eai-taxonomy-stem-w-dclm.6 likes2.4k downloads1y agoHugging FaceLogics-MLLM /Logics-STEM-SFT-Dataset-Open-1.6M Logics-STEM-SFT-Dataset-2.2M 📰 News [2026.01.05]🔥 Release of our Techinical Report. [2026.01.05]🔥 Release the first version of Logics-STEM-8B-SFT, Logics-STEM-8B-RL, /Logics-STEM-SFT-Dataset-Open-1.6M. Overview What is this dataset? Logics-STEM-SFT-Dataset-2.2M is a curated long Chain-of-Thought (CoT) SFT dataset for STEM reasoning, built on top of high-quality open-source data and enhanced through a rigorous curation and distillation… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M.text1M<n<10M33 likes2.1k downloads9mo agoHugging FaceIsaac105 /melodix-stemsaudio1M<n<10M0 likes1.4k downloads5d agoHugging Facestemdataset /STEM STEM Dataset 📃 [Paper] • 💻 [Github] • 🤗 [Dataset] • 🏆 [Leaderboard] • 📽 [Slides] • 📋 [Poster] This dataset is proposed in the ICLR 2024 paper: Measuring Vision-Language STEM Skills of Neural Models. We introduce a new challenge to test the STEM skills of neural models. The problems in the real world often require solutions, combining knowledge from STEM (science, technology, engineering, and math). Unlike existing datasets, our dataset requires the understanding of… See the full description on the dataset page: https://huggingface.co/datasets/stemdataset/STEM.text1M<n<10M6 likes1.2k downloads2y agoHugging Faceguhanthecagen /MASTER-STEM-DATASET0 likes955 downloads8d agoHugging Face