Team Ai
21 results

map

m-a-p /FineFineWeb FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb.tabulartext-classification1B<n<10B199 likes4.4m downloads2y agoHugging FaceSKPark1 /ngii-map-full-light ngii-map-full-light Light point/line extract from NGII 1/1000 topographic data for Korea. Not for shipping into GitHub — use this Hugging Face dataset instead. CRS Korea_2000_Central_Belt_2010 projected meters [x, y] Layers (per region under by_region/<region>/) Layer Description C023 poles (전주/통신주) C022 lights (가로등·보안등) A002 roads (도로 중심선) B001_tiny building footprints <25 m² as centroids B002 lines (구분/재질 라인) Also:… See the full description on the dataset page: https://huggingface.co/datasets/SKPark1/ngii-map-full-light.geospatialother1M<n<10M0 likes303k downloads28d agoHugging Facem-a-p /PIN-200M PIN-200M A mini version of "PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents" Paper: https://arxiv.org/abs/2406.13923 This dataset contains around 200M samples in PIN format, with around 312 TB storage. 🚀 News [ 2025.09.22 ] !NEW! 🔥 We have completed the final version of the PIN-200M dataset and conducted some simple statistics on it. [ 2024.12.06 ] !NEW! 🔥 We have updated the quality signals, enabling a swift assessment of whether a sample meets… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/PIN-200M.text10K<n<100K28 likes196k downloads6mo agoHugging Facem-a-p /FineFineWeb-sample FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.tabulartext-classification100M<n<1B4 likes56k downloads2y agoHugging Facem-a-p /Matrix Matrix An open-source pretraining dataset containing 4690 billion tokens, this bilingual dataset with both English and Chinese texts is used for training neo models. Dataset Composition The dataset consists of several components, each originating from different sources and serving various purposes in language modeling and processing. Below is a brief overview of each component: Common Crawl Extracts from the Common Crawl project, featuring a rich diversity of… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/Matrix.texttext-generation1B<n<10B176 likes34k downloads2y agoHugging Facem-a-p /CodeFeedback-Filtered-Instruction OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement [🏠Homepage] | [🛠️Code] OpenCodeInterpreter OpenCodeInterpreter is a family of open-source code generation systems designed to bridge the gap between large language models and advanced proprietary systems like the GPT-4 Code Interpreter. It significantly advances code generation capabilities by integrating execution and iterative refinement functionalities. For further information and… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/CodeFeedback-Filtered-Instruction.textquestion-answering100K<n<1M209 likes25k downloads3y agoHugging Face