Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Maple728 /Time-300B Dataset Card for Time-300B This repository contains the Time-300B dataset of the paper Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts. For details on how to use this dataset, please visit our GitHub page. time-series-forecastingn>1T34 likes7.8k downloads2y agoHugging Face02Maple222 /llmtcl ⚡ LitGPT 20+ high-performance LLMs with recipes to pretrain, finetune, and deploy at scale. ✅ From scratch implementations ✅ No abstractions ✅ Beginner friendly ✅ Flash attention ✅ FSDP ✅ LoRA, QLoRA, Adapter ✅ Reduce GPU memory (fp4/8/16/32) ✅ 1-1000+ GPUs/TPUs ✅ 20+ LLMs Quick start • Models • Finetune • Deploy • All workflows • Features • Recipes (YAML) • Lightning AI • Tutorials… See the full description on the dataset page: https://huggingface.co/datasets/Maple222/llmtcl.text1 likes2.4k downloads11mo agoHugging Face03kiendo82086 /maple0 likes1.5k downloads59m agoHugging Face04kai-02 /MAPLE MAPLE: Multi-Aspect Full-Paper Scientific Retrieval Benchmark MAPLE is an expert-validated benchmark for multi-aspect full-paper retrieval. It contains 2,095 fine-grained queries derived from 210 recent machine learning papers, together with a retrieval corpus of 73,973 candidate papers. Each target paper is paired with multiple queries grounded in textual or multimodal evidence and covering different aspects of the paper, including motivation, method, and experimental findings.… See the full description on the dataset page: https://huggingface.co/datasets/kai-02/MAPLE.text-retrieval10K<n<100K2 likes1.2k downloads2mo agoHugging Face05maplebb /UniREditBench-Results0 likes766 downloads9mo agoHugging Face06MAPLE-WestLake-AIGC /OpenstoryPlusPlus Openstory++: A Large-scale Dataset and Benchmark for Instance-aware Open-domain Visual Storytelling We introduce OpenStory++, a large-scale open-domain dataset contains focusing on enabling MLLMs to perform storytelling generation tasks. related resorcce paper: https://arxiv.org/abs/2408.03695 code: https://github.com/YeLuoSuiYou/openstorypp News 2024/7/31 We have reorganized and distributed the high-quality subset and released most of the story data collected… See the full description on the dataset page: https://huggingface.co/datasets/MAPLE-WestLake-AIGC/OpenstoryPlusPlus.image100K<n<1M5 likes726 downloads2y agoHugging Face07maplebb /UniREdit-Data-100KUniREditBench: A Unified Reasoning-based Image Editing Benchmark text10K<n<100K3 likes670 downloads11mo agoHugging Face08zhenghuayu /MAPLE_40K MAPLE_40K — 1024 × 1024 40,000 synthetic object-editing pairs derived from OBJect-3DIT / allenai/object-edit: 20,000 rotation and 20,000 translation pairs. All images are 1024 × 1024. This version replaces the previous 256 × 256 release on main. Source and target RGB images were enhanced from native 256 × 256 using SeedVR2-7B FP16 weights, fixed seed 42, LAB color correction, and three low-frequency consistency iterations. This is model-based super-resolution. Fine details are… See the full description on the dataset page: https://huggingface.co/datasets/zhenghuayu/MAPLE_40K.10K<n<100K0 likes451 downloads2d agoHugging Face09nuxsh /maple-collections-hackathon Maple Bank Collections Hackathon dataset (fully synthetic) Dataset for the CIBC Collections Hackathon build phase. One fictional bank ("Maple Bank"), 1,000,000 customers (1,020,000 CRM records), October 2016 to September 2026, snapshot date 2026-09-28. Every person, account, call, recording and document is synthetic. Files File Size What maple_collections_release.zip see file list Start here. 31 tables (CSV + Parquet), transcripts (JSON), policy… See the full description on the dataset page: https://huggingface.co/datasets/nuxsh/maple-collections-hackathon.tabular1M<n<10M2 likes306 downloads8d agoHugging Face10maplebb /UniREditBenchUniREditBench: A Unified Reasoning-based Image Editing Benchmark image1K<n<10K3 likes139 downloads11mo agoHugging Face11tudor-iustin22 /maple Overview Maple is an open-source full-stack code dataset developed and released by Tudor Iustin. It is designed to support code generation, web development, supervised fine-tuning, instruction tuning, post-training, dataset research, and evaluation workflows for code-capable AI systems. Maple contains 16,000 full-stack code samples totaling approximately 102 million tokens. It focuses on realistic software-building tasks, including web applications, product interfaces… See the full description on the dataset page: https://huggingface.co/datasets/tudor-iustin22/maple.texttext-generation10K<n<100K0 likes128 downloads3mo agoHugging Face12anchovy /maple728-time_300B Dataset Card for Time-300B This repository contains the Time-300B dataset of the paper Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts. For details on how to use this dataset, please visit our GitHub page. time-series-forecastingn>1T0 likes126 downloads2y agoHugging Face13x0me /maple-preview-cuda-benchmarks Maple Preview TQ2_0 CUDA Benchmarks Reproducibility data for the TQ2_0 CUDA patches in PascalAI2024/maple-preview-windows-cuda. This repository contains benchmark data, patch files, hashes, and raw validation evidence. It does not duplicate the Maple model weights. Result The fresh local A/B/B/A validation on an RTX 4080 SUPER reproduced the fused-MMQ prompt-processing gain: Variant pp512 mean pp512 median tg128 mean tg128 median Correctness MMQ enabled… See the full description on the dataset page: https://huggingface.co/datasets/x0me/maple-preview-cuda-benchmarks.tabularn<1K0 likes106 downloads2mo agoHugging Face14bartholomort /MAPLE-Lua-Corpus0 likes93 downloads5mo agoHugging Face15ayang903 /maple MAPLE (Bill Summarization, Tagging, Explanation) In this project, we generate summaries and category tags for of Massachusetts bills for MAPLE Platform. The goal is to simplify the legal language and content to make it comprehensible for a broader audience (9th-grade comprehension level) by exploring different ML and LLM services. This repository contains a pipeline from taking bills from Massachusetts legislature, generating summaries and category tags leveraging different the… See the full description on the dataset page: https://huggingface.co/datasets/ayang903/maple.textsummarization1K<n<10K0 likes76 downloads3y agoHugging Face16maplestang /escarpment-lab-data Escarpment Retreat Evolution Teaching Dataset This public dataset contains the precomputed display assets for an educational web platform about fluvial-erosion-driven escarpment retreat. Contents 81 MATLAB/TopoToolbox/TTLEM parameter scenarios 161 time steps per scenario from 0 to 32 Myr Corrected top-down plan views Fixed-perspective three-dimensional views Basin masks, river-profile data, knickpoints, and scenario metadata The web release contains 39,368 files… See the full description on the dataset page: https://huggingface.co/datasets/maplestang/escarpment-lab-data.image10K<n<100K0 likes67 downloads2mo agoHugging Face17lihVerma /MAPLE-bench MAPLE Benchmark Test Splits This repository contains the test splits for the MAPLE benchmark introduced in the paper MAPLE: Modality-Aware Post-training and Learning Ecosystem (https://arxiv.org/pdf/2602.11596). The benchmark is designed for modality-aware multimodal evaluation under different required-signal settings, where each sample is annotated with the minimal modality subset needed to solve the task. Dataset Overview MAPLE-bench evaluates multimodal reasoning… See the full description on the dataset page: https://huggingface.co/datasets/lihVerma/MAPLE-bench.visual-question-answering1K<n<10K0 likes63 downloads6mo agoHugging Face18irotem98 /maplestory_characters_hdimage10K<n<100K2 likes58 downloads2y agoHugging Face19Mapleyuchen /MME-RealWorld 2024.11.14 🌟 MME-RealWorld now has a lite version (50 samples per task) for inference acceleration, which is also supported by VLMEvalKit and Lmms-eval. 2024.10.27 🌟 LLaVA-OV currently ranks first on our leaderboard, but its overall accuracy remains below 55%, see our leaderboard for the detail. 2024.09.03 🌟 MME-RealWorld is now supported in the VLMEvalKit and Lmms-eval repository, enabling one-click evaluation—give it a try!" 2024.08.20 🌟 We are very proud to launch MME-RealWorld, which… See the full description on the dataset page: https://huggingface.co/datasets/Mapleyuchen/MME-RealWorld.multiple-choice100B<n<1T0 likes57 downloads7mo agoHugging Face20kylealexander6644 /maple0 likes48 downloads1h agoHugging Face21msw-ai-tf /maplestory-worlds-creator-qa MapleStory Worlds Creator QA Synthetic question-answer dataset built from the official MapleStory Worlds Creator Center documentation. Questions are generated to be self-contained and grounded in the source docs; answers avoid source/meta references so they read like an expert explanation. Some QA pairs are composed from multiple related documents (see combo_sources). Parallel Korean/English. Intended for instruction tuning, QA, and retrieval. Composition… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-qa.textquestion-answering100K<n<1M2 likes46 downloads3mo agoHugging Face22prdeepakbabu /maple-personas MAPLE-Personas: A Benchmark for Evaluating Personalized Conversational AI A dataset for evaluating how well conversational AI systems learn and apply user preferences from natural dialogue. This benchmark accompanies the MAPLE (Memory-Adaptive Personalized LEarning) framework. Dataset Description This dataset tests an AI assistant's ability to implicitly learn user traits from conversation context and apply that knowledge to personalize responses to open-ended… See the full description on the dataset page: https://huggingface.co/datasets/prdeepakbabu/maple-personas.texttext-generation1K<n<10K0 likes45 downloads8mo agoHugging Face23Davd-b01 /maple-analyst-cap-sft-data maple-analyst-cap-sft-data Dataset de SFT para fine-tune de maple-analyst-cap-bf16 (Qwen3.5-MoE 20.2B ternario). 4,956 trazas de razonamiento (pseudothinking + answer) en formato TC (ThinkingCap). Composición Fuente Filas thinkingcap (curriculum, trazas bigbang) 1,782 openmle-condensed (FrontisAI OpenMLE-SFT-Traces, condensadas con distiller LFM2.5-2.6B q8_0) 702 bigbang_mmlu 508 bigbang_bbh 441 hermes_function_calling 360 aya_dataset 342… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/maple-analyst-cap-sft-data.texttext-generation1K<n<10K0 likes44 downloads2mo agoHugging Face24MaplesWCT /DynaSolidGeo-SamplePaper: https://arxiv.org/abs/2510.22340 Github Repo: https://github.com/ChangtiWu/DynaSolidGeo In the "Appendix.E: DynaSolidGeo as a Training Dataset" in our paper, we sample K = 10 batches of instances using random seeds from 0 to 9, resulting in a total of 5,030 samples. These samples are divided into a training set (3,627 samples), a validation set (403 samples), and a test set (1,000 samples). text1K<n<10K1 likes40 downloads9mo agoHugging Face25SOTAagi2030 /Maple-Routing-Observations0 likes38 downloads15d agoHugging Face26verify-ppt /marin-starcoderdata_maple0 likes37 downloads6mo agoHugging Face27FabricAI /maple Overview Maple is an open-source full-stack code dataset developed and released by Fabric AI. It is designed to support code generation, web development, supervised fine-tuning, instruction tuning, post-training, dataset research, and evaluation workflows for code-capable AI systems. Maple contains 16,000 full-stack code samples totaling approximately 102 million tokens. It focuses on realistic software-building tasks, including web applications, product interfaces, dashboards… See the full description on the dataset page: https://huggingface.co/datasets/FabricAI/maple.text-generation10K<n<100K0 likes33 downloads2mo agoHugging Face28lastbattle /maplestory_captchaA huge collection of English MapleStory's captcha text in jpg that I have collected over the years. ENJOY!! It us used by pre-Big Bang MapleStory, throughout the game from Lie-Detector (anti-macro item), logins, to NPC conversations. Up till version 190 when they have switched using Runes (Up, Down, Left, Right arrow keys) for most of the time for detection of macros and bots. These images are not labelled, I'm releasing this for anyone that wants the dataset to be able to train a model… See the full description on the dataset page: https://huggingface.co/datasets/lastbattle/maplestory_captcha.imagetext-classification1K<n<10K3 likes28 downloads2y agoHugging Face29maplebridge /canada-china-trade Canada-China B2B Trade Dataset Dataset Description A curated dataset of Canada-China bilateral trade statistics, commodity breakdowns, provincial data, and B2B sourcing knowledge for use in AI/LLM research and applications. Maintained by: MapleBridge.io — AI-powered B2B matching platform for Canada-China trade. Dataset Contents File Description Rows canada_china_trade_annual.csv Annual bilateral trade volume 2015-2024 (CAD billions) 10… See the full description on the dataset page: https://huggingface.co/datasets/maplebridge/canada-china-trade.text-generationn<1K0 likes28 downloads7mo agoHugging Face30kemi-adekanbi /MapleStory_Monsters_Datasetimagen<1K1 likes27 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.