Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Maple222 /llmtcl ⚡ LitGPT 20+ high-performance LLMs with recipes to pretrain, finetune, and deploy at scale. ✅ From scratch implementations ✅ No abstractions ✅ Beginner friendly ✅ Flash attention ✅ FSDP ✅ LoRA, QLoRA, Adapter ✅ Reduce GPU memory (fp4/8/16/32) ✅ 1-1000+ GPUs/TPUs ✅ 20+ LLMs Quick start • Models • Finetune • Deploy • All workflows • Features • Recipes (YAML) • Lightning AI • Tutorials… See the full description on the dataset page: https://huggingface.co/datasets/Maple222/llmtcl.text1 likes2.4k downloads11mo agoHugging Face02MAPLE-WestLake-AIGC /OpenstoryPlusPlus Openstory++: A Large-scale Dataset and Benchmark for Instance-aware Open-domain Visual Storytelling We introduce OpenStory++, a large-scale open-domain dataset contains focusing on enabling MLLMs to perform storytelling generation tasks. related resorcce paper: https://arxiv.org/abs/2408.03695 code: https://github.com/YeLuoSuiYou/openstorypp News 2024/7/31 We have reorganized and distributed the high-quality subset and released most of the story data collected… See the full description on the dataset page: https://huggingface.co/datasets/MAPLE-WestLake-AIGC/OpenstoryPlusPlus.image100K<n<1M5 likes726 downloads2y agoHugging Face03maplebb /UniREdit-Data-100KUniREditBench: A Unified Reasoning-based Image Editing Benchmark text10K<n<100K3 likes670 downloads11mo agoHugging Face04tudor-iustin22 /maple Overview Maple is an open-source full-stack code dataset developed and released by Tudor Iustin. It is designed to support code generation, web development, supervised fine-tuning, instruction tuning, post-training, dataset research, and evaluation workflows for code-capable AI systems. Maple contains 16,000 full-stack code samples totaling approximately 102 million tokens. It focuses on realistic software-building tasks, including web applications, product interfaces… See the full description on the dataset page: https://huggingface.co/datasets/tudor-iustin22/maple.texttext-generation10K<n<100K0 likes128 downloads3mo agoHugging Face05x0me /maple-preview-cuda-benchmarks Maple Preview TQ2_0 CUDA Benchmarks Reproducibility data for the TQ2_0 CUDA patches in PascalAI2024/maple-preview-windows-cuda. This repository contains benchmark data, patch files, hashes, and raw validation evidence. It does not duplicate the Maple model weights. Result The fresh local A/B/B/A validation on an RTX 4080 SUPER reproduced the fused-MMQ prompt-processing gain: Variant pp512 mean pp512 median tg128 mean tg128 median Correctness MMQ enabled… See the full description on the dataset page: https://huggingface.co/datasets/x0me/maple-preview-cuda-benchmarks.tabularn<1K0 likes106 downloads2mo agoHugging Face06ayang903 /maple MAPLE (Bill Summarization, Tagging, Explanation) In this project, we generate summaries and category tags for of Massachusetts bills for MAPLE Platform. The goal is to simplify the legal language and content to make it comprehensible for a broader audience (9th-grade comprehension level) by exploring different ML and LLM services. This repository contains a pipeline from taking bills from Massachusetts legislature, generating summaries and category tags leveraging different the… See the full description on the dataset page: https://huggingface.co/datasets/ayang903/maple.textsummarization1K<n<10K0 likes76 downloads3y agoHugging Face07irotem98 /maplestory_characters_hdimage10K<n<100K2 likes58 downloads2y agoHugging Face08msw-ai-tf /maplestory-worlds-creator-qa MapleStory Worlds Creator QA Synthetic question-answer dataset built from the official MapleStory Worlds Creator Center documentation. Questions are generated to be self-contained and grounded in the source docs; answers avoid source/meta references so they read like an expert explanation. Some QA pairs are composed from multiple related documents (see combo_sources). Parallel Korean/English. Intended for instruction tuning, QA, and retrieval. Composition… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-qa.textquestion-answering100K<n<1M2 likes46 downloads3mo agoHugging Face09prdeepakbabu /maple-personas MAPLE-Personas: A Benchmark for Evaluating Personalized Conversational AI A dataset for evaluating how well conversational AI systems learn and apply user preferences from natural dialogue. This benchmark accompanies the MAPLE (Memory-Adaptive Personalized LEarning) framework. Dataset Description This dataset tests an AI assistant's ability to implicitly learn user traits from conversation context and apply that knowledge to personalize responses to open-ended… See the full description on the dataset page: https://huggingface.co/datasets/prdeepakbabu/maple-personas.texttext-generation1K<n<10K0 likes45 downloads8mo agoHugging Face10Davd-b01 /maple-analyst-cap-sft-data maple-analyst-cap-sft-data Dataset de SFT para fine-tune de maple-analyst-cap-bf16 (Qwen3.5-MoE 20.2B ternario). 4,956 trazas de razonamiento (pseudothinking + answer) en formato TC (ThinkingCap). Composición Fuente Filas thinkingcap (curriculum, trazas bigbang) 1,782 openmle-condensed (FrontisAI OpenMLE-SFT-Traces, condensadas con distiller LFM2.5-2.6B q8_0) 702 bigbang_mmlu 508 bigbang_bbh 441 hermes_function_calling 360 aya_dataset 342… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/maple-analyst-cap-sft-data.texttext-generation1K<n<10K0 likes44 downloads2mo agoHugging Face11MaplesWCT /DynaSolidGeo-SamplePaper: https://arxiv.org/abs/2510.22340 Github Repo: https://github.com/ChangtiWu/DynaSolidGeo In the "Appendix.E: DynaSolidGeo as a Training Dataset" in our paper, we sample K = 10 batches of instances using random seeds from 0 to 9, resulting in a total of 5,030 samples. These samples are divided into a training set (3,627 samples), a validation set (403 samples), and a test set (1,000 samples). text1K<n<10K1 likes40 downloads9mo agoHugging Face12msw-ai-tf /maplestory-worlds-creator-docs MapleStory Worlds Creator Center Documentation A curated dataset built from the official documentation of the MapleStory Worlds Creator Center. It is a parallel Korean/English documentation corpus intended for RAG, search, embeddings, and domain language-model training. The dataset covers all three Creator Center content types — guide documents (doc), API Reference (api), and resources (res). Composition Document counts by type and language: type Description… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-docs.texttext-generation1K<n<10K2 likes26 downloads4mo agoHugging Face13MapleBi /MetaRAG_Cross-Issue_OSSQA MetaRAG Cross-Issue OSSQA Dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA MetaRAG Cross-Issue OSSQA is an English open-source software issue question-answering and retrieval benchmark. Each example asks a question grounded in one GitHub issue and requires evidence from a related issue. The data contains explicit cross-issue references and a three-document silver evidence path. Dataset configurations Configuration Splits Rows… See the full description on the dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA.tabularquestion-answering10K<n<100K0 likes25 downloads2mo agoHugging Face14msw-ai-tf /maplestory-worlds-creator-code-instruct MapleStory Worlds Creator Code (mlua) Instruction-style code dataset for mlua, the scripting language of MapleStory Worlds. Built from the official Creator Center example code: each example is grounded in its source document and paired with a natural-language task, reasoning, a self-contained explanation, and commented mlua code. Intended to teach LLMs to write mlua game scripts. The example code is preserved from the official source (a code-preservation check rejects any record… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-code-instruct.texttext-generation1K<n<10K2 likes24 downloads3mo agoHugging Face15maple138 /naver-economy-news2stocktext1K<n<10K0 likes23 downloads1mo agoHugging Face16maplerxyz1 /yamltext10K<n<100K0 likes18 downloads2y agoHugging Face17zzhb /maple-umi-datatext0 likes17 downloads2mo agoHugging Face18MapleBi /CoastAdapt-KB CoastAdapt-KB Zero-Shot Hierarchical Events Dataset Dataset Summary This dataset is prepared from consolidated climate change solution extraction results. It is designed for zero-shot hierarchical multi-label text classification over climate adaptation and mitigation event records. Each example contains a natural-language input text plus one or more hierarchical label paths. The labels organize climate-related solution details into a taxonomy with phase, domain, and… See the full description on the dataset page: https://huggingface.co/datasets/MapleBi/CoastAdapt-KB.texttext-classification1K<n<10K0 likes15 downloads5mo agoHugging Face19Maple222 /pickplacenewtext1K<n<10K0 likes11 downloads4mo agoHugging Face20canxp-ai /maplept2-coder-corpustext1K<n<10K0 likes10 downloads4mo agoHugging Face21canxp-ai /maplept2-reasoning-corpustextn<1K0 likes10 downloads4mo agoHugging Face22mpstorys /maplestory-resource-index MapleStory Resource Index A structured, searchable, and deduplicated metadata index for useful MapleStory resources. The dataset covers six active series: MapleStory MapleStory Classic MapleStory M MapleStory Worlds MapleStory N MapleStory Idle Project website This dataset is maintained by MPStorys, a MapleStory resource discovery platform. Dataset contents The current export contains 93 resource records. Fields may include: Resource ID, name… See the full description on the dataset page: https://huggingface.co/datasets/mpstorys/maplestory-resource-index.textn<1K0 likes10 downloads3mo agoHugging Face23MapleLeavesKrish /short_selling Short Selling Data Notice: This dataset provides academic research access with a 6-month data lag. For real-time data access, please visit sov.ai to subscribe. For market insights and additional subscription options, check out our newsletter at blog.sov.ai. from datasets import load_dataset df_over_shorted = load_dataset("sovai/short_selling", split="train").to_pandas().set_index(["ticker","date"]) Data is updated weekly as data arrives after market close US-EST time. Tutorials… See the full description on the dataset page: https://huggingface.co/datasets/MapleLeavesKrish/short_selling.tabular1M<n<10M0 likes8 downloads9mo agoHugging Face24maple-matrix /sn38r5-u70-subtextn<1K0 likes8 downloads2mo agoHugging Face25supergoose /buzz_sources_410_mapletextn<1K0 likes7 downloads2y agoHugging Face26shyanchen /mapleautofarm MapleAutoFarm · 冒险岛自动打怪 Python 版 仅用于单机 / 离线 / 个人学习,不用于联网游戏。 Python 实现的冒险岛自动巡逻打怪工具,带 Tkinter 可视化面板,预留 OpenCV 视觉找怪能力。 功能 方向键移动、左右巡逻 跳跃键(默认 A)、攻击键(默认 D) 基础自动巡逻打怪 视觉找怪骨架(OpenCV 模板匹配) 可视化面板 + 日志 / 状态 / 循环次数 全局快捷键启动 / 停止 / 退出 宠物自动药水交给游戏内宠物,无需脚本处理 目录结构 MapleAutoFarm\ ├─ main.py 主程序 + 可视化面板 ├─ bot.py 自动打怪状态机 ├─ vision.py OpenCV 图像识别模块 ├─ config.json 配置文件(首次保存后生成) ├─ requirements.txt 依赖 ├─ install.bat 一键装环境… See the full description on the dataset page: https://huggingface.co/datasets/shyanchen/mapleautofarm.textn<1K1 likes7 downloads2mo agoHugging Face27MapleSage /msxgpt-dataset msxgpt Description This dataset, "msxgpt," is designed for training the GPT-3.5-turbo/ GPT-4 based language model for a task. The data consists of JSON lines, each representing an individual example for the model. The dataset has been created with an emphasis on encoding, which is pivotal to the functionality of Memory Features, Security, and API Endpoints. It is designed to process and store documents from various data sources continuously, using incoming webhooks to the… See the full description on the dataset page: https://huggingface.co/datasets/MapleSage/msxgpt-dataset.text1K<n<10K0 likes6 downloads3y agoHugging Face28canxp-ai /maplept-reasoning-corpustextn<1K0 likes6 downloads4mo agoHugging Face29maple-matrix /sn38r3-u170-subtextn<1K0 likes6 downloads3mo agoHugging Face30maple-matrix /sn38r4-u70-subtextn<1K0 likes6 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.