datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
maple-collections-hackathon
Maple Bank Collections Hackathon dataset (fully synthetic)
Dataset for the CIBC Collections Hackathon build phase. One fictional bank ("Maple Bank"), 1,000,000 customers
(1,020,000 CRM records), October 2016 to September 2026, snapshot date 2026-09-28. Every person, account, call,
recording and document is synthetic.
Files
File
Size
What
maple_collections_release.zip
see file list
Start here. 31 tables (CSV + Parquet), transcripts (JSON), policy… See the full description on the dataset page: https://huggingface.co/datasets/nuxsh/maple-collections-hackathon.maple-preview-cuda-benchmarks
Maple Preview TQ2_0 CUDA Benchmarks
Reproducibility data for the TQ2_0 CUDA patches in
PascalAI2024/maple-preview-windows-cuda.
This repository contains benchmark data, patch files, hashes, and raw validation
evidence. It does not duplicate the Maple model weights.
Result
The fresh local A/B/B/A validation on an RTX 4080 SUPER reproduced the fused-MMQ
prompt-processing gain:
Variant
pp512 mean
pp512 median
tg128 mean
tg128 median
Correctness
MMQ enabled… See the full description on the dataset page: https://huggingface.co/datasets/x0me/maple-preview-cuda-benchmarks.MetaRAG_Cross-Issue_OSSQA
MetaRAG Cross-Issue OSSQA
Dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA
MetaRAG Cross-Issue OSSQA is an English open-source software issue question-answering and retrieval benchmark. Each example asks a question grounded in one GitHub issue and requires evidence from a related issue. The data contains explicit cross-issue references and a three-document silver evidence path.
Dataset configurations
Configuration
Splits
Rows… See the full description on the dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA.short_selling
Short Selling
Data Notice: This dataset provides academic research access with a 6-month data lag.
For real-time data access, please visit sov.ai to subscribe.
For market insights and additional subscription options, check out our newsletter at blog.sov.ai.
from datasets import load_dataset
df_over_shorted = load_dataset("sovai/short_selling", split="train").to_pandas().set_index(["ticker","date"])
Data is updated weekly as data arrives after market close US-EST time.
Tutorials… See the full description on the dataset page: https://huggingface.co/datasets/MapleLeavesKrish/short_selling.
