datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineFineWeb-sample
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.MAPBench-V2For more details, please check our project page.
Paper: https://arxiv.org/abs/2601.05432
Repository: https://github.com/AMAP-ML/Thinking-with-Map
war3-maps
Warcraft III Community Map Archive
This public dataset preserves community-created Warcraft III maps and campaigns
for interoperability testing, search, research, and long-term access. Files are
deduplicated by SHA-256. Titles and other metadata are extracted with
war3-manager where the format permits.
Search and download individual maps: https://war3-archive.github.io/war3-maps/
Source and issue tracker: https://github.com/war3-archive/war3-maps
Layout… See the full description on the dataset page: https://huggingface.co/datasets/magicwenli/war3-maps.FineFineWeb-test
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-test.MAPBench-V1For more details, please check our project page.
Paper: https://arxiv.org/abs/2601.05432
Repository: https://github.com/AMAP-ML/Thinking-with-Map
GSFineFineWeb-validation
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-validation.EMOmaple-preview-cuda-benchmarks
Maple Preview TQ2_0 CUDA Benchmarks
Reproducibility data for the TQ2_0 CUDA patches in
PascalAI2024/maple-preview-windows-cuda.
This repository contains benchmark data, patch files, hashes, and raw validation
evidence. It does not duplicate the Maple model weights.
Result
The fresh local A/B/B/A validation on an RTX 4080 SUPER reproduced the fused-MMQ
prompt-processing gain:
Variant
pp512 mean
pp512 median
tg128 mean
tg128 median
Correctness
MMQ enabled… See the full description on the dataset page: https://huggingface.co/datasets/x0me/maple-preview-cuda-benchmarks.m-a-p-FineFineWeb-sample
Unofficial m-a-p/FineFineWeb Sample
This dataset is a processed, lightweight sample of the original m-a-p/FineFineWeb, a comprehensive corpus designed for fine-grained domain web text studies.
Sampling Methodology
To create this subset, the following processing steps were taken:
Selection: 100 random .jsonl files were chosen from the original dataset.
Extraction: 10,000 rows were downloaded per selected file.
Processing: The extracted rows were combined and… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/m-a-p-FineFineWeb-sample.Chords1217OpenO1_SFT_ultra_BoN_positvie_reward_v3_N-sampleMapEval-API
MapEval-API
MapEval-API is created using MapQaTor.
Usage
from datasets import load_dataset
# Load dataset
ds = load_dataset("MapEval/MapEval-API", name="benchmark")
# Generate better prompts
for item in ds["test"]:
# Start with a clear task description
prompt = (
"You are a highly intelligent assistant. "
"Answer the multiple-choice question by selecting the correct option.\n\n"
"Question:\n" + item["question"] + "\n\n"… See the full description on the dataset page: https://huggingface.co/datasets/MapEval/MapEval-API.MapEval-Textual
MapEval-Textual
MapEval-Textual is created using MapQaTor.
Usage
from datasets import load_dataset
# Load dataset
ds = load_dataset("MapEval/MapEval-Textual", name="benchmark")
# Generate better prompts
for item in ds["test"]:
# Start with a clear task description
prompt = (
"You are a highly intelligent assistant. "
"Based on the given context, answer the multiple-choice question by selecting the correct option.\n\n"
"Context:\n" +… See the full description on the dataset page: https://huggingface.co/datasets/MapEval/MapEval-Textual.repro-time-series-saliency-maps-explaining-models-across-multiple-domains-traces
Agent traces
Agent sessions published from a Trackio Logbook.
ai-datacenter-map-2026
ai-datacenter-map-2026
AI datacenter locations and capacity. 20 facilities
Records: 24 | Updated: 2026-10-07
API: GET https://api.legion-api.com/datacenter-map
Bundle: gemmo.gumroad.com/l/mdevxu
Gated — auto-approved.
MapRoute-Baselineucf-anomaly-detection-mapped
UCF-Crime: Precomputed I3D Features with Temporal Annotations
This dataset provides pre-extracted 1024-dimensional I3D RGB features along with frame-level temporal anomaly labels for videos from the UCF-Crime dataset.
Dataset Characteristics
Features
1024-dimensional I3D RGB feature vectors
Extracted from 64 uniformly sampled frames per video
Feature tensor shape: [64, 1024]
Temporal Annotations
Mapped from original anomaly intervals… See the full description on the dataset page: https://huggingface.co/datasets/Rahima411/ucf-anomaly-detection-mapped.MetaRAG_Cross-Issue_OSSQA
MetaRAG Cross-Issue OSSQA
Dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA
MetaRAG Cross-Issue OSSQA is an English open-source software issue question-answering and retrieval benchmark. Each example asks a question grounded in one GitHub issue and requires evidence from a related issue. The data contains explicit cross-issue references and a three-document silver evidence path.
Dataset configurations
Configuration
Splits
Rows… See the full description on the dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA.thirawat-mapper-demo-indexadaption-ip-to-port-mapping-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
ip_to_port_mapping
This dataset consists of prompt-completion pairs mapping IPv4 addresses to their associated network port numbers. The samples include both public and private IP addresses linked to common service ports such as 80, 53, and 8080. It is structured for tasks involving network configuration prediction or port identification based on IP inputs.
Dataset size
There are 138… See the full description on the dataset page: https://huggingface.co/datasets/NikitaSirotkin/adaption-ip-to-port-mapping-v1.adaption-ip-to-port-mapping
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
ip_to_port_mapping
This dataset consists of prompt-completion pairs mapping IPv4 addresses to their associated network port numbers. The samples include both public and private IP addresses linked to common service ports such as 80, 53, and 8080. It is structured for tasks involving network configuration prediction or port identification based on IP inputs.
Dataset size
There are 138… See the full description on the dataset page: https://huggingface.co/datasets/NikitaSirotkin/adaption-ip-to-port-mapping.mapform-gbm-dataFineLeanCorpusProofinstitutionsSRB_instance_key_mappingmedical-entity-code-mapper-indices
FAISS Indices Directory
This directory contains the pre-built FAISS indices for all medical ontologies.
Included Indices
indices/
├── icd10_bge_m3/ # ICD-10-CM diagnosis codes (4.4GB)
│ ├── faiss.index
│ └── metadata.pkl
├── snomed_bge_m3/ # SNOMED CT clinical concepts (3.1GB)
│ ├── faiss.index
│ └── metadata.pkl
├── loinc_bge_m3/ # LOINC laboratory codes (896MB)
│ ├── faiss.index
│ └── metadata.pkl
├── rxnorm_bge_m3/ # RxNorm medication… See the full description on the dataset page: https://huggingface.co/datasets/docdailey/medical-entity-code-mapper-indices.drugbank_drug_target_label_mapping_amino_acid_pairobstacle-map-annotations-json
Obstacle Map Annotations Dataset
Dataset annotating obstacles in a 2D environment
for robot navigation and mapping.
MapEval-Textual
MapEval-Textual
MapEval-Textual is created using MapQaTor.
Usage
from datasets import load_dataset
# Load dataset
ds = load_dataset("MapEval/MapEval-Textual", name="benchmark")
# Generate better prompts
for item in ds["test"]:
# Start with a clear task description
prompt = (
"You are a highly intelligent assistant. "
"Based on the given context, answer the multiple-choice question by selecting the correct option.\n\n"
"Context:\n"… See the full description on the dataset page: https://huggingface.co/datasets/TankVincent/MapEval-Textual.
