Benchmarking
benchmarking_sbi_runs
Benchmarking SBI Runs
This dataset contains the raw, per-run results underlying the manuscript
"Benchmarking Simulation-Based Inference"
(Lueckmann, Boelts, Greenberg, Goncalves & Macke, AISTATS 2021).
It is a direct migration of the Git LFS data from
mackelab/benchmarking_sbi_runs on GitHub.
For compiled, ready-to-use dataframes built from these raw results (and the code that produced
them), see the companion repository:… See the full description on the dataset page: https://huggingface.co/datasets/mackelab/benchmarking_sbi_runs.PDE_Inverse_Problem_Benchmarking
PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems
This is the official dataset for the paper PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems.
Code: GitHub - ASK-Berkeley/PDEInvBench
Sample Usage
You can use the provided script from the codebase to batch download the data:
pip install huggingface_hub
python3 huggingface_pdeinv_download.py --dataset… See the full description on the dataset page: https://huggingface.co/datasets/DabbyOWL/PDE_Inverse_Problem_Benchmarking.map-anything-benchmarking
MapAnything Benchmarking Dataset
Dataset Description
This dataset contains the WAI format data used for benchmarking feed-forward 3D reconstruction models in the MapAnything codebase.
Please see our Data Processing README for more details.
Citation
If you use this dataset in your research, please cite our paper:
@inproceedings{keetha2026mapanything,
title={{MapAnything}: Universal Feed-Forward Metric {3D} Reconstruction},
author={Nikhil Keetha and Norman… See the full description on the dataset page: https://huggingface.co/datasets/facebook/map-anything-benchmarking.PDB-Results
PDB-Results: model outputs and scores
📄 Paper ·
💻 Code ·
🌐 Project page ·
🏆 Leaderboard
Raw debugging outputs and evaluator scores for every model evaluated on the PDB
(Precise Debugging Benchmarking) suite, so that every reported number can be
inspected and recomputed.
Evaluation sets
Filename tag
Set
Tasks
Models
pdb_single_hard
PDB-Single
5,751 (BigCodeBench 2,525 + LiveCodeBench 3,226)
4
pdb_single… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Results.external-benchmarking
Vector Search Benchmarks
This repo contains datasets for benchmarking vector search performance, to help Superlinked prioritize integration partners.
For performing actual benchmarking on this dataset, see the github repository README.
Overview
We reviewed a number of publicly available datasets and noted 3 core problems + here is how this dataset fixes them:
Problems of other vector search benchmarks
How this dataset solves it
Not enough metadata of… See the full description on the dataset page: https://huggingface.co/datasets/superlinked/external-benchmarking.groundtruth-dynamic-benchmarking-submissions
Groundtruth Dynamic Benchmarking — Geology — Submissions
Community-submitted evaluation runs against the groundtruth-dynamic-benchmarking geology rubrics, feeding the leaderboard. We are currently running two tracks: model benchmarking (comparing different models with no special harness) and harness benchmarking (comparing different harnesses using a single standard model - GLM 4.7).
Each submission is a pointwise rubric score: one model, scored 0–10 per question against a… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking-submissions.
