Team Ai
20 results

Benchmarking

mackelab /benchmarking_sbi_runs Benchmarking SBI Runs This dataset contains the raw, per-run results underlying the manuscript "Benchmarking Simulation-Based Inference" (Lueckmann, Boelts, Greenberg, Goncalves & Macke, AISTATS 2021). It is a direct migration of the Git LFS data from mackelab/benchmarking_sbi_runs on GitHub. For compiled, ready-to-use dataframes built from these raw results (and the code that produced them), see the companion repository:… See the full description on the dataset page: https://huggingface.co/datasets/mackelab/benchmarking_sbi_runs.100K<n<1M0 likes14k downloads2mo agoHugging FaceDabbyOWL /PDE_Inverse_Problem_Benchmarking PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems This is the official dataset for the paper PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems. Code: GitHub - ASK-Berkeley/PDEInvBench Sample Usage You can use the provided script from the codebase to batch download the data: pip install huggingface_hub python3 huggingface_pdeinv_download.py --dataset… See the full description on the dataset page: https://huggingface.co/datasets/DabbyOWL/PDE_Inverse_Problem_Benchmarking.imageother100M<n<1B3 likes8.6k downloads5mo agoHugging Facefacebook /map-anything-benchmarking MapAnything Benchmarking Dataset Dataset Description This dataset contains the WAI format data used for benchmarking feed-forward 3D reconstruction models in the MapAnything codebase. Please see our Data Processing README for more details. Citation If you use this dataset in your research, please cite our paper: @inproceedings{keetha2026mapanything, title={{MapAnything}: Universal Feed-Forward Metric {3D} Reconstruction}, author={Nikhil Keetha and Norman… See the full description on the dataset page: https://huggingface.co/datasets/facebook/map-anything-benchmarking.image-to-3d100B<n<1T8 likes3.8k downloads9mo agoHugging FacePrecise-Debugging-Benchmarking /PDB-Results PDB-Results: model outputs and scores 📄 Paper &nbsp;·&nbsp; 💻 Code &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 🏆 Leaderboard Raw debugging outputs and evaluator scores for every model evaluated on the PDB (Precise Debugging Benchmarking) suite, so that every reported number can be inspected and recomputed. Evaluation sets Filename tag Set Tasks Models pdb_single_hard PDB-Single 5,751 (BigCodeBench 2,525 + LiveCodeBench 3,226) 4 pdb_single… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Results.0 likes741 downloads5d agoHugging Facesuperlinked /external-benchmarking Vector Search Benchmarks This repo contains datasets for benchmarking vector search performance, to help Superlinked prioritize integration partners. For performing actual benchmarking on this dataset, see the github repository README. Overview We reviewed a number of publicly available datasets and noted 3 core problems + here is how this dataset fixes them: Problems of other vector search benchmarks How this dataset solves it Not enough metadata of… See the full description on the dataset page: https://huggingface.co/datasets/superlinked/external-benchmarking.image10M<n<100M0 likes716 downloads1y agoHugging FaceEigenformAI /groundtruth-dynamic-benchmarking-submissions Groundtruth Dynamic Benchmarking — Geology — Submissions Community-submitted evaluation runs against the groundtruth-dynamic-benchmarking geology rubrics, feeding the leaderboard. We are currently running two tracks: model benchmarking (comparing different models with no special harness) and harness benchmarking (comparing different harnesses using a single standard model - GLM 4.7). Each submission is a pointwise rubric score: one model, scored 0–10 per question against a… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking-submissions.0 likes448 downloads1mo agoHugging Face