datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jevbench
JevBench
Cases and predictions for JevBench: An Open Evaluation Framework for Typed
Decision Models. The evaluation code is on GitHub: https://github.com/Leanmcp/jevbench JevBench evaluates models that return typed decisions
(a yes/no probability, a distribution over a set of choices, or an expected
level on an ordered rubric) on identical inputs built from public datasets.
Contents
Path
What it holds
cases/
One JSONL file per slice. Each row is the… See the full description on the dataset page: https://huggingface.co/datasets/Leanmcp/jevbench.craft-benchmark-lean
CRAFT Benchmark Dataset
Trajectory logs from the CRAFT benchmark — a multi-agent evaluation of pragmatic communication in LLMs under strict partial information. - TL;DR
Dataset Structure
Each row is one turn from a CRAFT game, with fields for:
Identity: structure_id, director_model, builder_model, model_type (base/frontier), turn_number
Director responses: D1_thinking, D1_message, D2_thinking, D2_message, D3_thinking, D3_message
Builder: builder_action, builder_block… See the full description on the dataset page: https://huggingface.co/datasets/Abhijnan/craft-benchmark-lean.leanflow-jhtdb-benchmark
🌊 LeanFlow — Formally Verified Dual-Scale Navier-Stokes Solver
LeanFlow is the next generation of Navier-Stokes solvers — combining formally verified mathematics (Lean 4), AI-native bare-metal execution (Runux AI runtime), and pseudo-spectral accuracy validated on real DNS turbulence data.
🏆 Key Results at a Glance
Metric
LeanFlow ETD-RK4
OpenFOAM icoFoam
FDM-PISO (Python)
Max Divergence $|\nabla\cdot u|_\infty$
2.994e-14
4.102e-07
N/A… See the full description on the dataset page: https://huggingface.co/datasets/callensxavier/leanflow-jhtdb-benchmark.imagenette-320px-resplit
Imagenette 320px with Fixed Validation and Test Splits
Dataset Description
This dataset is a reproducible, Parquet-based version of the 320px configuration of frgfm/imagenette. Imagenette is a subset of ten readily classified ImageNet classes created for fast experimentation with image-classification methods.
This version preserves the source images, numeric labels, and label metadata. Its only data change is a fixed, stratified division of the original validation… See the full description on the dataset page: https://huggingface.co/datasets/leandrodevai/imagenette-320px-resplit.Pixio_base_blind_spotsThe dataset now contains 990 images in addition to the previous 10 diverse samples.
Analysis of Blind Spots in Pixio (ViT-H/16) – A Vision Transformer
1. Model Selection
Model: facebook/pixio-vith16
Release Date: 17 Dec 2025
Parameters: 631M
Modality: Vision
Type: Base model
The model is a Vision Transformer (ViT) with a patch size of 16 and a hidden size of 1280, using 32 layers. A distinctive feature of
Pixio is that it uses 8 class tokens instead of a single… See the full description on the dataset page: https://huggingface.co/datasets/leandrehonore/Pixio_base_blind_spots.imagewoof-320px-resplit
ImageWoof 320px with Fixed Validation and Test Splits
Dataset Description
This dataset is a reproducible, Parquet-based version of the 320px configuration of frgfm/imagewoof. ImageWoof is a subset of ten dog-breed classes from ImageNet designed to be more difficult than broad-category image-classification benchmarks.
This version is intended for image classification and confidence-calibration experiments. It introduces two changes to the source dataset:
It… See the full description on the dataset page: https://huggingface.co/datasets/leandrodevai/imagewoof-320px-resplit.leanderjohanneskahrensI have been fighting schizophrenia and hoping for a cure.
facades2026_07_19_collect_leandojo_gemma3_12b_gemma4_31bDeepfakes-QA-Leaning
Deepfake Quality Assessment
Deepfake QA is a Deepfake Quality Assessment model designed to analyze the quality of deepfake images & videos. It evaluates whether a deepfake is of good or bad quality, where:
0 represents a bad-quality deepfake
1 represents a good-quality deepfake
This classification serves as the foundation for training models on deepfake quality assessment, helping improve deepfake detection and enhancement techniques.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Deepfakes-QA-Leaning.PneumoniaClassificationdatafacades_DSPneumoniaClassificationMultimodalDeepfakes-QA-Leaning
Deepfake Quality Assessment
Deepfake QA is a Deepfake Quality Assessment model designed to analyze the quality of deepfake images & videos. It evaluates whether a deepfake is of good or bad quality, where:
0 represents a bad-quality deepfake
1 represents a good-quality deepfake
This classification serves as the foundation for training models on deepfake quality assessment, helping improve deepfake detection and enhancement techniques.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/strangerguardhf/Deepfakes-QA-Leaning.
