Team Ai
Datasetpublic

WJHuang/calibration

Post-hoc Calibration Dataset This repository contains datasets designed for evaluating and developing post-hoc calibration methods for deep neural network classifiers. Each dataset includes precomputed logits and labels, divided into clear training and test splits. Dataset Overview Datasets provided here cover popular benchmark tasks, including CIFAR-10, CIFAR-100, SVHN, Stanford Cars (CARS), CUB-200 Birds (BIRDS), and ImageNet. Dataset composition here for… See the full description on the dataset page: https://huggingface.co/datasets/WJHuang/calibration.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
1likes482downloads
Dataset Card

Post-hoc Calibration Dataset

This repository contains datasets designed for evaluating and developing post-hoc calibration methods for deep neural network classifiers. Each dataset includes precomputed logits and labels, divided into clear training and test splits.

Dataset Overview

Datasets provided here cover popular benchmark tasks, including CIFAR-10, CIFAR-100, SVHN, Stanford Cars (CARS), CUB-200 Birds (BIRDS), and ImageNet. Dataset composition here for post-hoc calibration are listed below:

Dataset# ClassesTraining Set SizeTest Set Size
CIFAR-1010500010000
SVHN10600026032
CIFAR-100100500010000
CARS19640204020
BIRDS20028972897
ImageNet10002500025000

Included Pre-trained Networks on Classification Datasets

Each .p file, which represents one calibration task, contains ground-truth labels and predicted logits from specific pre-trained neural network architectures, as listed below:

  • —`probs_resnet110_c10_logits.p`: ResNet110 on CIFAR-10
  • —`probs_resnet_wide32_c10_logits.p`: WideResNet32 on CIFAR-10
  • —`probs_densenet40_c10_logits.p`: DenseNet40 on CIFAR-10
  • —`probs_resnet110_c100_logits.p`: ResNet110 on CIFAR-100
  • —`probs_resnet_wide32_c100_logits.p`: WideResNet32 on CIFAR-100
  • —`probs_densenet40_c100_logits.p`: DenseNet40 on CIFAR-100
  • —`probs_resnet152_SD_SVHN_logits.p`: ResNet152 SD on SVHN
  • —`probs_resnet50NTSNet_birds_logits.p`: ResNet50NTSNet on BIRDS
  • —`probs_resnet50_cars_logits.p`: ResNet50 on CARS
  • —`probs_resnet101scratch_cars_logits.p`: ResNet101 from scratch on CARS
  • —`probs_resnet101_cars_logits.p`: ResNet101 on CARS (initialized with ImageNet weights)
  • —`probs_densenet161_imgnet_logits.p`: DenseNet161 on ImageNet
  • —`probs_pnasnet5large_imgnet_logits.p`: PNASNet5large on ImageNet
  • —`probs_resnet152_imgnet_logits.p`: ResNet152 on ImageNet
  • —`probs_swintiny_imgnet_logits.p`: Swin Transformer (tiny) on ImageNet

Data Loading

Each dataset is stored as a Python pickle file (.p). Load the datasets with the following Python snippet:

python
import pickle

with open('path_to_dataset.p', 'rb') as f:
    (x_logits_train, y_train), (x_logits_test, y_test) = pickle.load(f)
  • —x_logits_train, x_logits_test: The logits (raw neural network outputs).
  • —y_train, y_test: Ground truth labels.

Reference

More detailed description for the dataset can be found in the following paper:

Huang W, Cao G, Xia J, Chen J, Wang H, Zhang J. h-calibration: Rethinking Classifier Recalibration with Probabilistic Error-Bounded Objective[J].
IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(10): 9023–9042, 2025. arXiv: arxiv.org/abs/2506.17968.

The official GitHub implementation of the above H-Calibration study using the dataset is accessible here: