datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qiskit-calibration-drift
Qiskit Calibration Drift
Calibration parameters from IBM Quantum Heron processors joined to ambient and space-weather conditions at the time of each measurement. Designed for time-series forecasting of qubit drift and for studying environmental coupling to superconducting calibrations.
A GitHub Action (source) polls backend.properties() on every available Heron device every 30 minutes and appends new calibration events keyed on (backend, property, qubit_a, qubit_b… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/qiskit-calibration-drift.imatrix-calibration
Importance Matrix Calibration Datasets
This repository provides calibration datasets used to generate importance matrices (imatrix), which are required to minimize errors when quantizing models with LLaMA C++.
The llama-imatrix program cannot handle parquet files directly and thus requires them to be converted into text format first. There are many ways to do this but a simple approach is to use DuckDB with the following command: duckdb -noheader -ascii -c "SELECT content FROM… See the full description on the dataset page: https://huggingface.co/datasets/eaddario/imatrix-calibration.LLM_calibrationmechanical-head-mouth-calibration
Mechanical Head Mouth Calibration Dataset
This dataset accompanies the paper Automated Synthesis of Facial Mechanisms for Conversational Animatronic Robots.
It contains 4993 aligned samples captured from a mechanical head mouth system. Each sample consists of:
one monocular camera image
raw paired PWM motor commands
normalized motor commands
MediaPipe mouth/jaw blendshape coefficients
normalized mouth-region landmark coordinates
Dataset Structure
.
+-- README.md… See the full description on the dataset page: https://huggingface.co/datasets/Zzz0918/mechanical-head-mouth-calibration.pa-warm-start-sft-xl-calibrationGLM-5.3-Flash-calibration-activations-v1
GLM-5.3-Flash calibration activations v1 (BF16, natural routing)
Per-layer block-input activations of zai-org/GLM-5.3-Flash-BF16 @ b1967181 over 92x2048
tokens of the exllamav3 standard_cal_data corpus (pinned): per context, layer_NNN.attn_in
and layer_NNN.mlp_in (bf16, post-norm linear inputs; mlp_in is the router + expert gate/up
input) and layer_NNN.router_logits (fp32, natural top-8 routing ground truth).
Per-expert Hessians E[xx^T], routing statistics and down-proj inputs… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-calibration-activations-v1.LLM_compression_calibration
LLM Compression Calibration dataset
This dataset is the default calibration dataset used by Neural Magic for one-shot compression of Large Language Models (LLMs).
Note: This dataset is the result of active research and subject to change without notice.
Dataset Details
Dataset Sources
The current version of this dataset is compiled from data from these datasets:
garage-bAInd/Open-Platypus: 10,000 samples
Data Fields
The dataset contains 2 data… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/LLM_compression_calibration.Inkling-Small-Multimodal-Calibration
Inkling-Small Multimodal Calibration
The exact 1,663 samples used for BF16 routed-expert importance collection
for Inkling-Small Mixed Quant GGUF.
This is calibration material, not a held-out evaluation benchmark.
The primary balanced pass is:
Category
Samples
Valid decoder tokens
Share
Text / reasoning
462
471,858
44.976%
Code / tool-oriented source text
205
209,715
19.989%
Real image / document
486
262,476
25.018%
Real speech audio
309
105,080
10.016%
Total… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/Inkling-Small-Multimodal-Calibration.robotwin2-pi05-calibration-clean
RoboTwin 2.0 pi0.5 Calibration Subset (clean)
Sampled from TianxingChen/RoboTwin2.0 on HuggingFace.
Embodiment: aloha-agilex
Setting: clean (no domain randomization)
Tasks: 50
Episodes per task: 2
Total episodes: 100
Sampling seed (numpy default_rng): 42
Layout
<task>/
scene_info.json # task scene metadata
episode{N}.hdf5 # sensor data (rgb / state / action)
episode{N}.pkl # trajectory metadata
episode{N}.json # per-episode instruction pool
See… See the full description on the dataset page: https://huggingface.co/datasets/JingwuLuo/robotwin2-pi05-calibration-clean.calibration-scorecards
Prediction Market Calibration Scorecards
Monthly Brier + log-loss calibration breakdowns for Kalshi + Polymarket. Each month provides mean Brier, mean log-loss, per-venue and per-category breakdowns, and a 10-bucket calibration histogram (actual vs predicted). Published with a 14-day delay after month-end to capture late resolutions.
License and Use
This dataset is released under Creative Commons Attribution 4.0 International
(CC-BY-4.0;… See the full description on the dataset page: https://huggingface.co/datasets/SimpleFunctions/calibration-scorecards.calibrationcalibration
Post-hoc Calibration Dataset
This repository contains datasets designed for evaluating and developing post-hoc calibration methods for deep neural network classifiers. Each dataset includes precomputed logits and labels, divided into clear training and test splits.
Dataset Overview
Datasets provided here cover popular benchmark tasks, including CIFAR-10, CIFAR-100, SVHN, Stanford Cars (CARS), CUB-200 Birds (BIRDS), and ImageNet. Dataset composition here for… See the full description on the dataset page: https://huggingface.co/datasets/WJHuang/calibration.HareSkip-calibration
HareSkip Calibration Dataset
1. Summary
This dataset contains all images and measurements generated during a July–September 2026 recalibration and method-comparison experiment for HareSkip, a step-skipping inference-acceleration extension for DiT-based image diffusion, implemented on top of Forge neo (upstream repository Haoming02/sd-webui-forge-classic, branch neo; the link is pinned to the commit used throughout the experiment). The HareSkip extension itself is… See the full description on the dataset page: https://huggingface.co/datasets/Rootport/HareSkip-calibration.SN-Calibration-2023opens2v-calibration
OpenS2V calibration manifest
This directory contains a small normalized calibration subset sourced from
BestWishYsh/OpenS2V-5M.
It is intended to prototype AutoRound diffusion T2V and I2V calibration without
requiring the full 11 TB upstream dataset.
Format
output/opens2v_calibration.tsv uses AutoRound's normalized schema:
id: stable upstream sample identifier
caption: metadata.face_cap_qwen, falling back to metadata.cap[0]
image: relative path to a reference… See the full description on the dataset page: https://huggingface.co/datasets/changwangss/opens2v-calibration.rl-natural-self-calibration
RL & Natural Self-Calibration
Does RL induce a model's self-knowledge of its own capacity/size? Experiments on Qwen2.5
(0.5–14B) and OLMo-3-7B (base / SFT / RL-Zero{Math,Code,General} / Think / Instruct).
Scripts, probe outputs, and findings. (Research scaffold — read FINDINGS_*.md in results/.)
TL;DR findings
Nobody knows their size explicitly. Base & instruct models across 0.5–14B confabulate
"GPT-3.5 / 175B" when asked their parameter count. The DV is… See the full description on the dataset page: https://huggingface.co/datasets/Jordine/rl-natural-self-calibration.imatrix-calibration
Importance Matrix Calibration Datasets
This repository provides calibration datasets used to generate importance matrices (imatrix), which are required to minimize errors when quantizing models with LLaMA C++.
The llama-imatrix program cannot handle parquet files directly and thus requires them to be converted into text format first. There are many ways to do this but a simple approach is to use DuckDB with the following command: duckdb -noheader -ascii -c "SELECT content FROM… See the full description on the dataset page: https://huggingface.co/datasets/ClonedGizzards/imatrix-calibration.mcqa_calibration_datasetDeepSeek-V4-Flash-0731-REAM-calibration-stats
DeepSeek-V4-Flash-0731 — expert calibration statistics (REAM line)
Layerwise routed-expert statistics of
deepseek-ai/DeepSeek-V4-Flash-0731
(43 MoE layers × 256 experts), collected by running the full model over a
~4.9M-token multi-domain calibration mix (multi-turn dialogs, thinking and
direct modes, rendered with the model's own chat encoder). These are the
statistics behind the REAM144/96 release line — published so that expert
selection, pruning, merging and routing research… See the full description on the dataset page: https://huggingface.co/datasets/WaveCut/DeepSeek-V4-Flash-0731-REAM-calibration-stats.insight9-calibration-20260923
Looper Insight 9 Calibration Capture
Two calibration-oriented recordings captured with a Looper Insight 9 on 2026-09-23. The release contains native-orientation stereo grayscale images, RGB JPEG images, device timestamps, raw IMU measurements, auxiliary device-estimated poses, capture metadata, quality reports, review videos, and validated ROS1 bags.
This is a raw sensor-data release, not a calibrated benchmark or ground-truth trajectory dataset. No intrinsics, extrinsics, or… See the full description on the dataset page: https://huggingface.co/datasets/nerako/insight9-calibration-20260923.safety-calibration-cases
Safety Calibration Cases
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/rmems/safety-calibration-cases.calibration-mixspecdec-calibration
Speculative-decoding calibration banks
Per-round speculative-decoding acceptance and speculator banks, plus
MoE expert-routing captures, collected by driving SGLang and logging every
draft round. Used to drive the discrete-event simulator in
inference-lab (see
examples/specdec/README.md for figure reproduction).
Supersedes
Doubleword/qwen3.6-specdec-calibration:
this dataset adds a model level to the path, the per-category SPEED-Bench
routing captures, and DeepSeek-V4-Flash.… See the full description on the dataset page: https://huggingface.co/datasets/Doubleword/specdec-calibration.agentic-vbench-calibration-trajectories
AgenticVBench volleyball calibration trajectories
Native raw agent traces from the calibration of two AgenticVBench understanding
tasks, published so a reviewer can audit turn counts, prompt parity and the
no-lookup rule independently rather than taking a summary on trust.
usc-wsu-2023-volleyball-block-timeline — 23 block points, two attributions each
byu-wsu-2023-volleyball-block-timeline — 18 block points, three attributions each
The tasks themselves, the answer keys, the… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/agentic-vbench-calibration-trajectories.jev-calibration-statistics
Confidence statistics for Jev and a self-judging Gemma 4 E2B
Aggregate statistics on the confidence scores from two judges in a retrieval benchmark: TypeSafe's Jev, pinned to jev-1.13.0, and Gemma 4 E2B judging its own work.
End to end, the pipeline with Jev making every decision did not beat the same pipeline with no judge: it scored 0.612 against 0.740, missed its main pre-registered bar, made about the same number of mistakes on questions both answered, and lost because it… See the full description on the dataset page: https://huggingface.co/datasets/clduab11/jev-calibration-statistics.qiskit_calibration_drift
Qiskit Calibration Drift (TsFile)
Apache TsFile version of
phanerozoic/qiskit-calibration-drift.
Overview
Calibration parameters from IBM Quantum Heron processors joined to ambient
and space-weather conditions at the time of each measurement, collected by a
GitHub Action that polls backend.properties() every 30 minutes. The pinned
revision covers 10,680,076 calibration events (2026-01-31 .. 2026-05-20) on
one backend across 3 calibration properties (1- and… See the full description on the dataset page: https://huggingface.co/datasets/THULab/qiskit_calibration_drift.llm-forecast-calibration
LLM Forecast Calibration Study — GLM-5.3 on resolved Manifold Markets questions
Raw generation data for the study "Does sampling K times beat thinking harder?
A controlled study of LLM forecast calibration on resolved binary questions."
Source repo: EzraStone/llm-forecast-calibration.
Data mirrored from GitHub commit 0f12f71a2c2ec8c54cafeb4231fecb87e705e660.
All eight JSONL files match the source data byte for byte. The source repository
remains canonical for analysis code… See the full description on the dataset page: https://huggingface.co/datasets/ezra77/llm-forecast-calibration.so101_pangyo_pickplace_cube_260121_calibrationDeepSeek-R1-Distill-Qwen-1.5B-Self-CalibrationThis dataset contains data for the paper Efficient Test-Time Scaling via Self-Calibration.
We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model uncertainty on different tasks. For example, we can incorporate the model’s confidence into self-consistency by assigning each sampled response $y_i$ a confidence score $c_i$. Instead of treating all responses… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/DeepSeek-R1-Distill-Qwen-1.5B-Self-Calibration.calibration
