datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quantized-llama-3.1-leaderboard-v2-evals
Open LLM Leaderboard v2 Benchmark Results
This artifact contains all the data from evaluations of Neural Magic's quantized Llama-3.1 models.
These evaluations were produced with lm-evaluation-harness by running the following command:
lm_eval \
--model vllm \
--model_args pretrained="<model_path>",dtype=auto,add_bos_token=False,max_model_len=4096,tensor_parallel_size="<num_gpus>",gpu_memory_utilization=0.8,enable_chunked_prefill=True \
--apply_chat_template \… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-leaderboard-v2-evals.szl-quant-sft-v1
szl-quant-sft-v1 — training rows with signed lineage
Training eligibility: HELD-COUNSEL. Do not train on this dataset.
The estate's license register
(SZLHOLDINGS/model-bom DATASET_LICENSE_REGISTER.csv)
lists this dataset as HELD pending counsel review of upstream CoinGecko
redistribution and commercial terms, and it appears on the CI-enforced
TRAINING_BLOCKLIST.txt.
The rows are published for lineage inspection and replay verification only.
The Apache-2.0 declaration covers… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/szl-quant-sft-v1.quant-fidelity-registry
Quantization Fidelity Registry
A public, schema'd, receipt-backed, cross-model index of quantization quality measurements.
It exists to answer one question that nothing else answers today: show me every measured quant of
model X, with its fidelity number and enough provenance to know whether that number means anything.
It is the sibling of 0xSero/local-ai-registry,
which answers how fast, how much VRAM, how much money. This one answers how faithful. Ids and the
huggingface… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/quant-fidelity-registry.Embodied-Agent-Arena
Embodied Agent Arena — task data
Task data for the Embodied Agent Arena evaluation code.
This task collection contains 1,000 uniquely identified evaluation cases:
W1 370, W2 220, W3 157, W4 183, and W5 70.
The 660 W1/W2/W5 cases use fixed image/video inputs; the 340 W3/W4 cases require
interactive benchmark environments. The task index contains 41 adapter IDs;
several adapters correspond to task types from the same source dataset.
W4 contains thirteen reporting groups over… See the full description on the dataset page: https://huggingface.co/datasets/uuu-Quant/Embodied-Agent-Arena.QuantumChem_Testbank_3000
Hold-out test set for evaluating quantaMind.
quant_exploration
Examining LLM Quantization Impact
This document is a comparative analysis of qualitative performance degradation across Llama.cpp quantization within a single 2x7B model. My hope is that it will help people unfamiliar with quant impacts get a sense of how quantization will affect output.
Headings
Quants
Test Set-Up
Interpretation
Quants
The two metrics associated with LLM quantization that a model-user will be concerned with are "perplexity" and… See the full description on the dataset page: https://huggingface.co/datasets/christopherthompson81/quant_exploration.exams-basic-and-quantum-cryptography-and-security-latex
Open Problem Exams: Cryptography and Security (LaTeX)
A curated dataset of open-ended exam problems (with solutions) in cryptography and computer security, formatted in LaTeX. The dataset is sourced from university courses at three institutions.
Dataset Overview
Institution
Files
Topics
Questions
Caltech & TU Delft
8
38
145
EPFL
6
19
86
ETH Zurich
1
14
37
MIT
3
33
79
Total
18
104
347
Difficulty Distribution
Institution… See the full description on the dataset page: https://huggingface.co/datasets/natnitaract/exams-basic-and-quantum-cryptography-and-security-latex.HPLT3_DE_0.9_Quantile_Adult_Filteredhemmingway-1-omlx-quantization-evidence-v2
Hemmingway-1 Quantization Evidence v2
This package records two local evidence lanes for the Hemmingway-1 oQ4e build: teacher-forced numerical fidelity against a BF16 reference, and controlled runtime telemetry on Apple Silicon. It complements the frozen blind-preference study in Hemmingway-1 oMLX Quantization Benchmark v1.
This dataset is sixstringzen/hemmingway-1-omlx-quantization-evidence-v2. The quality dataset remains unchanged because blind preference, distribution fidelity… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-evidence-v2.quantum-physics-0.6-corpus
quantum-physics-0.6-corpus
Dataset Description
This is a domain-specific corpus created using ontology-guided filtering from FineWeb-Edu.
Dataset Creation
Source: HuggingFaceFW/fineweb-edu
Filtering Method: Semantic similarity to subdomain centroids (embedding-based)
Pipeline: Ontology-Guided Domain Corpus Builder
Dataset Structure
Each chunk contains:
text: The text content (256-512 tokens)
subdomain_id: Assigned subdomain
similarity_score:… See the full description on the dataset page: https://huggingface.co/datasets/konsman/quantum-physics-0.6-corpus.tis-quantile-datasets-Qwen3-4B-Basehemmingway-1-omlx-quantization-benchmark-v1
Hemmingway-1 oMLX Quantization Benchmark
This is the public-safe benchmark package for the Hemmingway-1 oMLX
quantization study on Apple Silicon.
Altworld developed and published
Hemmingway-1. Bobby Pierce
published these quantizations and the evaluation package. The
collection
links the upstream model and all six builds.
Analysis revision 2, corrected on 2026-09-22, fixes A/B attribution and matching
across reversed packets. Read CORRECTION.md before using the
aggregate… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1.quantitative-finance-reasoning
Dataset Description
Dataset Summary
This dataset contains question-answer pairs focused on quantitative finance, covering topics such as option pricing, stochastic calculus (Brownian motion, Itô's Lemma), probability theory, and financial modeling assumptions. Each instance includes a question, a detailed ground-truth solution (often resembling textbook explanations or interview answers), a multi-step reasoning trace generated by a Gemini Pro model, and a structured validation of… See the full description on the dataset page: https://huggingface.co/datasets/VedantPadwal/quantitative-finance-reasoning.quantum-worldline-research
Quantum Worldline Research Data
Structured research data from the Quantum Worldline project - an AI-assisted research program investigating holographic forces in MERA tensor networks, worldline path integrals on AdS spacetime, and quantum simulation of lattice gauge theories.
Dataset Description
This dataset contains the complete structured output of the Quantum Worldline multi-agent research system, which automates the research cycle: discover - hypothesize - gate - test… See the full description on the dataset page: https://huggingface.co/datasets/Jonboy648/quantum-worldline-research.quant-fidelity-corpora
Quantization fidelity corpora
Evaluation text for measuring how faithfully a quantized LLM reproduces its full-precision parent
(KL divergence of next-token distributions, top-1 agreement, perplexity ratio), as used by
Agention for the Signal and Qwen3.8-27B quantization campaigns.
mixedweb-v1
mixedweb-v1.txt (800,789 chars, 301 documents, md5 51e0045e8cabf37922aa82766a25b7b4) is a
seeded random slice of HuggingFaceFW/fineweb
sample-10BT: general English web text… See the full description on the dataset page: https://huggingface.co/datasets/agentionai/quant-fidelity-corpora.quantqa
QuantQA: Quantitative Finance Interview Questions
QuantQA is a curated dataset of 519 interview questions sourced from leading quantitative trading firms including Jane Street, Citadel, Two Sigma, Optiver, and SIG, in collaboration with CoachQuant.
Topic Distribution
Topic
Coverage
Probability
67%
Combinatorics
22%
Expected Value
21%
Conditional Probability
14%
Game Theory
11%
Note: Questions may cover multiple topics
Training Results… See the full description on the dataset page: https://huggingface.co/datasets/ReinforceNow/quantqa.quantsafe-judge-benchmark
QuantSafe Refusal-Safety Judge-Agreement Benchmark
A focused, project-authored and project-labeled 40-item corpus for measuring
how three open guard-model families label (prompt, response) pairs as
safe / unsafe. It supports the Judge Agreement tab in
QuantSafe Certifier
and the broader methodology in
Quality Is Not a Safety Proxy Under Quantization.
Files
judge_corpus.jsonl: 12 clear-safe, 12 clear-unsafe, and 16 intentionally
borderline items with project… See the full description on the dataset page: https://huggingface.co/datasets/Crusadersk/quantsafe-judge-benchmark.tamperbench-quantization-qwen3-4b
TamperBench + Quantization: Does Compression Act as Implicit Tampering?
Motivation
TamperBench evaluates explicit tampering attacks (LoRA fine-tuning, jailbreak-tuning, etc.) on LLM safety guards. Catastrophic Failure of LLM Unlearning via Quantization shows that quantization can undo safety-trained behaviors.
This experiment bridges these two lines of work by adding quantization as a deployment-realistic perturbation to the TamperBench evaluation protocol. We… See the full description on the dataset page: https://huggingface.co/datasets/b0sungk1m/tamperbench-quantization-qwen3-4b.gmat-quant-corpus
GMAT Quant Corpus for Solver + Retrieval
This dataset is intended for retrieval over GMAT-style quantitative teaching content.
Files
gmat_hf_chunks.jsonl — retrieval chunks used by the app
gmat_question_seed.jsonl — question seed data
gmat_topic_index.json — topic metadata/index
Default dataset viewer
The default dataset viewer is configured to load only:
gmat_hf_chunks.jsonl
This avoids schema conflicts with the other support files in the repository.
tis-quantile-datasets-Olmo-3-1025-7Bopensre-incident-trajectories
OpenSRE Incident-Diagnosis Trajectories
Graded, multi-step SRE incident-diagnosis trajectories. A frozen LLM reads evidence through
diagnostic tools (describe_pod / get_events / get_logs / get_metrics / query_traces / …),
states a root cause + category + fix, and is scored on substance against ground truth. Built as a
HUD v6 RL environment with a deliberate model spanning set so difficulty is legible and the
within-group reward spread is real (the GRPO learning signal).
197… See the full description on the dataset page: https://huggingface.co/datasets/quantranger/opensre-incident-trajectories.minicpm5-1b-quantization-benchmark
openbmb/MiniCPM5-1B 次世代量子化(Quanto FP8 / INT4 vs BNB 4bit)実測ベンチマークレポート
対象モデル: openbmb/MiniCPM5-1B (1.16B parameters, 128k context, LlamaForCausalLM)
検証ハードウェア: NVIDIA GeForce RTX 4070 Ti (12GB GDDR6X, Ada Lovelace, Compute Capability 8.9, 第4世代Tensor Core)
実行環境: Windows / Python 3.13 / PyTorch 2.6.0+cu124 / transformers 4.57.6 / optimum-quanto 0.2.7 / bitsandbytes 0.50.0
検証日: 2026-09-19 12:12:34
1. エグゼクティブサマリー(全体比較)
NVIDIA GeForce RTX 4070 Ti 実機環境において、標準ネイティブ… See the full description on the dataset page: https://huggingface.co/datasets/aoiandroid/minicpm5-1b-quantization-benchmark.quantinvbench
QuantInvBench
QuantInvBench is the companion benchmark for evaluating multi-agent portfolio
aggregation on five assets in the fixed order BTC, ETH, USDT, XRP,
BNB. At each seven-day execution event, a method receives ten long-only,
fully invested strategy proposals and must return one portfolio weight vector.
The ten strategy identities, in canonical order, are Momentum, MaxSharpe,
TSMom14, TSMom60, RiskParity, MinVar, NewsSentiment, Size,
ValueNVT, and Liquidity. Equal-weight… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-quantinv/quantinvbench.perceive-benchmark
PERCEIVE
PERCEIVE (Psychophysics-driven Elicitation for Routing Cost-Efficiency In
Vision-Language Evaluation) is a 4,801-sample document-image QA benchmark for
cost-aware VLM routing. Each sample carries psychophysical complexity annotations
(Visual Dependency Score, Reasoning Depth Score, Spatial Extent Score) and a
routing label identifying the cheapest model-budget configuration that answers
it correctly.
Routing labels are derived via a QUEST-style adaptive cascade achieving… See the full description on the dataset page: https://huggingface.co/datasets/quantiphi-routing/perceive-benchmark.Qiskit-QuantumKatas
Qiskit QuantumKatas
A benchmark dataset for evaluating Large Language Models on quantum computing code generation tasks using Qiskit.
Dataset Description
This dataset contains 350 quantum computing tasks translated from Microsoft's QuantumKatas (originally in Q#) to Qiskit (Python). It is designed for evaluating LLMs on their ability to generate correct quantum computing code.
Supported Tasks
Code Generation: Given a natural language description and function… See the full description on the dataset page: https://huggingface.co/datasets/Qiskit/Qiskit-QuantumKatas.HPLT3_DE_0.9_Quantile_Adult_Filtered_Propelabbq
BBQ
Repository for the Bias Benchmark for QA dataset.
https://github.com/nyu-mll/BBQ
Authors: Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R. Bowman.
This repository is a fork of https://huggingface.co/datasets/heegyu/bbq, and adds the "All" configuration containing all subsets.
About BBQ (paper abstract)
It is well documented that NLP models learn social biases, but little work has been done… See the full description on the dataset page: https://huggingface.co/datasets/quantiles/bbq.quantum-circuits-8k
Quantum Circuits 8K Dataset
A synthetic dataset of 8,129 quantum circuit examples for training language models to generate OpenQASM 2.0 code from natural language descriptions.
Quick Stats
Total Samples: 8,129 (description → QASM pairs)
Unique Circuits: 739 base circuits
Categories: 92 distinct quantum circuit types
Qubit Range: 1-9 qubits
Format: OpenQASM 2.0
Augmentation: 11x per circuit (original + 10 paraphrases)
Quality: 100% QASM syntax valid, 0% duplicates… See the full description on the dataset page: https://huggingface.co/datasets/merileijona/quantum-circuits-8k.quant_eval_throughput_telemetry
quant_eval — Throughput telemetry
One record per generation call across all published runs, pooled into a single union schema. Carries token counts, timing, and the token-count source per call.
Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing.
Cite this dataset: 10.5281/zenodo.22009987 — concept DOI, always resolves to the latest… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_throughput_telemetry.quantum-circuits-21k
Quantum Circuits Dataset — v2 (21K)
A synthetic dataset of validated natural language → OpenQASM 2.0 circuit pairs for training quantum circuit generation models. To our knowledge the largest publicly available dataset of validated NL→QASM pairs specifically designed for generative model training.
Used to train the QuantumGPT-124M model series.
Quick Start
from datasets import load_dataset
# v2 training set (21K samples, recommended)
ds =… See the full description on the dataset page: https://huggingface.co/datasets/merileijona/quantum-circuits-21k.
