datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task_data
QuantCodeEval
A benchmark for evaluating LLM coding agents on quantitative-strategy code
reproduction from finance research papers.
Status: Anonymous artifact for the 30-task benchmark.
Release mirrors
The release is mirrored at two anonymous locations:
Hugging Face Datasets — complete anonymous release:
https://huggingface.co/datasets/quantcodeeval/task_data
anonymous.4open.science — browseable mirror:
https://anonymous.4open.science/r/QuantCodeEval-Anonymous… See the full description on the dataset page: https://huggingface.co/datasets/quantcodeeval/task_data.ths-quant-factor-dictionary
THS Quant Factor Dictionary (同花顺量化因子字典)
Quantitative factor dictionaries from THS (同花顺/Tonghuashun), covering A-share and overseas markets. Includes alpha factors, Barra risk factors, sell-side consensus estimates, and real-time news factors.
These dictionaries describe the schema and metadata of THS's quantitative factor database — they do not contain actual factor values, but serve as essential references for anyone working with THS quant data.
Files… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/ths-quant-factor-dictionary.ko-quant-loss-v0
Chipsnug koqloss (v0.2.0)
This dataset measures how much more information Korean loses than English when an open model is quantized, using the same content in both languages. It is a mirror of the result tables in https://github.com/chipsnug/koqloss. The full report, in English and then Korean, is in koqloss-public.md.
Q1 — KL vs Q8_0 on parallel text (FLORES-101 devtest, sentences 1–300). For the same content, Korean loses 1.09–3.45× more than English; the 95% interval is… See the full description on the dataset page: https://huggingface.co/datasets/chipsnug/ko-quant-loss-v0.quantized-llama-3.1-humaneval-evals
Coding Benchmark Results
The coding benchmark results were obtained with the EvalPlus library.
HumanEvalpass@1
HumanEval+pass@1
meta-llama_Meta-Llama-3.1-405B-Instruct
67.3
67.5
neuralmagic_Meta-Llama-3.1-405B-Instruct-W8A8-FP8
66.7
66.6
neuralmagic_Meta-Llama-3.1-405B-Instruct-W4A16
66.5
66.4
neuralmagic_Meta-Llama-3.1-405B-Instruct-W8A8-INT8
64.3
64.8
neuralmagic_Meta-Llama-3.1-70B-Instruct-W8A8-FP8
58.1
57.7
neuralmagic_Meta-Llama-3.1-70B-Instruct-W4A16
57.1… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-humaneval-evals.quant_eval_v7_21_per_case_results_and_run_provenance
quant_eval v7.21 — Per-Case Evaluation Results and Run Provenance
Supplementary evidence for the whitepaper quant_eval: A Behavioral Evaluation Harness for
Full-Weight and Quantized Large Language Models.
Author: Patrick Hill, PBH Applied Systems, LLC
ORCID: 0009-0008-3662-1681
Licence: CC BY 4.0
Concept DOI (all versions): 10.5281/zenodo.22851375
Version DOI (this deposit): 10.5281/zenodo.22851376
What this deposit is
Every quantitative result reported in the… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_v7_21_per_case_results_and_run_provenance.p2026-002-quantization-context-compression-results
Deployed Quantization Tier and Lossy Context Compression in Extractive QA
This result dataset mirrors the version-1.0.0 reproducibility artifact:
10.5281/zenodo.22847291.
The versioned report and full replication sources
are maintained together in the research-artifacts repository. Cite the exact
Zenodo version for the frozen evidence; this Hugging Face copy is a discovery
mirror.
Matthew Schwartz — ORCID 0009-0009-4171-7247
This dataset is the aggregate-only evidence for "No… See the full description on the dataset page: https://huggingface.co/datasets/mv1137/p2026-002-quantization-context-compression-results.quantization-benchmarksquant_eval_efficiency_and_footprint
quant_eval — Efficiency and footprint
One row per published run: stored weight artifact bytes before and after quantization, compression ratio, observed evaluation wall-time ratio with an explicit direction label, the accelerator used on each lane, and token throughput.
Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing.
Cite this… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_efficiency_and_footprint.gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation… See the full description on the dataset page: https://huggingface.co/datasets/quantiles/gpqa.quant_eval_run_provenance
quant_eval — Run provenance
One row per published run: model identity, contract identifiers, fixture hash, decoding conditions, licence, and the SHA-256 and byte size of both weight artifacts. Accompanied by the calibration lineage that informed each published run.
Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing.
Cite this dataset:… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_run_provenance.african-loss-damage-quantification
African Loss and Damage Quantification | Africa (original)
Size category: 10K<n<100K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/african-loss-damage-quantification.quant_eval_paired_degradation_statistics
quant_eval — Paired degradation statistics
One row per run per task family: the paired pass-rate difference with a 95% confidence interval, the two-sided exact McNemar test, the full discordance breakdown, and a semantic-cluster-adjusted delta and interval.
Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing.
Cite this dataset:… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_paired_degradation_statistics.ml-quant-trading-synthetic
ml-quant-trading deterministic synthetic panel
This dataset is the zero-account smoke-test panel generated by
ml-quant-trading. It contains
no real instruments, proprietary market data, or investment signals.
Run the full pipeline
Install the latest verified release and run the end-to-end synthetic demo:
python -m pip install --upgrade mlquantx
mlquant demo
The PyPI distribution is mlquantx; the import package and CLI remain mlquant. The demo covers data… See the full description on the dataset page: https://huggingface.co/datasets/dddyym/ml-quant-trading-synthetic.quant_eval_behavioral_per_case_results
quant_eval — Per-case behavioral results
One row per evaluation case per runner: the raw model output, every scored signal, per-case timing, expected/got pairs, the oracle trace, the fuzz audit envelope, and the decoding conditions under which the row was produced. Every aggregate statistic published in the other datasets is recomputable from this file.
Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_behavioral_per_case_results.quantarena-artifacts
QuantArena Artifact Bundle
Reproducibility artifacts for the paper QuantArena: Beat the Market or Be the
Market? A Live-Market Evaluation of Investment Paradigms (NeurIPS 2026
Evaluations & Datasets Track submission).
Summary
QuantArena is a controlled live-market evaluation protocol that holds the LLM
backend, market data stream, analyst workflow, capital, and execution harness
fixed across runs and varies only the investment doctrine (the policy
module). This bundle… See the full description on the dataset page: https://huggingface.co/datasets/NIPS26Repo/quantarena-artifacts.llm-quant-degradation
Capability-Specific Degradation of Quantized Small Language Models — raw evaluation outputs
Raw and aggregated evaluation results for the study Capability-Specific Degradation Patterns
in Quantized Small Language Models: seven open instruction-tuned small LLMs (1–4B parameters,
five architecture families) evaluated at FP16 and 4-bit (bitsandbytes NF4) across
six capabilities, for 84 controlled model × precision × benchmark evaluations.
Paper: Rahimov, E. (2026).… See the full description on the dataset page: https://huggingface.co/datasets/Emil-7/llm-quant-degradation.quant_eval_golden_oracle_fixtures
quant_eval — Golden oracle fixtures
The locked evaluation fixture set — 1,600 cases across eight agent task families with their deterministic ground truth — plus the crosswalk mapping every run's recorded fixture hash and version label to the published file.
Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing.
Cite this dataset:… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_golden_oracle_fixtures.quantum-cyber-attack-synthetic-logs-100kdiffusers-quantization-benchmarksquantum-machine-learninga continuous data scrape of arxiv and google scholar papers of quantum machine learning papers particularly regarding climate.
quantum-gate-sequence-instability-v0.1
quantum-gate-sequence-instability-v0.1
What this dataset does
This dataset evaluates whether models can detect instability in quantum gate sequences.
Each row represents a simplified quantum circuit execution scenario described through observable device and circuit proxies.
The task is to determine whether the gate sequence remains executable inside a stable coherence window or becomes unstable.
Core stability idea
Quantum gate sequences become unstable when… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/quantum-gate-sequence-instability-v0.1.quantum-error-correction-failure-v0.1
quantum-error-correction-failure-v0.1
What this dataset does
This dataset evaluates whether models can detect instability in quantum error correction regimes.
Each row represents a simplified quantum computing scenario where logical qubits are protected using error correction.
The task is to determine whether the correction mechanism remains stable or fails due to noise and correction latency.
Core stability idea
Quantum error correction works by detecting and… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/quantum-error-correction-failure-v0.1.qna-quantum-information
Q&A Quantum Information Dataset
This dataset is created by digesting 500 different papers from the quantum information directory on arXiv,
the papers are located based on their relevance to the quantum information keyword.
Data retrieval
The data is extracted from the .pdf files using PyMuPDF package proxied from langchain.
Then the Q&A pair is generated by:
Generate N questions per page of the .pdf document based on its content.
We will feed each question to an LLM… See the full description on the dataset page: https://huggingface.co/datasets/CoAILab/qna-quantum-information.simpleqa-verified
SimpleQA Verified
A 1,000-prompt factuality benchmark from Google DeepMind and Google Research, designed to reliably evaluate LLM parametric knowledge.
▶ SimpleQA Verified Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code
Benchmark
SimpleQA Verified is a 1,000-prompt benchmark for reliably evaluating Large Language Models (LLMs) on short-form factuality
and parametric knowledge. The authors from Google DeepMind and Google Research build on… See the full description on the dataset page: https://huggingface.co/datasets/quantiles/simpleqa-verified.BioPharmaCatalystsquantum-entanglement-decay-instability-v0.1
quantum-entanglement-decay-instability-v0.1
What this dataset does
This dataset evaluates whether models can detect instability in entangled quantum states.
Each row represents a simplified quantum system described through observable device and interaction proxies.
The task is to determine whether the entangled state remains stable or collapses due to noise and interaction instability.
Core stability idea
Entanglement stability depends on maintaining coherent… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/quantum-entanglement-decay-instability-v0.1.Quantum_Gate_Performance_Evaluation
🧪 Quantum Gate Performance Dataset
📘 Title:
Comprehensive Quantum Gate Performance Analysis: A Comparative Study of Noise and No-Noise Effects
📂 Dataset Description:
This repository contains benchmarking results for 13 quantum gates (e.g., H, CNOT, Toffoli) tested under noisy and noise-free conditions, based on 1000 simulation runs per gate configuration. Total 26000 rows and 13 columns.
📊 Features include:
Gate Type
Execution Time
Error Rate
Fidelity… See the full description on the dataset page: https://huggingface.co/datasets/ismielabir/Quantum_Gate_Performance_Evaluation.crows_pairs
Dataset Card for CrowS-Pairs
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/quantiles/crows_pairs.h200-quantization-benchmarks
H200 Quantization Benchmarks
Benchmark results for 40 quantized and non-quantized instruction-tuned LLMs evaluated on an NVIDIA H200 MIG (Multi-Instance GPU) setup. This dataset supports reproducible comparison of quantization methods (AWQ, GPTQ, fp8, bf16) across accuracy and throughput dimensions.
Dataset Configs
Config
Description
Rows
accuracy
Per-task accuracy results from lm-eval across all models
~240
accuracy_leaderboard
Aggregated accuracy… See the full description on the dataset page: https://huggingface.co/datasets/ssakethch/h200-quantization-benchmarks.Quantum-Programming🧑💻 Overview
This dataset focuses on Quantum Programming and contains curated information that can be used for research, education, and model training. Quantum programming is an emerging field that leverages the principles of quantum mechanics to develop new algorithms and computational techniques. This dataset aims to provide structured information that can help both beginners and advanced users explore concepts, applications, and trends in quantum computing.
📂 Dataset Contents
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/as-krn/Quantum-Programming.
