datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quantum-representations
Epsilon-Transformers Belief Analysis Dataset
This dataset contains trained neural network models and their corresponding belief state regression analysis from the Epsilon-Transformers project. The models were trained on four different stochastic processes and analyzed for their ability to learn and represent belief states.
See https://github.com/adamimos/epsilon-transformers/tree/quantum-public for codebase which generated this data.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/SimplexAI/quantum-representations.open-economic-quant-research-data
Open Economic & Quant Research Data
Versioned research content for CasualLab, Macroeconomics, Mortgage Rate Lock-In and Housing Market Dynamics, Tariff Incidence, and Order Flow to Price Impact, including project code, publishable data, fixtures, reports, tests, and reproducibility documentation.
Repository structure
CasualLab/: causal inference and policy-simulation research content.
Macroeconomics/: vintage-aware forecasting and public-source adapter research… See the full description on the dataset page: https://huggingface.co/datasets/ShawnChamberlain/open-economic-quant-research-data.quantbench-leaderboard-data
QuantBench leaderboard data
Raw benchmark data behind the QuantBench leaderboard:
calibration-quality GPTQ/AWQ quantization results across model sizes, calibration
corpora, and GPU tiers. 331 rows (239 ok / 92 failed — failed
runs are published too; a documented failure is a finding, not noise).
Models
Qwen/Qwen2.5-1.5B-Instruct (1.5B)
HuggingFaceTB/SmolLM2-1.7B-Instruct (1.7B)
deepgrove/Bonsai (0.5B)
Qwen/Qwen2.5-3B-Instruct (3B) — licence pending, rows only, no… See the full description on the dataset page: https://huggingface.co/datasets/Mohaaxa/quantbench-leaderboard-data.QuantiPhy-validation
QuantiPhy (Validation Set)
Dataset Summary
QuantiPhy is a benchmark for evaluating whether vision–language models (VLMs) can perform quantitative physical inference from visual evidence, rather than producing plausible but ungrounded numerical guesses.
This repository contains the official validation set of QuantiPhy, released to support model development, ablation studies, and preliminary evaluation.The validation set represents approximately 4% of the full… See the full description on the dataset page: https://huggingface.co/datasets/PaulineLi/QuantiPhy-validation.binance-future-orderbookquantized-llama-3.1-leaderboard-v2-evals
Open LLM Leaderboard v2 Benchmark Results
This artifact contains all the data from evaluations of Neural Magic's quantized Llama-3.1 models.
These evaluations were produced with lm-evaluation-harness by running the following command:
lm_eval \
--model vllm \
--model_args pretrained="<model_path>",dtype=auto,add_bos_token=False,max_model_len=4096,tensor_parallel_size="<num_gpus>",gpu_memory_utilization=0.8,enable_chunked_prefill=True \
--apply_chat_template \… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-leaderboard-v2-evals.task_data
QuantCodeEval
A benchmark for evaluating LLM coding agents on quantitative-strategy code
reproduction from finance research papers.
Status: Anonymous artifact for the 30-task benchmark.
Release mirrors
The release is mirrored at two anonymous locations:
Hugging Face Datasets — complete anonymous release:
https://huggingface.co/datasets/quantcodeeval/task_data
anonymous.4open.science — browseable mirror:
https://anonymous.4open.science/r/QuantCodeEval-Anonymous… See the full description on the dataset page: https://huggingface.co/datasets/quantcodeeval/task_data.quant-a-share-prices
A-share OHLCV Snapshot
Exported at: 2026-04-24T05:26:11.824110Z
Files: 5351 parquet files (*.SH.parquet, *.SZ.parquet, *.BJ.parquet)
Source pipeline: quant-agent download_data.py (akshare primary + yfinance fallback)
tcga-gene-expression-quantification-open
TCGA Gene Expression Quantification — Open Access
Every open-access TCGA RNA-Seq gene expression file from the NCI Genomic Data Commons (GDC), stacked into matrices: one row per sequenced tube of RNA, called an aliquot, and one column per gene. Every value is exactly as GDC published it.
Size: 11,505 aliquots × 60,660 genes, from 10,517 patients in 33 TCGA projects
Source: GDC Data Release 46.0 (2026-08-10)
Built: 2026-09-23 04:56:55 UTC
The source files
GDC… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-gene-expression-quantification-open.quant-h-share-prices
quant-h-share-prices
市场: 港股
格式: parquet
目录: 按哈希分片到子目录(00-ff)
说明: 由本地下载任务持续补齐,仓库支持断点续传更新。
quantum-simulation-chemistry-materials
Neura Parse — Quantum Simulation of Chemistry & Materials: Encodings, VQE/QPE & Dynamics
An application-deep, code-backed vertical on simulating quantum matter: electronic-structure problems, fermion-to-qubit encodings, Hamiltonian factorizations, ground/excited-state and real-time-dynamics algorithms, and analog simulation, with end-to-end resource estimates and honest classical-competitor accounting. Built with Qiskit Nature, OpenFermion, PennyLane-QChem, and PySCF — far… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-simulation-chemistry-materials.qmof_quantum
Dataset Details
Dataset Description
QMOF is a database of electronic properties of MOFs, assembled by Rosen et al.
Jablonka et al. added gas adsorption properties.
Curated by:
License: CC-BY-4.0
Dataset Sources
No links provided
Citation
BibTeX:
@article{Rosen_2021,
doi = {10.1016/j.matt.2021.02.015},
url = {https://doi.org/10.1016%2Fj.matt.2021.02.015},
year = 2021,
month = {may},
publisher = {Elsevier {BV}},
volume = {4},
number =… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/qmof_quantum.quant-data
Quant Lab · 共享行情数据
配套代码:yanlutao-scut/quant-lab。
本仓库保存可下载的行情快照,数据源为 NumCat;下载接口返回值经过脚本解析、校验并转成 Parquet。
本仓库不授予原始行情的额外使用或转售许可,使用及共享范围以数据源授权为准。
内容与覆盖范围
本次快照:15,063 个数据文件,6.016 GB(十进制)。
数据
日期范围
有数据的交易日
行数
全市场日 K
2024-07-15 — 2026-09-11
527
2,858,821
复权因子
2024-07-15 — 2026-09-11
527
2,931,150
两融汇总
2025-08-27 — 2026-09-07
250
750
两融明细
2025-08-27 — 2026-09-07
250
1,079,096
逐笔
20260105 — 20260724
80
401,835,082
逐笔日期并不连续,也不代表每只股票都有上述全部日期。
目录… See the full description on the dataset page: https://huggingface.co/datasets/lutaoyan/quant-data.quants
QuAnTS: Question Answering on Time Series – Dataset
QuAnTS is a challenging dataset designed to bridge the gap in question-answering research on time series data.
The dataset features a wide variety of questions and answers concerning human movements, presented as tracked skeleton trajectories.
QuAnTS also includes human reference performance to benchmark the practical usability of models trained on this dataset.
At present, there is no official leaderboard for this… See the full description on the dataset page: https://huggingface.co/datasets/dasyd/quants.polymarket-quant-bench
Polymarket Quant Bench — OHLCV bars for high-liquidity resolved markets
A teaching-grade derivative of the Polymarket
on-chain trade history, designed for indicator-engineering and
walk-forward back-testing assignments. Polymarket is a decentralised
prediction market on the Polygon blockchain; binary contracts settle to
one USDC if the underlying event occurs and zero otherwise.
This release replaces the per-trade fill stream of the upstream dataset
with per-token OHLCV bars at… See the full description on the dataset page: https://huggingface.co/datasets/smf-ulm/polymarket-quant-bench.quant-ai-trade-historyKAC-QuantFin-1M
Kreasof AI Capital (KAC) QuantFin-1M
We employ GARCH-based synthetic price generation, combined with our proprietary trading algorithm. The synthetic price path originally length in 28800 steps with 1-minute interval, but subsampled to 1028 steps with 28-minutes interval to save storage.
License: cc-by-nc-sa-4.0 (non-commercial)
HPLT3_DE_0.9_Quantile_Adult_Filteredstock-valuation-quantile-cn-hk-us
A股 / 港股 / 美股 估值分位数数据集
A-share / Hong Kong / US Stock Valuation Quantile Dataset
由 股查查(StockCheck) 数据团队整理产出,覆盖沪深A股、港股通及美股主要标的的 PE / PB / PS / 股息率 / 回购率 五项估值因子,并给出每只个股在「全市场」「所属行业」「自身历史 3/5/10/20 年分布」六个维度下的分位数排名。
Produced by the StockCheck (股查查) data team. Covers PE, PB, PS, dividend yield and buyback rate for Mainland China A-shares, Hong Kong-listed stocks and major US equities, with percentile rankings against the whole market, the stock's own industry, and its… See the full description on the dataset page: https://huggingface.co/datasets/stockcheck/stock-valuation-quantile-cn-hk-us.quant-pricesquantum-finance-risk-benchmark
Quantum vs Classical Kernels on Portfolio-Risk Structure — A Synthetic Benchmark
A synthetic benchmark for learning systemic portfolio-risk structure from a quantum
representation. As the portfolio grows from 8 to 16 assets the classical kernel degrades toward
chance (0.99) while the quantum representation keeps learning (0.69): the result is a persistent
sample-efficiency gap that widens with portfolio size.
Each example is a correlated-asset risk regime encoded as a quantum… See the full description on the dataset page: https://huggingface.co/datasets/SiriusQuantum/quantum-finance-risk-benchmark.hemmingway-1-omlx-quantization-evidence-v2
Hemmingway-1 Quantization Evidence v2
This package records two local evidence lanes for the Hemmingway-1 oQ4e build: teacher-forced numerical fidelity against a BF16 reference, and controlled runtime telemetry on Apple Silicon. It complements the frozen blind-preference study in Hemmingway-1 oMLX Quantization Benchmark v1.
This dataset is sixstringzen/hemmingway-1-omlx-quantization-evidence-v2. The quality dataset remains unchanged because blind preference, distribution fidelity… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-evidence-v2.quantum-physics-0.6-corpus
quantum-physics-0.6-corpus
Dataset Description
This is a domain-specific corpus created using ontology-guided filtering from FineWeb-Edu.
Dataset Creation
Source: HuggingFaceFW/fineweb-edu
Filtering Method: Semantic similarity to subdomain centroids (embedding-based)
Pipeline: Ontology-Guided Domain Corpus Builder
Dataset Structure
Each chunk contains:
text: The text content (256-512 tokens)
subdomain_id: Assigned subdomain
similarity_score:… See the full description on the dataset page: https://huggingface.co/datasets/konsman/quantum-physics-0.6-corpus.quantum-computing
Neura Parse — Quantum Computing
A multi-format quantum computing dataset spanning theory and hardware — from qubits, gates, and algorithms to QPUs, error correction, quantum software (Qiskit/Cirq/PennyLane), and quantum machine learning. Records come as instruction/response pairs, open and multiple-choice Q&A, runnable code tasks, encyclopedic concepts, and pretraining-style text, so the dataset supports SFT, evaluation, and continued pretraining under one schema.
Part of… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-computing.hemmingway-1-omlx-quantization-benchmark-v1
Hemmingway-1 oMLX Quantization Benchmark
This is the public-safe benchmark package for the Hemmingway-1 oMLX
quantization study on Apple Silicon.
Altworld developed and published
Hemmingway-1. Bobby Pierce
published these quantizations and the evaluation package. The
collection
links the upstream model and all six builds.
Analysis revision 2, corrected on 2026-09-22, fixes A/B attribution and matching
across reversed packets. Read CORRECTION.md before using the
aggregate… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1.Quant_Market_Dataths-quant-factor-dictionary
THS Quant Factor Dictionary (同花顺量化因子字典)
Quantitative factor dictionaries from THS (同花顺/Tonghuashun), covering A-share and overseas markets. Includes alpha factors, Barra risk factors, sell-side consensus estimates, and real-time news factors.
These dictionaries describe the schema and metadata of THS's quantitative factor database — they do not contain actual factor values, but serve as essential references for anyone working with THS quant data.
Files… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/ths-quant-factor-dictionary.quantum-worldline-research
Quantum Worldline Research Data
Structured research data from the Quantum Worldline project - an AI-assisted research program investigating holographic forces in MERA tensor networks, worldline path integrals on AdS spacetime, and quantum simulation of lattice gauge theories.
Dataset Description
This dataset contains the complete structured output of the Quantum Worldline multi-agent research system, which automates the research cycle: discover - hypothesize - gate - test… See the full description on the dataset page: https://huggingface.co/datasets/Jonboy648/quantum-worldline-research.quant-fidelity-corpora
Quantization fidelity corpora
Evaluation text for measuring how faithfully a quantized LLM reproduces its full-precision parent
(KL divergence of next-token distributions, top-1 agreement, perplexity ratio), as used by
Agention for the Signal and Qwen3.8-27B quantization campaigns.
mixedweb-v1
mixedweb-v1.txt (800,789 chars, 301 documents, md5 51e0045e8cabf37922aa82766a25b7b4) is a
seeded random slice of HuggingFaceFW/fineweb
sample-10BT: general English web text… See the full description on the dataset page: https://huggingface.co/datasets/agentionai/quant-fidelity-corpora.ko-quant-loss-v0
Chipsnug koqloss (v0.2.0)
This dataset measures how much more information Korean loses than English when an open model is quantized, using the same content in both languages. It is a mirror of the result tables in https://github.com/chipsnug/koqloss. The full report, in English and then Korean, is in koqloss-public.md.
Q1 — KL vs Q8_0 on parallel text (FLORES-101 devtest, sentences 1–300). For the same content, Korean loses 1.09–3.45× more than English; the 95% interval is… See the full description on the dataset page: https://huggingface.co/datasets/chipsnug/ko-quant-loss-v0.
