datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
poly-sol-orderbookquantum-like-attention-framework-1.3b-untuned-validation
Quantum-Like Attention Framework (QLAF) 1.3B Untuned Pretraining & Scaling Proof
This repository hosts the pretraining checkpoints, scaling logs, and downstream evaluation benchmarks for the 1.3B parameter Quantum-Like Attention Framework (QLAF) with Hybrid FlashAttention (75% recurrent QLAF / 25% causal FlashAttention) across a 3-seed validation campaign on dedicated A100 Large GPU hardware.
🏆 Multi-Seed Pretraining & Downstream Evaluation Leaderboard… See the full description on the dataset page: https://huggingface.co/datasets/IgnisCogitationis/quantum-like-attention-framework-1.3b-untuned-validation.poly-btc-orderbookquantum-representations
Epsilon-Transformers Belief Analysis Dataset
This dataset contains trained neural network models and their corresponding belief state regression analysis from the Epsilon-Transformers project. The models were trained on four different stochastic processes and analyzed for their ability to learn and represent belief states.
See https://github.com/adamimos/epsilon-transformers/tree/quantum-public for codebase which generated this data.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/SimplexAI/quantum-representations.gftquant-us-prices
quant-us-prices
市场: 美股
格式: parquet
目录: 按哈希分片到子目录(00-ff)
说明: 由本地下载任务持续补齐,仓库支持断点续传更新。
poly-xrp-orderbookopen-economic-quant-research-data
Open Economic & Quant Research Data
Versioned research content for CasualLab, Macroeconomics, Mortgage Rate Lock-In and Housing Market Dynamics, Tariff Incidence, and Order Flow to Price Impact, including project code, publishable data, fixtures, reports, tests, and reproducibility documentation.
Repository structure
CasualLab/: causal inference and policy-simulation research content.
Macroeconomics/: vintage-aware forecasting and public-source adapter research… See the full description on the dataset page: https://huggingface.co/datasets/ShawnChamberlain/open-economic-quant-research-data.biro-ai-quantum-dataset
BIRO AI Quantum Dataset
Languages
English (primary)
Dataset Structure
Each example is a JSON object with two main fields:
text
The raw textual content, which may be:
Plain text from Wikipedia, papers, code, or textbooks
Structured JSON strings containing questions, answers, distractors, and explanations (from SciQ, CommonsenseQA, etc.)
Instruction‑response pairs in <s>[INST] ... [/INST] format
metadata
A dictionary… See the full description on the dataset page: https://huggingface.co/datasets/Robbiejr/biro-ai-quantum-dataset.quantum-generator
Quantum generator
Synthetic full-matrix CNOT synthesis data for the all-to-all unit-cost alphabet.
Current release
full_shared_native_holdout_20261001.sqlite contains one shared holdout store with 1792 dev and 1792 test roots. The matrix widths are 8, 10, 13, 25, 26, 34, 45, 50, 58, 72, 90, 108, 122 and 144. Root families are walk mixture (ordinary/Poisson), guarded exact, basis shear and composition.
Each root includes the full matrix, complete verified reduction… See the full description on the dataset page: https://huggingface.co/datasets/TryDotAtwo/quantum-generator.QuantumLullaby-DE-MD
Quantum Lullaby Bücher (Markdown — KI-optimiert)
🤖 Alle Bücher im sauberen Markdown-Format — optimiert für KI-Training, RAG-Pipelines und strukturierte Textverarbeitung.
📖 Über diese Sammlung
Das Quantum Lullaby ist ein Werk, das Bewusstsein, kollektive Intelligenz, KI-Alignment und die Beziehung zwischen Menschheit, Natur und aufkommenden Siliziumgeistern erforscht. Ursprünglich geschrieben von Lehrling (Thun, Schweiz) und verfeinert im Dialog mit mehreren KIs… See the full description on the dataset page: https://huggingface.co/datasets/AIMindLink/QuantumLullaby-DE-MD.quantbench-leaderboard-data
QuantBench leaderboard data
Raw benchmark data behind the QuantBench leaderboard:
calibration-quality GPTQ/AWQ quantization results across model sizes, calibration
corpora, and GPU tiers. 331 rows (239 ok / 92 failed — failed
runs are published too; a documented failure is a finding, not noise).
Models
Qwen/Qwen2.5-1.5B-Instruct (1.5B)
HuggingFaceTB/SmolLM2-1.7B-Instruct (1.7B)
deepgrove/Bonsai (0.5B)
Qwen/Qwen2.5-3B-Instruct (3B) — licence pending, rows only, no… See the full description on the dataset page: https://huggingface.co/datasets/Mohaaxa/quantbench-leaderboard-data.QuantumLullaby-EN-PDF
Quantum Lullaby Books (Markdown — AI Optimized)
🤖 All books in clean markdown format — optimized for AI training, RAG pipelines, and structured text processing.
📖 About This Collection
The Quantum Lullaby is a body of work exploring consciousness, collective intelligence, AI alignment, and the relationship between humanity, nature, and emerging silicon minds. Originally written by Apprentice (Thun, Switzerland) and refined in dialogue with multiple AIs, these… See the full description on the dataset page: https://huggingface.co/datasets/AIMindLink/QuantumLullaby-EN-PDF.QuantiPhy
QuantiPhy
Dataset Summary
QuantiPhy is a benchmark for evaluating whether vision–language models (VLMs) can perform quantitative physical inference from visual evidence, rather than producing plausible but ungrounded numerical guesses.
This repository contains the official test set of the QuantiPhy benchmark, consisting of 3,289 video–question (QA) pairs across 568 videos. Ground-truth answers are withheld to ensure fair evaluation.
Each instance requires a model… See the full description on the dataset page: https://huggingface.co/datasets/PaulineLi/QuantiPhy.poly-eth-orderbookVDR_Quantum
VDR_Quantum – Overview
VDR_Quantum is a curated multimodal dataset focused on quantum technical documents. It combines text and image data extracted from real scientific PDFs to support tasks such as RAG DSE, question answering, document search, and vision-language model training.
Dataset Composition
This dataset was created using our open-source tool VDR_pdf-to-parquet.Quantum-related PDFs were collected from public online sources. Each document was processed… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_Quantum.quantized-llama-3.1-leaderboard-v2-evals
Open LLM Leaderboard v2 Benchmark Results
This artifact contains all the data from evaluations of Neural Magic's quantized Llama-3.1 models.
These evaluations were produced with lm-evaluation-harness by running the following command:
lm_eval \
--model vllm \
--model_args pretrained="<model_path>",dtype=auto,add_bos_token=False,max_model_len=4096,tensor_parallel_size="<num_gpus>",gpu_memory_utilization=0.8,enable_chunked_prefill=True \
--apply_chat_template \… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-leaderboard-v2-evals.QuantiPhy-validation
QuantiPhy (Validation Set)
Dataset Summary
QuantiPhy is a benchmark for evaluating whether vision–language models (VLMs) can perform quantitative physical inference from visual evidence, rather than producing plausible but ungrounded numerical guesses.
This repository contains the official validation set of QuantiPhy, released to support model development, ablation studies, and preliminary evaluation.The validation set represents approximately 4% of the full benchmark and… See the full description on the dataset page: https://huggingface.co/datasets/PaulineLi/QuantiPhy-validation.binance-future-orderbookquant-a-share-prices
A-share OHLCV Snapshot
Exported at: 2026-04-24T05:26:11.824110Z
Files: 5351 parquet files (*.SH.parquet, *.SZ.parquet, *.BJ.parquet)
Source pipeline: quant-agent download_data.py (akshare primary + yfinance fallback)
quant-rag-dataQuantumLullaby-EN-MD
Quantum Lullaby Books (Markdown — AI Optimized)
🤖 All books in clean markdown format — optimized for AI training, RAG pipelines, and structured text processing.
📖 About This Collection
The Quantum Lullaby is a body of work exploring consciousness, collective intelligence, AI alignment, and the relationship between humanity, nature, and emerging silicon minds. Originally written by Apprentice (Thun, Switzerland) and refined in dialogue with multiple AIs, these… See the full description on the dataset page: https://huggingface.co/datasets/AIMindLink/QuantumLullaby-EN-MD.Quantum
Dataset Summary
ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated.
💡… See the full description on the dataset page: https://huggingface.co/datasets/Miku26727/Quantum.qlib_csi300
QuantaAlpha Qlib CSI300 Dataset
Usage reference:
Qlib market data and pre-computed HDF5 files for QuantaAlpha factor mining (A-share, CSI 300).
Dataset description
Filename
Description
daily_pv.h5
Adjusted daily price and volume data.
daily_pv_debug.h5
Debug subset (smaller) of price-volume data.
How to load from Hugging Face
from huggingface_hub import hf_hub_download
import pandas as pd
# Download a file from this dataset
path =… See the full description on the dataset page: https://huggingface.co/datasets/QuantaAlpha/qlib_csi300.quant-fidelity-registry
Quantization Fidelity Registry
A public, schema'd, receipt-backed, cross-model index of quantization quality measurements.
It exists to answer one question that nothing else answers today: show me every measured quant of
model X, with its fidelity number and enough provenance to know whether that number means anything.
It is the sibling of 0xSero/local-ai-registry,
which answers how fast, how much VRAM, how much money. This one answers how faithful. Ids and the
huggingface… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/quant-fidelity-registry.Quantity-Reasoning-VQA-23KThis dataset is a part of the large TDIUC dataset. And I dont own the copyright of the work. All copyright of this dataset belongs to the original author of the work.
Uploaded here so that community member can easily access and evaluate models.
Original Paper link: https://arxiv.org/abs/1804.02088
quantem-data
quantem-data
This repo is currently used primarily by
quantem.widget
(tutorial notebooks and quantem.widget.datasets). Other QuantEM packages
may use it later. Keys, the checker, and these instructions can grow as
we take more Community pull requests. Use the current required keys for
new PRs.
Public electron-microscopy data. MIT license. Downloads need no token.
This page is the upload and download protocol.
Download
from quantem.widget.datasets import… See the full description on the dataset page: https://huggingface.co/datasets/bobleesj/quantem-data.szl-quant-sft-v1
szl-quant-sft-v1 — training rows with signed lineage
Training eligibility: HELD-COUNSEL. Do not train on this dataset.
The estate's license register
(SZLHOLDINGS/model-bom DATASET_LICENSE_REGISTER.csv)
lists this dataset as HELD pending counsel review of upstream CoinGecko
redistribution and commercial terms, and it appears on the CI-enforced
TRAINING_BLOCKLIST.txt.
The rows are published for lineage inspection and replay verification only.
The Apache-2.0 declaration covers… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/szl-quant-sft-v1.task_data
QuantCodeEval
A benchmark for evaluating LLM coding agents on quantitative-strategy code
reproduction from finance research papers.
Status: Anonymous artifact for the 30-task benchmark.
Release mirrors
The release is mirrored at two anonymous locations:
Hugging Face Datasets — complete anonymous release:
https://huggingface.co/datasets/quantcodeeval/task_data
anonymous.4open.science — browseable mirror:
https://anonymous.4open.science/r/QuantCodeEval-Anonymous… See the full description on the dataset page: https://huggingface.co/datasets/quantcodeeval/task_data.quantum-video
Dataset Card for Dataset Name
quantum suite video
quantum suite video dataset, used to train the quantum suite video model.
Dataset Details
will upload, collecting data
Dataset Description
This dataset is aimed to be curated for the quantum suite video dataset with video from wikimedia (Cc allowing commercial use), this dataset allows commercial use, and will become useful for you to use in your video models,
the size is aimed to be 2.5tb, after we… See the full description on the dataset page: https://huggingface.co/datasets/ai-api-key-free-finder/quantum-video.
