Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01quantcodeeval /task_data QuantCodeEval A benchmark for evaluating LLM coding agents on quantitative-strategy code reproduction from finance research papers. Status: Anonymous artifact for the 30-task benchmark. Release mirrors The release is mirrored at two anonymous locations: Hugging Face Datasets — complete anonymous release: https://huggingface.co/datasets/quantcodeeval/task_data anonymous.4open.science — browseable mirror: https://anonymous.4open.science/r/QuantCodeEval-Anonymous… See the full description on the dataset page: https://huggingface.co/datasets/quantcodeeval/task_data.tabulartext-generationn<1K2 likes906 downloads2mo agoHugging Face02SZLHOLDINGS /szl-quant-sft-v1 szl-quant-sft-v1 — training rows with signed lineage Training eligibility: HELD-COUNSEL. Do not train on this dataset. The estate's license register (SZLHOLDINGS/model-bom DATASET_LICENSE_REGISTER.csv) lists this dataset as HELD pending counsel review of upstream CoinGecko redistribution and commercial terms, and it appears on the CI-enforced TRAINING_BLOCKLIST.txt. The rows are published for lineage inspection and replay verification only. The Apache-2.0 declaration covers… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/szl-quant-sft-v1.texttext-generation1K<n<10K0 likes901 downloads12d agoHugging Face03burakaydinofficial /Quantuzo Quantuzo: KV Cache Quantization Benchmark Does KV cache quantization in llama.cpp hurt coding ability? Quantuzo measures the impact of KV cache quantization levels on real-world software engineering tasks using SWE-bench. Instead of synthetic benchmarks, models must actually browse repositories, understand code, write patches, and pass test suites. Motivation KV cache quantization (q8_0, q5_0, q4_0, etc.) significantly reduces VRAM usage during inference, making… See the full description on the dataset page: https://huggingface.co/datasets/burakaydinofficial/Quantuzo.text-generation1 likes507 downloads2mo agoHugging Face04Neura-parse /quantum-simulation-chemistry-materials Neura Parse — Quantum Simulation of Chemistry & Materials: Encodings, VQE/QPE & Dynamics An application-deep, code-backed vertical on simulating quantum matter: electronic-structure problems, fermion-to-qubit encodings, Hamiltonian factorizations, ground/excited-state and real-time-dynamics algorithms, and analog simulation, with end-to-end resource estimates and honest classical-competitor accounting. Built with Qiskit Nature, OpenFermion, PennyLane-QChem, and PySCF — far… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-simulation-chemistry-materials.tabulartext-generation100K<n<1M0 likes394 downloads3mo agoHugging Face05CyberNative-AI /qwen36-27b-gguf-bfcl-v4-quantization-pilot-corrected-v3 Qwen3.6-27B GGUF quantization on a bounded BFCL V4 pilot Q4_K_M matched Q8_0 on both tested categories: each scored 94 of 100 selected cases correct. Q5_K_M also scored 94/100; Q3_K_M scored 92/100. Read the results page · Inspect all 400 scored rows This is a post-result-corrected exploratory analysis of two selected non-live BFCL V4 categories, not a full leaderboard result. Inspect the scored rows without cloning The Hub Dataset Viewer does not render this… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative-AI/qwen36-27b-gguf-bfcl-v4-quantization-pilot-corrected-v3.text-generationn<1K1 likes385 downloads25d agoHugging Face06christopherthompson81 /quant_exploration Examining LLM Quantization Impact This document is a comparative analysis of qualitative performance degradation across Llama.cpp quantization within a single 2x7B model. My hope is that it will help people unfamiliar with quant impacts get a sense of how quantization will affect output. Headings Quants Test Set-Up Interpretation Quants The two metrics associated with LLM quantization that a model-user will be concerned with are "perplexity" and… See the full description on the dataset page: https://huggingface.co/datasets/christopherthompson81/quant_exploration.texttext-generationn<1K18 likes231 downloads3y agoHugging Face07sixstringzen /hemmingway-1-omlx-quantization-evidence-v2 Hemmingway-1 Quantization Evidence v2 This package records two local evidence lanes for the Hemmingway-1 oQ4e build: teacher-forced numerical fidelity against a BF16 reference, and controlled runtime telemetry on Apple Silicon. It complements the frozen blind-preference study in Hemmingway-1 oMLX Quantization Benchmark v1. This dataset is sixstringzen/hemmingway-1-omlx-quantization-evidence-v2. The quality dataset remains unchanged because blind preference, distribution fidelity… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-evidence-v2.tabulartext-generationn<1K0 likes153 downloads17d agoHugging Face08konsman /quantum-physics-0.6-corpus quantum-physics-0.6-corpus Dataset Description This is a domain-specific corpus created using ontology-guided filtering from FineWeb-Edu. Dataset Creation Source: HuggingFaceFW/fineweb-edu Filtering Method: Semantic similarity to subdomain centroids (embedding-based) Pipeline: Ontology-Guided Domain Corpus Builder Dataset Structure Each chunk contains: text: The text content (256-512 tokens) subdomain_id: Assigned subdomain similarity_score:… See the full description on the dataset page: https://huggingface.co/datasets/konsman/quantum-physics-0.6-corpus.tabulartext-generation1K<n<10K0 likes145 downloads10mo agoHugging Face09Neura-parse /quantum-computing Neura Parse — Quantum Computing A multi-format quantum computing dataset spanning theory and hardware — from qubits, gates, and algorithms to QPUs, error correction, quantum software (Qiskit/Cirq/PennyLane), and quantum machine learning. Records come as instruction/response pairs, open and multiple-choice Q&A, runnable code tasks, encyclopedic concepts, and pretraining-style text, so the dataset supports SFT, evaluation, and continued pretraining under one schema. Part of… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-computing.tabulartext-generation100K<n<1M1 likes135 downloads3mo agoHugging Face10sixstringzen /hemmingway-1-omlx-quantization-benchmark-v1 Hemmingway-1 oMLX Quantization Benchmark This is the public-safe benchmark package for the Hemmingway-1 oMLX quantization study on Apple Silicon. Altworld developed and published Hemmingway-1. Bobby Pierce published these quantizations and the evaluation package. The collection links the upstream model and all six builds. Analysis revision 2, corrected on 2026-09-22, fixes A/B attribution and matching across reversed packets. Read CORRECTION.md before using the aggregate… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1.tabulartext-generationn<1K0 likes129 downloads18d agoHugging Face11ReinforceNow /quantqa QuantQA: Quantitative Finance Interview Questions QuantQA is a curated dataset of 519 interview questions sourced from leading quantitative trading firms including Jane Street, Citadel, Two Sigma, Optiver, and SIG, in collaboration with CoachQuant. Topic Distribution Topic Coverage Probability 67% Combinatorics 22% Expected Value 21% Conditional Probability 14% Game Theory 11% Note: Questions may cover multiple topics Training Results… See the full description on the dataset page: https://huggingface.co/datasets/ReinforceNow/quantqa.documentquestion-answeringn<1K0 likes114 downloads9mo agoHugging Face12quantranger /opensre-incident-trajectories OpenSRE Incident-Diagnosis Trajectories Graded, multi-step SRE incident-diagnosis trajectories. A frozen LLM reads evidence through diagnostic tools (describe_pod / get_events / get_logs / get_metrics / query_traces / …), states a root cause + category + fix, and is scored on substance against ground truth. Built as a HUD v6 RL environment with a deliberate model spanning set so difficulty is legible and the within-group reward spread is real (the GRPO learning signal). 197… See the full description on the dataset page: https://huggingface.co/datasets/quantranger/opensre-incident-trajectories.tabulartext-generationn<1K0 likes90 downloads4mo agoHugging Face13Neura-parse /quantum-compilation-and-programming Neura Parse — Quantum Compilation & Programming A code-heavy vertical on the quantum software/compilation stack: turning abstract quantum circuits and unitaries into device-executable programs. Covers unitary decomposition and circuit synthesis (Euler/ZYZ, KAK/Cartan, Solovay-Kitaev, Ross-Selinger gridsynth, numerical synthesis with BQSKit), gate-set/basis transpilation to native gate sets, qubit layout/mapping and routing under connectivity constraints (SABRE, VF2, SWAP… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-compilation-and-programming.tabulartext-generation100K<n<1M0 likes87 downloads3mo agoHugging Face14emilioferrara /quantibias QuantiBias Benchmarking quantization-induced bias in large language models. QuantiBias measures a specific, under-audited failure mode: post-training quantization can leave a model's short-form safety behavior almost untouched while the bias it volunteers in open-ended generation rises. A standard audit that reads refusal rates and multiple-choice bias scores reports the compressed model unchanged; QuantiBias shows what that audit misses. Content warning. QuantiBias evaluates… See the full description on the dataset page: https://huggingface.co/datasets/emilioferrara/quantibias.text-generation1 likes85 downloads3mo agoHugging Face15aoiandroid /minicpm5-1b-quantization-benchmark openbmb/MiniCPM5-1B 次世代量子化(Quanto FP8 / INT4 vs BNB 4bit)実測ベンチマークレポート 対象モデル: openbmb/MiniCPM5-1B (1.16B parameters, 128k context, LlamaForCausalLM) 検証ハードウェア: NVIDIA GeForce RTX 4070 Ti (12GB GDDR6X, Ada Lovelace, Compute Capability 8.9, 第4世代Tensor Core) 実行環境: Windows / Python 3.13 / PyTorch 2.6.0+cu124 / transformers 4.57.6 / optimum-quanto 0.2.7 / bitsandbytes 0.50.0 検証日: 2026-09-19 12:12:34 1. エグゼクティブサマリー(全体比較) NVIDIA GeForce RTX 4070 Ti 実機環境において、標準ネイティブ… See the full description on the dataset page: https://huggingface.co/datasets/aoiandroid/minicpm5-1b-quantization-benchmark.texttext-generationn<1K0 likes83 downloads21d agoHugging Face16Neura-parse /quantum-machine-learning-models Neura Parse — Quantum Machine Learning Models: Encodings, Kernels, QNNs & Generative/Deep Architectures A hands-on, code-first vertical on quantum models that learn from data. Spans data encodings/feature maps, variational classifiers, quantum kernels/QSVMs, and quantum neural networks through modern generative and deep architectures (quantum GANs, circuit Born machines, quantum Boltzmann machines, QCNNs, quantum autoencoders, quantum RL, and quantum… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-machine-learning-models.tabulartext-generation100K<n<1M1 likes79 downloads3mo agoHugging Face17Neura-parse /quantum-information-and-complexity-theory Neura Parse — Quantum Information & Complexity Theory: Channels, Entropies, Classes & the Structure of Advantage A proof-based theoretical-foundations vertical uniting quantum information theory (channels, entropies, entanglement measures, distinguishability, capacities, Shannon theory) with quantum complexity theory and the structure of quantum advantage (classes, Hamiltonian complexity, sampling-based advantage and its verification, pseudorandomness, dequantization).… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-information-and-complexity-theory.tabulartext-generation100K<n<1M0 likes77 downloads3mo agoHugging Face18beatsprom /quant-finance-hft-trading-2026 ⚡ Quantitative Finance & High-Frequency Trading (HFT) SFT/DPO Suite (2026) Institutional-grade instruction fine-tuning and preference alignment dataset for training domain-expert Large Language Models in Quantitative Finance, Algorithmic Execution, and Ultra-Low-Latency HFT Systems. Engineered to the Mandatory Tier-1 Quality Standard: 80–150 lines of dense, production-grade C++20 and Rust per code snippet. Zero stubs, zero toy snippets, zero heap allocations on the critical… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/quant-finance-hft-trading-2026.texttext-generation1K<n<10K0 likes74 downloads1mo agoHugging Face19Neura-parse /fault-tolerant-quantum-computing Neura Parse — Fault-Tolerant Quantum Computing: QEC Codes, Decoders, Magic States & Resource Estimation A deep, Stim-informed vertical on fault tolerance — QEC code families, decoders, fault-tolerant gate constructions, and the full physical-to-logical resource-estimation pipeline. Expands the general dataset's handful of error-correction topics into research-grade coverage including the 2024-2026 milestones: surface-code below threshold, qLDPC/bivariate-bicycle memories… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/fault-tolerant-quantum-computing.tabulartext-generation100K<n<1M1 likes71 downloads3mo agoHugging Face20OO-LD /oold-quantity-schemas OO-LD quantity schemas 944 OO-LD schemas, one per QUDT quantity kind, each narrowing a shared QuantityValue base with the units that kind admits. Generated from the QUDT vocabulary. Every quantity figure in oold-llm-bench was measured against this generation, which is why it is published rather than regenerated. oold-bench fetch-corpus quantities --dest ./schemas The download is checked against a digest pinned in the benchmark:… See the full description on the dataset page: https://huggingface.co/datasets/OO-LD/oold-quantity-schemas.text-generation0 likes71 downloads6d agoHugging Face21quantcalc /WildChat-1M Dataset Card for WildChat Dataset Description Paper: https://arxiv.org/abs/2405.01470 Interactive Search Tool: https://wildvisualizer.com (paper) License: ODC-BY Language(s) (NLP): multi-lingual Point of Contact: Yuntian Deng Dataset Summary WildChat is a collection of 1 million conversations between human users and ChatGPT, alongside demographic data, including state, country, hashed IP addresses, and request headers. We collected WildChat by… See the full description on the dataset page: https://huggingface.co/datasets/quantcalc/WildChat-1M.texttext-generation100K<n<1M0 likes70 downloads10mo agoHugging Face22Qiskit /Qiskit-QuantumKatas Qiskit QuantumKatas A benchmark dataset for evaluating Large Language Models on quantum computing code generation tasks using Qiskit. Dataset Description This dataset contains 350 quantum computing tasks translated from Microsoft's QuantumKatas (originally in Q#) to Qiskit (Python). It is designed for evaluating LLMs on their ability to generate correct quantum computing code. Supported Tasks Code Generation: Given a natural language description and function… See the full description on the dataset page: https://huggingface.co/datasets/Qiskit/Qiskit-QuantumKatas.texttext-generationn<1K3 likes70 downloads5mo agoHugging Face23Neura-parse /quantum-networking-and-distributed Neura Parse — Quantum Networking, Repeaters & Distributed Quantum Computing A systems-frontier vertical on connecting quantum devices: entanglement distribution and distillation, quantum repeaters, quantum-internet protocol stacks, quantum memories/transduction, and modular/distributed quantum computing (nonlocal gates, circuit knitting across nodes, blind/verifiable delegated computation). Covers protocol and simulation methods used with tools such as NetSquid and SeQUeNCe… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-networking-and-distributed.tabulartext-generation100K<n<1M0 likes67 downloads3mo agoHugging Face24Neura-parse /advanced-quantum-algorithms Neura Parse — Advanced Quantum Algorithms: Derivations, QSVT/Block-Encoding & Hamiltonian Simulation A derivation- and resource-analyzed algorithms vertical spanning the canonical fault-tolerant canon (with full proofs, complexity, and worked traces) and the modern QSVT/block-encoding toolkit through Hamiltonian simulation, amplitude estimation, and quantum linear systems. Turns the general dataset's one-topic-per-algorithm summaries into line-by-line derivations, lower… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/advanced-quantum-algorithms.tabulartext-generation100K<n<1M1 likes66 downloads3mo agoHugging Face25Groovy-123 /QuantumAItext-classificationn>1T0 likes64 downloads1y agoHugging Face26merileijona /quantum-circuits-8k Quantum Circuits 8K Dataset A synthetic dataset of 8,129 quantum circuit examples for training language models to generate OpenQASM 2.0 code from natural language descriptions. Quick Stats Total Samples: 8,129 (description → QASM pairs) Unique Circuits: 739 base circuits Categories: 92 distinct quantum circuit types Qubit Range: 1-9 qubits Format: OpenQASM 2.0 Augmentation: 11x per circuit (original + 10 paraphrases) Quality: 100% QASM syntax valid, 0% duplicates… See the full description on the dataset page: https://huggingface.co/datasets/merileijona/quantum-circuits-8k.texttext-generation10K<n<100K1 likes64 downloads7mo agoHugging Face27quantiles /gpqa Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation… See the full description on the dataset page: https://huggingface.co/datasets/quantiles/gpqa.tabularquestion-answering1K<n<10K0 likes64 downloads3mo agoHugging Face28Lots-of-LoRAs /task740_lhoestq_answer_generation_quantity Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task740_lhoestq_answer_generation_quantity Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task740_lhoestq_answer_generation_quantity.texttext-generationn<1K0 likes62 downloads2y agoHugging Face29merileijona /quantum-circuits-21k Quantum Circuits Dataset — v2 (21K) A synthetic dataset of validated natural language → OpenQASM 2.0 circuit pairs for training quantum circuit generation models. To our knowledge the largest publicly available dataset of validated NL→QASM pairs specifically designed for generative model training. Used to train the QuantumGPT-124M model series. Quick Start from datasets import load_dataset # v2 training set (21K samples, recommended) ds =… See the full description on the dataset page: https://huggingface.co/datasets/merileijona/quantum-circuits-21k.texttext-generation10K<n<100K1 likes60 downloads7mo agoHugging Face30Neura-parse /quantum-error-mitigation-and-benchmarking Neura Parse — Quantum Error Mitigation, Characterization & Benchmarking A pre-fault-tolerance, code-backed vertical on getting trustworthy answers from noisy hardware and rigorously measuring device quality: error-mitigation techniques, characterization/tomography protocols, and benchmarking suites. Runnable Mitiq, pyGSTi, and Qiskit Experiments pipelines with honest sampling-overhead and bias/variance accounting — the practitioner and research toolkit the general dataset… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-error-mitigation-and-benchmarking.tabulartext-generation100K<n<1M0 likes60 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.