Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CyberNative-AI /qwen36-27b-gguf-bfcl-v4-quantization-pilot-corrected-v3 Qwen3.6-27B GGUF quantization on a bounded BFCL V4 pilot Q4_K_M matched Q8_0 on both tested categories: each scored 94 of 100 selected cases correct. Q5_K_M also scored 94/100; Q3_K_M scored 92/100. Read the results page · Inspect all 400 scored rows This is a post-result-corrected exploratory analysis of two selected non-live BFCL V4 categories, not a full leaderboard result. Inspect the scored rows without cloning The Hub Dataset Viewer does not render this… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative-AI/qwen36-27b-gguf-bfcl-v4-quantization-pilot-corrected-v3.text-generationn<1K1 likes378 downloads21d agoHugging Face02Tribunus-dev /tessera-quantization-research-evidence Tessera Quantization Research Evidence This dataset is the primary-source measurement evidence from an ongoing research program studying calibrated low-bit quantization (ternary, int4, vector-quantized codebooks) for LLM inference on heterogeneous AMD hardware (RDNA3 iGPU, XDNA1/2 NPU, Zen 4/5 CPU). The work is done in a fork of llama.cpp (project name "Tessera") that adds calibrated per-tensor ternary/payload4/VQ quantization, NPU offload, and RDNA3-native GPU kernels. This is… See the full description on the dataset page: https://huggingface.co/datasets/Tribunus-dev/tessera-quantization-research-evidence.audion<1K0 likes230 downloads1mo agoHugging Face03sixstringzen /hemmingway-1-omlx-quantization-evidence-v2 Hemmingway-1 Quantization Evidence v2 This package records two local evidence lanes for the Hemmingway-1 oQ4e build: teacher-forced numerical fidelity against a BF16 reference, and controlled runtime telemetry on Apple Silicon. It complements the frozen blind-preference study in Hemmingway-1 oMLX Quantization Benchmark v1. This dataset is sixstringzen/hemmingway-1-omlx-quantization-evidence-v2. The quality dataset remains unchanged because blind preference, distribution fidelity… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-evidence-v2.tabulartext-generationn<1K0 likes147 downloads12d agoHugging Face04thaki-AI /daily-paper-2026-09-17-tool-call-quantization-cliff The Tool-Call Cliff: Measuring the Accuracy Decay of Agentic Structured Output Under Low-Bit Quantization in Self-Hosted H200 Serving TL;DR — We formalize the tool-call cliff - the hypothesis that agentic structured tool calls decay faster than free-form prose under NVFP4/8-bit quantization - as an accuracy tax and a cliff ratio against a free-form control, derive two falsifiable predictions (a per-category failure-mode composition and a superlinear 8-to-4-bit tax jump), and… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-09-17-tool-call-quantization-cliff.0 likes141 downloads20d agoHugging Face05sixstringzen /hemmingway-1-omlx-quantization-benchmark-v1 Hemmingway-1 oMLX Quantization Benchmark This is the public-safe benchmark package for the Hemmingway-1 oMLX quantization study on Apple Silicon. Altworld developed and published Hemmingway-1. Bobby Pierce published these quantizations and the evaluation package. The collection links the upstream model and all six builds. Analysis revision 2, corrected on 2026-09-22, fixes A/B attribution and matching across reversed packets. Read CORRECTION.md before using the aggregate… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1.tabulartext-generationn<1K0 likes125 downloads13d agoHugging Face06ManPla /gauge-orbit-quantization-results Gauge equivalence does not survive quantisation: sweeps and results Raw sweeps, end-to-end vectors and reading tables behind Gauge equivalence does not survive quantisation: value-output orbit, norm-product bound, and consequences for rotational methods (Manuel Muñoz Plá, 2026, v2.1). The layout mirrors the code repository, so this dataset can be dropped over a clone of it and every reading script runs unchanged: https://github.com/mmunozpl/gauge-orbit-quantization (tag v2.1).… See the full description on the dataset page: https://huggingface.co/datasets/ManPla/gauge-orbit-quantization-results.1M<n<10M0 likes115 downloads6d agoHugging Face07b0sungk1m /tamperbench-quantization-qwen3-4b TamperBench + Quantization: Does Compression Act as Implicit Tampering? Motivation TamperBench evaluates explicit tampering attacks (LoRA fine-tuning, jailbreak-tuning, etc.) on LLM safety guards. Catastrophic Failure of LLM Unlearning via Quantization shows that quantization can undo safety-trained behaviors. This experiment bridges these two lines of work by adding quantization as a deployment-realistic perturbation to the TamperBench evaluation protocol. We… See the full description on the dataset page: https://huggingface.co/datasets/b0sungk1m/tamperbench-quantization-qwen3-4b.textn<1K4 likes108 downloads5mo agoHugging Face08kleinnner /camus-10-kv-cache-quantization We attempted KV cache quantization to Q4 — and documented why it fails on the qwen2vl architecture. KV Cache Quantization Attempt: type_k/type_v on qwen2vl Architecture The Problem The KV cache consumes significant memory bandwidth during autoregressive generation. On CPU-bound systems, memory bandwidth is the primary bottleneck. Quantizing the KV cache from FP16 to Q4 theoretically halves memory bandwidth requirements. What We Built We attempted… See the full description on the dataset page: https://huggingface.co/datasets/kleinnner/camus-10-kv-cache-quantization.8 likes100 downloads3mo agoHugging Face09apanferovnvidia /disaggregated-quantization-blog-assetsimagen<1K0 likes96 downloads15d agoHugging Face10jaytonde05 /blackwell-quantization-results0 likes94 downloads3mo agoHugging Face11aoiandroid /minicpm5-1b-quantization-benchmark openbmb/MiniCPM5-1B 次世代量子化(Quanto FP8 / INT4 vs BNB 4bit)実測ベンチマークレポート 対象モデル: openbmb/MiniCPM5-1B (1.16B parameters, 128k context, LlamaForCausalLM) 検証ハードウェア: NVIDIA GeForce RTX 4070 Ti (12GB GDDR6X, Ada Lovelace, Compute Capability 8.9, 第4世代Tensor Core) 実行環境: Windows / Python 3.13 / PyTorch 2.6.0+cu124 / transformers 4.57.6 / optimum-quanto 0.2.7 / bitsandbytes 0.50.0 検証日: 2026-09-19 12:12:34 1. エグゼクティブサマリー(全体比較) NVIDIA GeForce RTX 4070 Ti 実機環境において、標準ネイティブ… See the full description on the dataset page: https://huggingface.co/datasets/aoiandroid/minicpm5-1b-quantization-benchmark.texttext-generationn<1K0 likes82 downloads17d agoHugging Face12Anticloud /camus-10-kv-cache-quantization We attempted KV cache quantization to Q4 — and documented why it fails on the qwen2vl architecture. KV Cache Quantization Attempt: type_k/type_v on qwen2vl Architecture The Problem The KV cache consumes significant memory bandwidth during autoregressive generation. On CPU-bound systems, memory bandwidth is the primary bottleneck. Quantizing the KV cache from FP16 to Q4 theoretically halves memory bandwidth requirements. What We Built We attempted… See the full description on the dataset page: https://huggingface.co/datasets/Anticloud/camus-10-kv-cache-quantization.9 likes81 downloads3mo agoHugging Face13ThakiCloud /quantization-rebuild-noise-floor Quantization Rebuild Noise Floor Running the same quantization recipe twice produces two checkpoints that differ by more than most papers' reported deltas. This dataset is the measurement. We quantized Qwen3.8-27B to W4A16 with GPTQ, then ran the exact same recipe a second time — same model, same settings, same calibration set, only a different quantization run. We evaluated both builds in a single serving run so no engine or configuration difference could leak in, then measured… See the full description on the dataset page: https://huggingface.co/datasets/ThakiCloud/quantization-rebuild-noise-floor.text-generationn<1K0 likes81 downloads27d agoHugging Face14mv1137 /p2026-002-quantization-context-compression-results Deployed Quantization Tier and Lossy Context Compression in Extractive QA This result dataset mirrors the version-1.0.0 reproducibility artifact: 10.5281/zenodo.22847291. The versioned report and full replication sources are maintained together in the research-artifacts repository. Cite the exact Zenodo version for the frozen evidence; this Hugging Face copy is a discovery mirror. Matthew Schwartz — ORCID 0009-0009-4171-7247 This dataset is the aggregate-only evidence for "No… See the full description on the dataset page: https://huggingface.co/datasets/mv1137/p2026-002-quantization-context-compression-results.tabularquestion-answeringn<1K0 likes74 downloads7d agoHugging Face15derekl35 /quantization-benchmarkstabularn<1K3 likes68 downloads1y agoHugging Face16nielsr /r3al-vit-quantization-codex-trace R3AL ViT Quantization — Codex Agent Trace Codex session trace for installing the R3AL CLI and agent skill, exporting google/vit-base-patch16-224 to ONNX, performing dynamic INT8 post-training quantization on R3AL, and evaluating model size, Apple-arm64 CPU latency, and prediction fidelity on a 100-image ImageNet validation sample. The original Codex JSONL format is preserved for Hugging Face's native Agent Trace viewer. Credential values, email addresses, unrelated Gmail/Slack… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/r3al-vit-quantization-codex-trace.tabularn<1K0 likes65 downloads2mo agoHugging Face17hchautran /SAM_Quantization1 likes53 downloads1y agoHugging Face18KwabsHug /repro-robuq-pushing-dits-to-w1-58a2-via-robust-activation-quantization-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes46 downloads2mo agoHugging Face19mkvn /quantization-cache-amplification Quantization as Cache Amplification Trillion-Parameter Mixture-of-Experts Inference on a Commodity Laptop Kavin Kumar, Neural Metrics 📄 Read the paper — 11 pages What this is Weight quantization is usually justified as footprint reduction. This work argues that for offloaded mixture-of-experts inference that framing misses the leverage. The binding resource is not storage capacity but the fraction of expert slots resident in DRAM — and storage traffic depends on… See the full description on the dataset page: https://huggingface.co/datasets/mkvn/quantization-cache-amplification.documentn<1K0 likes41 downloads2mo agoHugging Face20bundaumamtom /kvcache-quantization-logs-qwen7b0 likes36 downloads2mo agoHugging Face21daksh-neo /qwen35-1bit-quantization-study This project was autonomously built using NEO — Your autonomous AI Agent. Try NEO → Extreme Quantization Feasibility Study: FP16 → 1-bit Model Under Test: Qwen/Qwen3.5-2B (proxy for production 31B deployment) Quantization Range: FP16 → INT8 → INT4 → 1-bit (W1.58A8) Hardware: NVIDIA RTX A6000 48GB Benchmark: WikiText-2 perplexity + layer sensitivity analysis Overview This study investigates whether extreme quantization down to 1-bit precision is viable for… See the full description on the dataset page: https://huggingface.co/datasets/daksh-neo/qwen35-1bit-quantization-study.0 likes26 downloads6mo agoHugging Face22droussis /quantization_calibration_data0 likes25 downloads7mo agoHugging Face23Kikinoking /Eval_dataset_quantizationtext1K<n<10K1 likes23 downloads1y agoHugging Face24openerotica /multi-turn-aware-quantization-llama-3.3-rp-testI added role headers and tokens for each turn in the LLaMA 3 Instruct format. The purpose is to test whether formatted multi-turn data can improve multi-turn performance after quantization. text100K<n<1M4 likes22 downloads2y agoHugging Face25beatsprom /llm-quantization-fine-tuning-2026 ⚡ LLM Fine-Tuning, Quantization & Model Optimization Dataset (2023–2026) This dataset contains 100 sample audit-verified research papers focusing on Large Language Model (LLM) quantization (GPTQ, AWQ, GGUF), fine-tuning (LoRA, QLoRA, PEFT), pruning, distillation, and speculative decoding. 📊 Features: 384-dimensional PyTorch Embeddings (all-MiniLM-L6-v2) for instant Vector Search NLP Sentence Extraction: Real extracted core problems & key technical innovations… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/llm-quantization-fine-tuning-2026.tabulartext-classificationn<1K0 likes21 downloads2mo agoHugging Face26derekl35 /diffusers-quantization-benchmarkstabularn<1K0 likes20 downloads1y agoHugging Face27ssakethch /h200-quantization-benchmarks H200 Quantization Benchmarks Benchmark results for 40 quantized and non-quantized instruction-tuned LLMs evaluated on an NVIDIA H200 MIG (Multi-Instance GPU) setup. This dataset supports reproducible comparison of quantization methods (AWQ, GPTQ, fp8, bf16) across accuracy and throughput dimensions. Dataset Configs Config Description Rows accuracy Per-task accuracy results from lm-eval across all models ~240 accuracy_leaderboard Aggregated accuracy… See the full description on the dataset page: https://huggingface.co/datasets/ssakethch/h200-quantization-benchmarks.tabularn<1K1 likes20 downloads6mo agoHugging Face28taozi555 /fp8-quantizationtabular1K<n<10K0 likes18 downloads2y agoHugging Face29dispatchAI /quantization-guide Quantization Guide Reference for choosing the right GGUF quantization level for mobile deployment. Q4_K_M is the recommended sweet spot — 40% of FP16 size, 92% quality. 🚀 dispatchAI textn<1K0 likes17 downloads3mo agoHugging Face30xedro98 /quantization-as-a-transfer-constraint Quantization as a Transfer Constraint: Zero-Shot Learning-Rate Transfer Survives Low Precision, but muP's Stability Margin Collapses Author: Shubhankar Kahali - Trumbo Labs, Inc - shubhankar@trumbo.dev License: CC BY 4.0 Paper: paper/quant_transfer_arxiv.pdf Abstract Maximal update parametrization (muP) licenses zero-shot hyperparameter transfer in exact arithmetic, but low-precision training perturbs precisely the coordinate magnitudes muP is designed to keep… See the full description on the dataset page: https://huggingface.co/datasets/xedro98/quantization-as-a-transfer-constraint.tabularother1K<n<10K1 likes17 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.