Team Ai
Datasetpublic

sixstringzen/hemmingway-1-omlx-quantization-evidence-v2

Hemmingway-1 Quantization Evidence v2 This package records two local evidence lanes for the Hemmingway-1 oQ4e build: teacher-forced numerical fidelity against a BF16 reference, and controlled runtime telemetry on Apple Silicon. It complements the frozen blind-preference study in Hemmingway-1 oMLX Quantization Benchmark v1. This dataset is sixstringzen/hemmingway-1-omlx-quantization-evidence-v2. The quality dataset remains unchanged because blind preference, distribution fidelity… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-evidence-v2.

sourceHugging Faceupdated 12d agoView on Hugging Face
0likes147downloads
Dataset Card

Hemmingway-1 Quantization Evidence v2

This package records two local evidence lanes for the Hemmingway-1 oQ4e build: teacher-forced numerical fidelity against a BF16 reference, and controlled runtime telemetry on Apple Silicon. It complements the frozen blind-preference study in Hemmingway-1 oMLX Quantization Benchmark v1.

This dataset is sixstringzen/hemmingway-1-omlx-quantization-evidence-v2. The quality dataset remains unchanged because blind preference, distribution fidelity, and runtime performance answer different questions.

Contents

FileEvidence laneContents
data/hemmingway-bf16-oq4e-fidelity-v1.jsonTeacher-forced fidelityModel and tokenizer provenance, aggregate and per-prompt KLD, top-k, and KV-cache summaries.
data/hemmingway-oq4e-runtime-v1.jsonControlled runtimeAggregate generation metrics, telemetry, and run provenance for the oQ4e condition.
viewer/fidelity/test.jsonlViewer tableFifteen prompt-level scalar fidelity rows derived from the fidelity JSON.
viewer/runtime/test.jsonlViewer tableOne controlled-runtime condition row derived from the runtime JSON.
RELEASE_MANIFEST.jsonRelease controlSource and table hashes, field-level privacy decisions, and the dataset identifier.

The two Viewer configurations use a test split because these rows report evaluation results. They do not contain training examples. The original JSON files remain the complete public evidence; the row tables make selected scalar results browsable. The manifest binds each table to its unchanged source artifact by SHA-256.

Study design

The fidelity run compared local BF16 and oQ4e model artifacts on 15 fixed prompts. Both tokenizers had to agree on special-token metadata, the full token-to-ID vocabulary mapping, and every prompt token sequence before a row could be emitted. KLD was calculated over the full next-token vocabulary. Top-k comparisons used k=10 and deterministic token-ID tie handling.

The model has a hybrid cache layout. The analysis compared 16 matched key/value cache layers and retained structural metadata for 48 excluded non-KV state-space layers. It did not reinterpret those state-space layers as attention KV tensors.

The runtime run used a temporary loopback-only oMLX server that discovered one model, Hemmingway-1-oQ4e-mtp. The workload contained five fixed prompts, three repetitions per prompt, a 512-token cap, one concurrent request, disabled oMLX paged caching, and AC power. The server unloaded the model at the end of the run. The host was an Apple M5 Max with 128 GiB of unified memory.

Aggregate fidelity results

MeasureValue
Fixed prompts15
Prediction positions566
Mean KL(PBF16 || PoQ4e)14.0982
Top-1 agreement0.3534%
Top-10 overlap1.3958%
Top-10 Jaccard overlap0.7753%
BF16 top-1 in oQ4e top-102.4735%
Mean KV cosine similarity0.2963
Mean KV relative L2 distance1.9011

The Viewer stores agreement and overlap values as fractions. For example, 0.0035335689 is 0.3534%. The KL values are in nats. These values describe numerical agreement on the fixed teacher-forced inputs. They do not measure writing quality, benchmark task completion, preference, safety, or usefulness in an agent workflow.

Aggregate runtime results

MeasureValue
Successful generations15 of 15
Generation errors0
Mean completion tokens208.6
Mean decode throughput23.19 tokens/s
Mean prefill throughput62.34 tokens/s
Mean time to first token0.829 s
Median end-to-end latency4.997 s
Telemetry samples288

The telemetry count consists of 144 oMLX status samples and 144 macmon samples. This is a single-condition runtime measurement. It does not rank the six quantizations or compare oQ4e runtime with BF16 runtime.

Privacy and provenance

The package omits raw prompt text, generated responses, token IDs, logits, KV tensors, local paths, credentials, and machine identifiers. The fidelity JSON preserves prompt hashes, tokenizer and model fingerprints, scalar metrics, and cache-shape metadata. The Viewer tables omit prompt hashes and expose scalar metrics, model labels, and source-artifact hashes only.

Each evidence JSON and Viewer table has a SHA-256 value in RELEASE_MANIFEST.json. The package supports an audit without exposing the private evaluation corpus or raw execution records.

Limits

This package does not make a cross-quant quality claim. The v1 blind-preference study remains the relevant evidence for that question. The fidelity values are not a quality score, and the runtime values are not a hardware-independent benchmark. An independent second capture of the large BF16-to-oQ4e divergence has not been completed. The withheld prompts mean that the public package supports provenance and aggregate review rather than an exact public rerun.