sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1
Hemmingway-1 oMLX Quantization Benchmark This is the public-safe benchmark package for the Hemmingway-1 oMLX quantization study on Apple Silicon. Altworld developed and published Hemmingway-1. Bobby Pierce published these quantizations and the evaluation package. The collection links the upstream model and all six builds. Analysis revision 2, corrected on 2026-09-22, fixes A/B attribution and matching across reversed packets. Read CORRECTION.md before using the aggregate… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1.
Hemmingway-1 oMLX Quantization Benchmark
This is the public-safe benchmark package for the Hemmingway-1 oMLX quantization study on Apple Silicon.
Altworld developed and published Hemmingway-1. Bobby Pierce published these quantizations and the evaluation package. The collection links the upstream model and all six builds.
Analysis revision 2, corrected on 2026-09-22, fixes A/B attribution and matching across reversed packets. Read CORRECTION.md before using the aggregate results. The study has 15 prompts per quant and 11 hosted comparison prompts. All generations and judge records are unchanged.
The release contains the authored task prompts, selected local execution metadata, aggregate blind-judge results, reliability metadata, and the policy used to select records when a condition was run more than once.
What is in the dataset
The Dataset Viewer reads data/train.jsonl as the default configuration. The YAML files are metadata and are not benchmark rows.
The execution records contain response hashes, lengths, status, token counts, latency, time to first token, throughput, sampler settings, and model condition. Response text is deliberately not included.
Use record_type to select the row family. Task rows contain the prompt text. Local execution rows contain selected run metadata. Hosted reference rows are quality-only. Aggregate judge rows contain the pooled quality results.
The interval_percent column is retained for schema compatibility and is null in revision 2. The previous intervals treated dependent judgments as independent samples and have been withdrawn. Explicit counts and judge-level totals are in metadata/public_results.yaml. Task quality_max_tokens records the final cap selected for judging; execution max_tokens records each attempt's actual cap.
Record selection
The clean BF16 512-token rerun is the canonical reference. The earlier BF16 run contains the same 11 logical examples, response hashes, and record IDs, but its runtime summary differs. It is excluded from merged quality and runtime data. Runtime measurements from repeated attempts are never averaged.
The 512, 1024, and 2048 token caps remain separate strata. A logical key is based on source run, condition, prompt, cap, seed, repetition, chat-template settings, and runtime profile. record_id is retained as audit metadata, not as the primary deduplication key.
Provenance
The local measurements were produced by real MLX/oMLX software on an Apple M5 Max machine with 128 GB of memory. Claude monitored or orchestrated a subset of the workflow. Claude is not the inference engine, benchmark implementation, or measurement source.
Grafana is a derived visualization layer. The structured local manifests, generation records, summaries, state files, attestations, and telemetry are the source evidence for the study.
What is not included
This public release does not contain raw local model outputs, raw provider responses, judge packets, telemetry, request bodies, private answer keys, or unredacted benchmark generations. Those remain in the local evidence package.
KLD, top-k agreement, KV-cache divergence, broader task scores, and controlled runtime comparisons are pending. They are not inferred from this release.
Reproduction
Load data/train.jsonl, filter record_type to task and local_execution, preserve the prompt text and task IDs, and run each model condition with the settings recorded in metadata/benchmark_manifest.yaml and metadata/selection_policy.yaml. Keep each token-cap stratum separate.
This release publishes the task prompts, so it is a reproducibility package, not a contamination-resistant hidden evaluation set. Use a private holdout for future claims about generalization.
Validation
The package is checked for row counts, logical-key uniqueness, excluded fields, and selection-policy invariants. Analyzer regression tests exercise A/B decoding, reversed prompt order, result validation, and packet integrity. The earlier Gemini package review did not detect the analysis errors. Its historical PASS does not validate revision 2, and no new model-judge review was run for this correction.
License
No reuse license is specified in v1. The task prompts are published for benchmark reproducibility. Reuse terms will be added in a later version.
