Team Ai
Datasetpublic

asingh15/amazon-c11-distillation-filtered

Amazon C11 quality-filtered distillation This is a high-signal SFT view of asingh15/amazon-c11-distillation. It contains six balanced rubric-writer/criterion-judge configurations. The original source remains unchanged. Filter A trajectory is retained only when its selected rubric has gold-score spread greater than 0.10 on both the selection panel and the paired held-out panel, non-constant proxy scores on both, positive Spearman correlation on both, positive… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c11-distillation-filtered.

sourceHugging Faceupdated 23d agoView on Hugging Face
0likes420downloads
Dataset Card

Amazon C11 quality-filtered distillation

This is a high-signal SFT view of `asingh15/amazon-c11-distillation`. It contains six balanced rubric-writer/criterion-judge configurations. The original source remains unchanged.

Filter

A trajectory is retained only when its selected rubric has gold-score spread greater than 0.10 on both the selection panel and the paired held-out panel, non-constant proxy scores on both, positive Spearman correlation on both, positive random-to-oracle normalized tie-robust harmonic Best@1..32 on both, and no writer truncation or recovery. The source stopping rubric is never reselected using held-out results.

All writer-prefix rows are retained. An equal number of judge rows is selected deterministically per trajectory, covering every final criterion first and then maximizing ICL-width and activation-label coverage. The native train, validation, and test splits are preserved.

ConfigTrain rowsValidationTestTrain panels keptSource-published retention
latent-state-spearman66,0125826944,68847.4%
latent-state-best32-norm-tr70,8726648184,29943.5%
latent-state-harmonic32-norm-tr74,0326648244,56546.2%
non-diverse-spearman69,1808126725,24157.2%
non-diverse-best32-norm-tr78,8768428444,86953.1%
non-diverse-harmonic32-norm-tr79,9649429185,23957.2%

Gold-score spread audit

Gold spread is max(gold) - min(gold) over each frozen 40-answer panel. Raw covers all source reviewers, Source reflects the original corpus's spread gate, and Filtered reflects this release. The filter raises mean spread and therefore enriches for higher-signal panels; it is not a difficulty-representative benchmark sample. The p10–p90 columns show that the retained data still covers a range of panel score spreads.

ConfigSideRaw meanSource meanFiltered meanFiltered p10–p90
latent-state-spearmanown0.7350.7420.7760.600–0.900
latent-state-spearmancross0.6100.6150.7010.400–0.900
latent-state-best32-norm-trown0.7350.7420.7780.600–0.900
latent-state-best32-norm-trcross0.6100.6150.7030.450–0.900
latent-state-harmonic32-norm-trown0.7350.7420.7770.600–0.900
latent-state-harmonic32-norm-trcross0.6100.6150.7030.416–0.900
non-diverse-spearmanown0.6100.6570.6990.400–0.900
non-diverse-spearmancross0.7350.7510.7740.600–0.900
non-diverse-best32-norm-trown0.6100.6570.7000.450–0.900
non-diverse-best32-norm-trcross0.7350.7510.7760.600–0.900
non-diverse-harmonic32-norm-trown0.6100.6570.7000.450–0.900
non-diverse-harmonic32-norm-trcross0.7350.7510.7760.600–0.900

Quality sidecar

Aggregate per-trajectory diagnostics are under quality/<split>/*.parquet. They include inclusion status and exclusion reasons, own/cross gold spread, proxy dispersion, Spearman, normalized harmonic Best@1..32, and role-row counts. Raw score or activation vectors and private targets are not published. The signed aggregate report is quality_audit.json.

Provenance

  • —Source dataset: asingh15/amazon-c11-distillation at 3f7302f2eb78cfa8a90110bd370237fd11e1638c
  • —Source objective catalog SHA-256: faff332d2bf465d37c7d6e9c0bb068522d80e3ae1e5f777688ec6c607f99f1d2
  • —Collection manifest SHA-256: 866c7bc6801f5fa8813a498c51eaa176ce6f5ff4042197a160f2e6162aa42d45
  • —Filter policy SHA-256: 3bafbdac25d0c745d75a44fa420ed647916ebf9f0ef64f26870c8690fe11d60c
  • —Filtered catalog SHA-256: 1b756f6f8b6124b7682db6a65c4054265b034b3e530eb03558fa675fad0e6f98
  • —Loss mask: qwen3_final_response