Team Ai
Datasetpublic

ken-jo/qev-data

QEV data Historical training snapshots for QEV, the LAYA-inspired Qwen3.5-2B decision model. Model. Configurations overlap. Do not concatenate them or assume independent test sets. ZIPs in corpora/ use qev-<stage>.zip and contain eligible original records, images and license notices. Extracted stage folders and original record/source IDs preserve the recorded training provenance. Viewer rows expose request/target schemas as JSON strings; parse with json.loads. IDs, group IDs… See the full description on the dataset page: https://huggingface.co/datasets/ken-jo/qev-data.

sourceHugging Faceotherupdated 6d agoView on Hugging Face
0likes1kdownloads
Dataset Card

QEV data

Historical training snapshots for QEV, the LAYA-inspired Qwen3.5-2B decision model. Model.

Configurations overlap. Do not concatenate them or assume independent test sets.

ZIPs in corpora/ use qev-<stage>.zip and contain eligible original records, images and license notices. Extracted stage folders and original record/source IDs preserve the recorded training provenance. Viewer rows expose request/target schemas as JSON strings; parse with json.loads. IDs, group IDs, language and original split assignments are preserved.

ConfigurationRowsTrainDevCalibrationTest
starter-v289727008771714479
policy-v341922400448448896
policy-v665924800448448896
policy-v8167349582179017863576
foundation-v111927910841244021323866
workflow-v121910011212271223442832
workflow-v12-cohort2378811212271270322832
workflow-v12b1910011212271223442832
workflow-v12c1910011212271223442832
workflow-v12d1910011212271223442832
workflow-v12e1910011212271223442832
workflow-v12f1910011212271223442832
workflow-v13g1910011212271223442832
workflow-v13h1910011212271223442832
workflow-v13-final1910011212271223442832

Counts include repeated observations across stages; the sum is not a unique count.

Licenses and exclusions

This collection is not uniformly Apache-2.0. Read LICENSES.md, per-record source/license fields, and the original notices in licenses/. Authored data and typed-decisions use Apache-2.0; Beans/TrashNet use MIT; BANKING77 uses CC-BY-4.0; SNLI derivatives retain CC-BY-SA-4.0. Original observations were selected and converted into typed decision views. Synthetic labels are not observed outcomes; typed-decisions uses teacher agreement.

CIFAR-10 evaluation requests/images are excluded because redistribution rights are unresolved. excluded/ records references and hashes for direct upstream reconstruction. No selected training rows were removed by that exclusion.

Use and limitations

Extract one corpus ZIP under data/ in the source repository. Image paths are relative to that corpus. Manifests record original and exported record/image hashes. Group and image-byte split checks apply within each corpus, not across historical stages. They do not prove distinct physical objects or absence from Qwen pretraining. Korean generated examples are retained; personal project notes are not included. Old test groups have been inspected. Future models need untouched final populations. Full retraining from all historical intermediate checkpoints is not turnkey.

Training details.

GitHub: ken-jo/qev