ken-jo/qev-data
QEV data Historical training snapshots for QEV, the LAYA-inspired Qwen3.5-2B decision model. Model. Configurations overlap. Do not concatenate them or assume independent test sets. ZIPs in corpora/ use qev-<stage>.zip and contain eligible original records, images and license notices. Extracted stage folders and original record/source IDs preserve the recorded training provenance. Viewer rows expose request/target schemas as JSON strings; parse with json.loads. IDs, group IDs… See the full description on the dataset page: https://huggingface.co/datasets/ken-jo/qev-data.
QEV data
Historical training snapshots for QEV, the LAYA-inspired Qwen3.5-2B decision model. Model.
Configurations overlap. Do not concatenate them or assume independent test sets.
ZIPs in corpora/ use qev-<stage>.zip and contain eligible original records, images and license notices. Extracted stage folders and original record/source IDs preserve the recorded training provenance. Viewer rows expose request/target schemas as JSON strings; parse with json.loads. IDs, group IDs, language and original split assignments are preserved.
Counts include repeated observations across stages; the sum is not a unique count.
Licenses and exclusions
This collection is not uniformly Apache-2.0. Read LICENSES.md, per-record source/license fields, and the original notices in licenses/. Authored data and typed-decisions use Apache-2.0; Beans/TrashNet use MIT; BANKING77 uses CC-BY-4.0; SNLI derivatives retain CC-BY-SA-4.0. Original observations were selected and converted into typed decision views. Synthetic labels are not observed outcomes; typed-decisions uses teacher agreement.
CIFAR-10 evaluation requests/images are excluded because redistribution rights are unresolved. excluded/ records references and hashes for direct upstream reconstruction. No selected training rows were removed by that exclusion.
Use and limitations
Extract one corpus ZIP under data/ in the source repository. Image paths are relative to that corpus. Manifests record original and exported record/image hashes. Group and image-byte split checks apply within each corpus, not across historical stages. They do not prove distinct physical objects or absence from Qwen pretraining. Korean generated examples are retained; personal project notes are not included. Old test groups have been inspected. Future models need untouched final populations. Full retraining from all historical intermediate checkpoints is not turnkey.
