KeenForgeAI/PCB_AoI_KubeEdge-corrected
PCB_AoI_KubeEdge-corrected PCB automated optical inspection (AOI) defect detection — cleaned version of the KubeEdge-Ianvs PCB-AoI public dataset. PCB 自动光学检测(AOI)缺陷检测 —— KubeEdge-Ianvs PCB-AoI 公开数据集的清理修正版。 English · 中文 English What is this? A cleaned, de-duplicated and consistently packaged version of the PCB-AoI Public Dataset, released by KubeEdge SIG AI members from China Telecom and Raisecom Technology as an industrial defect-detection… See the full description on the dataset page: https://huggingface.co/datasets/KeenForgeAI/PCB_AoI_KubeEdge-corrected.
PCBAoIKubeEdge-corrected

PCB automated optical inspection (AOI) defect detection — cleaned version of the KubeEdge-Ianvs PCB-AoI public dataset. PCB 自动光学检测(AOI)缺陷检测 —— KubeEdge-Ianvs PCB-AoI 公开数据集的清理修正版。
English
What is this?
A cleaned, de-duplicated and consistently packaged version of the PCB-AoI Public Dataset, released by KubeEdge SIG AI members from China Telecom and Raisecom Technology as an industrial defect-detection scenario of the KubeEdge-Ianvs distributed synergy AI benchmarking project.
The upstream release ships as a flat, unsplit train_data/ + test_data/ + train_data_augmentation/ trio (Pascal VOC), where the augmentation folder silently mixes the original tiles with their geometric transforms. This release keeps every image pixel unchanged, separates originals from augmentations, removes a leakage-prone slice of the augmentation set, and packages the data in both Pascal VOC and YOLO format with a reproducible train/val/test split.
⚠️ Upstream license: none stated. PCB-AoI ships no LICENSE file and no terms of use. See License & attribution — verify before commercial use.
Corrections vs. the official release
Image pixels are never modified. All 1,115 image files (233 core + 882 augmented) are byte-for-byte identical to the upstream files.
Scope — structural clean-up only. This is not a re-annotation. We separate, de-duplicate and re-package the data; we do not re-verify individual boxes (missed defects, wrong class labels, box tightness) — that requires domain expertise and is left untouched.
Structural checks performed on the released set:
Dataset at a glance
Core split (recommended evaluation)
Boxes per class (core): Bad_podu 830 · Bad_qiaojiao 181
Boxes per split: train 604 (Bad_podu 477 / Bad_qiaojiao 127) · val 75 (58 / 17) · test 332 (295 / 37)
Augmented set (training-only)
Boxes per class (augmented): Bad_podu 2,862 · Bad_qiaojiao 762
Structure
PCB_AoI_KubeEdge-corrected/
├── images/{train,val,test}/ # 600x600 RGB .jpeg (core split)
├── labels/{train,val,test}/ # YOLO: cls cx cy w h (normalized)
├── Annotations/ # Pascal VOC XML (flat, cleaned)
├── JPEGImages/ # copy of the core images (VOC layout)
├── augmented/
│ ├── images/ # 882 training-only transforms
│ ├── labels/ # YOLO labels for the augmented set
│ ├── Annotations/ # Pascal VOC XML for the augmented set
│ └── data.yaml # train-only config (originals excluded)
├── classes.txt
├── data.yaml # core: images/{train,val,test}
├── data_with_augment.yaml # core train + augmented (recommended recipe)
├── LICENSE
└── README.mdAnnotations/**/*.xml and labels/**/*.txt are two representations of the same cleaned annotations and are verified to agree. data.yaml is the clean evaluation recipe; data_with_augment.yaml is the recommended training recipe (core train + augmented).
Quick start
from ultralytics import YOLO
model = YOLO("yolov8n.pt")
model.train(data="data_with_augment.yaml", epochs=100, imgsz=640) # train with augmentation
model.val(data="data.yaml") # evaluate on clean val/testCitation
1. The original dataset — please always cite this.
@misc{pcbaoikubeedge,
title = {PCB-AoI Public Dataset},
author = {{KubeEdge SIG AI} and {China Telecom} and {Raisecom Technology}},
year = {2021},
howpublished = {KubeEdge-Ianvs distributed synergy AI benchmarking project},
url = {https://www.kaggle.com/datasets/kubeedgeianvs/pcb-aoi},
note = {Industrial defect-detection scenario of the Ianvs benchmark}
}2. This corrected release — please cite it as well. It is not identical to the official release: the augmentation set was separated and leakage-pruned, a reproducible split was added, and the data was re-packaged, so a citation to the original alone does not describe this version.
@misc{pcb_aoi_kubeedge_corrected,
author = {KeenForgeAI},
title = {PCB_AoI_KubeEdge-corrected: a cleaned release of the KubeEdge-Ianvs PCB-AoI dataset},
year = {2026},
version = {1.0},
publisher = {KeenForgeAI},
url = {https://huggingface.co/datasets/KeenForgeAI/PCB_AoI_KubeEdge-corrected},
note = {Curated by Lu Gan and Sam Li. Derived from the KubeEdge-Ianvs PCB-AoI
Public Dataset (China Telecom / Raisecom Technology). Upstream licence unstated.}
}3. The annotation tool (optional).
@software{keenforge,
author = {KeenForgeAI},
title = {KeenForge: a local-first, offline image annotation and model-training desktop tool},
year = {2026},
publisher = {KeenForgeAI},
url = {https://github.com/KeenForgeAI/KeenForge},
note = {MIT licensed. Developed by Lu Gan and Sam Li.}
}License & attribution
The upstream PCB-AoI dataset states no license. The archive ships no LICENSE file and no terms of use, and the KubeEdge-Ianvs project license covers the tooling, not the image data. The copyright status of the original images and annotations is therefore undetermined — verify it before any commercial use. See LICENSE in this repository.
Our modifications (separation, de-duplication, split, format conversion, documentation) are released under CC BY 4.0. Image pixels are unmodified and remain subject to the upstream terms.
Provenance
- Source: KubeEdge-Ianvs PCB-AoI public dataset (Kaggle
kubeedgeianvs/pcb-aoi/ KubeEdge OBS), retrieved 2026-09-29. - Upstream:
train_data173 img / 679 box ·test_data60 img / 332 box ·train_data_augmentation1,211 img / 4,753 box (173 originals + 1,038 transforms). - Verified: 233 core images (1,011 boxes, byte-identical) + 882 training-only transforms (3,624 boxes); 0 leakage, 0 invalid boxes.
中文
这是什么?
PCB-AoI 公开数据集的清理、去重、规范化打包版。该数据集由 中国电信 与 瑞斯康达 (Raisecom Technology) 的 KubeEdge SIG AI 成员,作为 KubeEdge-Ianvs 分布式协同 AI 基准测试的工业缺陷检测 场景发布。
上游以扁平的、无划分的 train_data/ + test_data/ + train_data_augmentation/ 三目录 (Pascal VOC)发布,其中增强目录把原始图与几何变换图混在一起。本版不修改任何图像像素, 将原始图与增强图分离、剔除易泄漏的增强切片,并以 Pascal VOC + YOLO 双格式打包, 附带可复现的 train/val/test 划分。
⚠️ 上游未声明许可证。 PCB-AoI 未附带 LICENSE 文件,也无使用条款。详见 许可证与署名——商用前请自行确认。
相对官方版的修正
图像像素从不修改——全部 1,115 张图(233 核心 + 882 增强)与上游文件逐字节一致。
本版范围——仅结构化清理。 这是结构性清理,不是重新标注。我们只做分离、去重、重新打包; 不重新核对单个标注框(漏标、类别标错、框松紧)——那需要领域专业知识,一律保持原样。
已执行的结构化检查(针对发布的集合):
数据集概览
核心划分(推荐用于评估)
各类框数(核心):Bad_podu 830 · Bad_qiaojiao 181
各划分框数:train 604(Bad_podu 477 / Bad_qiaojiao 127)· val 75(58 / 17)· test 332(295 / 37)
增强集(仅用于训练)
各类框数(增强):Bad_podu 2,862 · Bad_qiaojiao 762
目录结构
PCB_AoI_KubeEdge-corrected/
├── images/{train,val,test}/ # 600x600 RGB .jpeg(核心划分)
├── labels/{train,val,test}/ # YOLO:cls cx cy w h(归一化)
├── Annotations/ # Pascal VOC XML(扁平,已清理)
├── JPEGImages/ # 核心图片副本(VOC 布局)
├── augmented/
│ ├── images/ # 882 张仅训练用的变换图
│ ├── labels/ # 增强集的 YOLO 标签
│ ├── Annotations/ # 增强集的 Pascal VOC XML
│ └── data.yaml # 仅训练配置(不含原始图)
├── classes.txt
├── data.yaml # 核心:images/{train,val,test}
├── data_with_augment.yaml # 核心训练 + 增强(推荐配方)
├── LICENSE
└── README.mdAnnotations/**/*.xml 与 labels/**/*.txt 是同一份清理后标注的两种表示,已校验一致。 data.yaml 为干净的评估配方;data_with_augment.yaml 为推荐的训练配方(核心训练 + 增强)。
快速开始
from ultralytics import YOLO
model = YOLO("yolov8n.pt")
model.train(data="data_with_augment.yaml", epochs=100, imgsz=640) # 带增强训练
model.val(data="data.yaml") # 在干净的 val/test 上评估引用
使用本数据集请同时引用原始工作与本修正版:
1. 原始数据集(请务必引用)
@misc{pcbaoikubeedge,
title = {PCB-AoI Public Dataset},
author = {{KubeEdge SIG AI} and {China Telecom} and {Raisecom Technology}},
year = {2021},
howpublished = {KubeEdge-Ianvs distributed synergy AI benchmarking project},
url = {https://www.kaggle.com/datasets/kubeedgeianvs/pcb-aoi},
note = {Industrial defect-detection scenario of the Ianvs benchmark}
}2. 本修正版(请一并引用) —— 本版与官方发布并不相同:分离并对增强集做了防泄漏处理, 补充了可复现划分,并重新打包,因此只引用原数据集无法描述本版本。
@misc{pcb_aoi_kubeedge_corrected,
author = {KeenForgeAI},
title = {PCB_AoI_KubeEdge-corrected: a cleaned release of the KubeEdge-Ianvs PCB-AoI dataset},
year = {2026},
version = {1.0},
publisher = {KeenForgeAI},
url = {https://huggingface.co/datasets/KeenForgeAI/PCB_AoI_KubeEdge-corrected},
note = {Curated by Lu Gan and Sam Li. Derived from the KubeEdge-Ianvs PCB-AoI
Public Dataset (China Telecom / Raisecom Technology). Upstream licence unstated.}
}3. 标注工具(可选)
@software{keenforge,
author = {KeenForgeAI},
title = {KeenForge: a local-first, offline image annotation and model-training desktop tool},
year = {2026},
publisher = {KeenForgeAI},
url = {https://github.com/KeenForgeAI/KeenForge},
note = {MIT licensed. Developed by Lu Gan and Sam Li.}
}许可证与署名
上游 PCB-AoI 数据集未声明许可证。 压缩包内无 LICENSE 文件、无使用条款,且 KubeEdge-Ianvs 项目许可证覆盖的是工具而非图像数据。因此原始图像与标注的版权状态未定——商用前请自行确认。 详见本仓库 LICENSE。
我们对数据集的修改部分(分离、去重、划分、格式转换、文档)以 CC BY 4.0 发布。 图像像素未修改,仍受上游条款约束。
来源
- 来源:KubeEdge-Ianvs PCB-AoI 公开数据集(Kaggle
kubeedgeianvs/pcb-aoi/ KubeEdge OBS),获取于 2026-09-29 - 上游:
train_data173 图 / 679 框 ·test_data60 图 / 332 框 ·train_data_augmentation1,211 图 / 4,753 框(173 原始图 + 1,038 变换图) - 已校验:233 核心图(1,011 框,逐字节一致)+ 882 仅训练变换图(3,624 框);0 泄漏、0 非法框
