Team Ai
Datasetpublic

KeenForgeAI/PCB_AoI_KubeEdge-corrected

PCB_AoI_KubeEdge-corrected PCB automated optical inspection (AOI) defect detection — cleaned version of the KubeEdge-Ianvs PCB-AoI public dataset. PCB 自动光学检测(AOI)缺陷检测 —— KubeEdge-Ianvs PCB-AoI 公开数据集的清理修正版。 English · 中文 English What is this? A cleaned, de-duplicated and consistently packaged version of the PCB-AoI Public Dataset, released by KubeEdge SIG AI members from China Telecom and Raisecom Technology as an industrial defect-detection… See the full description on the dataset page: https://huggingface.co/datasets/KeenForgeAI/PCB_AoI_KubeEdge-corrected.

sourceHugging Faceotherupdated 11d agoView on Hugging Face
0likes211downloads
Dataset Card

PCBAoIKubeEdge-corrected

![DOI](https://doi.org/10.57967/hf/10677)

PCB automated optical inspection (AOI) defect detection — cleaned version of the KubeEdge-Ianvs PCB-AoI public dataset. PCB 自动光学检测(AOI)缺陷检测 —— KubeEdge-Ianvs PCB-AoI 公开数据集的清理修正版。

English · 中文


English

What is this?

A cleaned, de-duplicated and consistently packaged version of the PCB-AoI Public Dataset, released by KubeEdge SIG AI members from China Telecom and Raisecom Technology as an industrial defect-detection scenario of the KubeEdge-Ianvs distributed synergy AI benchmarking project.

The upstream release ships as a flat, unsplit train_data/ + test_data/ + train_data_augmentation/ trio (Pascal VOC), where the augmentation folder silently mixes the original tiles with their geometric transforms. This release keeps every image pixel unchanged, separates originals from augmentations, removes a leakage-prone slice of the augmentation set, and packages the data in both Pascal VOC and YOLO format with a reproducible train/val/test split.

⚠️ Upstream license: none stated. PCB-AoI ships no LICENSE file and no terms of use. See License & attribution — verify before commercial use.

Corrections vs. the official release

#CorrectionDetail
1Separated originals from augmentationsUpstream train_data_augmentation/ stored the 173 original tiles together with 1,038 geometric transforms of them. Here the originals are dropped (they already live in the core split) and the transforms moved to a dedicated augmented/ folder.
2Removed a validation-leaking augmentation slice156 transforms were derived from tiles that now belong to the val split. They are removed, leaving 882 transforms generated from the 147 core-train tiles only — so training on augmented/ can never leak validation content.
3Added a reproducible splitUpstream has no train/val split (only train_data 173 / test_data 60). This release provides a seeded, class-stratified 85/15 split of train_data into train/val, keeping the upstream test set intact.
4Dual formatUpstream is Pascal VOC only. This release ships VOC + YOLO from a single verified conversion.
5VOC metadata clean-upThe upstream XML carried <folder>UAV_data</folder>, <source>…UAV autolanding…</source> and <owner>ChaojieZhu</owner> boilerplate copied from an unrelated "UAV autolanding" dataset, and named .jpeg files as .jpg. All fixed.
6Packagingclasses.txt, data.yaml, augmented/data.yaml, data_with_augment.yaml added.

Image pixels are never modified. All 1,115 image files (233 core + 882 augmented) are byte-for-byte identical to the upstream files.

Scope — structural clean-up only. This is not a re-annotation. We separate, de-duplicate and re-package the data; we do not re-verify individual boxes (missed defects, wrong class labels, box tightness) — that requires domain expertise and is left untouched.

Structural checks performed on the released set:

CheckMethodResult
Filename ↔ image ↔ XML pairingbidirectional0 mismatches / orphans
Image size vs. XML <size>pixel dimensions vs. header0 discrepancies (600×600 core; 600×600 / 900×900 aug)
Out-of-range boxesxmin/ymin ≥ 0, xmax ≤ W, ymax ≤ H0
Degenerate boxeszero / negative width or height0
Image ↔ label ↔ XML box agreementVOC ↔ YOLO round-trip0 inconsistencies
Train/val/test board leakagebusiness-date disjointness + byte-identical pairs0
Augmentation ↔ val/test overlapbase-tile provenance check0 (leaking slice removed)
Class-name consistencyunique label set2 classes, no anomalies

Dataset at a glance

Core split (recommended evaluation)
PropertyValue
Images233 (600 × 600, RGB .jpeg)
Classes2 — Bad_podu (solder-paste insufficient / 少锡), Bad_qiaojiao (solder bridge / 桥连)
Bounding boxes1,011 (≈ 4.34 per image)
FormatsPascal VOC XML · YOLO TXT
Splittrain 147 / val 26 / test 60 (seeded, class-stratified; test = upstream test_data)

Boxes per class (core): Bad_podu 830 · Bad_qiaojiao 181

Boxes per split: train 604 (Bad_podu 477 / Bad_qiaojiao 127) · val 75 (58 / 17) · test 332 (295 / 37)

Augmented set (training-only)
PropertyValue
Images882 (600 × 600 or 900 × 900, RGB .jpeg)
Bounding boxes3,624
Construction6 geometric transforms of each of the 147 core-train tiles only (0 originals, 0 val/test-derived)
Transformsrotate 90° / 180° / 270° · horizontal flip · vertical flip · scale
UsageTraining only. Never use as a validation or test source.

Boxes per class (augmented): Bad_podu 2,862 · Bad_qiaojiao 762

Structure

PCB_AoI_KubeEdge-corrected/
├── images/{train,val,test}/      # 600x600 RGB .jpeg (core split)
├── labels/{train,val,test}/      # YOLO: cls cx cy w h (normalized)
├── Annotations/                  # Pascal VOC XML (flat, cleaned)
├── JPEGImages/                   # copy of the core images (VOC layout)
├── augmented/
│   ├── images/                   # 882 training-only transforms
│   ├── labels/                   # YOLO labels for the augmented set
│   ├── Annotations/              # Pascal VOC XML for the augmented set
│   └── data.yaml                 # train-only config (originals excluded)
├── classes.txt
├── data.yaml                     # core: images/{train,val,test}
├── data_with_augment.yaml        # core train + augmented (recommended recipe)
├── LICENSE
└── README.md

Annotations/**/*.xml and labels/**/*.txt are two representations of the same cleaned annotations and are verified to agree. data.yaml is the clean evaluation recipe; data_with_augment.yaml is the recommended training recipe (core train + augmented).

Quick start

python
from ultralytics import YOLO

model = YOLO("yolov8n.pt")
model.train(data="data_with_augment.yaml", epochs=100, imgsz=640)  # train with augmentation
model.val(data="data.yaml")                                        # evaluate on clean val/test

Citation

1. The original dataset — please always cite this.

bibtex
@misc{pcbaoikubeedge,
  title        = {PCB-AoI Public Dataset},
  author       = {{KubeEdge SIG AI} and {China Telecom} and {Raisecom Technology}},
  year         = {2021},
  howpublished = {KubeEdge-Ianvs distributed synergy AI benchmarking project},
  url          = {https://www.kaggle.com/datasets/kubeedgeianvs/pcb-aoi},
  note         = {Industrial defect-detection scenario of the Ianvs benchmark}
}

2. This corrected release — please cite it as well. It is not identical to the official release: the augmentation set was separated and leakage-pruned, a reproducible split was added, and the data was re-packaged, so a citation to the original alone does not describe this version.

bibtex
@misc{pcb_aoi_kubeedge_corrected,
  author    = {KeenForgeAI},
  title     = {PCB_AoI_KubeEdge-corrected: a cleaned release of the KubeEdge-Ianvs PCB-AoI dataset},
  year      = {2026},
  version   = {1.0},
  publisher = {KeenForgeAI},
  url       = {https://huggingface.co/datasets/KeenForgeAI/PCB_AoI_KubeEdge-corrected},
  note      = {Curated by Lu Gan and Sam Li. Derived from the KubeEdge-Ianvs PCB-AoI
               Public Dataset (China Telecom / Raisecom Technology). Upstream licence unstated.}
}

3. The annotation tool (optional).

bibtex
@software{keenforge,
  author    = {KeenForgeAI},
  title     = {KeenForge: a local-first, offline image annotation and model-training desktop tool},
  year      = {2026},
  publisher = {KeenForgeAI},
  url       = {https://github.com/KeenForgeAI/KeenForge},
  note      = {MIT licensed. Developed by Lu Gan and Sam Li.}
}

License & attribution

The upstream PCB-AoI dataset states no license. The archive ships no LICENSE file and no terms of use, and the KubeEdge-Ianvs project license covers the tooling, not the image data. The copyright status of the original images and annotations is therefore undetermined — verify it before any commercial use. See LICENSE in this repository.

Our modifications (separation, de-duplication, split, format conversion, documentation) are released under CC BY 4.0. Image pixels are unmodified and remain subject to the upstream terms.

Provenance

  • —Source: KubeEdge-Ianvs PCB-AoI public dataset (Kaggle kubeedgeianvs/pcb-aoi / KubeEdge OBS), retrieved 2026-09-29.
  • —Upstream: train_data 173 img / 679 box · test_data 60 img / 332 box · train_data_augmentation 1,211 img / 4,753 box (173 originals + 1,038 transforms).
  • —Verified: 233 core images (1,011 boxes, byte-identical) + 882 training-only transforms (3,624 boxes); 0 leakage, 0 invalid boxes.

中文

这是什么?

PCB-AoI 公开数据集的清理、去重、规范化打包版。该数据集由 中国电信 与 瑞斯康达 (Raisecom Technology) 的 KubeEdge SIG AI 成员,作为 KubeEdge-Ianvs 分布式协同 AI 基准测试的工业缺陷检测 场景发布。

上游以扁平的、无划分的 train_data/ + test_data/ + train_data_augmentation/ 三目录 (Pascal VOC)发布,其中增强目录把原始图与几何变换图混在一起。本版不修改任何图像像素, 将原始图与增强图分离、剔除易泄漏的增强切片,并以 Pascal VOC + YOLO 双格式打包, 附带可复现的 train/val/test 划分。

⚠️ 上游未声明许可证。 PCB-AoI 未附带 LICENSE 文件,也无使用条款。详见 许可证与署名——商用前请自行确认。

相对官方版的修正

#修正说明
1分离原始图与增强图上游 train_data_augmentation/ 把 173 张原始图与 1,038 张几何变换图混存。本版删除其中的原始图(核心划分中已有),并将变换图单独放入 augmented/。
2剔除会泄漏验证集的增强切片有 156 张变换图源自现已归入 val 的图块,本版将其删除,仅保留由 147 张核心训练图生成的 882 张变换图,因此用 augmented/ 训练绝不会泄漏验证集内容。
3补充可复现划分上游无 train/val 划分(仅 train_data 173 / test_data 60)。本版对 train_data 做固定随机种子、按类别分层的 85/15 划分,并保持上游测试集不变。
4双格式上游仅有 Pascal VOC;本版同时提供 VOC + YOLO(由同一次转换生成并校验一致)。
5VOC 元数据清理上游 XML 带有从无关的 “UAV autolanding” 数据集复制来的 <folder>UAV_data</folder>、<source>…UAV autolanding…</source>、<owner>ChaojieZhu</owner> 样板,并把 .jpeg 文件写成 .jpg。均已修正。
6规范化打包新增 classes.txt、data.yaml、augmented/data.yaml、data_with_augment.yaml。

图像像素从不修改——全部 1,115 张图(233 核心 + 882 增强)与上游文件逐字节一致。

本版范围——仅结构化清理。 这是结构性清理,不是重新标注。我们只做分离、去重、重新打包; 不重新核对单个标注框(漏标、类别标错、框松紧)——那需要领域专业知识,一律保持原样。

已执行的结构化检查(针对发布的集合):

检查方法结果
文件名↔图片↔XML 配对双向0 错配 / 孤儿
图片尺寸 vs. XML <size>像素尺寸 vs. 头信息0 不一致(核心 600×600;增强 600×600 / 900×900)
越界框xmin/ymin ≥ 0,xmax ≤ W,ymax ≤ H0
退化框宽或高为零/负0
图片↔标签↔XML 框一致性VOC ↔ YOLO 往返校验0 不一致
train/val/test 板级泄漏业务日期互斥 + 逐字节相同对0
增强集↔验证/测试集重叠基图来源检查0(泄漏切片已删)
类别名一致性唯一类别名集合2 类,无杂名

数据集概览

核心划分(推荐用于评估)
属性值
图片233(600 × 600,RGB .jpeg)
类别2 —— Bad_podu(锡膏不足/少锡)、Bad_qiaojiao(焊点桥连/桥连)
标注框1,011(约 4.34 框/图)
格式Pascal VOC XML · YOLO TXT
划分train 147 / val 26 / test 60(固定种子、类别分层;test = 上游 test_data)

各类框数(核心):Bad_podu 830 · Bad_qiaojiao 181

各划分框数:train 604(Bad_podu 477 / Bad_qiaojiao 127)· val 75(58 / 17)· test 332(295 / 37)

增强集(仅用于训练)
属性值
图片882(600 × 600 或 900 × 900,RGB .jpeg)
标注框3,624
构成由 147 张核心训练图各生成的 6 种几何变换(0 原始图,0 验证/测试来源)
变换旋转 90° / 180° / 270° · 水平翻转 · 垂直翻转 · 缩放
用途仅用于训练。 切勿作为验证或测试来源。

各类框数(增强):Bad_podu 2,862 · Bad_qiaojiao 762

目录结构

PCB_AoI_KubeEdge-corrected/
├── images/{train,val,test}/      # 600x600 RGB .jpeg(核心划分)
├── labels/{train,val,test}/      # YOLO:cls cx cy w h(归一化)
├── Annotations/                  # Pascal VOC XML(扁平,已清理)
├── JPEGImages/                   # 核心图片副本(VOC 布局)
├── augmented/
│   ├── images/                   # 882 张仅训练用的变换图
│   ├── labels/                   # 增强集的 YOLO 标签
│   ├── Annotations/              # 增强集的 Pascal VOC XML
│   └── data.yaml                 # 仅训练配置(不含原始图)
├── classes.txt
├── data.yaml                     # 核心:images/{train,val,test}
├── data_with_augment.yaml        # 核心训练 + 增强(推荐配方)
├── LICENSE
└── README.md

Annotations/**/*.xml 与 labels/**/*.txt 是同一份清理后标注的两种表示,已校验一致。 data.yaml 为干净的评估配方;data_with_augment.yaml 为推荐的训练配方(核心训练 + 增强)。

快速开始

python
from ultralytics import YOLO

model = YOLO("yolov8n.pt")
model.train(data="data_with_augment.yaml", epochs=100, imgsz=640)  # 带增强训练
model.val(data="data.yaml")                                        # 在干净的 val/test 上评估

引用

使用本数据集请同时引用原始工作与本修正版:

1. 原始数据集(请务必引用)

bibtex
@misc{pcbaoikubeedge,
  title        = {PCB-AoI Public Dataset},
  author       = {{KubeEdge SIG AI} and {China Telecom} and {Raisecom Technology}},
  year         = {2021},
  howpublished = {KubeEdge-Ianvs distributed synergy AI benchmarking project},
  url          = {https://www.kaggle.com/datasets/kubeedgeianvs/pcb-aoi},
  note         = {Industrial defect-detection scenario of the Ianvs benchmark}
}

2. 本修正版(请一并引用) —— 本版与官方发布并不相同:分离并对增强集做了防泄漏处理, 补充了可复现划分,并重新打包,因此只引用原数据集无法描述本版本。

bibtex
@misc{pcb_aoi_kubeedge_corrected,
  author    = {KeenForgeAI},
  title     = {PCB_AoI_KubeEdge-corrected: a cleaned release of the KubeEdge-Ianvs PCB-AoI dataset},
  year      = {2026},
  version   = {1.0},
  publisher = {KeenForgeAI},
  url       = {https://huggingface.co/datasets/KeenForgeAI/PCB_AoI_KubeEdge-corrected},
  note      = {Curated by Lu Gan and Sam Li. Derived from the KubeEdge-Ianvs PCB-AoI
               Public Dataset (China Telecom / Raisecom Technology). Upstream licence unstated.}
}

3. 标注工具(可选)

bibtex
@software{keenforge,
  author    = {KeenForgeAI},
  title     = {KeenForge: a local-first, offline image annotation and model-training desktop tool},
  year      = {2026},
  publisher = {KeenForgeAI},
  url       = {https://github.com/KeenForgeAI/KeenForge},
  note      = {MIT licensed. Developed by Lu Gan and Sam Li.}
}

许可证与署名

上游 PCB-AoI 数据集未声明许可证。 压缩包内无 LICENSE 文件、无使用条款,且 KubeEdge-Ianvs 项目许可证覆盖的是工具而非图像数据。因此原始图像与标注的版权状态未定——商用前请自行确认。 详见本仓库 LICENSE。

我们对数据集的修改部分(分离、去重、划分、格式转换、文档)以 CC BY 4.0 发布。 图像像素未修改,仍受上游条款约束。

来源

  • —来源:KubeEdge-Ianvs PCB-AoI 公开数据集(Kaggle kubeedgeianvs/pcb-aoi / KubeEdge OBS),获取于 2026-09-29
  • —上游:train_data 173 图 / 679 框 · test_data 60 图 / 332 框 · train_data_augmentation 1,211 图 / 4,753 框(173 原始图 + 1,038 变换图)
  • —已校验:233 核心图(1,011 框,逐字节一致)+ 882 仅训练变换图(3,624 框);0 泄漏、0 非法框