ApyHTML19/Medication_Boxes_Arabe_Latin
Medication Boxes — Arabic / Latin This dataset contains photos of medication boxes, with packaging text in Arabic and Latin script (French and others). It is annotated for COCO instance segmentation and was built for MediSeG, which segments each box so the text on it can be read afterwards. One class: 1 = medicine_box (0 = background) Format: COCO JSON (polygons + bbox) Images: original resolution, never resized Structure train/ 540 images val/… See the full description on the dataset page: https://huggingface.co/datasets/ApyHTML19/Medication_Boxes_Arabe_Latin.
Medication Boxes — Arabic / Latin
This dataset contains photos of medication boxes, with packaging text in Arabic and Latin script (French and others). It is annotated for COCO instance segmentation and was built for MediSeG, which segments each box so the text on it can be read afterwards.
- One class:
1 = medicine_box(0 = background) - Format: COCO JSON (polygons + bbox)
- Images: original resolution, never resized
Structure
train/ 540 images
val/ 68 images
test/ 68 images
annotations/
instances_train.json
instances_val.json
instances_test.jsonEach images entry also has three extra fields: source, original_split and original_file.
Sources
The drugs source (1,068 images) was left out because it has bounding boxes only and no segmentation.
Preparation
- 80/10/10 split (seed 42), grouped by source photo. Near-duplicates (dHash of the object crop) go into the same split, which fixes the train/val/test leakage in the original splits (128 images affected).
- Checks: no corrupted images, no invalid annotations, no empty masks, no exact duplicates (md5).
- Validation: every split loads with
torchvisionCocoDetection+wrap_dataset_for_transforms_v2, and amaskrcnn_resnet50_fpn(num_classes=2)forward pass gives finite losses.
Usage
from huggingface_hub import snapshot_download
from torchvision.datasets import CocoDetection, wrap_dataset_for_transforms_v2
root = snapshot_download("ApyHTML19/Medication_Boxes_AR_Latin", repo_type="dataset")
ds = CocoDetection(f"{root}/train", f"{root}/annotations/instances_train.json")
ds = wrap_dataset_for_transforms_v2(ds, target_keys=("boxes", "labels", "masks"))
img, target = ds[0]Limitations
- Small dataset (676 images), a single class, no per-drug labels.
- The
main_ar_frpolygons have 4 points, so masks are quadrilaterals that approximate the box outline.
