AlYldz/Gradio
Sheep Identification Dataset (172 Classes) This dataset is specifically curated for individual sheep identification tasks. It contains processed images of 172 different sheep IDs, derived from a Mendeley Data source and optimized for deep learning models. 1. Project Overview Individual identification is a critical task in precision livestock farming. This dataset provides a multi-class classification structure where each class represents a unique sheep ID (labeled… See the full description on the dataset page: https://huggingface.co/datasets/AlYldz/Gradio.
Sheep Identification Dataset (172 Classes)
This dataset is specifically curated for individual sheep identification tasks. It contains processed images of 172 different sheep IDs, derived from a Mendeley Data source and optimized for deep learning models.
1. Project Overview
Individual identification is a critical task in precision livestock farming. This dataset provides a multi-class classification structure where each class represents a unique sheep ID (labeled from koyun1 to koyun172).
2. Dataset Statistics
The dataset is structured to support robust training and objective evaluation:
- Number of Classes: 172 unique sheep identities.
- Total Images: 14,810 images.
- Training Set: 14,443 samples (including augmented data).
- Validation Set: 367 samples.
3. Preprocessing & Engineering (The "Pipeline")
To handle the complexity of 172 different classes, the following preprocessing steps were implemented:
- Image Standardization: All images were resized and normalized to ensure consistent input for CNN/Transformer architectures.
- Data Augmentation: Given the high number of classes, we applied augmentation to the training set to prevent overfitting and improve the model's ability to recognize sheep from different angles:
- Rotation & Flips: To simulate different camera perspectives in the field.
- Brightness/Contrast: To account for varying lighting conditions in farm environments.
- Data Splitting: A strict separation between training and validation sets was maintained to ensure the integrity of the performance metrics.
4. Source & Citation
- Original Data Source: https://data.mendeley.com/datasets/cstwjgtxfd/
- Original Authors: Sanabel Abu Jwade,Andrew Guzzomi,Ajmal Mian,Galib Muhammad Shahriar Himel,Md Masudul Islam
- Modifications: Data re-structured into ImageFolder format, augmented, and split into train/val by @azaliyasli.
5. How to Load
from datasets import load_dataset
dataset = load_dataset("AlYldz/Gradio")