datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Plant_Detection_Classification
Plant Species Classification Dataset
A comprehensive dataset containing 64 different plant species with high-quality images for machine learning and computer vision applications.
Model Trainig code and trainned model with detailed performance analysis is present on the Github
https://github.com/jameelkhalidawan/Plant-Detection-Model-using-Yolo
Dataset Overview
This dataset is designed for plant species classification tasks and contains images of various plant… See the full description on the dataset page: https://huggingface.co/datasets/jameelkhalidawan/Plant_Detection_Classification.intel-image-classification
Intel Image Classification
The Intel Image Classification dataset contains images of natural scenes categorized into six classes:
Buildings
Forest
Glacier
Mountain
Sea
Street
📆 Content
The dataset contains ~25,000 images of size 150x150 pixels.
Images are evenly distributed across 6 categories:
{'buildings' -> 0,
'forest' -> 1,
'glacier' -> 2,
'mountain' -> 3,
'sea' -> 4,
'street' -> 5 }
It is divided into three parts:
Training set: ~14… See the full description on the dataset page: https://huggingface.co/datasets/sfarrukhm/intel-image-classification.autotrain-data-galaxy_classification
AutoTrain Dataset for project: galaxy_classification
Dataset Description
This dataset has been automatically processed by AutoTrain for project galaxy_classification.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<256x256 RGB PIL image>",
"target": 0
},
{
"image": "<256x256 RGB PIL image>",
"target": 0
}]… See the full description on the dataset page: https://huggingface.co/datasets/Xanadu00/autotrain-data-galaxy_classification.aidovecl-vehicle-detection-classification-localization
AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization
We introduce an annotated AI-generated dataset of eye-level vehicle images using outpainting, offering versatile generation of diverse vehicle classes in varied contexts with pretrained models.
Citation Notice
Please ensure that all publications and presentations using this data reference the following paper:
Kazemi, A., Fatima, Q. ul A., Kindratenko, V., & Tessum, C. W.… See the full description on the dataset page: https://huggingface.co/datasets/amir-kazemi/aidovecl-vehicle-detection-classification-localization.AfriMCQA-category-classification
Afri-MCQA cross-modal cultural category classification (MTEB)
Classify the cultural category of an entry from its photograph and the question
about it spoken by a native speaker, across 16 African languages.
Labels index this list:
geography, building, and landmarks
public figure and pop culture
cooking and food
objects, materials, clothing
tranditions, art, and history
brands, products, and companies
plants and animals
people, and everyday life
vehicles and transportation… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/AfriMCQA-category-classification.data-csgo-weapon-classification
Dataset for project: csgo-weapon-classification
Dataset Description
This dataset has for project csgo-weapon-classification was collected with the help of a bulk google image downloader.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<1768x718 RGB PIL image>",
"target": 0
},
{
"image": "<716x375 RGBA PIL image>"… See the full description on the dataset page: https://huggingface.co/datasets/Kaludi/data-csgo-weapon-classification.multimodal_meme_classification_singapore
Dataset Card for Offensive Memes in Singapore Context
Dataset Details
Dataset Description
This dataset is a collection of memes from various existing datasets, online forums, and freshly scrapped contents. It contains both global-context memes and Singapore-context memes, in different splits. It has textual description and a label stating if it is offensive under Singapore society's standards.
Curated by: Cao Yuxuan, Wu Jiayang, Alistair Cheong, Theodore Lee… See the full description on the dataset page: https://huggingface.co/datasets/aliencaocao/multimodal_meme_classification_singapore.plant_classification_v11food-category-classification-v2.0
Dataset for project: food-category-classification-v2.0
Dataset Description
This dataset for project food-category-classification-v2.0 was scraped with the help of a bulk google image downloader.
Dataset Structure
Dataset Fields
The dataset has the following fields (also called "features"):
{
"image": "Image(decode=True, id=None)",
"target": "ClassLabel(names=['Bread', 'Dairy', 'Dessert', 'Egg', 'Fried Food', 'Fruit', 'Meat', 'Noodles', 'Rice'… See the full description on the dataset page: https://huggingface.co/datasets/Kaludi/food-category-classification-v2.0.chest-xray-classification
Dataset Labels
['NORMAL', 'PNEUMONIA']
Number of Images
{'train': 4077, 'test': 582, 'valid': 1165}
How to Use
Install datasets:
pip install datasets
Load the dataset:
from datasets import load_dataset
ds = load_dataset("keremberke/chest-xray-classification", name="full")
example = ds['train'][0]
Roboflow Dataset Page
https://universe.roboflow.com/mohamed-traore-2ekkp/chest-x-rays-qjmia/dataset/2
Citation… See the full description on the dataset page: https://huggingface.co/datasets/keremberke/chest-xray-classification.OpenWhistle-Classification-Finetuning
OpenWhistle Classification Finetuning Dataset
dolphinteam/OpenWhistle-Classification-Finetuning is the public
classification finetuning dataset used for dolphin whistle identity
classification. It contains short whistle clips, whistle-level metadata,
fundamental-frequency tracks, rendered F0 spectrograms, and integer class
labels.
The main reviewer-facing subset is the balanced balanced config. It contains
six classes:
NSW_1 (label=0)
SW_Luna (label=1)
SW_Nana (label=2)
SW_Neo… See the full description on the dataset page: https://huggingface.co/datasets/dolphinteam/OpenWhistle-Classification-Finetuning.data-food-classification
Dataset for project: food-classification
Dataset Description
This dataset has been processed for project food-classification.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<308x512 RGB PIL image>",
"target": 0
},
{
"image": "<512x512 RGB PIL image>",
"target": 0
}]
Dataset Fields
The dataset has the… See the full description on the dataset page: https://huggingface.co/datasets/Kaludi/data-food-classification.autotrain-data-image-classification
AutoTrain Dataset for project: image-classification
Dataset Description
This dataset has been automatically processed by AutoTrain for project image-classification.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<79x80 RGBA PIL image>",
"target": 1
},
{
"image": "<547x108 RGBA PIL image>",
"target": 1
}]… See the full description on the dataset page: https://huggingface.co/datasets/fsuarez/autotrain-data-image-classification.hagrid-classification-512p-dataset
Dataset Card for "hagrid-classification-512p-dataset"
More Information needed
smoking_classificationhappy-whale-dolphin-classificationautotrain-data-weather-classification
AutoTrain Dataset for project: weather-classification
Dataset Description
This dataset has been automatically processed by AutoTrain for project weather-classification.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<771x514 RGB PIL image>",
"target": 2
},
{
"image": "<269x254 RGB PIL image>",
"target": 8
}]… See the full description on the dataset page: https://huggingface.co/datasets/dazzle-nu/autotrain-data-weather-classification.medicinal-leaf-classification-dataset
🌿 Medicinal Leaf Classification Dataset
An image dataset of 3 medicinal plant leaves — Aloe Vera, Neem, and Tulsi — used to train and evaluate deep learning classifiers.
Dataset Summary
Property
Value
Total Images
~7,380 (train + val)
Classes
3
Image Format
JPEG / PNG
Task
Image Classification
Classes
Index
Class
Train+Val Images
Test Images
0
Aloe Vera
—
183
1
Neem
—
453
2
Tulsi
—
284
Splits… See the full description on the dataset page: https://huggingface.co/datasets/poojan-s/medicinal-leaf-classification-dataset.indoor-scene-classification
Dataset Labels
['meeting_room', 'cloister', 'stairscase', 'restaurant', 'hairsalon', 'children_room', 'dining_room', 'lobby', 'museum', 'laundromat', 'computerroom', 'grocerystore', 'hospitalroom', 'buffet', 'office', 'warehouse', 'garage', 'bookstore', 'florist', 'locker_room', 'inside_bus', 'subway', 'fastfood_restaurant', 'auditorium', 'studiomusic', 'airport_inside', 'pantry', 'restaurant_kitchen', 'casino', 'movietheater', 'kitchen', 'waitingroom', 'artstudio', 'toystore'… See the full description on the dataset page: https://huggingface.co/datasets/keremberke/indoor-scene-classification.Garbage_Classification_YOLONotice: train set include 80% of original dataset, test and val sets have 10%.
data-food-category-classification
Dataset for project: food-category-classification
Dataset Description
This dataset is for project food-category-classification.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<512x512 RGB PIL image>",
"target": 0
},
{
"image": "<512x512 RGB PIL image>",
"target": 0
}]
Dataset Fields
The dataset has… See the full description on the dataset page: https://huggingface.co/datasets/Kaludi/data-food-category-classification.fresh_rotten_fruit_classification
Fresh Rotten Fruit Classification
A dataset for quality classification of 8 types of fruit. The dataset contains raw and augmented versions.The raw dataset contains 3,200 images.Images per class:
Fresh: 1,600
Rotten: 1,600
The augmented dataset contains 12,335 images.Images per class:
Fresh: 6,194
Rotten: 6,141
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{SULTANA2022108552,
title = {An… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/fresh_rotten_fruit_classification.Uddessho-Bangla-Multimodal-Intent-Classification
📊 Uddessho Dataset — Multimodal Author Intent Classification
Uddessho (meaning "Intent" in English) is a multimodal dataset created for author intent classification in the low-resource Bangla language.It contains 3,048 social media posts (text + images) labeled into six distinct intent types.
🏷️ Intent Categories & Label Mapping
Label ID
Class Name
0
Advocative
1
Controversial
2
Exhibitionist
3
Expressive
4
Informative
5
Promotive
📂… See the full description on the dataset page: https://huggingface.co/datasets/Mukaffi28/Uddessho-Bangla-Multimodal-Intent-Classification.bus_uc_classification-ultrasound-datasetusvc-classificationbirds-525-species-image-classificationOriginal dataset is https://www.kaggle.com/datasets/gpiosenka/100-bird-species
TN5000-thyroid-nodule-classification
TN5000 Thyroid Nodule Classification Dataset
A preprocessed, CNN-ready version of the TN5000 thyroid ultrasound dataset, cropped to individual nodule regions of interest and standardized to 224×224 PNG images for binary classification (Benign vs. Malignant).
Source Dataset
This dataset is derived from:
TN5000: An Ultrasound Image Dataset for Thyroid Nodule Detection and ClassificationXiaoxian Yu et al., Scientific Data (Nature Publishing Group), 2025DOI:… See the full description on the dataset page: https://huggingface.co/datasets/Johnyquest7/TN5000-thyroid-nodule-classification.bruised_vegetable_classification
Bruised Vegetable Classification
A dataset for classification of Bruised Vegetable Classification. The dataset contains 4,464 images across 3 classes.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{samanta2025nature,
title={Nature's best vs. bruised: A veggie edibility evaluation database},
author={Samanta, Bidisha and Banerjee, Sriparna and Das, Ranadhir and Chaudhuri, Sheli Sinha and Djemal… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/bruised_vegetable_classification.Crop_Weed_classification
Crop & Weed Classification Dataset
© 2026 Rishi. All rights reserved.
This dataset is provided under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license.
Usage
You are free to share and adapt this dataset under the following terms:
Attribution: You must give appropriate credit to the author.
Non-Commercial: You may not use this material for commercial purposes.
Citation & Paper
A research paper utilizing this… See the full description on the dataset page: https://huggingface.co/datasets/Rishi210904/Crop_Weed_classification.spark-plug-classification
Spark Plug Condition Classification
Labeled images of spark plugs for training a multi-class visual classifier. Includes four categories:
normal
carbon_fouled
oil_fouled
mechanical_damage
Optimized for use in Edge Impulse with the Transfer Learning (MobileNetV2) block.
Input: Images (96x96)
Task: Multi-class classification
Use case: Visual diagnostics and condition monitoring
Model trained and demonstrated on Edge Impulse.
Looking for anomaly detection?See the… See the full description on the dataset page: https://huggingface.co/datasets/eoinedge/spark-plug-classification.
