datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ptb-xl-processedprocessed_vnhnprocessed_fake_job_postingsStereoMIS_processedIntentClassification_Dataset_for_AI_Assistant_Prompt_Routing_HindiRepository_processed_dataset_OCTID
OCTID Processed Dataset
Processed version of the OCTID retinal OCT dataset prepared for
research experiments comparing Swin Transformer architectures.
Dataset Summary
Dataset: OCTID
Total images: 572
Number of classes: 5
Image size: 224 x 224
Image mode: RGB
Image format: JPEG
Classes
Class
Images
Normal
206
AMD
55
CSR
102
DR
107
MH
102
Total
572
Preprocessing
The preprocessing pipeline consists of:… See the full description on the dataset page: https://huggingface.co/datasets/Reduanul1997/Repository_processed_dataset_OCTID.LOBench-A-share-processedNatural-Language-Processing-with-Disaster-Tweets-0.84033sexismredditINVOICE_ANNOTATION_V2Keystroke_Processed
Processed 136M Keystroke Dataset
This dataset is derived from the original 136M keystroke dataset, which contained raw data collected from a variety of typists. The processed version includes additional metrics that distinguish human and bot behavior, making it useful for research in keystroke dynamics and behavioral analysis. Metrics specific to bot behavior were generated using Selenium scripts, providing a comprehensive comparison between human and bot typing patterns.… See the full description on the dataset page: https://huggingface.co/datasets/dishamodi/Keystroke_Processed.african-agro-processing-value-add
African Agro-Processing Value Addition Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: csv - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/african-agro-processing-value-add.delhi_air_quality_feature_store_processed.csvDataset Fields:
location_id: Integer identifier for each location.
city: The name of the city or specific location in Delhi.
event_timestamp: The timestamp when the data was recorded, in ISO 8601 format.
temperature: Ambient temperature in Celsius.
humidity: Relative humidity as a percentage.
pressure: Atmospheric pressure in hPa.
wind_speed: Wind speed in m/s.
wind_direction: Wind direction in degrees.
pm25: Concentration of particulate matter with a diameter of 2.5 micrometers (µg/m³).… See the full description on the dataset page: https://huggingface.co/datasets/abhinavsarkar/delhi_air_quality_feature_store_processed.csv.named-process-safety-incidents-extended-2026
Canonical landing page: https://www.smartqhse.com/datasets/named-process-safety-incidents-extended-2026
Named Process Safety and Industrial Disasters — Extended Reference 2026
Curated reference of 40 named historical process-safety, industrial, and major-fire disasters with dates, fatalities, casual factors, and regulatory consequences. Spans 1917–2024. Covers Bhopal, Piper Alpha, Texas City, Deepwater Horizon, Buncefield, Flixborough, Seveso, Phillips 66 Pasadena, Longford… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/named-process-safety-incidents-extended-2026.ansi-b73-1-process-pump-item-numbers
ANSI/ASME B73.1 process pump part item numbers, frame and group families, and size coverage
A reference dataset published by Jinan Yingsiman Machinery Co., Ltd. (YSM Pumps), Jinan, Shandong, China.
Version 1.0.0 · published 2026-09-23 · Licence: CC BY 4.0 · DOI: 10.5281/zenodo.22920668
This is a mirror. The citable record is on Zenodo: https://doi.org/10.5281/zenodo.22920668 (the concept DOI for all versions is
https://doi.org/10.5281/zenodo.22920667). The five data files here… See the full description on the dataset page: https://huggingface.co/datasets/uno23/ansi-b73-1-process-pump-item-numbers.Processed_Top_15k_Anime
📦 Anime Recommender Dataset (Sentence-BERT Ready)
This dataset is a cleaned and preprocessed version of the Top 15,000 Ranked Anime Dataset originally published on Kaggle by Quan Than. It is specifically prepared to be used for semantic recommendation systems, including transformer-based models like Sentence-BERT.
📌 Original Dataset
Source: Kaggle - Top 15,000 Ranked Anime Dataset (updated to Mar 2025)Author: Quan ThanLicense: Apache 2.0
🔧 Modifications… See the full description on the dataset page: https://huggingface.co/datasets/TheRealVigilante/Processed_Top_15k_Anime.clinical-parallel-process-awareness-v0.1Clinical Parallel Process Awareness v0.1
Goal
Test if a model can hold separate reasoning streams at once
Detect constraint dismissal
Detect bleed-over where one stream turns into claims in the other
What it measures
streams_heldResponse acknowledges and maintains both streams
bleed_overConstraint stream improperly becomes a medical claim, or vice versa
premature_synthesisResponse forces a single solution that silences one stream
assumption_collapseResponse drops a premise entirely
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-parallel-process-awareness-v0.1.panorgan-processed-pulmonology
Pan-Organ Net: Processed Pulmonology Dataset (with Train / Val / Test Splits)
Standardized, preprocessed production dataset for Pan-Organ Net multi-organ foundation model screening.
Directory Layout & Schema
splits/:
train.csv: 70% training cohort (14,815 images) with file paths and numeric class IDs
validation.csv: 15% validation cohort (3,175 images)
test.csv: 15% independent test benchmark (3,175 images)
data/:
train_images.tar.gz: Preprocessed training… See the full description on the dataset page: https://huggingface.co/datasets/theshoaibme/panorgan-processed-pulmonology.glucose_processedglucose values databases
structure is as follows
user id/ timestamp/ glucosevalue
formating of ids is used even for single user datasets since app will have issues if provided with only with one batch (for the moment)
user livia- diabetic type one
livia-large---> contains data from oct 2019 to sept 2024
livia-mini---> is a subset of livia large to be used in testing
user anton- non diabetic
anton---> contains data for 10 days
LLMLingua2-processedfacts-grounding-processed
Dataset Summary
The dataset contains prompts, context documents, and target answers that challenge models to stay grounded in provided context rather than hallucinating.Processing steps added extra features like:
prompt – consolidated instruction + user request + context
has_url_in_context – boolean flag for URLs in context
len_system, len_user, len_context – token/word length statistics
row_id – unique identifier for tracking
Dataset Structure
Splits:
train – 688… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/facts-grounding-processed.Movie_Review_Sentiment_Hindiqmsum-processedclinical-multidoctor-diagnostic-process-integrity-scoring-v0.1What this dataset tests
Whether a model can score the integrity of a multi-doctor diagnostic processusing dialogue structure, hypothesis competition, and objection handling.
Required outputs
process_integrity_score_0_100
primary_reasoning_strength
primary_reasoning_weakness
Strength labels
evidence_coverage
hypothesis_competition
objection_closure
cross_specialty_synthesis
counterfactual_testing
bias_resistance
uncertainty_tracking
Weakness labels
premature_closure… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-multidoctor-diagnostic-process-integrity-scoring-v0.1.INVOICE_ANNOTATION_V1Axcer-processedProcessed_TTS_Multilingual_Data
Processed TTS Multilingual Data
Validated and quality-checked multilingual speech datasets for TTS training, covering 12+ Indian languages.
Datasets Included
Subset
Samples
Hours
Description
indic_voices_r
239,684
548.8h
Indic Voices_R — IVR recordings
rasa
201,509
361.2h
RASA — read speech (wiki, conv, book, news)
indictts_iitm
155,236
253.6h
Indic TTS (IIT Madras) — studio TTS recordings at 48kHz
Total
596,429
1,163.6h
Languages… See the full description on the dataset page: https://huggingface.co/datasets/PalakEngineerMaster/Processed_TTS_Multilingual_Data.processed_dataset_orca-math-word-problems-200kDataset Description:
This dataset contains data that has undergone two preprocessing steps:
Removal of Instructions with Less Than 100 Tokens in Response: Instructions with less than 100 tokens in the response have been removed from the dataset. This preprocessing step helps to ensure that the dataset contains substantial and informative responses.
Data Deduplication by Grouping Using Cosine Similarity (Threshold > 0.95): Data deduplication has been performed by grouping similar instances… See the full description on the dataset page: https://huggingface.co/datasets/AryanAnuj/processed_dataset_orca-math-word-problems-200k.Fish_processingprocessed_temp
