datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ptb-xl-processedprocessed_vnhnprocessed_fake_job_postingsStereoMIS_processedLOBench-A-share-processedNatural-Language-Processing-with-Disaster-Tweets-0.84033INVOICE_ANNOTATION_V2Keystroke_Processed
Processed 136M Keystroke Dataset
This dataset is derived from the original 136M keystroke dataset, which contained raw data collected from a variety of typists. The processed version includes additional metrics that distinguish human and bot behavior, making it useful for research in keystroke dynamics and behavioral analysis. Metrics specific to bot behavior were generated using Selenium scripts, providing a comprehensive comparison between human and bot typing patterns.… See the full description on the dataset page: https://huggingface.co/datasets/dishamodi/Keystroke_Processed.african-agro-processing-value-add
African Agro-Processing Value Addition Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: csv - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/african-agro-processing-value-add.delhi_air_quality_feature_store_processed.csvDataset Fields:
location_id: Integer identifier for each location.
city: The name of the city or specific location in Delhi.
event_timestamp: The timestamp when the data was recorded, in ISO 8601 format.
temperature: Ambient temperature in Celsius.
humidity: Relative humidity as a percentage.
pressure: Atmospheric pressure in hPa.
wind_speed: Wind speed in m/s.
wind_direction: Wind direction in degrees.
pm25: Concentration of particulate matter with a diameter of 2.5 micrometers (µg/m³).… See the full description on the dataset page: https://huggingface.co/datasets/abhinavsarkar/delhi_air_quality_feature_store_processed.csv.Processed_Top_15k_Anime
📦 Anime Recommender Dataset (Sentence-BERT Ready)
This dataset is a cleaned and preprocessed version of the Top 15,000 Ranked Anime Dataset originally published on Kaggle by Quan Than. It is specifically prepared to be used for semantic recommendation systems, including transformer-based models like Sentence-BERT.
📌 Original Dataset
Source: Kaggle - Top 15,000 Ranked Anime Dataset (updated to Mar 2025)Author: Quan ThanLicense: Apache 2.0
🔧 Modifications… See the full description on the dataset page: https://huggingface.co/datasets/TheRealVigilante/Processed_Top_15k_Anime.panorgan-processed-pulmonology
Pan-Organ Net: Processed Pulmonology Dataset (with Train / Val / Test Splits)
Standardized, preprocessed production dataset for Pan-Organ Net multi-organ foundation model screening.
Directory Layout & Schema
splits/:
train.csv: 70% training cohort (14,815 images) with file paths and numeric class IDs
validation.csv: 15% validation cohort (3,175 images)
test.csv: 15% independent test benchmark (3,175 images)
data/:
train_images.tar.gz: Preprocessed training… See the full description on the dataset page: https://huggingface.co/datasets/theshoaibme/panorgan-processed-pulmonology.glucose_processedglucose values databases
structure is as follows
user id/ timestamp/ glucosevalue
formating of ids is used even for single user datasets since app will have issues if provided with only with one batch (for the moment)
user livia- diabetic type one
livia-large---> contains data from oct 2019 to sept 2024
livia-mini---> is a subset of livia large to be used in testing
user anton- non diabetic
anton---> contains data for 10 days
LLMLingua2-processedfacts-grounding-processed
Dataset Summary
The dataset contains prompts, context documents, and target answers that challenge models to stay grounded in provided context rather than hallucinating.Processing steps added extra features like:
prompt – consolidated instruction + user request + context
has_url_in_context – boolean flag for URLs in context
len_system, len_user, len_context – token/word length statistics
row_id – unique identifier for tracking
Dataset Structure
Splits:
train – 688… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/facts-grounding-processed.INVOICE_ANNOTATION_V1Axcer-processedProcessed_TTS_Multilingual_Data
Processed TTS Multilingual Data
Validated and quality-checked multilingual speech datasets for TTS training, covering 12+ Indian languages.
Datasets Included
Subset
Samples
Hours
Description
indic_voices_r
239,684
548.8h
Indic Voices_R — IVR recordings
rasa
201,509
361.2h
RASA — read speech (wiki, conv, book, news)
indictts_iitm
155,236
253.6h
Indic TTS (IIT Madras) — studio TTS recordings at 48kHz
Total
596,429
1,163.6h
Languages… See the full description on the dataset page: https://huggingface.co/datasets/PalakEngineerMaster/Processed_TTS_Multilingual_Data.Fish_processingprocessed_tempprocessed-hotel-reviewsGo-Emotions-Processedipa-pharma-compactor-platform
IPA Pharmaceutical Roller Compactor Platform: Scale-Up & Performance (Synthetic)
Version: 1.0
Publisher: Innovative Process Applications (IPA)
License: Creative Commons Attribution 4.0 International (CC BY 4.0)
⚠️ This dataset is 100% synthetic and intended for educational use only.
Generated from IPA's published CL-series specifications and standard
compaction physics — not real production or clinical batch data.
What's in this dataset
3,000 simulated… See the full description on the dataset page: https://huggingface.co/datasets/Innovative-Process-Applications/ipa-pharma-compactor-platform.Paul_RNA_Sequence_Processed_Datasetroller-compaction-ribbon-density
Roller Compaction: Ribbon Density vs. Process Parameters (Synthetic)
Version: 1.0
Publisher: Innovative Process Applications (IPA)
License: Creative Commons Attribution 4.0 International (CC BY 4.0)
Contact: Crestwood, IL, USA
⚠️ This dataset is 100% synthetic and intended for educational use only.
It was generated from a published physical model (Johanson rolling theory + Heckel densification) — not measured on any real equipment, customer, or production batch. Do not use it for… See the full description on the dataset page: https://huggingface.co/datasets/Innovative-Process-Applications/roller-compaction-ribbon-density.mimic_hosp_processed_100_resampled_v2red_team_agent_analysis_rl_csvs_diverse1_processed
red_team_agent_analysis_rl_csvs_diverse1_processed
This dataset was automatically uploaded from the red-team-agent repository.
Dataset Information
Original file: diverse1_processed.csv
Source path: /home/ubuntu/red-team-agent/red_team_agent/analysis_rl/csvs/diverse1_processed.csv
Validation: Valid CSV with 1 rows, 12 columns (0.0MB) - Loaded with strategy 1
Usage
import pandas as pd
from datasets import load_dataset
# Load using datasets library
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/aq1048576/red_team_agent_analysis_rl_csvs_diverse1_processed.Selective-Context-processedred_team_agent_analysis_rl_csvs_diverse1_processed_streaming
red_team_agent_analysis_rl_csvs_diverse1_processed_streaming
This dataset was automatically uploaded from the red-team-agent repository.
Dataset Information
Original file: diverse1_processed_streaming.csv
Source path: /home/ubuntu/red-team-agent/red_team_agent/analysis_rl/csvs/diverse1_processed_streaming.csv
Validation: Valid CSV with 703 rows, 12 columns (2.5MB) - Loaded with strategy 1
Usage
import pandas as pd
from datasets import load_dataset
# Load… See the full description on the dataset page: https://huggingface.co/datasets/aq1048576/red_team_agent_analysis_rl_csvs_diverse1_processed_streaming.EmoPillars-Contextless-Processed
