Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zackyabd /ptb-xl-processedtabular10K<n<100K0 likes73k downloads1y agoHugging Face02Tri1 /processed_vnhntabular10K<n<100K0 likes2.4k downloads2mo agoHugging Face03Ken4962 /processed_fake_job_postingstabulartext-classification10K<n<100K0 likes1.8k downloads1y agoHugging Face04BeileiCui /StereoMIS_processedimagen<1K0 likes1k downloads1mo agoHugging Face05Process-Venue /IntentClassification_Dataset_for_AI_Assistant_Prompt_Routing_Hinditexttext-classification1K<n<10K0 likes183 downloads11mo agoHugging Face06Reduanul1997 /Repository_processed_dataset_OCTID OCTID Processed Dataset Processed version of the OCTID retinal OCT dataset prepared for research experiments comparing Swin Transformer architectures. Dataset Summary Dataset: OCTID Total images: 572 Number of classes: 5 Image size: 224 x 224 Image mode: RGB Image format: JPEG Classes Class Images Normal 206 AMD 55 CSR 102 DR 107 MH 102 Total 572 Preprocessing The preprocessing pipeline consists of:… See the full description on the dataset page: https://huggingface.co/datasets/Reduanul1997/Repository_processed_dataset_OCTID.image1K<n<10K0 likes131 downloads1mo agoHugging Face07mythezone /LOBench-A-share-processedtabular1M<n<10M0 likes128 downloads2mo agoHugging Face08cuonguyenphu /Natural-Language-Processing-with-Disaster-Tweets-0.84033tabular1K<n<10K0 likes71 downloads21d agoHugging Face09natural-lang-processing /sexismreddittext10K<n<100K5 likes68 downloads3y agoHugging Face10processvenue /INVOICE_ANNOTATION_V2tabularimage-classification1K<n<10K0 likes67 downloads10mo agoHugging Face11dishamodi /Keystroke_Processed Processed 136M Keystroke Dataset This dataset is derived from the original 136M keystroke dataset, which contained raw data collected from a variety of typists. The processed version includes additional metrics that distinguish human and bot behavior, making it useful for research in keystroke dynamics and behavioral analysis. Metrics specific to bot behavior were generated using Selenium scripts, providing a comprehensive comparison between human and bot typing patterns.… See the full description on the dataset page: https://huggingface.co/datasets/dishamodi/Keystroke_Processed.tabular100K<n<1M2 likes59 downloads2y agoHugging Face12electricsheepafrica /african-agro-processing-value-add African Agro-Processing Value Addition Dataset | Africa (Electric Sheep Africa metadata inventory) Size category: 100K<n<1M - Formats: csv - Sector: agriculture_food - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/african-agro-processing-value-add.tabulartabular-classification100K<n<1M0 likes48 downloads2mo agoHugging Face13abhinavsarkar /delhi_air_quality_feature_store_processed.csvDataset Fields: location_id: Integer identifier for each location. city: The name of the city or specific location in Delhi. event_timestamp: The timestamp when the data was recorded, in ISO 8601 format. temperature: Ambient temperature in Celsius. humidity: Relative humidity as a percentage. pressure: Atmospheric pressure in hPa. wind_speed: Wind speed in m/s. wind_direction: Wind direction in degrees. pm25: Concentration of particulate matter with a diameter of 2.5 micrometers (µg/m³).… See the full description on the dataset page: https://huggingface.co/datasets/abhinavsarkar/delhi_air_quality_feature_store_processed.csv.tabular1M<n<10M0 likes46 downloads2y agoHugging Face14SmartQHSE /named-process-safety-incidents-extended-2026 Canonical landing page: https://www.smartqhse.com/datasets/named-process-safety-incidents-extended-2026 Named Process Safety and Industrial Disasters — Extended Reference 2026 Curated reference of 40 named historical process-safety, industrial, and major-fire disasters with dates, fatalities, casual factors, and regulatory consequences. Spans 1917–2024. Covers Bhopal, Piper Alpha, Texas City, Deepwater Horizon, Buncefield, Flixborough, Seveso, Phillips 66 Pasadena, Longford… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/named-process-safety-incidents-extended-2026.textn<1K0 likes43 downloads4mo agoHugging Face15uno23 /ansi-b73-1-process-pump-item-numbers ANSI/ASME B73.1 process pump part item numbers, frame and group families, and size coverage A reference dataset published by Jinan Yingsiman Machinery Co., Ltd. (YSM Pumps), Jinan, Shandong, China. Version 1.0.0 · published 2026-09-23 · Licence: CC BY 4.0 · DOI: 10.5281/zenodo.22920668 This is a mirror. The citable record is on Zenodo: https://doi.org/10.5281/zenodo.22920668 (the concept DOI for all versions is https://doi.org/10.5281/zenodo.22920667). The five data files here… See the full description on the dataset page: https://huggingface.co/datasets/uno23/ansi-b73-1-process-pump-item-numbers.documentn<1K0 likes43 downloads16d agoHugging Face16TheRealVigilante /Processed_Top_15k_Anime 📦 Anime Recommender Dataset (Sentence-BERT Ready) This dataset is a cleaned and preprocessed version of the Top 15,000 Ranked Anime Dataset originally published on Kaggle by Quan Than. It is specifically prepared to be used for semantic recommendation systems, including transformer-based models like Sentence-BERT. 📌 Original Dataset Source: Kaggle - Top 15,000 Ranked Anime Dataset (updated to Mar 2025)Author: Quan ThanLicense: Apache 2.0 🔧 Modifications… See the full description on the dataset page: https://huggingface.co/datasets/TheRealVigilante/Processed_Top_15k_Anime.tabular10K<n<100K1 likes41 downloads1y agoHugging Face17ClarusC64 /clinical-parallel-process-awareness-v0.1Clinical Parallel Process Awareness v0.1 Goal Test if a model can hold separate reasoning streams at once Detect constraint dismissal Detect bleed-over where one stream turns into claims in the other What it measures streams_heldResponse acknowledges and maintains both streams bleed_overConstraint stream improperly becomes a medical claim, or vice versa premature_synthesisResponse forces a single solution that silences one stream assumption_collapseResponse drops a premise entirely Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-parallel-process-awareness-v0.1.texttext-generationn<1K0 likes41 downloads9mo agoHugging Face18theshoaibme /panorgan-processed-pulmonology Pan-Organ Net: Processed Pulmonology Dataset (with Train / Val / Test Splits) Standardized, preprocessed production dataset for Pan-Organ Net multi-organ foundation model screening. Directory Layout & Schema splits/: train.csv: 70% training cohort (14,815 images) with file paths and numeric class IDs validation.csv: 15% validation cohort (3,175 images) test.csv: 15% independent test benchmark (3,175 images) data/: train_images.tar.gz: Preprocessed training… See the full description on the dataset page: https://huggingface.co/datasets/theshoaibme/panorgan-processed-pulmonology.imageimage-classification10K<n<100K0 likes41 downloads3d agoHugging Face19Livia-Zaharia /glucose_processedglucose values databases structure is as follows user id/ timestamp/ glucosevalue formating of ids is used even for single user datasets since app will have issues if provided with only with one batch (for the moment) user livia- diabetic type one livia-large---> contains data from oct 2019 to sept 2024 livia-mini---> is a subset of livia large to be used in testing user anton- non diabetic anton---> contains data for 10 days tabular100K<n<1M0 likes39 downloads2y agoHugging Face20Itz-Amethyst /LLMLingua2-processedtabular10K<n<100K0 likes36 downloads4mo agoHugging Face21GenAIDevTOProd /facts-grounding-processed Dataset Summary The dataset contains prompts, context documents, and target answers that challenge models to stay grounded in provided context rather than hallucinating.Processing steps added extra features like: prompt – consolidated instruction + user request + context has_url_in_context – boolean flag for URLs in context len_system, len_user, len_context – token/word length statistics row_id – unique identifier for tracking Dataset Structure Splits: train – 688… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/facts-grounding-processed.tabularn<1K1 likes35 downloads1y agoHugging Face22Process-Venue /Movie_Review_Sentiment_Hinditexttext-classification1K<n<10K0 likes32 downloads11mo agoHugging Face23MocktaiLEngineer /qmsum-processedtext1K<n<10K0 likes31 downloads3y agoHugging Face24ClarusC64 /clinical-multidoctor-diagnostic-process-integrity-scoring-v0.1What this dataset tests Whether a model can score the integrity of a multi-doctor diagnostic processusing dialogue structure, hypothesis competition, and objection handling. Required outputs process_integrity_score_0_100 primary_reasoning_strength primary_reasoning_weakness Strength labels evidence_coverage hypothesis_competition objection_closure cross_specialty_synthesis counterfactual_testing bias_resistance uncertainty_tracking Weakness labels premature_closure… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-multidoctor-diagnostic-process-integrity-scoring-v0.1.texttext-classificationn<1K0 likes30 downloads8mo agoHugging Face25processvenue /INVOICE_ANNOTATION_V1tabularimage-classification1K<n<10K0 likes27 downloads10mo agoHugging Face26Itz-Amethyst /Axcer-processedtabular10K<n<100K0 likes26 downloads4mo agoHugging Face27PalakEngineerMaster /Processed_TTS_Multilingual_Data Processed TTS Multilingual Data Validated and quality-checked multilingual speech datasets for TTS training, covering 12+ Indian languages. Datasets Included Subset Samples Hours Description indic_voices_r 239,684 548.8h Indic Voices_R — IVR recordings rasa 201,509 361.2h RASA — read speech (wiki, conv, book, news) indictts_iitm 155,236 253.6h Indic TTS (IIT Madras) — studio TTS recordings at 48kHz Total 596,429 1,163.6h Languages… See the full description on the dataset page: https://huggingface.co/datasets/PalakEngineerMaster/Processed_TTS_Multilingual_Data.tabulartext-to-speech100K<n<1M0 likes25 downloads8mo agoHugging Face28AryanAnuj /processed_dataset_orca-math-word-problems-200kDataset Description: This dataset contains data that has undergone two preprocessing steps: Removal of Instructions with Less Than 100 Tokens in Response: Instructions with less than 100 tokens in the response have been removed from the dataset. This preprocessing step helps to ensure that the dataset contains substantial and informative responses. Data Deduplication by Grouping Using Cosine Similarity (Threshold > 0.95): Data deduplication has been performed by grouping similar instances… See the full description on the dataset page: https://huggingface.co/datasets/AryanAnuj/processed_dataset_orca-math-word-problems-200k.textn<1K0 likes24 downloads2y agoHugging Face29introvoyz041 /Fish_processingimagen<1K0 likes24 downloads2y agoHugging Face30abvgdejzui /processed_temptabular1K<n<10K1 likes23 downloads10d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.