Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kammartina /MA_Query_Expansion_MLT26tabular10K<n<100K0 likes2.5k downloads16d agoHugging Face02CShorten /ML-ArXiv-PapersThis dataset contains the subset of ArXiv papers with the "cs.LG" tag to indicate the paper is about Machine Learning. The core dataset is filtered from the full ArXiv dataset hosted on Kaggle: https://www.kaggle.com/datasets/Cornell-University/arxiv. The original dataset contains roughly 2 million papers. This dataset contains roughly 100,000 papers following the category filtering. The dataset is maintained by with requests to the ArXiv API. The current iteration of the dataset only contains… See the full description on the dataset page: https://huggingface.co/datasets/CShorten/ML-ArXiv-Papers.tabular100K<n<1M73 likes2.4k downloads4y agoHugging Face03gle3D /ML-Proto-Dataset3dn<1K1 likes807 downloads11mo agoHugging Face04daekeun-ml /naver-news-summarization-ko Naver-News-KO: A Korean News Summarization Dataset A Korean news summarization dataset of 27,400 (document, summary) pairs, crawled from Naver News over a ten-day window in July 2022. It was originally built for a Korean NLP hands-on lab and has been publicly hosted on the Hugging Face Hub since January 2023. A technical report documenting the collection protocol, corpus statistics, contamination analysis, and reproducible baselines is available on arXiv: arXiv:2607.20442.… See the full description on the dataset page: https://huggingface.co/datasets/daekeun-ml/naver-news-summarization-ko.textsummarization10K<n<100K68 likes571 downloads3mo agoHugging Face05GD-ML /CCN Towards Full Candidate Interaction: A Comprehensive Comparison Network for Better Route Recommendation This is the dataset for our paper. The following table contains the feature dimensions and key features of our dataset. Feature Type Interpretation Shape Some Key Features Route Features Used to describe each route, including static features, dynamic features, and trajectory statistical features N * 62 The estimated time of arrival for the routeThe total distance length… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/CCN.texttabular-classification100K<n<1M37 likes436 downloads7mo agoHugging Face06GD-ML /Taming-Hallucinations Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation CVPR 2026 Findings Project Page | Paper | Code Dataset Summary This repository hosts DualityVidQA, the large-scale paired video–QA dataset introduced in Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation. Taming Hallucinations introduces DualityForge, a controllable diffusion-based framework that turns real videos into… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/Taming-Hallucinations.textvideo-text-to-text1K<n<10K0 likes407 downloads4mo agoHugging Face07Kokoslocke /NACA_4_Digit_for_ML NACA 4-Digit Airfoil CFD Dataset Point-cloud CFD solutions for NACA 4-digit airfoils, generated with OpenFOAM v13 (k-ω SST). Intended for training surrogate models that predict steady-state flow fields from airfoil geometry and flow conditions. Dataset Summary ~850 converged in-distribution cases across 50 distinct NACA 4-digit profiles AoA range: −5° to +5° Reynolds number range: 100,000 – 500,000 129 out-of-distribution (OOD) probe cases at high Re (1–2 × 10⁶)… See the full description on the dataset page: https://huggingface.co/datasets/Kokoslocke/NACA_4_Digit_for_ML.tabularothern<1K0 likes360 downloads3mo agoHugging Face08GD-ML /TransitLM TransitLM: Dataset Release & Evaluation Protocol Dataset Description TransitLM is a dataset for public transit route planning in Chinese urban environments, designed to support training and evaluation of language models that generate structured transit routes from origin-destination information. The full dataset covers four cities: Beijing, Shanghai, Shenzhen, and Chengdu, and includes coordinates, station sequences, transfer structure, line information, and route… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/TransitLM.tabulartext-generation100K<n<1M82 likes318 downloads4mo agoHugging Face09ML-Owl /faang-engineered-time-series-features-2013-2025 FAANG Stocks Historical Raw and Engineered Time-Series Dataset (2013-2025) Since this is a comprehensive ReadMe file with multiple sections and crosslinks to other documents and images, I wanted to start by providing a ToC with hyperlinks to simplify navigation for the readers. (special thanks to @csavur for this very helpful suggestion!) DOCUMENT NAVIGATION GUIDE (ToC) 1 - Summary2 - Usage & Reproducability3 - Practical Uses of this Dataset 3.1 - A real-world ML… See the full description on the dataset page: https://huggingface.co/datasets/ML-Owl/faang-engineered-time-series-features-2013-2025.imagetabular-classification10K<n<100K2 likes315 downloads7mo agoHugging Face10nedjmaou /MLMA_hate_speech Disclaimer This is a hate speech dataset (in Arabic, French, and English). Offensive content that does not reflect the opinions of the authors. Dataset of our EMNLP 2019 Paper (Multilingual and Multi-Aspect Hate Speech Analysis) For more details about our dataset, please check our paper: @inproceedings{ousidhoum-etal-multilingual-hate-speech-2019, title = "Multilingual and Multi-Aspect Hate Speech Analysis", author = "Ousidhoum, Nedjma… See the full description on the dataset page: https://huggingface.co/datasets/nedjmaou/MLMA_hate_speech.text10K<n<100K5 likes285 downloads2y agoHugging Face11ARTeLab /mlsum-it Dataset Card for mlsum-it Dataset Summary The MLSum-it dataset is the translated version (Helsinki-NLP/opus-mt-es-it) of the spanish portion of MLSum, containing news articles taken from BBC/mundo. More informations on the official dataset page HuggingFace page. There are two features: source: Input news article. target: Summary of the article. Supported Tasks and Leaderboards abstractive-summarization, summarization Languages The text in… See the full description on the dataset page: https://huggingface.co/datasets/ARTeLab/mlsum-it.textsummarization10K<n<100K2 likes278 downloads4y agoHugging Face12GD-ML /SCASRec SCASRec: A Self-Correcting and Auto-Stopping Model for Generative Route List Recommendation This is the dataset for our paper. The following table contains the feature dimensions and key features of our dataset. Feature Type Interpretation Shape Some Key Features Route Features Used to describe each route, including static features, dynamic features, and trajectory statistical features N * 62 The estimated time of arrival for the routeThe total distance length of the… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/SCASRec.texttabular-classification100K<n<1M30 likes273 downloads8mo agoHugging Face13GD-ML /MobilityBench Note: This work is currently under review. The full dataset will be released progressively. MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios Paper | GitHub MobilityBench is a scalable benchmark for evaluating route-planning agents in real-world mobility scenarios. It is built from large-scale, anonymized mobility queries from Amap, organized with a comprehensive task taxonomy, and provides structured ground truth (required tool calls… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/MobilityBench.tabularquestion-answering10K<n<100K18 likes220 downloads7mo agoHugging Face14GD-ML /GenMRP GenMRP: A Generative Multi-Route Planning Framework for Efficient and Personalized Real-Time Industrial Navigation This is the dataset for our paper. The following table contains the feature dimensions and key features of our dataset. Feature Type Interpretation Shape Some Key Features Link Features Includes the road segment attributes K * 2 * N Link lengthLink Lane width Frequency Features Logs the user's travel history within the past three months K * 2 * 10 * 7 Delta… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/GenMRP.texttabular-classification100K<n<1M43 likes219 downloads7mo agoHugging Face15bep40 /ml-intern-kpis bep40/ml-intern-kpis Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub. Try ML Intern: https://smolagents-ml-intern.hf.space Source code: https://github.com/huggingface/ml-intern Usage from datasets import load_dataset dataset = load_dataset("bep40/ml-intern-kpis") tabularn<1K0 likes197 downloads9h agoHugging Face16scikit-fingerprints /ASAP_OpenADMET_MLM ASAP-OpenADMET MLM ASAP_OpenADMET_MLM dataset from the ASAP Discovery-OpenADMET Antiviral Drug Discovery Challenge [1] [2] [3]. It is intended to be used through scikit-fingerprints library. The task is to predict MLM (mouse liver microsomal intrinsic clearance in uL/min/mg) of molecules. Characteristic Description Tasks 1 Task type regression Total samples 425 Recommended splittime Recommended metric MAE References [1] ASAP Discovery "ASAP… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ASAP_OpenADMET_MLM.texttabular-regressionn<1K0 likes167 downloads6mo agoHugging Face17ai-eldorado /ML-KEM-SideChannel-Traces ML-KEM Side Channel Traces Dataset Description This dataset contains power traces captured from the Post-Quantum Cryptography (PQC) ML-KEM implementation of the PQM4[1] library (commit: a24bb4b), running on an STM32 Nucleo-L4R5ZI development board equipped with an ARM Cortex-M4 processor. The traces were collected using a Rohde & Schwarz RTC1002 100 MHz digital oscilloscope. The purpose of this dataset is to evaluate the ML-KEM implementation for side-channel… See the full description on the dataset page: https://huggingface.co/datasets/ai-eldorado/ML-KEM-SideChannel-Traces.textother100K<n<1M1 likes162 downloads1mo agoHugging Face18gbyuvd /coconut-chembl34-selfies-mlm Dataset Card for COCONUT+ChemBL34 SELFIES for MLM training (unmasked) This dataset is a collection of molecular structures represented as SELFIES (Self-Referencing Embedded Strings), created by combining and processing data from COCONUTDB and ChemBL34. It contains 2,700,462 unique molecules across 13 chunks. The dataset is specifically designed for pre-training language models on molecular representations using the Masked Language Model (MLM) approach. It consists of a single column… See the full description on the dataset page: https://huggingface.co/datasets/gbyuvd/coconut-chembl34-selfies-mlm.text100K<n<1M0 likes155 downloads1y agoHugging Face19MLLab-TS /energy_datasettabular1K<n<10K0 likes154 downloads8mo agoHugging Face20michaelmallari /mlb-statcast-batterstabular1K<n<10K0 likes151 downloads3y agoHugging Face21wwbrannon /ml-interview-examples-movielens-1mtabular1M<n<10M0 likes138 downloads10mo agoHugging Face22scikit-fingerprints /ExpansionRx_OpenADMET_MLM_CLint ExpansionRx-OpenADMET MLM CLint MLM CLint dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through scikit-fingerprints library. The task is to predict MLM CLint of molecules. Characteristic Description Tasks 1 Task type regression Total samples 5692 Recommended split time Recommended metric MAE References [1] OpenADMET team "Announcement 1: ExpansionRx-OpenADMET Blind Challenge"… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_MLM_CLint.texttabular-regression1K<n<10K0 likes135 downloads6mo agoHugging Face23MLBtrio /genz-slang-dataset Dataset Details This dataset contains a rich collection of popular slang terms and acronyms used primarily by Generation Z. It includes detailed descriptions of each term, its context of use, and practical examples that demonstrate how the slang is used in real-life conversations. The dataset is designed to capture the unique and evolving language patterns of GenZ, reflecting their communication style in digital spaces such as social media, text messaging, and online forums. Each… See the full description on the dataset page: https://huggingface.co/datasets/MLBtrio/genz-slang-dataset.texttext-generation1K<n<10K52 likes134 downloads2y agoHugging Face24blanchon /parler-tts_mls_eng_10k_snac_token_old Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.tabularautomatic-speech-recognition100K<n<1M1 likes119 downloads2y agoHugging Face25ShieldX /Rash-Driving-Detection-on-Bikes-for-ML Dataset Card for Rash Driving Detection on Bikes Using Mobile and Sensor Data This dataset is designed to aid the detection of rash driving behavior on bikes using data collected from mobile and sensor-based systems. It includes sensor readings such as accelerometer values, orientation (azimuth, pitch, roll), and speed, with labels indicating whether the riding behavior is classified as rash or not. Dataset Details Dataset Description This dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/ShieldX/Rash-Driving-Detection-on-Bikes-for-ML.tabular10K<n<100K0 likes116 downloads10mo agoHugging Face26dagmawi-ml /cwe-workshop-datasetimage1M<n<10M0 likes114 downloads4mo agoHugging Face27robro612 /lotte_pooled_dev_search_mlateon lotte_pooled_dev_search_mlateon Multi-vector (late-interaction) embeddings of LoTTE pooled/dev/search (lotte/pooled/dev/search), encoded with lightonai/mLateOn at revision edd378f99593c0ac8a15518b97ad89786b02685e. Source data: ir_datasets lotte/pooled/dev/search (ir_datasets 0.6.3), which downloads lotte.tar.gz (md5 3b2e88b1d66933627462950b4c3f5d0f). The ColBERTv2 authors also publish LoTTE on the Hub as colbertv2/lotte, whose card gives this dataset's license; the data here was… See the full description on the dataset page: https://huggingface.co/datasets/robro612/lotte_pooled_dev_search_mlateon.tabular1K<n<10K0 likes112 downloads10d agoHugging Face28robro612 /trec-covid_mlateon trec-covid_mlateon Multi-vector (late-interaction) embeddings of BEIR trec-covid (beir/trec-covid), encoded with lightonai/mLateOn at revision edd378f99593c0ac8a15518b97ad89786b02685e. Source data: ir_datasets beir/trec-covid (ir_datasets 0.6.3), which downloads trec-covid.zip (md5 ce62140cb23feb9becf6270d0d1fe6d1). BEIR also publishes this corpus on the Hub as BeIR/trec-covid, whose card gives this dataset's license; the data here was loaded through ir_datasets, not from that… See the full description on the dataset page: https://huggingface.co/datasets/robro612/trec-covid_mlateon.tabular10K<n<100K0 likes110 downloads11d agoHugging Face29mlexplorer008 /malayalam_news_classificationtext1K<n<10K0 likes101 downloads2y agoHugging Face30takarahomes /mlit-japan-real-estate-price-2005-2024q3tabular1M<n<10M4 likes100 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.