Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ada-datadruids /booksummaries_cleanedtext10K<n<100K0 likes1.5k downloads2y agoHugging Face02Finnish-NLP /mc4_fi_cleaned Dataset Card for mC4 Finnish Cleaned Dataset Summary mC4 Finnish cleaned is cleaned version of the original mC4 Finnish split. Supported Tasks and Leaderboards mC4 Finnish is mainly intended to pretrain Finnish language models and word representations. Languages Finnish Dataset Structure Data Instances [Needs More Information] Data Fields The data have several fields: url: url of the source as a string text: text… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/mc4_fi_cleaned.texttext-generation10M<n<100M4 likes280 downloads4y agoHugging Face03Sufiyan83 /Low-Carbon-London-Smart-Meter-Cleaned-FeatureReadytext100M<n<1B9 likes274 downloads11mo agoHugging Face04haipradana /indonesian-twitter-hate-speech-cleaned Dataset Card for indonesian-twitter-hate-speech-cleaned Dataset Summary Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories. The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/haipradana/indonesian-twitter-hate-speech-cleaned.texttext-classification10K<n<100K0 likes143 downloads1y agoHugging Face05arian81 /TEDS-Full-Cleanedtabular1M<n<10M0 likes129 downloads3y agoHugging Face06bogdanminko /wildguardmix-cleaned Dataset This dataset is cleaned from missing values original wildguardmix dataset: https://huggingface.co/datasets/allenai/wildguardmix text10K<n<100K0 likes122 downloads2y agoHugging Face07hugginglearners /reddit-depression-cleaned Dataset Card for Depression: Reddit Dataset (Cleaned) Dataset Summary The raw data is collected through web scrapping Subreddits and is cleaned using multiple NLP techniques. The data is only in English language. It mainly targets mental health classification. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/reddit-depression-cleaned.text1K<n<10K2 likes121 downloads4y agoHugging Face08Aipresso /10k_rows_cleaned_prompts 10K Rows Cleaned Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use You must provide attribution when using this data in publications, research, or commercial products. Dataset Overview A chunked collection of 2.7 million cleaned English prompts, organized into 200 files of 10,000 rows each for easy processing and distributed training of language models. 📊 Dataset Statistics Metric… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/10k_rows_cleaned_prompts.texttext-generation1M<n<10M0 likes121 downloads1y agoHugging Face09kellyhongg /cleaned-longmemeval-s Cleaned Version of LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory This dataset is a cleaned version of LongMemEval by Wu et al. (2024). This dataset was used in Context Rot. Modifications Removed ambiguous question-answers Fixed focused (orcale) version to ensure question can be fully answered by input Citation If you use this dataset, please cite the original authors: @article{wu2024longmemeval, title={LongMemEval:… See the full description on the dataset page: https://huggingface.co/datasets/kellyhongg/cleaned-longmemeval-s.tabularn<1K2 likes99 downloads1y agoHugging Face10Siddish /change-my-view-subreddit-cleaned Opinionated LLM texttext-generation1K<n<10K1 likes96 downloads3y agoHugging Face11blo05 /cleaned_wiki_enCleaned wikipedia dataset text1M<n<10M4 likes86 downloads5y agoHugging Face12ikekobby /40-percent-cleaned-preprocessed-fake-real-newsKaggle based dataset for text classification task. The data has been cleaned and processed for preparation into any model for classification based tasks. This is just 40% of the entire dataset. text10K<n<100K1 likes71 downloads5y agoHugging Face13ramachandrajoshi /english-kannada-cleaned English–Kannada Cleaned A cleaned parallel corpus of English–Kannada sentence pairs suitable for training and evaluating machine translation models. Languages: English -> Kannada License: Apache License 2.0 Dataset statistics Train: 8,00,000 sentence pairs Validation: 1,000 sentence pairs Test: 1,000 sentence pairs Total: 5,02,000 sentence pairs These counts exclude per-file CSV headers. Source and provenance The dataset is provided as UTF-8 CSV files with… See the full description on the dataset page: https://huggingface.co/datasets/ramachandrajoshi/english-kannada-cleaned.text100K<n<1M1 likes65 downloads6mo agoHugging Face14Iris314 /recipe-cleaned Recipe Cleaned Dataset Dataset Summary This dataset is a structured and cleaned collection of recipe data derived from the Food.com Recipes and Interactions dataset. It is designed for ingredient-based personalization, machine learning training, and interactive recommendation systems. The dataset integrates a hierarchical ingredient taxonomy, standardized nutrition information, and categorical metadata (e.g., diet tags, cuisine attributes, region) to support downstream… See the full description on the dataset page: https://huggingface.co/datasets/Iris314/recipe-cleaned.tabular100K<n<1M0 likes62 downloads1y agoHugging Face15Telugu-LLM-Labs /telugu_alpaca_yahma_cleaned_filtered_romanizedtext10K<n<100K19 likes59 downloads3y agoHugging Face16blo05 /cleaned_wiki_en_0-20text1M<n<10M2 likes50 downloads5y agoHugging Face17dharmam-stjude /CT-RATE-Dataset-cleanedtabular10K<n<100K0 likes47 downloads11mo agoHugging Face18egdrga /indonesian-twitter-hate-speech-cleaned Dataset Card for indonesian-twitter-hate-speech-cleaned Dataset Summary Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories. The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable… See the full description on the dataset page: https://huggingface.co/datasets/egdrga/indonesian-twitter-hate-speech-cleaned.texttext-classification10K<n<100K0 likes46 downloads23d agoHugging Face19starsofchance /MSR_data_cleaned MSR Data Cleaned - C/C++ Code Vulnerability Dataset 📌 Dataset Description A curated collection of C/C++ code vulnerabilities paired with: CVE details (scores, classifications, exploit status) Code changes (commit messages, added/deleted lines) File-level and function-level diffs 🔍 Sample Data Structure from original file +---------------+-----------------+----------------------+---------------------------+ | CVE ID | Attack Origin | Publish Date… See the full description on the dataset page: https://huggingface.co/datasets/starsofchance/MSR_data_cleaned.tabular100K<n<1M1 likes44 downloads1y agoHugging Face20Ryanna567 /cardiovascular-cleaned-datasettabularn<1K0 likes40 downloads1d agoHugging Face21iamshnoo /alpaca-cleaned-albaniantext10K<n<100K2 likes31 downloads3y agoHugging Face22Davidpm02 /spanish_housing_cleanedtabular10K<n<100K0 likes31 downloads19d agoHugging Face23REILX /cleaned-lmsys-arena-human-preference-55k original dataset https://huggingface.co/datasets/lmsys/lmsys-arena-human-preference-55k Use the following code to process the original data to obtain the cleaned data. import csv import random input_file = R'C:\Users\Downloads\train.csv' output_file = 'cleaned-lmsys-arena-human-preference-55k.csv' def clean_text(text): if text.startswith('["') and text.endswith('"]'): return text[2:-2] return text with open(input_file, mode='r', encoding='utf-8') as… See the full description on the dataset page: https://huggingface.co/datasets/REILX/cleaned-lmsys-arena-human-preference-55k.text10K<n<100K0 likes29 downloads2y agoHugging Face24SanaullahTareen07 /grab-safe-driver-telematics-cleaned-datasettabular1M<n<10M0 likes27 downloads2mo agoHugging Face25blo05 /cleaned_wiki_en_40-60text1M<n<10M1 likes26 downloads5y agoHugging Face26NLPC-UOM /nllb-top25k-enta-cleaned Licensing Information The dataset is released under the terms of ODC-BY. By using this, you are also bound to the respective Terms of Use and License of the original source. Citation Information @inproceedings{ranathunga-etal-2024-quality, title = "Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel Corpora", author = "Ranathunga, Surangika and De Silva, Nisansa and Menan, Velayuthan and Fernando, Aloka and… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/nllb-top25k-enta-cleaned.texttranslation10K<n<100K0 likes26 downloads2y agoHugging Face27arjunsunar /imdb-cleanedtabular10K<n<100K0 likes26 downloads2mo agoHugging Face28GaussianWorld /crowdsourced_3dgs_cleanedgatedtextothern<1K0 likes24 downloads1y agoHugging Face29MaxPrestige /CA_Weather_Fire_Dataset_Cleaned📦 Dataset Card: CA_Weather_Fire_Dataset_Cleaned Dataset Summary This dataset contains cleaned and preprocessed weather and fire incident data for California (1984–2025). The original dataset, California Weather and Fire Prediction Dataset (1984–2025) with Engineered Features, includes features such as temperature, humidity, wind speed, fire occurrence, and seasonal indicators. From the Original Dataset, I changed the data types to floats, rearranged the columns, removed… See the full description on the dataset page: https://huggingface.co/datasets/MaxPrestige/CA_Weather_Fire_Dataset_Cleaned.tabularreinforcement-learning10K<n<100K2 likes24 downloads1y agoHugging Face30Aipresso /cleaned-english-prompts Cleaned English Prompts Dataset Dataset Description A cleaned dataset containing English prompts and their corresponding responses. This dataset is designed for training conversational AI models and language models. Dataset Summary Columns: Questions and Response Language: English Size: 1,000-10,000 examples Format: CSV Cleaning: Data has been processed and cleaned for training Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/cleaned-english-prompts.textquestion-answering1M<n<10M0 likes24 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.