datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
booksummaries_cleanedmc4_fi_cleaned
Dataset Card for mC4 Finnish Cleaned
Dataset Summary
mC4 Finnish cleaned is cleaned version of the original mC4 Finnish split.
Supported Tasks and Leaderboards
mC4 Finnish is mainly intended to pretrain Finnish language models and word representations.
Languages
Finnish
Dataset Structure
Data Instances
[Needs More Information]
Data Fields
The data have several fields:
url: url of the source as a string
text: text… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/mc4_fi_cleaned.Low-Carbon-London-Smart-Meter-Cleaned-FeatureReadyindonesian-twitter-hate-speech-cleaned
Dataset Card for indonesian-twitter-hate-speech-cleaned
Dataset Summary
Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories.
The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/haipradana/indonesian-twitter-hate-speech-cleaned.TEDS-Full-Cleanedwildguardmix-cleaned
Dataset
This dataset is cleaned from missing values original wildguardmix dataset: https://huggingface.co/datasets/allenai/wildguardmix
reddit-depression-cleaned
Dataset Card for Depression: Reddit Dataset (Cleaned)
Dataset Summary
The raw data is collected through web scrapping Subreddits and is cleaned using multiple NLP techniques. The data is only in English language. It mainly targets mental health classification.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/reddit-depression-cleaned.10k_rows_cleaned_prompts
10K Rows Cleaned Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
You must provide attribution when using this data in publications, research, or commercial products.
Dataset Overview
A chunked collection of 2.7 million cleaned English prompts, organized into 200 files of 10,000 rows each for easy processing and distributed training of language models.
📊 Dataset Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/10k_rows_cleaned_prompts.cleaned-longmemeval-s
Cleaned Version of LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
This dataset is a cleaned version of LongMemEval by Wu et al. (2024).
This dataset was used in Context Rot.
Modifications
Removed ambiguous question-answers
Fixed focused (orcale) version to ensure question can be fully answered by input
Citation
If you use this dataset, please cite the original authors:
@article{wu2024longmemeval,
title={LongMemEval:… See the full description on the dataset page: https://huggingface.co/datasets/kellyhongg/cleaned-longmemeval-s.change-my-view-subreddit-cleaned
Opinionated LLM
cleaned_wiki_enCleaned wikipedia dataset
40-percent-cleaned-preprocessed-fake-real-newsKaggle based dataset for text classification task. The data has been cleaned and processed for preparation into any model for classification based tasks. This is just 40% of the entire dataset.
english-kannada-cleaned
English–Kannada Cleaned
A cleaned parallel corpus of English–Kannada sentence pairs suitable for training and evaluating machine translation models.
Languages: English -> Kannada
License: Apache License 2.0
Dataset statistics
Train: 8,00,000 sentence pairs
Validation: 1,000 sentence pairs
Test: 1,000 sentence pairs
Total: 5,02,000 sentence pairs
These counts exclude per-file CSV headers.
Source and provenance
The dataset is provided as UTF-8 CSV files with… See the full description on the dataset page: https://huggingface.co/datasets/ramachandrajoshi/english-kannada-cleaned.recipe-cleaned
Recipe Cleaned Dataset
Dataset Summary
This dataset is a structured and cleaned collection of recipe data derived from the Food.com Recipes and Interactions dataset. It is designed for ingredient-based personalization, machine learning training, and interactive recommendation systems. The dataset integrates a hierarchical ingredient taxonomy, standardized nutrition information, and categorical metadata (e.g., diet tags, cuisine attributes, region) to support downstream… See the full description on the dataset page: https://huggingface.co/datasets/Iris314/recipe-cleaned.telugu_alpaca_yahma_cleaned_filtered_romanizedcleaned_wiki_en_0-20CT-RATE-Dataset-cleanedindonesian-twitter-hate-speech-cleaned
Dataset Card for indonesian-twitter-hate-speech-cleaned
Dataset Summary
Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories.
The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable… See the full description on the dataset page: https://huggingface.co/datasets/egdrga/indonesian-twitter-hate-speech-cleaned.MSR_data_cleaned
MSR Data Cleaned - C/C++ Code Vulnerability Dataset
📌 Dataset Description
A curated collection of C/C++ code vulnerabilities paired with:
CVE details (scores, classifications, exploit status)
Code changes (commit messages, added/deleted lines)
File-level and function-level diffs
🔍 Sample Data Structure from original file
+---------------+-----------------+----------------------+---------------------------+
| CVE ID | Attack Origin | Publish Date… See the full description on the dataset page: https://huggingface.co/datasets/starsofchance/MSR_data_cleaned.cardiovascular-cleaned-datasetalpaca-cleaned-albanianspanish_housing_cleanedcleaned-lmsys-arena-human-preference-55k
original dataset
https://huggingface.co/datasets/lmsys/lmsys-arena-human-preference-55k
Use the following code to process the original data to obtain the cleaned data.
import csv
import random
input_file = R'C:\Users\Downloads\train.csv'
output_file = 'cleaned-lmsys-arena-human-preference-55k.csv'
def clean_text(text):
if text.startswith('["') and text.endswith('"]'):
return text[2:-2]
return text
with open(input_file, mode='r', encoding='utf-8') as… See the full description on the dataset page: https://huggingface.co/datasets/REILX/cleaned-lmsys-arena-human-preference-55k.grab-safe-driver-telematics-cleaned-datasetcleaned_wiki_en_40-60nllb-top25k-enta-cleaned
Licensing Information
The dataset is released under the terms of ODC-BY. By using this, you are also bound to the respective Terms of Use and License of the original source.
Citation Information
@inproceedings{ranathunga-etal-2024-quality,
title = "Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel Corpora",
author = "Ranathunga, Surangika and
De Silva, Nisansa and
Menan, Velayuthan and
Fernando, Aloka and… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/nllb-top25k-enta-cleaned.imdb-cleanedcrowdsourced_3dgs_cleanedCA_Weather_Fire_Dataset_Cleaned📦 Dataset Card: CA_Weather_Fire_Dataset_Cleaned
Dataset Summary
This dataset contains cleaned and preprocessed weather and fire incident data for California (1984–2025).
The original dataset, California Weather and Fire Prediction Dataset (1984–2025) with Engineered Features, includes features such as temperature, humidity, wind speed, fire occurrence, and seasonal indicators.
From the Original Dataset, I changed the data types to floats, rearranged the columns, removed… See the full description on the dataset page: https://huggingface.co/datasets/MaxPrestige/CA_Weather_Fire_Dataset_Cleaned.cleaned-english-prompts
Cleaned English Prompts Dataset
Dataset Description
A cleaned dataset containing English prompts and their corresponding responses. This dataset is designed for training conversational AI models and language models.
Dataset Summary
Columns: Questions and Response
Language: English
Size: 1,000-10,000 examples
Format: CSV
Cleaning: Data has been processed and cleaned for training
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/cleaned-english-prompts.
