Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HPLT /HPLT2.0_cleanedNB: HPLT2.0 is now superseded by a newer release: HPLT3.0 We recommed switching to v3.0, unless you have a compelling reason to stay on 2.0. This is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project. The source of the data is mostly Internet Archive with some additions from Common Crawl. For a detailed description of the dataset, please refer to our website and our pre-print. The Cleaned variant of HPLT Datasets v2.0 This is… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/HPLT2.0_cleaned.tabularfill-mask1B<n<10B46 likes187k downloads4mo agoHugging Face02Hula0401 /cad-corpus-cleanedtabular1M<n<10M4 likes27k downloads4mo agoHugging Face03argilla /ultrafeedback-binarized-preferences-cleaned UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences, and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback. Read more about Argilla's approach towards UltraFeedback binarization at argilla/ultrafeedback-binarized-preferences/README.md. Differences with argilla/ultrafeedback-binarized-preferences… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned.tabulartext-generation10K<n<100K165 likes25k downloads3y agoHugging Face04xiaowu0162 /longmemeval-cleanedThis dataset replaces the original LongMemEval dataset. The main difference is that this version removes noisy history sessions that interfere with the answer correctness. More detailed session processing information can be found here. 36 likes25k downloads1y agoHugging Face05yahma /alpaca-cleaned Dataset Card for Alpaca-Cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an answer. "instruction":"Summarize… See the full description on the dataset page: https://huggingface.co/datasets/yahma/alpaca-cleaned.texttext-generation10K<n<100K891 likes22k downloads4y agoHugging Face06argilla /ultrafeedback-binarized-preferences-cleaned-kto UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) KTO A KTO signal transformed version of the highly loved UltraFeedback Binarized Preferences Cleaned, the preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences, and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback. Read more about… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned-kto.texttext-generation100K<n<1M10 likes14k downloads3y agoHugging Face07ArtificialAnalysis /Earnings22-Cleaned-AA Earnings22-Cleaned-AA Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.audioautomatic-speech-recognitionn<1K7 likes8.8k downloads8mo agoHugging Face08ClementRomac /cleaned_deduplicated_oscar Dataset Card for "cleaned_deduplicated_oscar" More Information needed text100M<n<1B0 likes7.1k downloads3y agoHugging Face09Avelina /smollm-corpus-cleaned SmolLM-Corpus: Now shuffled and sharded (and Cleaned)! This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming! The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo. Dataset Structure The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.texttext-generation100M<n<1B2 likes6.6k downloads2y agoHugging Face10theelderemo /genius-lyrics-cleaned ◎ Genius Lyrics Dataset Cleaned & Deduplicated 🤗 Hugging Face 🤗 Hugging Face DOI: 10.57967/hf/7978 DOI: 10.57967/hf/7978 revision: 9742989 revision: 9742989 A heavily cleaned, English-only, genre-filtered subset of the Genius Song… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/genius-lyrics-cleaned.texttext-generation1M<n<10M19 likes5.6k downloads7mo agoHugging Face11AiAF /SCPWiki-Cleaned-PDF-Archivesdocumenttext-generationn<1K1 likes5.4k downloads1y agoHugging Face12dacorvo /funes-xiaowu0162-longmemeval-cleaned-s Funes recall store — LongMemEval_s cleaned corpus A funes recall store built by indexing the longmemeval_s_cleaned.json haystack of xiaowu0162/longmemeval-cleaned (LongMemEval, arXiv:2410.10813) — every unique chat session across all 500 questions' haystacks, in one corpus-wide store. What this is This is not a raw trace dataset — it is a pre-built funes index: the source sessions chunked into content blocks and embedded, stored as a Lance table (chunks.lance).… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-xiaowu0162-longmemeval-cleaned-s.tabular100K<n<1M0 likes3.7k downloads9d agoHugging Face13zicsx /mC4-Hindi-Cleaned-3.0 Dataset Card for "mC4-Hindi-Cleaned-3.0" More Information needed text1M<n<10M2 likes3.7k downloads3y agoHugging Face14Crownelius /Creative-Writing-Sonnet4.6-Cleaned Creative-Writing-Sonnet4.6-Cleaned Cleaned creative writing SFT dataset from Sonnet 4.6 (833 samples). Prompts cleaned, thinking traces preserved. Format Each line is a JSON object with: messages: list of message dicts with roles (system, user, assistant) System: writing quality instructions User: cleaned creative writing prompt Assistant: creative writing response (may include <think> traces) Stats Metric Value Total prompt tokens… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Sonnet4.6-Cleaned.texttext-generationn<1K3 likes3.6k downloads3mo agoHugging Face15vikp /starcoder_cleanedThis is starcoderdata, but with leading boilerplate text/license text removed, and with short sequences filtered out. It also removes the extra tags at the beginning of some of the files, like <reponame>. text10M<n<100M4 likes3.5k downloads3y agoHugging Face16devilkingpc /Telegram-Cleaned-DBtext100M<n<1B1 likes3.1k downloads1mo agoHugging Face17bayes-group-diffusion /OAS95-aligned-cleanedtext100M<n<1B3 likes2.9k downloads10mo agoHugging Face18unsloth /alpaca-cleaned Dataset Card for Alpaca-Cleaned Forked from https://huggingface.co/datasets/yahma/alpaca-cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused… See the full description on the dataset page: https://huggingface.co/datasets/unsloth/alpaca-cleaned.texttext-generation10K<n<100K30 likes2.8k downloads10mo agoHugging Face19Dr3dre /Genius-song-lyrics-cleaned 🎵 Genius Song Lyrics cleaned Dataset Dataset Description This dataset is originally taken from Genius Song Lyrics and it contains cleaned and normalized song lyrics for more than 5 million songs, designed for large-scale topic modeling, clustering, and semantic analysis. The dataset was specifically preprocessed to be compatible with embedding-based models (e.g. Sentence Transformers, BERTopic) while preserving lyrical meaning and thematic content. Repetitive structures… See the full description on the dataset page: https://huggingface.co/datasets/Dr3dre/Genius-song-lyrics-cleaned.tabulartext-classification1M<n<10M5 likes2.7k downloads9mo agoHugging Face20eKaiva /Telegram-Cleaned-DBtext100M<n<1B1 likes2.3k downloads1mo agoHugging Face21kejian /codeparrot-train-more-filter-3.3b-cleanedtabulartext-classification1M<n<10M2 likes2.1k downloads4y agoHugging Face22sunsau91 /Telegram-Cleaned-DBtext100M<n<1B0 likes2.1k downloads1mo agoHugging Face23Chat-Error /book2-lite-cleanedtext10K<n<100K2 likes1.8k downloads3y agoHugging Face24serda-dev /turkish-raw-text-cleaned Turkish Raw Text Cleaned turkish-raw-text-cleaned, Türkçe dil modeli çalışmaları için hazırlanmış temizlenmiş ham metin veri kümesidir. Veri kümesi, turkish-nlp-suite çatısı altında yayımlanan Türkçe metin kaynaklarının temizlenmesi, filtrelenmesi ve model eğitimine daha uygun hale getirilmesiyle oluşturulmuştur. Bu çalışma özellikle Türkçe LLM ön-eğitimi, continual pre-training (CPT), tokenizer analizi, embedding modeli eğitimi, alan bağımsız Türkçe metin modelleme ve veri… See the full description on the dataset page: https://huggingface.co/datasets/serda-dev/turkish-raw-text-cleaned.text-generation1M<n<10M0 likes1.6k downloads3mo agoHugging Face25ada-datadruids /booksummaries_cleanedtext10K<n<100K0 likes1.5k downloads2y agoHugging Face26jed351 /Traditional-Chinese-Common-Crawl-NOT-CleanedCommon Crawl Dumps that were briefly filtered by keywords to remove bad words and simplified Chinese. The hash based cleaned dataset can be found here. Files here are for future usage (downloading from Common Crawl and keyword filtering are very slow) text100M<n<1B0 likes1.4k downloads1y agoHugging Face27ArtificialAnalysis /Earnings22-Cleaned-AA-chunked Earnings22-Cleaned-AA-chunked Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation. The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.audioautomatic-speech-recognitionn<1K1 likes1.3k downloads3mo agoHugging Face28allenai /ultrafeedback_binarized_cleaned Dataset Card for "ultrafeedback_binarized_cleaned" Update 1/12/2023: I've removed examples identified as faulty by Argilla - see their awesome work for more details. This is a version of the UltraFeedback binarized dataset but with TruthfulQA prompts removed and source annotations added (so you can filter out samples from different sources yourself if you want!). Please see the binarized dataset card for more information, or the original UltraFeedback dataset card. tabular100K<n<1M73 likes1.3k downloads3y agoHugging Face29tuetschek /e2e_nlg_cleanedAn update release of E2E NLG Challenge data with cleaned MRs and scripts, accompanying the following paper: Ondřej Dušek, David M. Howcroft, and Verena Rieser (2019): Semantic Noise Matters for Neural Natural Language Generation. In INLG, Tokyo, Japan.10K<n<100K3 likes1.3k downloads3y agoHugging Face30niccogreek /nmr-canonical-cleaned Canonical NMR Dataset Collection — Data Card Dataset release: v4 Canonical schema: v2 Spectral modalities: 1H and 13C resonance-level peak lists Collection overview This release brings several of the largest openly available processed NMR corpora used by current deep-learning methods into one model-independent schema. It combines simulated and literature-derived spectra while preserving the provenance and annotation coverage of every source. The collection has… See the full description on the dataset page: https://huggingface.co/datasets/niccogreek/nmr-canonical-cleaned.feature-extraction100M<n<1B0 likes1.3k downloads8d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.