Team Ai
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yordanoswuletaw /amharic-pretraining-corpusAmharic Pretraining Corpus is a large-scale dataset (~103M) for general amharic language pretraining tasks. It consists of diverse text sources, including news articles, books, social media posts, government documents, and web content, all written in Amharic. You can load the dataset as follows from datasets import load_dataset ds = load_dataset("yordanoswuletaw/amharic-pretraining-corpus") texttext-generation100M<n<1B4 likes295 downloads2y agoHugging Face02MichelNivard /proteinLM-mixed-pretraining-v1 Pretraining mix for Protein language models In order to construct a solid pre-training data mixture for protein language models we sample a mix of proteins from 3 sources: MG_Prot50 (https://huggingface.co/datasets/tattabio/OMG_prot50): meta-genomic proteins created by clustering the Open MetaGenomic dataset (OMG) at 50% sequence identity. UniRef50: UniProt proteins from across all species clustered to 50% sequences identity, downloaded form UniProt on 31th of March 2025… See the full description on the dataset page: https://huggingface.co/datasets/MichelNivard/proteinLM-mixed-pretraining-v1.texttoken-classification100M<n<1B0 likes200 downloads2y agoHugging Face03Gunulhona /pretraining_datasettext100K<n<1M0 likes69 downloads3y agoHugging Face04CRUISEResearchGroup /CGM-JEPA-Pretraining CGM-JEPA Pretraining Corpus Continuous glucose monitor (CGM) time-series corpus used for self-supervised pretraining of CGM-JEPA, X-CGM-JEPA, GluFormer, and TS2Vec encoders in the paper CGM-JEPA: Learning Consistent Continuous Glucose Monitor Representations via Predictive Self-Supervised Pretraining. Code: https://github.com/cruiseresearchgroup/CGM-JEPA Pretraining-only corpus. For the labeled downstream-evaluation cohorts (insulin resistance and β-cell dysfunction classification)… See the full description on the dataset page: https://huggingface.co/datasets/CRUISEResearchGroup/CGM-JEPA-Pretraining.texttime-series-forecasting100K<n<1M0 likes62 downloads5mo agoHugging Face05open-paws /continued-pretraining-llama-format Open Paws Continued Pretraining Llama Format Overview This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Specialized Data Format: CSV (Comma-separated values) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/continued-pretraining-llama-format.texttext-generation10K<n<100K2 likes54 downloads1y agoHugging Face06ChengsenWang /GenoJEPA-Pretraining GenoJEPA-Pretraining This dataset provides the pre-training resources used for GenoJEPA, a genomic representation learning framework based on joint-embedding predictive architecture. GenoJEPA learns semantic representations of DNA sequences by shifting the optimization target from nucleotide-level reconstruction to latent-space semantic alignment. The pre-training data is used to construct global and local sequence views for self-supervised genomic representation learning.… See the full description on the dataset page: https://huggingface.co/datasets/ChengsenWang/GenoJEPA-Pretraining.tabularn<1K0 likes52 downloads1mo agoHugging Face07ChallengerSpaceShuttle /zulu-pretraining-datasetThis is IsiZulu Pretraining Dataset. The dataset was used to pre-train BafoGPT-3B Books: Zulu-English Dictionary – A dictionary offering Zulu terms with English definitions, ideal for teaching basic word mappings. Translation: South African Government Speeches – Official speeches in Zulu, which help the model understand structured Zulu sentences and phrases. Transcription: Zulu Community Corpus – A collection of transcriptions, exposing the model to real-life conversational Zulu. Document:… See the full description on the dataset page: https://huggingface.co/datasets/ChallengerSpaceShuttle/zulu-pretraining-dataset.texttext-generationn<1K3 likes26 downloads2y agoHugging Face08jonghyunlee1993 /ChEMBL_v33_pretrainingtext1M<n<10M0 likes21 downloads3y agoHugging Face09TannerGladson /chess-roberta-pretraining-sansconfigs: config_name: default data_files: split: train path: train/*.csv split: eval path: eval/*.csv tabular100M<n<1B0 likes18 downloads2y agoHugging Face10xin1997 /bugfix_pretrainingtext100K<n<1M2 likes14 downloads3y agoHugging Face11Disclosures-SSRC /Detecting-Access-Violations-in-a-LLMs-Pre-Training-Data Beyond Public Access in LLM Pre-Training Data The official HuggingFace repository for the paper "Beyond Public Access in LLM Pre-Training Data" by The AI Disclosures Project. Using a legally obtained dataset of 34 copyrighted O'Reilly Media books, we apply the DE-COP membership inference attack method to investigate whether OpenAI's large language models were trained on copyrighted content without consent. tabular100K<n<1M0 likes14 downloads11mo agoHugging Face12Gaoj124 /pretraining_synthetic_long_100textn<1K0 likes6 downloads3y agoHugging Face13Gaoj124 /pretraining_synthetic_longtextn<1K0 likes5 downloads3y agoHugging Face14Gaoj124 /pretraining_synthetic_shorttextn<1K0 likes4 downloads3y agoHugging Face15Gaoj124 /pretraining_synthetic_long_100_5_real_random_examplestextn<1K0 likes4 downloads3y agoHugging Face16Gaoj124 /pretraining_synthetic_long_100_1_real_random_examplestextn<1K0 likes4 downloads3y agoHugging Face17Haxirus /rasbt_pretrainingtext10K<n<100K0 likes4 downloads2y agoHugging Face18xin1997 /vrepair_pretraining_datatext100K<n<1M0 likes3 downloads3y agoHugging Face19Gaoj124 /pretraining_synthetic_long_100_5_real_examplestextn<1K0 likes3 downloads3y agoHugging Face20ananttrivedi /hinglish_pretraining_datasettext100K<n<1M0 likes3 downloads2y agoHugging Face21Gaoj124 /pretraining_synthetic_short_100_placetextn<1K0 likes2 downloads3y agoHugging Face22ThomasTheMaker /ikigai-pretraining-corpustext10K<n<100K0 likes1 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.