Team Ai
30 results

CL

HPLT /HPLT2.0_cleanedNB: HPLT2.0 is now superseded by a newer release: HPLT3.0 We recommed switching to v3.0, unless you have a compelling reason to stay on 2.0. This is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project. The source of the data is mostly Internet Archive with some additions from Common Crawl. For a detailed description of the dataset, please refer to our website and our pre-print. The Cleaned variant of HPLT Datasets v2.0 This is… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/HPLT2.0_cleaned.tabularfill-mask1B<n<10B46 likes187k downloads4mo agoHugging Faceclip-benchmark /wds_objectnetimage1K<n<10K4 likes81k downloads4y agoHugging Facejhu-clsp /ettin-pretraining-data Ettin Pre-training Data Phase 1 of 3: Diverse pre-training data mixture (1.7T tokens) used to train the Ettin model suite. This dataset contains the pre-training phase data used to train all Ettin encoder and decoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository. 📊 Data Composition Data Source Tokens (B) Percentage Description DCLM 837.2 49.1% High-quality web crawl data CC Head 356.6… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/ettin-pretraining-data.text-generation10 likes53k downloads1y agoHugging Facerobbyant /robotwin-clean-and-aug-lerobot Robotwin Dataset Robotwin dataset in Lerobot format, with video latents already extracted in WAN 2.2 format, ready for use in Lingbot-VA post-training. License Agreement This project is licensed under the CC BY-NC-SA 4.0. video10K<n<100K18 likes49k downloads8mo agoHugging Facekarpathy /climbmix-400b-shuffle73 likes45k downloads7mo agoHugging FaceBVRA /animal-clef-2026 AnimalCLEF26 Kaggle Competition Dataset This is a HuggingFace mirror of the official AnimalCLEF26 competition dataset. Images have been repackaged into one zipfile per split, which include additional metadata that makes the dataset easier to use with HuggingFace. Otherwise, no files have been changed. Loading from datasets import load_dataset dataset = load_dataset("BVRA/animal-clef-2026") print(dataset["train"][0]["image"]) Documentation For… See the full description on the dataset page: https://huggingface.co/datasets/BVRA/animal-clef-2026.imageimage-classification10K<n<100K0 likes43k downloads3mo agoHugging Face

People

Projects