Team Ai
Datasetpublic

OptimalScale/ClimbLab

ClimbLab is a high-quality pre-training corpus released by NVIDIA. Here is the description: ClimbLab is a filtered 1.2-trillion-token corpus with 20 clusters. Based on Nemotron-CC and SmolLM-Corpus, we employed our proposed CLIMB-clustering to semantically reorganize and filter this combined dataset into 20 distinct clusters, leading to a 1.2-trillion-token high-quality corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we applied two… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbLab.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
16likes6.8kdownloads
settings

This repository belongs to OptimalScale on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameClimbLab
visibilitypublic
licenceapache-2.0
gatedno
ownerOptimalScale
Account settings
OptimalScale/ClimbLab · Team Ai