Team Ai
Datasetpublic

OptimalScale/ClimbLab

ClimbLab is a high-quality pre-training corpus released by NVIDIA. Here is the description: ClimbLab is a filtered 1.2-trillion-token corpus with 20 clusters. Based on Nemotron-CC and SmolLM-Corpus, we employed our proposed CLIMB-clustering to semantically reorganize and filter this combined dataset into 20 distinct clusters, leading to a 1.2-trillion-token high-quality corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we applied two… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbLab.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
16likes6.8kdownloads

OptimalScale/ClimbLab · main · files are served by the source, never re-hosted here