Team Ai
Datasetpublic

KantaHayashiAI/ClimbLab-Ja

Japanese / 日本語版 ClimbLab-Ja ClimbLab-Ja is a high-quality 300-billion-token Japanese corpus with 20 clusters. It is a Japanese adaptation of the nvidia/Nemotron-ClimbLab approach. Based on LLM-jp Corpus v4, we semantically reorganized and filtered the dataset into 20 distinct clusters, resulting in a high-quality 300-billion-token corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we assigned six scores from 0 to 5 to each group… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbLab-Ja.

sourceHugging Faceodc-byupdated 16d agoView on Hugging Face
2likes2.8kdownloads
settings

This repository belongs to KantaHayashiAI on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameClimbLab-Ja
visibilitypublic
licenceodc-by
gatedno
ownerKantaHayashiAI
Account settings
KantaHayashiAI/ClimbLab-Ja · Team Ai